Integrating Supertag Features into Neural Discontinuous Constituent Parsing
–arXiv.org Artificial Intelligence
Syntactic parsing is essential in natural-language processing, with constituent structure being one widely used description of syntax. Traditional views of constituency demand that constituents consist of adjacent words, but this poses challenges in analysing syntax with non-local dependencies, common in languages like German. Therefore, in a number of treebanks like NeGra and TIGER for German and DPTB for English, long-range dependencies are represented by crossing edges. Various grammar formalisms have been used to describe discontinuous trees - often with high time complexities for parsing. Transition-based parsing aims at reducing this factor by eliminating the need for an explicit grammar. Instead, neural networks are trained to produce trees given raw text input using supervised learning on large annotated corpora. An elegant proposal for a stack-free transition-based parser developed by Coavoux and Cohen (2019) successfully allows for the derivation of any discontinuous constituent tree over a sentence in worst-case quadratic time. The purpose of this work is to explore the introduction of supertag information into transition-based discontinuous constituent parsing. In lexicalised grammar formalisms like CCG (Steedman, 1989) informative categories are assigned to the words in a sentence and act as the building blocks for composing the sentence's syntax. These supertags indicate a word's structural role and syntactic relationship with surrounding items. The study examines incorporating supertag information by using a dedicated supertagger as additional input for a neural parser (pipeline) and by jointly training a neural model for both parsing and supertagging (multi-task). In addition to CCG, several other frameworks (LTAG-spinal, LCFRS) and sequence labelling tasks (chunking, dependency parsing) will be compared in terms of their suitability as auxiliary tasks for parsing.
arXiv.org Artificial Intelligence
Oct-11-2024
- Country:
- North America
- Dominican Republic (0.04)
- United States
- Maryland > Baltimore (0.04)
- New Jersey (0.04)
- Texas > Travis County
- Austin (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.13)
- Nevada > Clark County
- Las Vegas (0.04)
- Ohio > Franklin County
- Columbus (0.04)
- New York
- Richmond County > New York City (0.04)
- Queens County > New York City (0.04)
- New York County > New York City (0.04)
- Kings County > New York City (0.04)
- Bronx County > New York City (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Pennsylvania > Philadelphia County
- Philadelphia (0.13)
- Oregon > Multnomah County
- Portland (0.04)
- Illinois > Cook County
- Chicago (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- Colorado
- Denver County > Denver (0.04)
- Boulder County > Boulder (0.04)
- Massachusetts
- Suffolk County > Boston (0.04)
- Middlesex County
- California
- San Diego County > San Diego (0.04)
- Monterey County > Pacific Grove (0.04)
- Santa Clara County
- Los Angeles County
- Los Angeles (0.13)
- Long Beach (0.04)
- Wisconsin > Milwaukee County
- Milwaukee (0.04)
- Canada > British Columbia
- Europe
- Switzerland (0.04)
- Czechia > Prague (0.04)
- Iceland > Capital Region
- Reykjavik (0.04)
- Germany
- Berlin (0.04)
- Hamburg (0.04)
- North Rhine-Westphalia > Düsseldorf Region
- Düsseldorf (0.13)
- Baden-Württemberg > Tübingen Region
- Tübingen (0.04)
- Spain
- Canary Islands (0.04)
- Valencian Community > Valencia Province
- Valencia (0.04)
- Catalonia > Barcelona Province
- Barcelona (0.04)
- Andalusia > Granada Province
- Granada (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- United Kingdom > England
- Oxfordshire > Oxford (0.13)
- Cambridgeshire > Cambridge (0.13)
- Greece > Attica
- Athens (0.04)
- Portugal > Lisbon
- Lisbon (0.04)
- France
- Île-de-France > Paris
- Paris (0.04)
- Provence-Alpes-Côte d'Azur > Bouches-du-Rhône
- Marseille (0.04)
- Île-de-France > Paris
- Italy
- Netherlands > South Holland
- Dordrecht (0.04)
- Finland > Uusimaa
- Helsinki (0.04)
- Sweden
- Vaestra Goetaland > Gothenburg (0.04)
- Uppsala County > Uppsala (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Asia
- North America
- Genre:
- Overview (1.00)
- Research Report
- New Finding (0.45)
- Experimental Study (0.33)
- Industry:
- Government (0.67)
- Technology: