From Lengthy to Lucid: A Systematic Literature Review on NLP Techniques for Taming Long Sentences

Passali, Tatiana, Chatzikyriakidis, Efstathios, Andreadis, Stelios, Stavropoulos, Thanos G., Matonaki, Anastasia, Fachantidis, Anestis, Tsoumakas, Grigorios

Dec-8-2023–arXiv.org Artificial Intelligence

Long sentences have been a persistent issue in written communication for many years since they make it challenging for readers to grasp the main points or follow the initial intention of the writer. This survey, conducted using the PRISMA guidelines, systematically reviews two main strategies for addressing the issue of long sentences: a) sentence compression and b) sentence splitting. An increased trend of interest in this area has been observed since 2005, with significant growth after 2017. Current research is dominated by supervised approaches for both sentence compression and splitting. Yet, there is a considerable gap in weakly and self-supervised techniques, suggesting an opportunity for further research, especially in domains with limited data. In this survey, we categorize and group the most representative methods into a comprehensive taxonomy. We also conduct a comparative evaluation analysis of these methods on common sentence compression and splitting datasets. Finally, we discuss the challenges and limitations of current methods, providing valuable insights for future research directions. This survey is meant to serve as a comprehensive resource for addressing the complexities of long sentences. We aim to enable researchers to make further advancements in the field until long sentences are no longer a barrier to effective communication.

compression, computational linguistic, sentence compression, (11 more...)

arXiv.org Artificial Intelligence

Dec-8-2023

arXiv.org PDF

Add feedback

Country:
- Oceania > Australia
  - New South Wales > Sydney (0.14)
  - Victoria > Melbourne (0.04)
- North America
  - Dominican Republic (0.04)
  - United States
    - Maryland > Baltimore (0.04)
    - Texas > Travis County
      - Austin (0.04)
    - Michigan > Washtenaw County
      - Ann Arbor (0.04)
    - Minnesota > Hennepin County
      - Minneapolis (0.28)
    - Ohio > Franklin County
      - Columbus (0.04)
    - New York
      - New York County > New York City (0.04)
      - Monroe County > Rochester (0.04)
    - Louisiana > Orleans Parish
      - New Orleans (0.04)
    - Oregon > Multnomah County
      - Portland (0.04)
    - New Mexico > Santa Fe County
      - Santa Fe (0.04)
    - Washington > King County
      - Seattle (0.14)
    - California
      - San Diego County > San Diego (0.04)
      - Los Angeles County
        Los Angeles (0.14)
        Long Beach (0.04)
  - Canada
    - Quebec > Montreal (0.04)
    - British Columbia > Metro Vancouver Regional District
      - Vancouver (0.04)
- Europe
  - Middle East > Malta (0.04)
  - Greece > Central Macedonia
    - Thessaloniki (0.05)
  - Lithuania > Vilnius County
    - Vilnius (0.04)
  - France > Provence-Alpes-Côte d'Azur
    - Bouches-du-Rhône > Marseille (0.04)
  - Denmark > Capital Region
    - Copenhagen (0.04)
  - Bulgaria > Sofia City Province
    - Sofia (0.04)
  - United Kingdom > England
    - Oxfordshire > Oxford (0.04)
    - Greater Manchester > Manchester (0.04)
  - Italy
    - Tuscany
      - Florence (0.04)
      - Pisa Province > Pisa (0.04)
    - Trentino-Alto Adige/Südtirol > Trentino Province
      - Trento (0.04)
  - Spain > Catalonia
    - Barcelona Province > Barcelona (0.04)
  - Sweden > Östergötland County
    - Linköping (0.04)
  - Czechia
    - South Moravian Region > Brno (0.04)
    - Prague (0.04)
  - Ireland > Leinster
    - County Dublin > Dublin (0.04)
  - Belgium > Brussels-Capital Region
    - Brussels (0.04)
  - Portugal
    - Lisbon > Lisbon (0.04)
    - Évora > Évora (0.04)
- Asia
  - South Korea (0.04)
  - Vietnam
    - Quảng Ninh Province > Hạ Long (0.04)
    - Hanoi > Hanoi (0.04)
  - Singapore > Central Region
    - Singapore (0.04)
  - Middle East > UAE
    - Abu Dhabi Emirate > Abu Dhabi (0.04)
  - Japan
    - Honshū > Kantō
      - Tokyo Metropolis Prefecture > Tokyo (0.14)
    - Hokkaidō > Hokkaidō Prefecture
      - Sapporo (0.04)
  - India > Maharashtra
    - Mumbai (0.04)
  - China
    - Hong Kong (0.04)
    - Hubei Province > Wuhan (0.04)
- Africa > Ethiopia
  - Addis Ababa > Addis Ababa (0.04)

Genre:
- Overview (1.00)
- Research Report > New Finding (0.67)

Technology:
- Information Technology > Artificial Intelligence
  - Representation & Reasoning
    - Optimization (1.00)
    - Constraint-Based Reasoning (1.00)
    - Expert Systems (0.67)
  - Natural Language
    - Grammars & Parsing (1.00)
    - Large Language Model (0.68)
    - Text Processing (0.67)
  - Machine Learning
    - Statistical Learning (1.00)
    - Neural Networks > Deep Learning (1.00)