Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
Lior, Gili, Goldberg, Yoav, Stanovsky, Gabriel
–arXiv.org Artificial Intelligence
Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both human users and structure-aware models. We propose to identify the typical structure of document within a collection, which requires to capture recurring topics across the collection, while abstracting over arbitrary header paraphrases, and ground each topic to respective document locations. These requirements pose several challenges: headers that mark recurring topics frequently differ in phrasing, certain section headers are unique to individual documents and do not reflect the typical structure, and the order of topics can vary between documents. Subsequently, we develop an unsupervised graph-based method which leverages both inter- and intra-document similarities, to extract the underlying collection-wide structure. Our evaluations on three diverse domains in both English and Hebrew indicate that our method extracts meaningful collection-wide structure, and we hope that future work will leverage our method for multi-document applications and structure-aware models.
arXiv.org Artificial Intelligence
Jun-20-2024
- Country:
- Asia
- China > Hong Kong (0.04)
- Middle East
- Israel > Jerusalem District
- Jerusalem (0.04)
- Jordan (0.04)
- UAE > Abu Dhabi Emirate
- Abu Dhabi (0.04)
- Israel > Jerusalem District
- Singapore (0.04)
- Europe
- Bulgaria (0.04)
- Croatia > Dubrovnik-Neretva County
- Dubrovnik (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Italy > Tuscany
- Florence (0.04)
- North America
- Canada > Ontario
- Toronto (0.04)
- Dominican Republic (0.04)
- United States > New York
- New York County > New York City (0.04)
- Canada > Ontario
- Asia
- Genre:
- Research Report (0.40)
- Industry:
- Health & Medicine > Therapeutic Area
- Psychiatry/Psychology (0.46)
- Law > Business Law (0.46)
- Health & Medicine > Therapeutic Area
- Technology: