Enhancing Keyphrase Extraction from Long Scientific Documents using Graph Embeddings
Martínez-Cruz, Roberto, Mahata, Debanjan, López-López, Alvaro J., Portela, José
–arXiv.org Artificial Intelligence
In this study, we investigate using graph neural network (GNN) representations to enhance contextualized representations of pre-trained language models (PLMs) for keyphrase extraction from lengthy documents. We show that augmenting a PLM with graph embeddings provides a more comprehensive semantic understanding of words in a document, particularly for long documents. We construct a co-occurrence graph of the text and embed it using a graph convolutional network (GCN) trained on the task of edge prediction. We propose a graph-enhanced sequence tagging architecture that augments contextualized PLM embeddings with graph representations. Evaluating on benchmark datasets, we demonstrate that enhancing PLMs with graph embeddings outperforms state-of-the-art models on long documents, showing significant improvements in F1 scores across all the datasets. Our study highlights the potential of GNN representations as a complementary approach to improve PLM performance for keyphrase extraction from long documents.
arXiv.org Artificial Intelligence
May-16-2023
- Country:
- North America > United States
- Washington > King County
- Seattle (0.04)
- New York > New York County
- New York City (0.04)
- Washington > King County
- Europe
- Spain
- Galicia > Madrid (0.04)
- Catalonia > Barcelona Province
- Barcelona (0.04)
- Norway > Western Norway
- Spain
- Asia
- North America > United States
- Genre:
- Research Report > New Finding (1.00)
- Technology: