HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction

Sarmah, Bhaskarjit, Hall, Benika, Rao, Rohan, Patel, Sunil, Pasquali, Stefano, Mehta, Dhagash

Aug-9-2024–arXiv.org Machine Learning

Although LLMs have substantial potential in financial applications, there are notable challenges in using pre-trained models to Extraction and interpretation of intricate information from unstructured extract information from financial documents outside their training text data arising in financial applications, such as earnings data while also reducing hallucination [7, 8]. Financial documents call transcripts, present substantial challenges to large language typically contain domain-specific language, multiple data formats, models (LLMs) even using the current best practices to use Retrieval and unique contextual relationships that general purpose-trained Augmented Generation (RAG) (referred to as VectorRAG LLMs do not handle well. In addition, extracting consistent and techniques which utilize vector databases for information retrieval) coherent information from multiple financial documents can be due to challenges such as domain specific terminology and complex challenging due to variations in terminology, format, and context formats of the documents. We introduce a novel approach based across different textual sources. The specialized terminology and on a combination, called HybridRAG, of the Knowledge Graphs complex data formats in financial documents make it difficult for (KGs) based RAG techniques (called GraphRAG) and VectorRAG models to extract meaningful insights, in turn, causing inaccurate techniques to enhance question-answer (Q&A) systems for information predictions, overlooked insights, and unreliable analysis, which extraction from financial documents that is shown to be ultimately hinder the ability to make well-informed decisions.

hybridrag, information, vectorrag, (11 more...)

arXiv.org Machine Learning

Aug-9-2024

arXiv.org PDF

Add feedback

Country:
- Asia > India (0.04)
- North America > United States
  - New York > New York County
    - New York City (0.04)
  - California > Santa Clara County
    - Santa Clara (0.04)

Genre:
- Overview (0.88)
- Research Report > Promising Solution (0.34)

Industry:
- Banking & Finance > Trading (1.00)

Technology:
- Information Technology > Artificial Intelligence
  - Representation & Reasoning (1.00)
  - Natural Language > Large Language Model (1.00)
  - Machine Learning > Neural Networks
    - Deep Learning (0.84)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found