TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
Khatun, Aisha, Brown, Daniel G.
–arXiv.org Artificial Intelligence
The typical benchmark evaluations have begun to However, it remains unclear if the model's fall short and do not cover the nuances of LLMs' responses bear useful meaning - whether the model abilities (Zoph et al., 2022). Did the model provide understands the topic or is responding probabilistically a certain answer simply because of the huge purely based on training data. TruthfulQA amount of similar text it saw during training? Or (Lin et al., 2021) comes close to assessing a did the model register a piece of knowledge and use model's understanding of the world but it is designed that to answer the question? It is impossible to tell to exploit the imitative weaknesses of models them apart without analyzing the training dataset, and relies on a model's elaborate response and which, given the current trend, is not available for text-matching metrics. In contrast, our work intends most models. Current RAG (Retrieval Augmented to extract knowledge and understanding from Generation) systems rely on LLM's prompt memory LLMs without intentionally tricking or confusing to register some facts and expect the model the model.
arXiv.org Artificial Intelligence
Jun-3-2024
- Country:
- Asia > China (0.04)
- North America
- United States > New York
- New York County > New York City (0.04)
- Canada > Ontario
- Waterloo Region > Waterloo (0.04)
- Toronto (0.04)
- United States > New York
- Europe
- Genre:
- Research Report (0.50)
- Industry:
- Technology: