Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition
Azin, Tara, Dumitrescu, Daniel, Inkpen, Diana, Singh, Raj
–arXiv.org Artificial Intelligence
Natural Language Inference (NLI) is the task of determining whether a sentence pair represents entailment, contradiction, or a neutral relationship. While NLI models perform well on many inference tasks, their ability to handle fine-grained pragmatic inferences, particularly presupposition in conditionals, remains underexplored. In this study, we introduce CONFER, a novel dataset designed to evaluate how NLI models process inference in conditional sentences. We assess the performance of four NLI models, including two pre-trained models, to examine their generalization to conditional reasoning. Additionally, we evaluate Large Language Models (LLMs), including GPT-4o, LLaMA, Gemma, and DeepSeek-R1, in zero-shot and few-shot prompting settings to analyze their ability to infer presuppositions with and without prior context. Our findings indicate that NLI models struggle with presuppositional reasoning in conditionals, and fine-tuning on existing NLI datasets does not necessarily improve their performance.
arXiv.org Artificial Intelligence
Jun-9-2025
- Country:
- Europe
- Denmark > Capital Region
- Copenhagen (0.04)
- Italy > Tuscany
- Florence (0.04)
- Denmark > Capital Region
- North America
- Canada > Ontario
- National Capital Region > Ottawa (0.04)
- Toronto (0.04)
- United States > Massachusetts
- Hampshire County > Amherst (0.04)
- Canada > Ontario
- Oceania > Australia
- Europe
- Genre:
- Research Report > New Finding (1.00)
- Technology: