BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoning
Nguyen, Ha-Thanh, Liu, Chaoran, Liu, Qianying, Tachibana, Hideyuki, Noe, Su Myat, Miyao, Yusuke, Takeda, Koichi, Kurohashi, Sadao
–arXiv.org Artificial Intelligence
We present BIS Reasoning 1.0, the first large-scale Japanese dataset of syllogistic reasoning problems explicitly designed to evaluate belief-inconsistent reasoning in large language models (LLMs). Unlike prior datasets such as NeuBAROCO and JFLD, which focus on general or belief-aligned reasoning, BIS Reasoning 1.0 introduces logically valid yet belief-inconsistent syllogisms to uncover reasoning biases in LLMs trained on human-aligned corpora. We benchmark state-of-the-art models - including GPT models, Claude models, and leading Japanese LLMs - revealing significant variance in performance, with GPT-4o achieving 79.54% accuracy. Our analysis identifies critical weaknesses in current LLMs when handling logically valid but belief-conflicting inputs. These findings have important implications for deploying LLMs in high-stakes domains such as law, healthcare, and scientific literature, where truth must override intuitive belief to ensure integrity and safety.
arXiv.org Artificial Intelligence
Jul-15-2025
- Country:
- Europe > France (0.28)
- Asia > Middle East
- UAE (0.28)
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Health & Medicine (0.48)
- Technology: