BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
Zhou, Jinan, Ghosh, Rajat, Bhargava, Vaishnavi, Dutta, Debojyoti, Singhal, Aryan
–arXiv.org Artificial Intelligence
When designing LLM services, practitioners care about three key properties: inference-time budget, factual authenticity, and reasoning capacity. However, our analysis shows that no model can simultaneously optimize for all three. We formally prove this trade-off and propose a principled framework named The BAR Theorem for LLM-application design. Large language models (LLMs) with the Transformer architecture (V aswani et al., 2017) are pre-trained with a massive number of tokens to be auto-regressive next token predictors, null While authenticity and reasoning are qualitative measures evaluating the capabilities of LLMs, inference overhead is a budget metric that measures the operational cost of deploying LLMs. Inference for large language models (LLMs) is primarily a memory-bound process due to several factors: these models typically contain billions or even trillions of parameters, requiring constant memory accesses during inference (Brown et al., 2020; Chowdhery et al., 2022).
arXiv.org Artificial Intelligence
Aug-5-2025