BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning

Zhou, Jinan, Ghosh, Rajat, Bhargava, Vaishnavi, Dutta, Debojyoti, Singhal, Aryan

arXiv.org Artificial Intelligence 

When designing LLM services, practitioners care about three key properties: inference-time budget, factual authenticity, and reasoning capacity. However, our analysis shows that no model can simultaneously optimize for all three. We formally prove this trade-off and propose a principled framework named The BAR Theorem for LLM-application design. Large language models (LLMs) with the Transformer architecture (V aswani et al., 2017) are pre-trained with a massive number of tokens to be auto-regressive next token predictors, null While authenticity and reasoning are qualitative measures evaluating the capabilities of LLMs, inference overhead is a budget metric that measures the operational cost of deploying LLMs. Inference for large language models (LLMs) is primarily a memory-bound process due to several factors: these models typically contain billions or even trillions of parameters, requiring constant memory accesses during inference (Brown et al., 2020; Chowdhery et al., 2022).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found