Towards Optimal Caching and Model Selection for Large Model Inference

May-27-2025, 09:14:42 GMT–Neural Information Processing Systems

Large Language Models (LLMs) and other large foundation models have achieved impressive results, but their size exacerbates existing resource consumption and latency challenges. In particular, the large-scale deployment of these models is hindered by the significant resource requirements during inference. In this paper, we study two approaches for mitigating these challenges: employing a cache to store previous queries and learning a model selector to choose from an ensemble of models for query processing.Theoretically, we provide an optimal algorithm for jointly optimizing both approaches to reduce the inference cost in both offline and online tabular settings. By combining a caching algorithm, namely Greedy Dual Size with Frequency (GDSF) or Least Expected Cost (LEC), with a model selector, we achieve optimal rates in both offline and online settings. Empirically, simulations show that our caching and model selection algorithm greatly improves over the baselines, with up to 50\times improvement over the baseline when the ratio between the maximum cost and minimum cost is 100 .

large language model, machine learning, natural language, (9 more...)

Neural Information Processing Systems

May-27-2025, 09:14:42 GMT

Conferences Web Page

Add feedback

Country:
- Asia > Middle East > Jordan (0.08)

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning > Statistical Learning (0.64)
  - Natural Language > Large Language Model (0.62)