Large Language Model
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
Mishra, Manit, Braham, Abderrahman, Marsom, Charles, Chung, Bryan, Griffin, Gavin, Sidnerlikar, Dakshesh, Sarin, Chatanya, Rajaram, Arjun
Conventional processes for analyzing datasets and extracting meaningful information are often time-consuming and laborious. Previous work has identified manual, repetitive coding and data collection as major obstacles that hinder data scientists from undertaking more nuanced labor and high-level projects. To combat this, we evaluated OpenAI's GPT-3.5 as a "Language Data Scientist" (LDS) that can extrapolate key findings, including correlations and basic information, from a given dataset. The model was tested on a diverse set of benchmark datasets to evaluate its performance across multiple standards, including data science code-generation based tasks involving libraries such as NumPy, Pandas, Scikit-Learn, and TensorFlow, and was broadly successful in correctly answering a given data science query related to the benchmark dataset. The LDS used various novel prompt engineering techniques to effectively answer a given question, including Chain-of-Thought reinforcement and SayCan prompt engineering. Our findings demonstrate great potential for leveraging Large Language Models for low-level, zero-shot data analysis.
LayerNorm: A key component in parameter-efficient fine-tuning
ValizadehAslani, Taha, Liang, Hualou
Fine-tuning a pre-trained model, such as Bidirectional Encoder Representations from Transformers (BERT), has been proven to be an effective method for solving many natural language processing (NLP) tasks. However, due to the large number of parameters in many state-of-the-art NLP models, including BERT, the process of fine-tuning is computationally expensive. One attractive solution to this issue is parameter-efficient fine-tuning, which involves modifying only a minimal segment of the model while keeping the remainder unchanged. Yet, it remains unclear which segment of the BERT model is crucial for fine-tuning. In this paper, we first analyze different components in the BERT model to pinpoint which one undergoes the most significant changes after fine-tuning. We find that output LayerNorm changes more than any other components when fine-tuned for different General Language Understanding Evaluation (GLUE) tasks. Then we show that only fine-tuning the LayerNorm can reach comparable, or in some cases better, performance to full fine-tuning and other parameter-efficient fine-tuning methods. Moreover, we use Fisher information to determine the most critical subset of LayerNorm and demonstrate that many NLP tasks in the GLUE benchmark can be solved by fine-tuning only a small portion of LayerNorm with negligible performance degradation.
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Stojkovic, Jovan, Choukse, Esha, Zhang, Chaojie, Goiri, Inigo, Torrellas, Josep
With the ubiquitous use of modern large language models (LLMs) across industries, the inference serving for these models is ever expanding. Given the high compute and memory requirements of modern LLMs, more and more top-of-the-line GPUs are being deployed to serve these models. Energy availability has come to the forefront as the biggest challenge for data center expansion to serve these models. In this paper, we present the trade-offs brought up by making energy efficiency the primary goal of LLM serving under performance SLOs. We show that depending on the inputs, the model, and the service-level agreements, there are several knobs available to the LLM inference provider to use for being energy efficient. We characterize the impact of these knobs on the latency, throughput, as well as the energy. By exploring these trade-offs, we offer valuable insights into optimizing energy usage without compromising on performance, thereby paving the way for sustainable and cost-effective LLM deployment in data center environments.
Incentivizing News Consumption on Social Media Platforms Using Large Language Models and Realistic Bot Accounts
Askari, Hadi, Chhabra, Anshuman, von Hohenberg, Bernhard Clemm, Heseltine, Michael, Wojcieszak, Magdalena
Polarization, declining trust, and wavering support for democratic norms are pressing threats to U.S. democracy. Exposure to verified and quality news may lower individual susceptibility to these threats and make citizens more resilient to misinformation, populism, and hyperpartisan rhetoric. This project examines how to enhance users' exposure to and engagement with verified and ideologically balanced news in an ecologically valid setting. We rely on a large-scale two-week long field experiment (from 1/19/2023 to 2/3/2023) on 28,457 Twitter users. We created 28 bots utilizing GPT-2 that replied to users tweeting about sports, entertainment, or lifestyle with a contextual reply containing two hardcoded elements: a URL to the topic-relevant section of quality news organization and an encouragement to follow its Twitter account. To further test differential effects by gender of the bots, treated users were randomly assigned to receive responses by bots presented as female or male. We examine whether our over-time intervention enhances the following of news media organization, the sharing and the liking of news content and the tweeting about politics and the liking of political content. We find that the treated users followed more news accounts and the users in the female bot treatment were more likely to like news content than the control. Most of these results, however, were small in magnitude and confined to the already politically interested Twitter users, as indicated by their pre-treatment tweeting about politics. These findings have implications for social media and news organizations, and also offer direction for future work on how Large Language Models and other computational interventions can effectively enhance individual on-platform engagement with quality news and public affairs.
Causal Inference for Human-Language Model Collaboration
Zhang, Bohan, Wang, Yixin, Dhillon, Paramveer S.
In this paper, we examine the collaborative dynamics between humans and language models (LMs), where the interactions typically involve LMs proposing text segments and humans editing or responding to these proposals. Productive engagement with LMs in such scenarios necessitates that humans discern effective text-based interaction strategies, such as editing and response styles, from historical human-LM interactions. This objective is inherently causal, driven by the counterfactual `what-if' question: how would the outcome of collaboration change if humans employed a different text editing/refinement strategy? A key challenge in answering this causal inference question is formulating an appropriate causal estimand: the conventional average treatment effect (ATE) estimand is inapplicable to text-based treatments due to their high dimensionality. To address this concern, we introduce a new causal estimand -- Incremental Stylistic Effect (ISE) -- which characterizes the average impact of infinitesimally shifting a text towards a specific style, such as increasing formality. We establish the conditions for the non-parametric identification of ISE. Building on this, we develop CausalCollab, an algorithm designed to estimate the ISE of various interaction strategies in dynamic human-LM collaborations. Our empirical investigations across three distinct human-LM collaboration scenarios reveal that CausalCollab effectively reduces confounding and significantly improves counterfactual estimation over a set of competitive baselines.
Data Quality May Be All You Need
History has a lesson for the development of artificial intelligence (AI): when in doubt, make it bigger. In "The Bitter Lesson," he argued that over its 70-year history, AI has succeeded when it has exploited available computing power. A series of papers published during the past decade that analyzed deep learning performance have confirmed the powerful effects of scaling up model size. This process accelerated in the wake of Google's development of the Transformer architecture for the BERT large language models (LLMs). Model size, measured by the number of stored neural weights, ballooned in just five years. From BERT's 340 million parameters, today's largest implementations, known as frontier models, such as OpenAI's GPT-4 have pushed beyond a trillion.
Google reverses course and brings its Gemini AI to the regular Pixel 8
Google will bring Gemini, the company's new large language model, to Pixel 8 smartphones after all. The phone will incorporate Gemini Nano, a version of the model built to run locally on personal devices. This follows a successful rollout to the Pixel 8 Pro late last year and the Samsung Galaxy S24 in January. The Pixel 8 features the same proprietary Tensor G3 chip as the Pro, which was designed to speed up AI performance. So the overall experience should be similar with both gadgets.
Microsoft Copilot AI will soon run locally on PCs
ChromeOS and macOS both use NPU power for more video and audio processing features, though, along with OCR, translation, live transcription and more, Ars Technica noted. So far, the processor with the fastest NPU speed is Apple M3, which offers 18 TOPS across the lineup (M3, M3 Pro and M3 Ultra). AMD's Ryzen 8040 and 7040 laptop chips are next with 16 and 10 TOPS respectively, while Intel's Meteor Lake laptop hits 10 TOPS as well. Qualcomm may offer the first processor with enough power for Copilot via the Snapdragon X Elite, which will offer 45 TOPS of AI compute speed. Intel's Lunar Lake chips, set to arrive in 2025, will ship with triple its current NPU speeds. Yesterday, the company introduced 300 new AI features optimized specifically for its own OpenVino platform. The chip giant also announced an AI PC development kit based on the the ASUS NUC Pro that uses its current Meteor Lake silicon. "From a desktop standpoint, we have plans on the desktop side, what we would say [is an] AI PC. And then there's also the next-gen AI PC, the 40 TOPS requirements; we have all of our different steps in our roadmap on how we cover all the different segments," the company told Tom's Hardware.
New AI test measures how fast robots can respond to user commands
WEHEAD connects to ChatGPT and displays a face, expressions and voice. Artificial intelligence benchmarking group MLCommons on Wednesday released a fresh set of tests and results that rate the speed at which top-of-the-line hardware can run AI applications and respond to users. The two new benchmarks added by MLCommons measure the speed at which the AI chips and systems can generate responses from the powerful AI models packed with data. The results roughly demonstrate to how quickly an AI application such as ChatGPT can deliver a response to a user query. One of the new benchmarks added the capability to measure the speediness of a question-and-answer scenario for large language models.
RouterBench: A Benchmark for Multi-LLM Routing System
Hu, Qitian Jason, Bieker, Jacob, Li, Xiuyu, Jiang, Nan, Keigwin, Benjamin, Ranganath, Gaurav, Keutzer, Kurt, Upadhyay, Shriyash Kaustubh
As the range of applications for Large Language Models (LLMs) continues to grow, the demand for effective serving solutions becomes increasingly critical. Despite the versatility of LLMs, no single model can optimally address all tasks and applications, particularly when balancing performance with cost. This limitation has led to the development of LLM routing systems, which combine the strengths of various models to overcome the constraints of individual LLMs. Yet, the absence of a standardized benchmark for evaluating the performance of LLM routers hinders progress in this area. To bridge this gap, we present RouterBench, a novel evaluation framework designed to systematically assess the efficacy of LLM routing systems, along with a comprehensive dataset comprising over 405k inference outcomes from representative LLMs to support the development of routing strategies. We further propose a theoretical framework for LLM routing, and deliver a comparative analysis of various routing approaches through RouterBench, highlighting their potentials and limitations within our evaluation framework. This work not only formalizes and advances the development of LLM routing systems but also sets a standard for their assessment, paving the way for more accessible and economically viable LLM deployments. The code and data are available at https://github.com/withmartian/routerbench.