Goto

Collaborating Authors

 Large Language Model


How Effective Is Self-Consistency for Long-Context Problems?

arXiv.org Artificial Intelligence

Self-consistency (SC) has been demonstrated to enhance the performance of large language models (LLMs) across various tasks and domains involving short content. However, does this evidence support its effectiveness for long-context problems? This study examines the role of SC in long-context scenarios, where LLMs often struggle with position bias, hindering their ability to utilize information effectively from all parts of their long input context. We examine a range of design parameters, including different models, context lengths, prompt formats, and types of datasets and tasks. Our findings demonstrate that SC, while effective for short-context problems, fundamentally fails for long-context tasks -- not only does it fail to mitigate position bias, but it can also actively degrade performance. We observe that the effectiveness of SC varies with context length and model size but remains mainly unaffected by prompt format or task type. These results provide valuable insight into the limitations of current LLMs in long-context understanding and highlight the need for more sophisticated approaches to address position bias in these models.


Generative Memesis: AI Mediates Political Memes in the 2024 USA Presidential Election

arXiv.org Artificial Intelligence

Visual content on social media has become increasingly influential in shaping political discourse and civic engagement. Using a dataset of 239,526 Instagram images, deep learning, and LLM-based workflows, we examine the impact of different content types on user engagement during the 2024 US presidential Elections, with a focus on synthetic visuals. Results show while synthetic content may not increase engagement alone, it mediates how political information is created through highly effective, often absurd, political memes. We define the notion of generative memesis, where memes are no longer shared person-to-person but mediated by AI through customized, generated images. We also find partisan divergences: Democrats use AI for in-group support whereas Republicans use it for out-group attacks. Non-traditional, left-leaning outlets are the primary creators of political memes; emphasis on different topics largely follows issue ownership.


Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement

arXiv.org Artificial Intelligence

Recent advancements in LLM-based agents have led to significant progress in automatic software engineering, particularly in software maintenance and evolution. Despite these encouraging advances, current research faces two major challenges. First, SOTA performance primarily depends on closed-source models, which significantly limits the technology's accessibility, and potential for customization in diverse SE tasks. Second, these models are predominantly trained on static code data, lacking a deep understanding of the dynamic interactions, iterative problem-solving processes, and evolutionary characteristics inherent in software development. To address these challenges, our study adopts a software engineering perspective. We recognize that real-world software maintenance and evolution processes encompass not only static code data but also developers' thought processes, utilization of external tools, and the interaction between different functional personnel. Consequently, we introduce the Lingma SWE-GPT series, comprising Lingma SWE-GPT 7B and 72B. By learning from and simulating real-world code submission activities, Lingma SWE-GPT systematically incorporates the dynamic interactions and iterative problem-solving inherent in software development process, thereby achieving a more comprehensive understanding of software improvement processes. We conducted experimental evaluations using SWE-bench Verified benchmark. The results demonstrate that Lingma SWE-GPT 72B successfully resolves 30.20% of the GitHub issues, marking a significant improvement in automatic issue resolution (22.76% relative improvement compared to Llama 3.1 405B), approaching the performance of closed-source models (31.80\% issues of GPT-4o resolved). Notably, Lingma SWE-GPT 7B resolves 18.20% of the issues, highlighting the potential for applying smaller models to ASE tasks.


On the Exploration of LM-Based Soft Modular Robot Design

arXiv.org Artificial Intelligence

Recent large language models (LLMs) have demonstrated promising capabilities in modeling real-world knowledge and enhancing knowledge-based generation tasks. In this paper, we further explore the potential of using LLMs to aid in the design of soft modular robots, taking into account both user instructions and physical laws, to reduce the reliance on extensive trial-and-error experiments typically needed to achieve robot designs that meet specific structural or task requirements. Specifically, we formulate the robot design process as a sequence generation task and find that LLMs are able to capture key requirements expressed in natural language and reflect them in the construction sequences of robots. To simplify, rather than conducting real-world experiments to assess design quality, we utilize a simulation tool to provide feedback to the generative model, allowing for iterative improvements without requiring extensive human annotations. Furthermore, we introduce five evaluation metrics to assess the quality of robot designs from multiple angles including task completion and adherence to instructions, supporting an automatic evaluation process. Our model performs well in evaluations for designing soft modular robots with uni- and bi-directional locomotion and stair-descending capabilities, highlighting the potential of using natural language and LLMs for robot design. However, we also observe certain limitations that suggest areas for further improvement.


LLM Chain Ensembles for Scalable and Accurate Data Annotation

arXiv.org Artificial Intelligence

Abstract--The ability of large language models (LLMs) to perform zero-shot classification makes them viable solutions for data annotation in rapidly evolving domains where quality labeled data is often scarce and costly to obtain. However, the large-scale deployment of LLMs can be prohibitively expensive. This paper introduces an LLM chain ensemble methodology that aligns multiple LLMs in a sequence, routing data subsets to subsequent models based on classification uncertainty. This approach leverages the strengths of individual LLMs within a broader system, allowing each model to handle data points where it exhibits the highest confidence, while forwarding more complex cases to potentially more robust models. Our results show that the chain ensemble method often exceeds the performance of the best individual model in the chain and achieves substantial cost savings, making LLM chain ensembles a practical and efficient solution for large-scale data annotation challenges.


Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers

arXiv.org Machine Learning

Per-example gradient norms are a vital ingredient for estimating gradient noise scale (GNS) with minimal variance. Observing the tensor contractions required to compute them, we propose a method with minimal FLOPs in 3D or greater tensor regimes by simultaneously computing the norms while computing the parameter gradients. Using this method we are able to observe the GNS of different layers at higher accuracy than previously possible. We find that the total GNS of contemporary transformer models is predicted well by the GNS of only the normalization layers. As a result, focusing only on the normalization layer, we develop a custom kernel to compute the per-example gradient norms while performing the LayerNorm backward pass with zero throughput overhead. Tracking GNS on only those layers, we are able to guide a practical batch size schedule that reduces training time by 18% on a Chinchilla-optimal language model.


Provable optimal transport with transformers: The essence of depth and prompt engineering

arXiv.org Machine Learning

Can we establish provable performance guarantees for transformers? Establishing such theoretical guarantees is a milestone in developing trustworthy generative AI. In this paper, we take a step toward addressing this question by focusing on optimal transport, a fundamental problem at the intersection of combinatorial and continuous optimization. Leveraging the computational power of attention layers, we prove that a transformer with fixed parameters can effectively solve the optimal transport problem in Wasserstein-2 with entropic regularization for an arbitrary number of points. Consequently, the transformer can sort lists of arbitrary sizes up to an approximation factor. Our results rely on an engineered prompt that enables the transformer to implement gradient descent with adaptive stepsizes on the dual optimal transport. Combining the convergence analysis of gradient descent with Sinkhorn dynamics, we establish an explicit approximation bound for optimal transport with transformers, which improves as depth increases. Our findings provide novel insights into the essence of prompt engineering and depth for solving optimal transport. In particular, prompt engineering boosts the algorithmic expressivity of transformers, allowing them implement an optimization method. With increasing depth, transformers can simulate several iterations of gradient descent.


Apple reports robust demand for iPhone 16 even as overall sales in China slow

The Guardian

Apple reported strong demand for the iPhone 16 in its quarterly earnings report on Thursday, though overall sales in China slightly decreased year-over-year. The company reported 94.9bn in revenue, up 6% year-over-year, and 1.64 in earnings per share (EPS). The company's earnings slightly beat Wall Street projections of 94.4bn in sales and an EPS of 1.60. The company saw 46.2bn in revenue from iPhone sales, up from 43.8bn year-over-year. Fourth-quarter revenue from its services division, which include subscriptions, increased from 22.31bn to 24.97bn year-over-year.


ChatGPT Search will do the legwork for you

Engadget

ChatGPT Search is here to try to combine the best of chatbots and web searches. OpenAI's latest feature searches the web in response to your natural language queries, delivering "fast, timely answers with links to relevant web sources." When using ChatGPT, the bot will search the web depending on what you ask. Or, if you want to manually override its decision-making, you can tap a new web search icon below the input bar. OpenAI says the feature looks for "original, high-quality content from the web," integrating it into its conversational answers.


This Is a Glimpse of the Future of AI Robot

WIRED

The idea of a robot that does a wide range of household chores, from unloading the dryer to folding laundry to cleaning up a messy table, has long seemed like pure science fiction--perhaps most famously embodied by the 1960s fantasy that was Rosey in The Jetsons. Physical Intelligence, a startup in San Francisco, has shown that such a dream might actually not be so far off, demonstrating a single artificial intelligence model that has learned to do a wide range of useful home chores--including all of the above--by being trained on an unprecedented amount of data. The feat raises the prospect of bringing something as magical and generally capable as other AI models like ChatGPT into the physical world. The advent of large language models (LLMs)--general-purpose learning algorithms fed vast swaths of text from books and the internet--has given chatbots vastly more general capabilities. Physical Intelligence aims to create something similarly capable in the physical world by training a similar kind of algorithm with enormous amounts of robotic data instead.