Goto

Collaborating Authors

 Large Language Model


Can LLMs Interpret and Leverage Structured Linguistic Representations? A Case Study with AMRs

arXiv.org Artificial Intelligence

This paper evaluates the ability of Large Language Models (LLMs) to leverage contextual information in the form of structured linguistic representations. Specifically, we examine the impact of encoding both short and long contexts using Abstract Meaning Representation (AMR) structures across a diverse set of language tasks. We perform our analysis using 8-bit quantized and instruction-tuned versions of Llama 3.1 (8B), Phi-3, and Mistral 7B. Our results indicate that, for tasks involving short contexts, augmenting the prompt with the AMR of the original language context often degrades the performance of the underlying LLM. However, for tasks that involve long contexts, such as dialogue summarization in the SAMSum dataset, this enhancement improves LLM performance, for example, by increasing the zero-shot cosine similarity score of Llama 3.1 from 66% to 76%. This improvement is more evident in the newer and larger LLMs, but does not extend to the older or smaller ones. In addition, we observe that LLMs can effectively reconstruct the original text from a linearized AMR, achieving a cosine similarity of 81% in the best-case scenario.


OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have transformed software development by enabling code generation, automated debugging, and complex reasoning. However, their continued advancement is constrained by the scarcity of high-quality, publicly available supervised fine-tuning (SFT) datasets tailored for coding tasks. To bridge this gap, we introduce OpenCodeInstruct, the largest open-access instruction tuning dataset, comprising 5 million diverse samples. Each sample includes a programming question, solution, test cases, execution feedback, and LLM-generated quality assessments. We fine-tune various base models, including LLaMA and Qwen, across multiple scales (1B+, 3B+, and 7B+) using our dataset. Comprehensive evaluations on popular benchmarks (HumanEval, MBPP, LiveCodeBench, and BigCodeBench) demonstrate substantial performance improvements achieved by SFT with OpenCodeInstruct. We also present a detailed methodology encompassing seed data curation, synthetic instruction and solution generation, and filtering.


Single-Pass Document Scanning for Question Answering

arXiv.org Artificial Intelligence

Handling extremely large documents for question answering is challenging: chunk-based embedding methods often lose track of important global context, while full-context transformers can be prohibitively expensive for hundreds of thousands of tokens. We propose a single-pass document scanning approach that processes the entire text in linear time, preserving global coherence while deciding which sentences are most relevant to the query. On 41 QA benchmarks, our single-pass scanner consistently outperforms chunk-based embedding methods and competes with large language models at a fraction of the computational cost. By conditioning on the entire preceding context without chunk breaks, the method preserves global coherence, which is especially important for long documents. Overall, single-pass document scanning offers a simple solution for question answering over massive text. All code, datasets, and model checkpoints are available at https://github.com/MambaRetriever/MambaRetriever


OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

arXiv.org Artificial Intelligence

Since the advent of reasoning-based large language models, many have found great success from distilling reasoning capabilities into student models. Such techniques have significantly bridged the gap between reasoning and standard LLMs on coding tasks. Despite this, much of the progress on distilling reasoning models remains locked behind proprietary datasets or lacks details on data curation, filtering and subsequent training. To address this, we construct a superior supervised fine-tuning (SFT) dataset that we use to achieve state-of-the-art coding capability results in models of various sizes. Our distilled models use only SFT to achieve 61.8% on LiveCodeBench and 24.6% on CodeContests, surpassing alternatives trained with reinforcement learning. We then perform analysis on the data sources used to construct our dataset, the impact of code execution filtering, and the importance of instruction/solution diversity. We observe that execution filtering negatively affected benchmark accuracy, leading us to prioritize instruction diversity over solution correctness. Finally, we also analyze the token efficiency and reasoning patterns utilized by these models. We will open-source these datasets and distilled models to the community.


It shocked the market but has China's DeepSeek changed AI?

BBC News

DeepSeek's arrival also marked a turning point in the US-China AI rivalry, some experts say. "China was seen as playing catch-up in large language models until this point, with competitive models but always trailing the best western ones," policy analyst Wendy Chang of the Mercator Institute for China Studies told the BBC. A large language model (LLM) is a reasoning system trained to predict the next word in a given sentence or phrase. DeepSeek changed perceptions when it claimed to have achieved a leading model for a fraction of the computational resources and costs common among its American counterparts. OpenAI had spent 5bn ( 3.7bn) in 2024 alone.


'It's missing something': AGI, superintelligence and a race for the future

The Guardian

That was how Sam Altman, chief executive of OpenAI, described the latest upgrade to ChatGPT this week. The race Altman was referring to was artificial general intelligence (AGI), a theoretical state of AI where, by OpenAI's definition, a highly autonomous system is able to do a human's job. Describing the new GPT-5 model, which will power ChatGPT, as a "significant step on the path to AGI", he nonetheless added a hefty caveat. "[It is] missing something quite important, many things quite important," said Altman, such as the model's inability to "continuously learn" even after its launch. In other words, these systems are impressive but they have yet to crack the autonomy that would allow them to do a full-time job.


OpenAI will not disclose GPT-5's energy use. It could be higher than past models

The Guardian

In mid-2023, if a user asked OpenAI's ChatGPT for a recipe for artichoke pasta or instructions on how to make a ritual offering to the ancient Canaanite deity Moloch, its response might have taken โ€“ very roughly โ€“ 2 watt-hours, or about as much electricity as an incandescent bulb consumes in 2 minutes. OpenAI released a model on Thursday that will underpin the popular chatbot โ€“ GPT-5. Ask that version of the AI for an artichoke recipe, and the same amount of pasta-related text could take several times โ€“ even 20 times โ€“ that amount of energy, experts say. As it rolled out GPT-5, the company highlighted the model's breakthrough capabilities: its ability to create websites, answer PhD-level science questions, and reason through difficult problems. But experts who have spent the past years working to benchmark the energy and resource usage of AI models say those new powers come at a cost: a response from GPT-5 may take a significantly larger amount of energy than a response from previous versions of ChatGPT.


Fox News AI Newsletter: OpenAI GPT-5 draws Musk eyeroll

FOX News

Open AI CEO Sam Altman, center, speaks with boxer Jake Paul and wrestler Logan Paul in Emancipation Hall at the 60th Presidential Inauguration, Monday, Jan. 20, 2025, at the U.S. Capitol in Washington. TECH TENSIONS: Elon Musk escalated tensions in the critical artificial intelligence race Thursday, asserting his most advanced AI model, Grok 4 Heavy, was already outperforming OpenAI's newly launched GPT-5 two weeks ago. BOT BOOM: Small business owners are rapidly adopting artificial intelligence to power their growth, with many saying it will lead to more job opportunities this year, according to a Goldman Sachs survey. POCKET GENIUS: OpenAI unveiled GPT-5 on Thursday, calling it a significant upgrade from its predecessors and a major step forward in building the capabilities of large language models. AI-DOCTORED PHOTOS: Airbnb has reportedly apologized to a woman after the host of a Manhattan apartment where she stayed used artificial intelligence to doctor images of the home, saying she caused thousands of dollars in damage.


Join Our Next Livestream: What GPT-5 Means for ChatGPT Users

WIRED

Few recent software releases have been as hyped as OpenAI's launch of its GPT-5 model. "GPT-5 is the first time that it really feels like talking to an expert in any topic, like a PhD level expert," said CEO Sam Altman in a recent press briefing. Is this new release as big of an upgrade as OpenAI claims? What do these changes actually mean for ChatGPT users? WIRED reporters are currently testing this newest drop from OpenAI, and seeing how GPT-5's ability to write, code, and perform other tasks compares to past releases.


The Download: GPT-5 is here, and Intel's CEO drama

MIT Technology Review

"I don't think we should think of them as the'new Google' yet." The arrhythmia of our current age Arrhythmia means the heart beats, but not in proper time--a critical rhythm of life suddenly going rogue and unpredictable. It's frightening to experience, but what if it's also a good metaphor for our current times? That a pulse once seemingly so steady is now less sure. Perhaps this wobbliness might be extrapolated into a broader sense of life in the 2020s.