Large Language Model
Two-way Evidence self-Alignment based Dual-Gated Reasoning Enhancement
Zhang, Kexin, Chen, Junlan, Li, Daifeng, Zhang, Yuxuan, Feng, Yangyang, Deng, Bowen, Chen, Weixu
Large language models (LLMs) encounter difficulties in knowledge-intensive multi-step reasoning (KIMSR) tasks. One challenge is how to effectively extract and represent rationale evidence. The current methods often extract semantically relevant but logically irrelevant evidence, resulting in flawed reasoning and inaccurate responses. We propose a two-way evidence self-alignment (TW-ESA) module, which utilizes the mutual alignment between strict reasoning and LLM reasoning to enhance its understanding of the causal logic of evidence, thereby addressing the first challenge. Another challenge is how to utilize the rationale evidence and LLM's intrinsic knowledge for accurate reasoning when the evidence contains uncertainty. We propose a dual-gated reasoning enhancement (DGR) module to gradually fuse useful knowledge of LLM within strict reasoning, which can enable the model to perform accurate reasoning by focusing on causal elements in the evidence and exhibit greater robustness. The two modules are collaboratively trained in a unified framework ESA-DGR. Extensive experiments on three diverse and challenging KIMSR datasets reveal that ESA-DGR significantly surpasses state-of-the-art LLM-based fine-tuning methods, with remarkable average improvements of 4% in exact match (EM) and 5% in F1 score. The implementation code is available at https://anonymous.4open.science/r/ESA-DGR-2BF8.
Sparse Activation Editing for Reliable Instruction Following in Narratives
Zhao, Runcong, Cao, Chengyu, Zhu, Qinglin, Lv, Xiucheng, Shao, Shun, Gui, Lin, Xu, Ruifeng, He, Yulan
Complex narrative contexts often challenge language models' ability to follow instructions, and existing benchmarks fail to capture these difficulties. To address this, we propose Concise-SAE, a training-free framework that improves instruction following by identifying and editing instruction-relevant neurons using only natural language instructions, without requiring labelled data. To thoroughly evaluate our method, we introduce FreeInstruct, a diverse and realistic benchmark of 1,212 examples that highlights the challenges of instruction following in narrative-rich settings. While initially motivated by complex narratives, Concise-SAE demonstrates state-of-the-art instruction adherence across varied tasks without compromising generation quality.
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI
Goodfellow, Martin, Booth, Robbie, Fagan, Andrew, Lambert, Alasdair
Students often do not fully understand the code they have written. This sometimes does not become evident until later in their education, which can mean it is harder to fix their incorrect knowledge or misunderstandings. In addition, being able to fully understand code is increasingly important in a world where students have access to generative artificial intelligence (GenAI) tools, such as GitHub Copilot. One effective solution is to utilise code comprehension questions, where a marker asks questions about a submission to gauge understanding, this can also have the side effect of helping to detect plagiarism. However, this approach is time consuming and can be difficult and/or expensive to scale. This paper introduces AutoMCQ, which uses GenAI for the automatic generation of multiple-choice code comprehension questions. This is integrated with the CodeRunner automated assessment platform.
Resource for Error Analysis in Text Simplification: New Taxonomy and Test Collection
Vendeville, Benjamin, Ermakova, Liana, De Loor, Pierre
The general public often encounters complex texts but does not have the time or expertise to fully understand them, leading to the spread of misinformation. Automatic Text Simplification (ATS) helps make information more accessible, but its evaluation methods have not kept up with advances in text generation, especially with Large Language Models (LLMs). In particular, recent studies have shown that current ATS metrics do not correlate with the presence of errors. Manual inspections have further revealed a variety of errors, underscoring the need for a more nuanced evaluation framework, which is currently lacking. This resource paper addresses this gap by introducing a test collection for detecting and classifying errors in simplified texts. First, we propose a taxonomy of errors, with a formal focus on information distortion. Next, we introduce a parallel dataset of automatically simplified scientific texts. This dataset has been human-annotated with labels based on our proposed taxonomy. Finally, we analyze the quality of the dataset, and we study the performance of existing models to detect and classify errors from that taxonomy. These contributions give researchers the tools to better evaluate errors in ATS, develop more reliable models, and ultimately improve the quality of automatically simplified texts.
iPhone design guru and OpenAI chief promise an AI device revolution
Everything over the last 30 years, according to Sir Jony Ive, has led to this moment: a partnership between the iPhone designer and the developer of ChatGPT. Ive has sold his hardware startup, io, to OpenAI and will take on creative and design leadership across the merged businesses. "I have a growing sense that everything I have learned over the last 30 years has led me to this place, to this moment," he says in a video announcing the 6.4bn ( 4.8bn) deal. The main aim will be to move on from Ive's signature achievement designing Apple's most successful product, as well as the iPod, iPad and Apple Watch. The British-born designer has already developed a prototype io device, and one of its users is OpenAI's chief executive, Sam Altman.
AI Is Eating Data Center Power Demand--and It's Only Getting Worse
AI's energy use already represents as much as 20 percent of global data-center power demand, research published Thursday in the journal Joule shows. That demand from AI, the research states, could double by the end of this year, comprising nearly half of all total data-center electricity consumption worldwide, excluding the electricity used for bitcoin mining. The new research is published in a commentary by Alex de Vries-Gao, the founder of Digiconomist, a research company that evaluates the environmental impact of technology. De Vries-Gao started Digiconomist in the late 2010s to explore the impact of bitcoin mining, another extremely energy-intensive activity, would have on the environment. Looking at AI, he says, has grown more urgent over the past few years because of the widespread adoption of ChatGPT and other large language models that use massive amounts of energy. According to his research, worldwide AI energy demand is now set to surpass demand from bitcoin mining by the end of this year.
A United Arab Emirates Lab Announces Frontier AI Projects--and a New Outpost in Silicon Valley
A United Arab Emirates (UAE) academic lab today launched an artificial intelligence world model and agent, two large language models (LLMs) and a new research center in Silicon Valley as it ramps up its investment in the cutting-edge field. The UAE's Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) revealed an AI world model called PAN, which can be used to build physically realistic simulations for testing and honing the performance of AI agents. Eric Xing, President and Professor of MBZUAI and a leading AI researcher, revealed the models and lab at the Computer History Museum in Mountain View, California today. The UAE has made big investments in AI in recent years under the guidance of Sheikh Tahnoun bin Zayed al Nahyan, the nation's tech-savvy national security advisor and younger brother of president Mohamed bin Zayed Al Nahyan. Xing says the UAE's new center in Sunnyvale, California, will help the nation tap into the world's most concentrated source of AI knowledge and talent.
DOGE Used Meta AI Model to Review Emails From Federal Workers
Elon Musk's so-called Department of Government Efficiency (DOGE) used artificial intelligence from Meta's Llama model to comb through and analyze emails from federal workers. Materials viewed by WIRED show that DOGE affiliates within the Office of Personnel Management (OPM) tested and used Meta's Llama 2 model to review and classify responses from federal workers to the infamous "Fork in the Road" email that was sent across the government in late January. The email offered deferred resignation to anyone opposed to changes the Trump administration was making to its federal workforce, including an enforced return to office policy, downsizing, and a requirement to be "loyal." To leave their position, recipients merely needed to reply with the word "resign." This email closely mirrored one that Musk sent to Twitter employees shortly after he took over the company in 2022.
Anthropic's New Model Excels at Reasoning and Planning--and Has the Pokémon Skills to Prove It
Anthropic announced two new models, Claude 4 Opus and Claude Sonnet 4, during its first developer conference in San Francisco on Thursday. The pair will be immediately available to paying Claude subscribers. The new models, which jump the naming convention from 3.7 straight to 4, have a number of strengths, including their ability to reason, plan, and remember the context of conversations over extended periods of time, the company says. Claude 4 Opus is also even better at playing Pokémon than its predecessor. "It was able to work agentically on Pokémon for 24 hours," says Anthropic's chief product officer Mike Krieger in an interview with WIRED.
Google's New AI Puts Breasts on Minors--And J. D. Vance
Sorry to tell you this, but Google's new AI shopping tool appears eager to give J. D. Vance breasts. This week, at its annual software conference, Google released an AI tool called Try It On, which acts as a virtual dressing room: Upload images of yourself while shopping for clothes online, and Google will show you what you might look like in a selected garment. Curious to play around with the tool, we began uploading images of famous men--Vance, Sam Altman, Abraham Lincoln, Michelangelo's David, Pope Leo XIV--and dressed them in linen shirts and three-piece suits. But when we tested a number of articles designed for women on these famous men, the tool quickly adapted: Whether it was a mesh shirt, a low-cut top, or even just a T-shirt, Google's AI rapidly spun up images of the vice president, the CEO of OpenAI, and the vicar of Christ with breasts. It's not just men: When we uploaded images of women, the tool repeatedly enhanced their décolletage or added breasts that were not visible in the original images.