Large Language Model
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
Wang, Dingzirui, Zhang, Xuangliang, Xu, Keyan, Zhu, Qingfu, Che, Wanxiang, Deng, Yang
Numerous studies have investigated the underlying mechanisms of in-context learning (ICL) effectiveness to inspire the design of related methods. However, existing work predominantly assumes the effectiveness of the demonstrations provided within ICL, while many research indicates that not all demonstrations are effective, failing to yielding any performance improvement during ICL. Therefore, in this paper, we investigate the reasons behind demonstration ineffectiveness. Our analysis is based on gradient flow and linear self-attention models. By setting the gradient flow to zero, we deduce that a demonstration becomes ineffective if its information has either been learned by the model or is irrelevant to the user query. Furthermore, we demonstrate that in multi-layer models, the disparity in effectiveness among demonstrations is amplified with layer increasing, causing the model to focus more on effective ones. Considering that current demonstration selection methods primarily focus on the relevance to the user query while overlooking the information that the model has already assimilated, we propose a novel method called GradS, which leverages gradient flow for demonstration selection. We use the magnitude of the gradient flow of the demonstration with respect to a given user query as the criterion, thereby ensuring the effectiveness of the chosen ones. We validate our derivation and GradS on four prominent LLMs across five mainstream datasets. The experimental results confirm that the disparity in effectiveness among demonstrations is magnified as the model layer increases, substantiating our derivations. Moreover, GradS achieves a relative improvement of $6.8\%$ on average over the strongest baselines, demonstrating its effectiveness.
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
Albert, Paul, Zhang, Frederic Z., Saratchandran, Hemanth, Hengel, Anton van den, Abbasnejad, Ehsan
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large pre-trained models. Amongst PEFT methods, low-rank adaptation (LoRA) has achieved notable success. However, recent studies have highlighted its limitations compared against full-rank alternatives, particularly when applied to multimodal and large language models. In this work, we present a quantitative comparison amongst full-rank and low-rank PEFT methods using a synthetic matrix approximation benchmark with controlled spectral properties. Our results confirm that LoRA struggles to approximate matrices with relatively flat spectrums or high frequency components -- signs of high effective ranks. To this end, we introduce KRAdapter, a novel PEFT algorithm that leverages the Khatri-Rao product to produce weight updates, which, by construction, tends to produce matrix product with a high effective rank. We demonstrate performance gains with KRAdapter on vision-language models up to 1B parameters and on large language models up to 8B parameters, particularly on unseen common-sense reasoning tasks. In addition, KRAdapter maintains the memory and compute efficiency of LoRA, making it a practical and robust alternative to fine-tune billion-scale parameter models.
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
Chen, Mingda, Li, Yang, Chen, Xilun, Williams, Adina, Ghosh, Gargi, Yih, Scott
Long-form factuality evaluation assesses the ability of models to generate accurate, comprehensive responses to short prompts. Existing benchmarks often lack human verification, leading to potential quality issues. To address this limitation, we introduce FACTORY, a large-scale, human-verified prompt set. Developed using a model-in-the-loop approach and refined by humans, FACTORY includes challenging prompts that are fact-seeking, answerable, and unambiguous. We conduct human evaluations on 6 state-of-the-art language models using FACTORY and existing datasets. Our results show that FACTORY is a challenging benchmark: approximately 40% of the claims made in the responses of SOTA models are not factual, compared to only 10% for other datasets. Our analysis identifies the strengths of FACTORY over prior benchmarks, emphasizing its reliability and the necessity for models to reason across long-tailed facts.
Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench
Mutisya, Fred, Gitau, Shikoh, Ongoma, Nasubo, Mbae, Keith, Wamicha, Elizabeth
HealthBench, a benchmark designed to measure the capabilities of AI systems for health better (Arora et al., 2025), has advanced medical language model evaluation through physician-crafted dialogues and transparent rubrics. However, its reliance on expert opinion, rather than high-tier clinical evidence, risks codifying regional biases and individual clinician idiosyncrasies, further compounded by potential biases in automated grading systems. These limitations are particularly magnified in low- and middle-income settings, where issues like sparse neglected tropical disease coverage and region-specific guideline mismatches are prevalent. The unique challenges of the African context, including data scarcity, inadequate infrastructure, and nascent regulatory frameworks, underscore the urgent need for more globally relevant and equitable benchmarks. To address these shortcomings, we propose anchoring reward functions in version-controlled Clinical Practice Guidelines (CPGs) that incorporate systematic reviews and GRADE evidence ratings. Our roadmap outlines "evidence-robust" reinforcement learning via rubric-to-guideline linkage, evidence-weighted scoring, and contextual override logic, complemented by a focus on ethical considerations and the integration of delayed outcome feedback. By re-grounding rewards in rigorously vetted CPGs, while preserving HealthBench's transparency and physician engagement, we aim to foster medical language models that are not only linguistically polished but also clinically trustworthy, ethically sound, and globally relevant.
The Second Machine Turn: From Checking Proofs to Creating Concepts
We identify a second machine turn in the process of mathematical discovery: after automating proof-checking, AI is now poised to automate the *creation* of mathematical concepts themselves. We discuss the current state of the art, obstacles and potential solutions as well as a preliminary attempt at mathematizing the creation of concepts itself. The paper ends with an assessment of how these capabilities could reshape mathematics and human-machine collaboration, and a few different futures we might find ourselves in.
The A.I. Data Center Push That Wasn't
OpenAI's Sam Altman, flanked by President Trump and Softbank's Masayoshi Son, announced a hugely ambitious investment in data centers across America to support all the artificial intelligence we're going to be using. Months in, the project has been scaled back to a single, power-hungry data center in Ohio. Subscribe to Slate Plus to access ad-free listening to the whole What Next family and all your favorite Slate podcasts. Subscribe today on Apple Podcasts by clicking "Try Free" at the top of our show page. Sign up now at slate.com/whatnextplus to get access wherever you listen.
The clanker of clankers: Get every major AI model for 85% off
TL;DR 1min.AI gives you access to GPT-4o, Claude 3, Gemini Pro, and more in one platform -- Lifetime access is just 79.97 (MSRP 540). Running a business, making content, and scaling a brand takes more than one bot. For writing, editing, designing, researching, and automating, you don't need a single assistant. You need a whole AI workforce. It's built for serious output across content, image, audio, and video, all from one browser-based dashboard.
Anthropic Revokes OpenAI's Access to Claude
Anthropic revoked OpenAI's API access to its models on Tuesday, multiple sources familiar with the matter tell WIRED. OpenAI was informed that its access was cut off due to violating the terms of service. "Claude Code has become the go-to choice for coders everywhere and so it was no surprise to learn OpenAI's own technical staff were also using our coding tools ahead of the launch of GPT-5," Anthropic spokesperson Christopher Nulty said in a statement to WIRED. "Unfortunately, this is a direct violation of our terms of service." According to Anthropic's commercial terms of service, customers are barred from using the service to "build a competing product or service, including to train competing AI models" or "reverse engineer or duplicate" the services.
ChatGPT is crushing all other AI chatbots, and the numbers prove it
It might seem that "ChatGPT" is all you ever hear about when discussing AI chatbots, also known as LLMs. As it turns out, that's reflected in the real world, too. Statcounter, which tracks the market share of operating systems, browsers, social media sites, and more, has begun tracking the number of sessions by users who visit artificial intelligence sites like ChatGPT, Google Gemini, Claude AI, Microsoft Copilot, and more. The winner, not surprisingly, is ChatGPT, by an enormous margin: over 80 percent and climbing, which coincides with our own chatbot tests. Statcounter began tracking the statistics in March, and they've roughly remained the same since then: ChatGPT absolutely dominates, with a cluster of smaller AI chatbots below.
WIRED Roundup: ChatGPT Goes Full Demon Mode
On today's episode, our host Zoë Schiffer is joined by WIRED's senior business editor Louise Matsakis to run through five of the most important stories we published this week, from Meta continuing its AI talent poaching spree to how much faster our brains have aged since the pandemic. Afterward, they dive into the surprising reason ChatGPT reportedly went full demon mode last week. Write to us at uncannyvalley@wired.com. Mentioned in this episode: The Real Demon Inside ChatGPT by Louise Matsakis Meta's AI Recruiting Campaign Finds a New Target by Kylie Robison The Pandemic Appears to Have Accelerated Brain Aging, Even in People Who Never Got Covid by Javier Carbajal Age Verification Laws Send VPN Use Soaring--and Threaten the Open Internet by Lily Hay Newman and Matt Burgess This Smart Basketball Tracks Data About Every Shot. You can always listen to this week's podcast through the audio player on this page, but if you want to subscribe for free to get every episode, here's how: If you're on an iPhone or iPad, open the app called Podcasts, or just tap this link.