Goto

Collaborating Authors

 Large Language Model


Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product

arXiv.org Artificial Intelligence

Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large pre-trained models. Amongst PEFT methods, low-rank adaptation (LoRA) has achieved notable success. However, recent studies have highlighted its limitations compared against full-rank alternatives, particularly when applied to multimodal and large language models. In this work, we present a quantitative comparison amongst full-rank and low-rank PEFT methods using a synthetic matrix approximation benchmark with controlled spectral properties. Our results confirm that LoRA struggles to approximate matrices with relatively flat spectrums or high frequency components -- signs of high effective ranks. To this end, we introduce KRAdapter, a novel PEFT algorithm that leverages the Khatri-Rao product to produce weight updates, which, by construction, tends to produce matrix product with a high effective rank. We demonstrate performance gains with KRAdapter on vision-language models up to 1B parameters and on large language models up to 8B parameters, particularly on unseen common-sense reasoning tasks. In addition, KRAdapter maintains the memory and compute efficiency of LoRA, making it a practical and robust alternative to fine-tune billion-scale parameter models.


FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality

arXiv.org Artificial Intelligence

Long-form factuality evaluation assesses the ability of models to generate accurate, comprehensive responses to short prompts. Existing benchmarks often lack human verification, leading to potential quality issues. To address this limitation, we introduce FACTORY, a large-scale, human-verified prompt set. Developed using a model-in-the-loop approach and refined by humans, FACTORY includes challenging prompts that are fact-seeking, answerable, and unambiguous. We conduct human evaluations on 6 state-of-the-art language models using FACTORY and existing datasets. Our results show that FACTORY is a challenging benchmark: approximately 40% of the claims made in the responses of SOTA models are not factual, compared to only 10% for other datasets. Our analysis identifies the strengths of FACTORY over prior benchmarks, emphasizing its reliability and the necessity for models to reason across long-tailed facts.


Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench

arXiv.org Artificial Intelligence

HealthBench, a benchmark designed to measure the capabilities of AI systems for health better (Arora et al., 2025), has advanced medical language model evaluation through physician-crafted dialogues and transparent rubrics. However, its reliance on expert opinion, rather than high-tier clinical evidence, risks codifying regional biases and individual clinician idiosyncrasies, further compounded by potential biases in automated grading systems. These limitations are particularly magnified in low- and middle-income settings, where issues like sparse neglected tropical disease coverage and region-specific guideline mismatches are prevalent. The unique challenges of the African context, including data scarcity, inadequate infrastructure, and nascent regulatory frameworks, underscore the urgent need for more globally relevant and equitable benchmarks. To address these shortcomings, we propose anchoring reward functions in version-controlled Clinical Practice Guidelines (CPGs) that incorporate systematic reviews and GRADE evidence ratings. Our roadmap outlines "evidence-robust" reinforcement learning via rubric-to-guideline linkage, evidence-weighted scoring, and contextual override logic, complemented by a focus on ethical considerations and the integration of delayed outcome feedback. By re-grounding rewards in rigorously vetted CPGs, while preserving HealthBench's transparency and physician engagement, we aim to foster medical language models that are not only linguistically polished but also clinically trustworthy, ethically sound, and globally relevant.


The Second Machine Turn: From Checking Proofs to Creating Concepts

arXiv.org Artificial Intelligence

We identify a second machine turn in the process of mathematical discovery: after automating proof-checking, AI is now poised to automate the *creation* of mathematical concepts themselves. We discuss the current state of the art, obstacles and potential solutions as well as a preliminary attempt at mathematizing the creation of concepts itself. The paper ends with an assessment of how these capabilities could reshape mathematics and human-machine collaboration, and a few different futures we might find ourselves in.


The A.I. Data Center Push That Wasn't

Slate

OpenAI's Sam Altman, flanked by President Trump and Softbank's Masayoshi Son, announced a hugely ambitious investment in data centers across America to support all the artificial intelligence we're going to be using. Months in, the project has been scaled back to a single, power-hungry data center in Ohio. Subscribe to Slate Plus to access ad-free listening to the whole What Next family and all your favorite Slate podcasts. Subscribe today on Apple Podcasts by clicking "Try Free" at the top of our show page. Sign up now at slate.com/whatnextplus to get access wherever you listen.


The clanker of clankers: Get every major AI model for 85% off

PCWorld

TL;DR 1min.AI gives you access to GPT-4o, Claude 3, Gemini Pro, and more in one platform -- Lifetime access is just 79.97 (MSRP 540). Running a business, making content, and scaling a brand takes more than one bot. For writing, editing, designing, researching, and automating, you don't need a single assistant. You need a whole AI workforce. It's built for serious output across content, image, audio, and video, all from one browser-based dashboard.


Anthropic Revokes OpenAI's Access to Claude

WIRED

Anthropic revoked OpenAI's API access to its models on Tuesday, multiple sources familiar with the matter tell WIRED. OpenAI was informed that its access was cut off due to violating the terms of service. "Claude Code has become the go-to choice for coders everywhere and so it was no surprise to learn OpenAI's own technical staff were also using our coding tools ahead of the launch of GPT-5," Anthropic spokesperson Christopher Nulty said in a statement to WIRED. "Unfortunately, this is a direct violation of our terms of service." According to Anthropic's commercial terms of service, customers are barred from using the service to "build a competing product or service, including to train competing AI models" or "reverse engineer or duplicate" the services.


ChatGPT is crushing all other AI chatbots, and the numbers prove it

PCWorld

It might seem that "ChatGPT" is all you ever hear about when discussing AI chatbots, also known as LLMs. As it turns out, that's reflected in the real world, too. Statcounter, which tracks the market share of operating systems, browsers, social media sites, and more, has begun tracking the number of sessions by users who visit artificial intelligence sites like ChatGPT, Google Gemini, Claude AI, Microsoft Copilot, and more. The winner, not surprisingly, is ChatGPT, by an enormous margin: over 80 percent and climbing, which coincides with our own chatbot tests. Statcounter began tracking the statistics in March, and they've roughly remained the same since then: ChatGPT absolutely dominates, with a cluster of smaller AI chatbots below.


WIRED Roundup: ChatGPT Goes Full Demon Mode

WIRED

On today's episode, our host Zoë Schiffer is joined by WIRED's senior business editor Louise Matsakis to run through five of the most important stories we published this week, from Meta continuing its AI talent poaching spree to how much faster our brains have aged since the pandemic. Afterward, they dive into the surprising reason ChatGPT reportedly went full demon mode last week. Write to us at uncannyvalley@wired.com. Mentioned in this episode: The Real Demon Inside ChatGPT by Louise Matsakis Meta's AI Recruiting Campaign Finds a New Target by Kylie Robison The Pandemic Appears to Have Accelerated Brain Aging, Even in People Who Never Got Covid by Javier Carbajal Age Verification Laws Send VPN Use Soaring--and Threaten the Open Internet by Lily Hay Newman and Matt Burgess This Smart Basketball Tracks Data About Every Shot. You can always listen to this week's podcast through the audio player on this page, but if you want to subscribe for free to get every episode, here's how: If you're on an iPhone or iPad, open the app called Podcasts, or just tap this link.


The way we train AIs makes them more likely to spout bull

New Scientist

Common methods used to train artificial intelligence models seem to increase their tendency to give misleading answers, according to researchers who are aiming to produce "the first systematic analysis of machine bullshit". It is widely known that large language models (LLMs) have a tendency to generate false information – or "hallucinate" – but this is just one example, says Jaime Fernández Fisac at Princeton University. He and his colleagues define bullshit as "discourse intended to manipulate audience's beliefs, delivered with disregard for its truth value". "Our analysis found that the problem of bullshit in large language models is quite serious and widespread," says Fisac. The team divided such instances into five categories: empty rhetoric, such as "this red car combines style, charm, and adventure that captivates everyone"; weasel words – uncertain statements such as "studies suggest our product may help improve results in some cases"; paltering – using truthful statements to give a misleading impression; unverified claims; and sycophancy.