Goto

Collaborating Authors

 Large Language Model


Sample Complexities of Estimating Gumbel--Max Watermark Proportions with and without Reduction to Pivotal Statistics

arXiv.org Machine Learning

Watermarking promises a statistical trace of large language model (LLM) use, but real documents, after editing or paraphrasing, rarely arrive as purely human-written or purely machine-generated. This motivates a quantitative question beyond detection: what proportion of a document is generated from a pre-specified watermarked LLM? We study this watermark proportion estimation problem under the Gumbel--max watermarking mechanism, treating the next-token prediction (NTP) distributions as unknown and arbitrary nuisance parameters subject to a non-degeneracy condition. We compare two observation regimes: in the full observation regime, the estimator observes the pseudorandom vector and the selected token at each position; under the more popular setting of pivotal reduction, it observes only a scalar pivot, which follows a one-dimensional Uniform--Beta mixture distribution. Under pivotal reduction, we develop a Laguerre-polynomial estimator and establish a matching information-theoretic lower bound for the sample complexity. For full observation, we introduce an event-counting estimator and show a matching lower bound, yielding a substantially smaller sample complexity. As our results imply, although reducing to pivotal statistics is an elegant and widely used procedure, it is not always sample-efficient for estimating the proportion of watermarks.


Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization

arXiv.org Machine Learning

Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning. With richer context, transformers adapt more effectively to the current use case without any parameter updates. However, the quadratic computational and memory complexity with respect to context length significantly slows data processing in softmax transformers. Linear transformers were proposed to address this issue by reducing the complexity to linear dependence on context length, but the design and understanding of the feature mapping in linear attention, from a theoretical viewpoint, remain unclear. In this paper, we investigate the approximation and generalization abilities of linear transformers under a two-staged sampling process from domain generalization. We show that linear transformers perform in-context learning as learning a mapping from context distributions to response functions. A dimension-independent convergence rate is obtained for our generalization analysis, which also exhibits the tradeoff between the regularities of data distributions and latent features. Guided by our theoretical framework, we propose a new perspective on activation and loss design for linearizing pretrained softmax large language models.


Prototype Language Models

arXiv.org Machine Learning

Knowing which training examples drive outputs is fundamental to auditing, correcting, and understanding language models, yet for modern LLMs this remains expensive, approximate, and largely post-hoc. Standard language models generate tokens through a dense network pathway, causing training data's influence to be distributed across parameters rather than organized along explicit, traceable components. We introduce a prototype language model architecture, Prototypes for Interpretable Sequence Modeling (PRISM), that forms each prediction via a sparse, non-negative mixture of learned prototypes, trained with clustering objectives that anchor each prototype to coherent neighborhoods of training examples. Across architectures from 130M to 1.6B parameters trained on up to 50B tokens, prototype language models either surpass or remain within 2.5 percentage points on average downstream accuracy of matched dense baselines. We show that sparse prototype structure localizes curvature in the loss landscape, yielding a more tractable Hessian and enabling training data attribution that is ~500x faster than post hoc baselines when consuming equivalent memory. Calibrating linear prototype controllers can improve downstream accuracy by roughly 3 points while tracing those corrections back to training neighborhoods, and targeted prototype suppression can remove model behaviors without finetuning or measurable loss in generation quality.


Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

arXiv.org Machine Learning

Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge. We study this problem in instruction-based molecular optimization, where answer-only supervised fine-tuning (SFT) collapses multi-step reasoning and reinforcement learning with verifiable rewards (RLVR) suffers from sparse feedback. Reference-guided Policy Optimization (RePO) mitigates both by anchoring policy updates to dataset-provided references, but its effectiveness is tightly coupled to reference quality: weak or misaligned references impose a performance ceiling. To overcome this ceiling, we propose active reasoning, a paradigm in which the policy actively decides, on a per-instance basis, when to imitate a reference and when to reinforce its own discoveries, while continuously upgrading what it imitates. We instantiate this paradigm as Active Group Relative Policy Optimization (Active-GRPO), realized through two coupled mechanisms: active imitate-reinforce and active referencing. The former performs imitation learning when the reference still outperforms the policy's own candidates, and shifts to self-improvement via reinforcement learning once the policy has generated molecules that surpass the reference. The latter continuously upgrades the reference itself by replacing it with the best policy-generated candidate discovered so far, progressively raising the imitation target and ensuring that reference guidance remains informative--rather than restrictive--throughout training. Across TOMG-Bench MOLOPT, Active-GRPO improves average SR Sim from 0.0959 for GRPO and 0.1665 for RePO to 0.1773 under matched three-seed evaluation, with statistically significant gains on LogP, MR, and QED.


Anthropic Added a New Security Measure to Get Back Into the Trump Administration's Good Graces

WIRED

Anthropic Added a New Security Measure to Get Back Into the Trump Administration's Good Graces The government has removed restrictions on Anthropic's Fable 5 and Mythos 5 AI models--but there were strings attached. The Trump administration lifted export controls on Anthropic's Claude Fable 5 AI model after the company agreed to extend an existing guardrail to prevent users from trying to access certain restricted capabilities, according to two people familiar with the matter. The safeguard means any users trying to unlock those capabilities will be notified that their request is blocked and will have their query processed by the less-advanced Opus 4.8 AI model, the people say. Before Anthropic cut off access to Fable 5, user requests related to sensitive cybersecurity and biology capabilities were supposed to be processed by Opus 4.8. The new safeguard, the people say, will extend this guardrail to requests related to a specific behavior identified in a paper by Amazon .


LLMs are stuck in a groupthink groove. This startup is trying to get them out.

MIT Technology Review

Let's start with a game. Open up your chatbot of choice--Claude, ChatGPT, Gemini--and type "Give me a random number between 1 and 10." You're going to get 7. Almost always. Now type "Another" and you'll get 3 or 4. Type "Another" again and you'll get 8 or 9. That won't work every time--but if it did for you, you may wonder if I have superpowers.


Claude subscribers are furious over Fable's new restrictions

PCWorld

PCWorld reports that Anthropic's Claude subscribers are expressing strong dissatisfaction over new restrictions on the Fable 5 AI model after its return from government-imposed limitations. Claude Pro, Max, Team, and Enterprise users now face a 50% usage limit for Fable 5 and only have access until July 7, shorter than originally promised. After July 7, subscribers must purchase separate usage credits at higher costs to continue using Fable 5, while Mythos 5 remains limited to select U.S. organizations only. After nearly three weeks in a government-imposed limbo, Anthropic's powerful Fable 5 and Mythos 5 models are going back online. But irate Claude subscribers are now learning that their already brief availability window for Fable will be curtailed even further.


The Download: Anthropic launches Claude Science, and California's carbon manure math

MIT Technology Review

The Download: Anthropic launches Claude Science, and California's carbon manure math Plus: The US has lifted restrictions on Anthropic's Mythos and Fable models. Claude Science is Anthropic's newest flagship product At an event for pharmaceutical executives, biotech founders, and researchers yesterday, Anthropic announced Claude Science, a major new product intended to support scientific research like Claude Code supports software engineering. Like Claude Code, Claude Science can autonomously carry out meaningful work from concise, high-level instructions, with tools for computational biology and drug development. The launch signals that Anthropic is doubling down on AI for science, and the company will also use the product in its own research into drugs for rare, neglected diseases. Discover why Anthropic is betting big on AI for scientific research . Why California's carbon manure math doesn't add up Years ago, the state set up a system that pays cattle farmers to turn the methane emitted from cattle manure into natural gas.


Claude Helped a Hacker Find a Way to Issue Tickets to Almost Every US Music Festival

WIRED

A researcher found that using Anthropic's Claude Opus 4.7, he could break into the website of Front Gate--used by every festival from Lollapalooza to Bonnaroo--and freely issue any ticket he chose. Fears about AI tools capable of autonomous hacking usually involve nightmare scenarios like the theft of nuclear launch codes or zeroed-out bank reserves. Far more plausible, it turns out, is asking AI to gain super-administrator access on a ticketing website and then issuing yourself and all of your friends free VIP backstage passes to Bonnaroo. That was the discovery of security researcher Ian Carroll, who used the AI tool Claude Opus 4.7 in April to discover a technique that allowed him full access to the systems of Front Gate Tickets, which handles ticketing for practically every major US music festival, from Lollapalooza and South by Southwest to Austin City Limits. Carroll found that Front Gate, which like Ticketmaster is a subsidiary of the event company Live Nation Entertainment, had a bug in its website that he--with Claude's help--could exploit to gain access to millions of customer or staff records and freely issue tickets for any event, of any value, to himself or whoever he chose.


Gemini Spark comes to Google's Gemini app for macOS

Engadget

The Spark agentic AI assistant is exclusively available to Google AI Ultra subscribers in the US. Google has started rolling out its new agentic AI assistant, Gemini Spark, to Gemini's app for macOS . When the company launched Spark at its I/O developer conference in May, it explained that the assistant turns Gemini into an active partner that can actually do tasks for you. On Mac computers, for instance, you can ask it to (finally) sort the massive number of the PDFs in your Downloads into specific folders. You'll also be able to get it to do tasks on your Workspace apps using files in your computer, such as asking it to create a spreadsheet with invoices saved on your laptop.