Goto

Collaborating Authors

 Deep Learning


Optimal Policy Minimum Bayesian Risk

arXiv.org Artificial Intelligence

Inference scaling helps LLMs solve complex reasoning problems through extended runtime computation. On top of long chain-of-thought (long-CoT) models, purely inference-time techniques such as best-of-N (BoN) sampling, majority voting, or more generally, minimum Bayes risk decoding (MBRD), can further improve LLM accuracy by generating multiple candidate solutions and aggregating over them. These methods typically leverage additional signals in the form of reward models and risk/similarity functions that compare generated samples, e.g., exact match in some normalized space or standard similarity metrics such as Rouge. Here we present a novel method for incorporating reward and risk/similarity signals into MBRD. Based on the concept of optimal policy in KL-controlled reinforcement learning, our framework provides a simple and well-defined mechanism for leveraging such signals, offering several advantages over traditional inference-time methods: higher robustness, improved accuracy, and well-understood asymptotic behavior. In addition, it allows for the development of a sample-efficient variant of MBRD that can adjust the number of samples to generate according to the difficulty of the problem, without relying on majority vote counts. We empirically demonstrate the advantages of our approach on math (MATH-$500$) and coding (HumanEval) tasks using recent open-source models. We also present a comprehensive analysis of its accuracy-compute trade-offs.


SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection

arXiv.org Artificial Intelligence

Predicting earnings surprises from financial documents, such as earnings conference calls, regulatory filings, and financial news, has become increasingly important in financial economics. However, these financial documents present significant analytical challenges, typically containing over 5,000 words with substantial redundancy and industry-specific terminology that creates obstacles for language models. In this work, we propose the SAE-FiRE (Sparse Autoencoder for Financial Representation Enhancement) framework to address these limitations by extracting key information while eliminating redundancy. SAE-FiRE employs Sparse Autoencoders (SAEs) to decompose dense neural representations from large language models into interpretable sparse components, then applies statistical feature selection methods, including ANOVA F-tests and tree-based importance scoring, to identify the top-k most discriminative dimensions for classification. By systematically filtering out noise that might otherwise lead to overfitting, we enable more robust and generalizable predictions. Experimental results across three financial datasets demonstrate that SAE-FiRE significantly outperforms baseline approaches.


FAID: Fine-Grained AI-Generated Text Detection Using Multi-Task Auxiliary and Multi-Level Contrastive Learning

arXiv.org Artificial Intelligence

The growing collaboration between humans and AI models in generative tasks has introduced new challenges in distinguishing between human-written, LLM-generated, and human--LLM collaborative texts. In this work, we collect a multilingual, multi-domain, multi-generator dataset FAIDSet. We further introduce a fine-grained detection framework FAID to classify text into these three categories, and also to identify the underlying LLM family of the generator. Unlike existing binary classifiers, FAID is built to capture both authorship and model-specific characteristics. Our method combines multi-level contrastive learning with multi-task auxiliary classification to learn subtle stylistic cues. By modeling LLM families as distinct stylistic entities, we incorporate an adaptation to address distributional shifts without retraining for unseen data. Our experimental results demonstrate that FAID outperforms several baselines, particularly enhancing the generalization accuracy on unseen domains and new LLMs, thus offering a potential solution for improving transparency and accountability in AI-assisted writing.


CottonSim: A vision-guided autonomous robotic system for cotton harvesting in Gazebo simulation

arXiv.org Artificial Intelligence

Cotton is a major cash crop in the United States, with the country being a leading global producer and exporter. Nearly all U.S. cotton is grown in the Cotton Belt, spanning 17 states in the southern region. Harvesting remains a critical yet challenging stage, impacted by the use of costly, environmentally harmful defoliants and heavy, expensive cotton pickers. These factors contribute to yield loss, reduced fiber quality, and soil compaction, which collectively threaten long-term sustainability. To address these issues, this study proposes a lightweight, small-scale, vision-guided autonomous robotic cotton picker as an alternative. An autonomous system, built on Clearpath's Husky platform and integrated with the CottonEye perception system, was developed and tested in the Gazebo simulation environment. A virtual cotton field was designed to facilitate autonomous navigation testing. The navigation system used Global Positioning System (GPS) and map-based guidance, assisted by an RGBdepth camera and a YOLOv8nseg instance segmentation model. The model achieved a mean Average Precision (mAP) of 85.2%, a recall of 88.9%, and a precision of 93.0%. The GPS-based approach reached a 100% completion rate (CR) within a $(5e-6)^{\circ}$ threshold, while the map-based method achieved a 96.7% CR within a 0.25 m threshold. The developed Robot Operating System (ROS) packages enable robust simulation of autonomous cotton picking, offering a scalable baseline for future agricultural robotics. CottonSim code and datasets are publicly available on GitHub: https://github.com/imtheva/CottonSim


A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility

arXiv.org Artificial Intelligence

Reasoning has emerged as the next major frontier for language models (LMs), with rapid advances from both academic and industrial labs. However, this progress often outpaces methodological rigor, with many evaluations relying on benchmarking practices that lack transparency, robustness, or statistical grounding. In this work, we conduct a comprehensive empirical study and find that current mathematical reasoning benchmarks are highly sensitive to subtle implementation choices--including decoding parameters, random seeds, prompt formatting, and even hardware and software configurations. Performance gains reported in recent studies frequently hinge on unclear comparisons or unreported sources of variance. To address these issues, we propose a standardized evaluation framework with clearly defined best practices and reporting standards. Using this framework, we reassess recent methods and find that most reinforcement learning (RL) approaches yield only modest improvements--far below prior claims--and are prone to overfitting, especially on small-scale benchmarks like AIME'24. In contrast, supervised finetuning (SFT) methods show consistently stronger generalization in the settings we study. To foster reproducibility, we release all code, prompts, and model outputs, for reasoning benchmarks, establishing more rigorous foundations for future work.


AI toys are all the rage in China--and now they're appearing on shelves in the US too

MIT Technology Review

AI toys are all the rage in China--and now they're appearing on shelves in the US too Competition is heating up, with Mattel and OpenAI expected to launch a product for kids this year. Kids have always played with and talked to stuffed animals. But now their toys can talk back, thanks to a wave of companies that are fitting children's playthings with chatbots and voice assistants. It's a trend that has particularly taken off in China: A recent report by the Shenzhen Toy Industry Association and JD.com predicts that the sector will surpass ¥100 billion ($14 billion) by 2030, growing faster than almost any other branch of consumer AI. According to the Chinese corporation registration database Qichamao, there are over 1,500 AI toy companies operating in China as of October 2025. One of the latest entrants to the market is a toy called BubblePal, a device the size of a Ping-Pong ball that clips onto a child's favorite stuffed animal and makes it "talk."


OpenAI Sneezes, and Software Firms Catch a Cold

WIRED

OpenAI revealed last week the custom AI tools it uses internally. The news sent some software companies into turmoil. Allan Thygesen, the CEO of Docusign, was not particularly concerned when he saw the news last week that OpenAI had created an internal tool called DocuGPT . He might have preferred that OpenAI choose a different name for its contracting tool. But still, he thought, DocuGPT barely scratched the surface of what Docusign can do.


The Download: extracting lithium, and what we still don't know about Sora

MIT Technology Review

The Download: extracting lithium, and what we still don't know about Sora On a bright afternoon in August, the shore of Utah's Great Salt Lake looks like something out of a science fiction film set in a scorching alien world. This otherworldly scene is the test site for a company called Lilac Solutions, which is developing a technology it says will shake up the United States' efforts to pry control over the global supply of lithium, the so-called "white gold" needed for electric vehicles and batteries, away from China. The startup is in a race to commercialize a new, less environmentally-damaging way to extract lithium from rocks. If everything pans out, it could significantly increase domestic supply at a crucial moment for the nation's lithium extraction industry. Last week OpenAI released Sora, a TikTok-style app that presents an endless feed of exclusively AI-generated videos, each up to 10 seconds long. The app allows you to create a "cameo" of yourself--a hyperrealistic avatar that mimics your appearance and voice--and insert other peoples' cameos into your own videos (depending on what permissions they set).


MrBeast says AI advance is scary for YouTube creators

BBC News

MrBeast: AI means it's'scary times' for YouTube creators The world's biggest YouTuber, MrBeast, says the rapid advance of generative artificial intelligence (AI) is scary for the millions of creators currently making content for a living. AI tools that can create fully-formed videos from simple text prompts by users have made rapid advances in recent years. On social media, MrBeast, real name Jimmy Donaldson, asked what would happen to people like him when AI videos are just as good as normal videos. Fears about the impact AI will have on the jobs market are widespread - but particularly acute in the creative industries. In the film and video game industries, there has been extensive industrial action over the use of AI.


The Future of AI Filmmaking Is a Parody of the Apocalypse, Made by a Guy Named Josh

WIRED

The filmmaker could not get Tiggy the alien to cooperate. He just needed the glistening brown creature to turn its head. But Tiggy, who was sitting in the passenger's seat of a cop car, kept disobeying. At first Tiggy rotated his gaze only slightly. Then he looked to the wrong side of the camera. Then his skin turned splotchy, like an overripe fruit. The filmmaker was not on a movie set, or Mars. He was sitting at his home computer in Los Angeles using a piece of AI software called FLUX Kontext to generate and regenerate images of the alien, waiting for a workable one to appear. He'd used a different AI tool, Midjourney, to generate the very first image of Tiggy (prompt: "fat blob alien with a tiny mouth and tiny lips"); one called ElevenLabs to create the timbre of Tiggy's voice (the filmmaker's voice overlaid with a synthetic one, then pitch-shifted way up); and yet another called Runway to describe the precise shot he wanted in this scene ("close up on the little alien as they ride in the passenger seat, shallow depth of field").