Large Language Model
Deep Research Bench: Evaluating AI Web Research Agents
FutureSearch, null, :, null, Bosse, Nikos I., Evans, Jon, Gambee, Robert G., Hnyk, Daniel, Mühlbacher, Peter, Phillips, Lawrence, Schwarz, Dan, Wildman, Jack
Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the continually-changing web. We introduce Deep Research Bench, consisting of 89 multi-step web research task instances of varying difficulty across 8 diverse task categories, with the answers carefully worked out by skilled humans. We provide a "RetroSearch" environment with a large frozen set of scraped web pages, and demonstrate that offline "RetroSearch" agents perform comparably to "live web" agents, enabling reliable evaluations of models over time. We provide robust agent tooling and scaffolding to benchmark major LLMs as they are released, including "thinking" models like o3 and Gemini 2.5 Pro. We include automated evaluations of the lengthy agent traces to report progress over time in hallucinations, tool use, and forgetting. Finally, we evaluate the major web research products branded as "Deep Research", "Deep Search", "Search", or "Research." Results are available on a public leaderboard at https://drb.futuresearch.ai/.
Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach
Deng, Shuangyan, Peng, Haizhou, Xu, Jiachen, Liu, Chunhou, Giurcuaneanu, Ciprian Doru, Liu, Jiamou
Effective financial reasoning demands not only textual understanding but also the ability to interpret complex visual data such as charts, tables, and trend graphs. This paper introduces a new benchmark designed to evaluate how well AI models - especially large language and multimodal models - reason in finance-specific contexts. Covering 3,200 expert-level question-answer pairs across 15 core financial topics, the benchmark integrates both textual and visual modalities to reflect authentic analytical challenges in finance. To address limitations in current reasoning approaches, we propose an error-aware learning framework that leverages historical model mistakes and feedback to guide inference, without requiring fine-tuning. Our experiments across state-of-the-art models show that multimodal inputs significantly enhance performance and that incorporating error feedback leads to consistent and measurable improvements. The results highlight persistent challenges in visual understanding and mathematical logic, while also demonstrating the promise of self-reflective reasoning in financial AI systems. Our code and data can be found at https://anonymous/FinMR/CodeData.
AI takes backseat as Apple unveils software revamp and new apps
Apple's artificial intelligence features took a backseat on Monday at its latest annual Worldwide Developers Conference. The company announced a revamped software design called Liquid Glass, new phone and camera apps as well as new features on Apple Watch and Vision Pro. But in spite of pressure to compete with firms that have gone all-in on AI, Apple's AI announcements were limited to incremental features and upgrades. Users will have a few new Apple Intelligence-powered features to look forward to including live translation, a real-time language translation feature that will be integrated into messages, FaceTime and the Phone app. The Android operating system has offered a similar feature for several years.
Advanced AI suffers 'complete accuracy collapse' in face of complex problems, study finds
Apple researchers have found "fundamental limitations" in cutting-edge artificial intelligence models, in a paper raising doubts about the technology industry's race to develop ever more powerful systems. Apple said in a paper published at the weekend that large reasoning models (LRMs) – an advanced form of AI – faced a "complete accuracy collapse" when presented with highly complex problems. It found that standard AI models outperformed LRMs in low-complexity tasks, while both types of model suffered "complete collapse" with high-complexity tasks. Large reasoning models attempt to solve complex queries by generating detailed thinking processes that break down the problem into smaller steps. The study, which tested the models' ability to solve puzzles, added that as LRMs neared performance collapse they began "reducing their reasoning effort".
I'm a Polite Person. But in This One Specific Situation, I Recommend Being a Total Jerk.
Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. Fairly recently, I started being verbally abusive to large language models. I highly recommend you experiment with doing so yourself. Over the past 30 days, I have called large language models (primarily OpenAI's paid product) the following names, among others that I won't repeat here because my mom might read this: Dipshit, fucknuts, shitstain, dummy, dumbass, dum-dum fucking dumbass dum-dum, numbnuts, hockey puck (thank you, Don Rickles), turdburger, lickspittle, cockroach, fucking cockroach (thank you, Tony Montana), idiot, fucking idiot, total fucking idiot, and fucking numbnuts dipshit. Ethan Mollick, author of Co-Intelligence: Living and Working With AI, and currently the reigning A.I. whisperer for the consultant class, says that anthropomorphizing A.I. is "a sin of necessity."
Meta set to throw billions at startup that leads AI data market
Three months after the Chinese artificial intelligence developer DeepSeek upended the tech world with a model that rivaled America's best, a 28-year-old AI executive named Alexandr Wang came to Capitol Hill to tell policymakers what they needed to do to maintain U.S. dominance. The U.S. needs to establish a "national AI data reserve," supply enough power for data centers and avoid an onerous patchwork of state-level rules, Wang said at the April hearing. "It's good to see you again here in Washington," Republican Representative Neal Dunn of Florida said. Wang, the chief executive officer of Scale AI, may not be a household name in the same way OpenAI's Sam Altman has become. But he and his company have gained significant influence in tech and policy circles in recent years.
Conformal Prediction Adaptive to Unknown Subpopulation Shifts
Wang, Nien-Shao, Yaldiz, Duygu Nur, Bakman, Yavuz Faruk, Karimireddy, Sai Praneeth
Conformal prediction is widely used to equip black-box machine learning models with uncertainty quantification enjoying formal coverage guarantees. However, these guarantees typically break down in the presence of distribution shifts, where the data distribution at test time differs from the training (or calibration-time) distribution. In this work, we address subpopulation shifts, where the test environment exhibits an unknown and differing mixture of subpopulations compared to the calibration data. We propose new methods that provably adapt conformal prediction to such shifts, ensuring valid coverage without requiring explicit knowledge of subpopulation structure. Our algorithms scale to high-dimensional settings and perform effectively in realistic machine learning tasks. Extensive experiments on vision (with vision transformers) and language (with large language models) benchmarks demonstrate that our methods reliably maintain coverage and controls risk in scenarios where standard conformal prediction fails.
Zero-shot protein stability prediction by inverse folding models: a free energy interpretation
Frellsen, Jes, Kassem, Maher M., Bengtsen, Tone, Olsen, Lars, Lindorff-Larsen, Kresten, Ferkinghoff-Borg, Jesper, Boomsma, Wouter
Inverse folding models have proven to be highly effective zero-shot predictors of protein stability. Despite this success, the link between the amino acid preferences of an inverse folding model and the free-energy considerations underlying thermodynamic stability remains incompletely understood. A better understanding would be of interest not only from a theoretical perspective, but also potentially provide the basis for stronger zero-shot stability prediction. In this paper, we take steps to clarify the free-energy foundations of inverse folding models. Our derivation reveals the standard practice of likelihood ratios as a simplistic approximation and suggests several paths towards better estimates of the relative stability. We empirically assess these approaches and demonstrate that considerable gains in zero-shot performance can be achieved with fairly simple means.
Preference Learning for AI Alignment: a Causal Perspective
Kobalczyk, Katarzyna, van der Schaar, Mihaela
Reward modelling from preference data is a crucial step in aligning large language models (LLMs) with human values, requiring robust generalisation to novel prompt-response pairs. In this work, we propose to frame this problem in a causal paradigm, providing the rich toolbox of causality to identify the persistent challenges, such as causal misidentification, preference heterogeneity, and confounding due to user-specific factors. Inheriting from the literature of causal inference, we identify key assumptions necessary for reliable generalisation and contrast them with common data collection practices. We illustrate failure modes of naive reward models and demonstrate how causally-inspired approaches can improve model robustness. Finally, we outline desiderata for future research and practices, advocating targeted interventions to address inherent limitations of observational data.
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Lei, Xuanyu, Li, Chenliang, Wu, Yuning, Liu, Kaiming, Shen, Weizhou, Li, Peng, Yan, Ming, Zhang, Ji, Huang, Fei, Liu, Yang
Recent advances in Large Language Models (LLMs) have enabled strong performance in long-form writing, yet existing supervised fine-tuning (SFT) approaches suffer from limitations such as data saturation and restricted learning capacity bounded by teacher signals. In this work, we present Writing-RL: an Adaptive Curriculum Reinforcement Learning framework to advance long-form writing capabilities beyond SFT. The framework consists of three key components: Margin-aware Data Selection strategy that prioritizes samples with high learning potential, Pairwise Comparison Reward mechanism that provides discriminative learning signals in the absence of verifiable rewards, and Dynamic Reference Scheduling approach, which plays a particularly critical role by adaptively adjusting task difficulty based on evolving model performance. Experiments on 7B-scale writer models show that our RL framework largely improves long-form writing performance over strong SFT baselines. Furthermore, we observe that models trained with long-output RL generalize surprisingly well to long-input reasoning tasks, potentially offering a promising perspective for rethinking long-context training.