Large Language Model
Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
Dedhia, Bhishma, Kansal, Yuval, Jha, Niraj K.
Language models traditionally used for cross-domain generalization have recently demonstrated task-specific reasoning. However, their top-down training approach on general corpora is insufficient for acquiring abstractions needed for deep domain expertise. This may require a bottom-up approach that acquires expertise by learning to compose simple domain concepts into more complex ones. A knowledge graph (KG) provides this compositional structure, where domain primitives are represented as head-relation-tail edges and their paths encode higher-level concepts. We present a task generation pipeline that synthesizes tasks directly from KG primitives, enabling models to acquire and compose them for reasoning. We fine-tune language models on the resultant KG-grounded curriculum to demonstrate domain-specific superintelligence. While broadly applicable, we validate our approach in medicine, where reliable KGs exist. Using a medical KG, we curate 24,000 reasoning tasks paired with thinking traces derived from diverse medical primitives. We fine-tune the QwQ-32B model on this curriculum to obtain QwQ-Med-3 that takes a step towards medical superintelligence. We also introduce ICD-Bench, an evaluation suite to quantify reasoning abilities across 15 medical domains. Our experiments demonstrate that QwQ-Med-3 significantly outperforms state-of-the-art reasoning models on ICD-Bench categories. Further analysis reveals that QwQ-Med-3 utilizes acquired primitives to widen the performance gap on the hardest tasks of ICD-Bench. Finally, evaluation on medical question-answer benchmarks shows that QwQ-Med-3 transfers acquired expertise to enhance the base model's performance. While the industry's approach to artificial general intelligence (AGI) emphasizes broad expertise, we envision a future in which AGI emerges from the composable interaction of efficient domain-specific superintelligent agents.
Agentic-R1: Distilled Dual-Strategy Reasoning
Du, Weihua, Aggarwal, Pranjal, Welleck, Sean, Yang, Yiming
Current long chain-of-thought (long-CoT) models excel at mathematical reasoning but rely on slow and error-prone natural language traces. Tool-augmented agents address arithmetic via code execution, but often falter on complex logical tasks. We introduce a fine-tuning framework, DualDistill, that distills complementary reasoning strategies from multiple teachers into a unified student model. Using this approach, we train Agentic-R1, which dynamically selects the optimal strategy for each query, invoking tools for arithmetic and algorithmic problems, and using text-based reasoning for abstract ones. Our method improves accuracy across a range of tasks, including both computation-intensive and standard benchmarks, demonstrating the effectiveness of multi-strategy distillation in achieving robust and efficient reasoning. Our project is available at https://github.com/StigLidu/DualDistill
Oversight Structures for Agentic AI in Public-Sector Organizations
Schmitz, Chris, Rystrรธm, Jonathan, Batzner, Jan
This paper finds that the introduction of agentic AI systems intensifies existing challenges to traditional public sector oversight mechanisms -- which rely on siloed compliance units and episodic approvals rather than continuous, integrated supervision. We identify five governance dimensions essential for responsible agent deployment: cross-departmental implementation, comprehensive evaluation, enhanced security protocols, operational visibility, and systematic auditing. We evaluate the capacity of existing oversight structures to meet these challenges, via a mixed-methods approach consisting of a literature review and interviews with civil servants in AI-related roles. We find that agent oversight poses intensified versions of three existing governance challenges: continuous oversight, deeper integration of governance and operational capabilities, and interdepartmental coordination. We propose approaches that both adapt institutional structures and design agent oversight compatible with public sector constraints.
FinS-Pilot: A Benchmark for Online Financial RAG System
Wang, Feng, Sun, Yiding, Mao, Jiaxin, Xue, Wei, Xu, Danqing
Large language models (LLMs) have demonstrated remarkable capabilities across various professional domains, with their performance typically evaluated through standardized benchmarks. In the financial field, the stringent demands for professional accuracy and real-time data processing often necessitate the use of retrieval-augmented generation (RAG) techniques. However, the development of financial RAG benchmarks has been constrained by data confidentiality issues and the lack of dynamic data integration. To address this issue, we introduce FinS-Pilot, a novel benchmark for evaluating RAG systems in online financial applications. Constructed from real-world financial assistant interactions, our benchmark incorporates both real-time API data and text data, organized through an intent classification framework covering critical financial domains. The benchmark enables comprehensive evaluation of financial assistants' capabilities in handling both static knowledge and time-sensitive market information.Through systematic experiments with multiple Chinese leading LLMs, we demonstrate FinS-Pilot's effectiveness in identifying models suitable for financial applications while addressing the current gap in specialized evaluation tools for the financial domain. Our work contributes both a practical evaluation framework and a curated dataset to advance research in financial NLP systems. The code and dataset are accessible on GitHub.
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
This study illustrates how incorporating feedback-oriented annotations into the scoring pipeline can enhance the accuracy of automated essay scoring (AES). This approach is demonstrated with the Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements (PERSUADE) corpus. We integrate two types of feedback-driven annotations: those that identify spelling and grammatical errors, and those that highlight argumentative components. To illustrate how this method could be applied in real-world scenarios, we employ two LLMs to generate annotations -- a generative language model used for spell correction and an encoder-based token-classifier trained to identify and mark argumentative elements. By incorporating annotations into the scoring process, we demonstrate improvements in performance using encoder-based large language models fine-tuned as classifiers.
Therapists are secretly using ChatGPT. Clients are triggered.
Declan was so shocked he didn't say anything, and for the rest of the session he was privy to a real-time stream of ChatGPT analysis rippling across his therapist's screen. The session became even more surreal when Declan began echoing ChatGPT in his own responses, preempting his therapist. "I became the best patient ever," he says, "because ChatGPT would be like, 'Well, do you consider that your way of thinking might be a little too black and white?' And I would be like, 'Huh, you know, I think my way of thinking might be too black and white,' and [my therapist would] be like, 'Exactly.' I'm sure it was his dream session."
Latam-GPT: The Free, Open Source, and Collaborative AI of Latin America
Latam-GPT is new large language model being developed in and for Latin America. The project, led by the nonprofit Chilean National Center for Artificial Intelligence (CENIA), aims to help the region achieve technological independence by developing an open source AI model trained on Latin American languages and contexts. "This work cannot be undertaken by just one group or one country in Latin America: It is a challenge that requires everyone's participation," says รlvaro Soto, director of CENIA, in an interview with WIRED en Espaรฑol. "Latam-GPT is a project that seeks to create an open, free, and, above all, collaborative AI model. We've been working for two years with a very bottom-up process, bringing together citizens from different countries who want to collaborate. Recently, it has also seen some more top-down initiatives, with governments taking an interest and beginning to participate in the project."
WIRED Roundup: Meta's AI Brain Drain
In today's episode, our host Zoรซ Schiffer is joined by WIRED's senior politics editor Leah Feiger to run through five of this week's best stories--from how AI is eliminating entry level jobs to why a secretive Democrat group is funding high-profile influencers. Then, Zoรซ and Leah dive into the scoop that AI researchers recently recruited to Meta Superintelligence Labs are already leaving--with some heading back to OpenAI. Join us live in San Francisco on September 9. Get your tickets here. Mentioned in this episode: Researchers Are Already Leaving Meta's New Superintelligence Lab by Zoรซ Schiffer and Will Knight AI Is Eliminating Jobs for Younger Workers by Will Knight Elon Musk's xAI Sues Apple and OpenAI Over App Store Rankings by Zoรซ Schiffer A Dark Money Group Is Secretly Funding High-Profile Democratic Influencers by Taylor Lorenz What It's Like Watching Dozens of Bodies Decompose (for Science) by Jess Thomson Write to us at uncannyvalley@wired.com. You can always listen to this week's podcast through the audio player on this page, but if you want to subscribe for free to get every episode, here's how: If you're on an iPhone or iPad, open the app called Podcasts, or just tap this link.
Free AI training comes to California colleges -- but at what cost?
As artificial intelligence replaces entry-level jobs, California's universities and community colleges are offering a glimmer of hope for students: free AI training that will help them master the new technology. "You're seeing in certain coding spaces significant declines in hiring for obvious reasons," Gov. Gavin Newsom said in early August from the seventh floor of Google's San Francisco office. Flanked by leadership from California's higher education systems, he called attention to the recent layoffs at Microsoft, Google's parent company, Alphabet, and at nearby Salesforce Tower, home to the tech company that is still the city's largest private employer. Now, some of those companies -- including Google and Microsoft -- will offer a suite of AI resources free to California schools and universities. In return, the companies could gain access to millions of new users.
Robustness is Important: Limitations of LLMs for Data Fitting
Liu, Hejia, Yang, Mochen, Adomavicius, Gediminas
Large Language Models (LLMs) are being applied in a wide array of settings, well beyond the typical language-oriented use cases. In particular, LLMs are increasingly used as a plug-and-play method for fitting data and generating predictions. Prior work has shown that LLMs, via in-context learning or supervised fine-tuning, can perform competitively with many tabular supervised learning techniques in terms of predictive performance. However, we identify a critical vulnerability of using LLMs for data fitting -- making changes to data representation that are completely irrelevant to the underlying learning task can drastically alter LLMs' predictions on the same data. For example, simply changing variable names can sway the size of prediction error by as much as 82% in certain settings. Such prediction sensitivity with respect to task-irrelevant variations manifests under both in-context learning and supervised fine-tuning, for both close-weight and open-weight general-purpose LLMs. Moreover, by examining the attention scores of an open-weight LLM, we discover a non-uniform attention pattern: training examples and variable names/values which happen to occupy certain positions in the prompt receive more attention when output tokens are generated, even though different positions are expected to receive roughly the same attention. This partially explains the sensitivity in the presence of task-irrelevant variations. We also consider a state-of-the-art tabular foundation model (TabPFN) trained specifically for data fitting. Despite being explicitly designed to achieve prediction robustness, TabPFN is still not immune to task-irrelevant variations. Overall, despite LLMs' impressive predictive capabilities, currently they lack even the basic level of robustness to be used as a principled data-fitting tool.