Goto

Collaborating Authors

 Large Language Model


Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning

arXiv.org Artificial Intelligence

Large language models have demonstrated remarkable capabilities in complex mathematical reasoning tasks, but they inevitably generate errors throughout multi-step solutions. Process-level Reward Models (PRMs) have shown great promise by providing supervision and evaluation at each intermediate step, thereby effectively improving the models' reasoning abilities. However, training effective PRMs requires high-quality process reward data, yet existing methods for constructing such data are often labour-intensive or inefficient. In this paper, we propose an uncertainty-driven framework for automated process reward data construction, encompassing both data generation and annotation processes for PRMs. Additionally, we identify the limitations of both majority vote and PRMs, and introduce two generic uncertainty-aware output aggregation methods: Hybrid Majority Reward Vote and Weighted Reward Frequency Vote, which combine the strengths of majority vote with PRMs. Extensive experiments on ProcessBench, MATH, and GSMPlus show the effectiveness and efficiency of the proposed PRM data construction framework, and demonstrate that the two output aggregation methods further improve the mathematical reasoning abilities across diverse PRMs. The code and data will be publicly available at https://github.com/Jiuzhouh/UnPRM.


Harnessing Collective Intelligence of LLMs for Robust Biomedical QA: A Multi-Model Approach

arXiv.org Artificial Intelligence

Biomedical text mining and question-answering are essential yet highly demanding tasks, particularly in the face of the exponential growth of biomedical literature. In this work, we present our participation in the 13th edition of the BioASQ challenge, which involves biomedical semantic question-answering for Task 13b and biomedical question-answering for developing topics for the Synergy task. We deploy a selection of open-source large language models (LLMs) as retrieval-augmented generators to answer biomedical questions. Various models are used to process the questions. A majority voting system combines their output to determine the final answer for Yes/No questions, while for list and factoid type questions, the union of their answers in used. We evaluated 13 state-of-the-art open source LLMs, exploring all possible model combinations to contribute to the final answer, resulting in tailored LLM pipelines for each question type. Our findings provide valuable insight into which combinations of LLMs consistently produce superior results for specific question types. In the four rounds of the 2025 BioASQ challenge, our system achieved notable results: in the Synergy task, we secured 1st place for ideal answers and 2nd place for exact answers in round 2, as well as two shared 1st places for exact answers in round 3 and 4.


TripTailor: A Real-World Benchmark for Personalized Travel Planning

arXiv.org Artificial Intelligence

The continuous evolution and enhanced reasoning capabilities of large language models (LLMs) have elevated their role in complex tasks, notably in travel planning, where demand for personalized, high-quality itineraries is rising. However, current benchmarks often rely on unrealistic simulated data, failing to reflect the differences between LLM-generated and real-world itineraries. Existing evaluation metrics, which primarily emphasize constraints, fall short of providing a comprehensive assessment of the overall quality of travel plans. To address these limitations, we introduce TripTailor, a benchmark designed specifically for personalized travel planning in real-world scenarios. This dataset features an extensive collection of over 500,000 real-world points of interest (POIs) and nearly 4,000 diverse travel itineraries, complete with detailed information, providing a more authentic evaluation framework. Experiments show that fewer than 10\% of the itineraries generated by the latest state-of-the-art LLMs achieve human-level performance. Moreover, we identify several critical challenges in travel planning, including the feasibility, rationality, and personalized customization of the proposed solutions. We hope that TripTailor will drive the development of travel planning agents capable of understanding and meeting user needs while generating practical itineraries. Our code and dataset are available at https://github.com/swxkfm/TripTailor


Win-k: Improved Membership Inference Attacks on Small Language Models

arXiv.org Artificial Intelligence

Small language models (SLMs) are increasingly valued for their efficiency and deployability in resource-constrained environments, making them useful for on-device, privacy-sensitive, and edge computing applications. On the other hand, membership inference attacks (MIAs), which aim to determine whether a given sample was used in a model's training, are an important threat with serious privacy and intellectual property implications. In this paper, we study MIAs on SLMs. Although MIAs were shown to be effective on large language models (LLMs), they are relatively less studied on emerging SLMs, and furthermore, their effectiveness decreases as models get smaller. Motivated by this finding, we propose a new MIA called win-k, which builds on top of a state-of-the-art attack (min-k). We experimentally evaluate win-k by comparing it with five existing MIAs using three datasets and eight SLMs.


WebDS: An End-to-End Benchmark for Web-based Data Science

arXiv.org Artificial Intelligence

A large portion of real-world data science tasks are complex and require multi-hop web-based interactions: finding appropriate data available on the internet, synthesizing real-time data of various modalities from different locations, and producing summarized analyses. Existing web benchmarks often focus on simplistic interactions, such as form submissions or e-commerce transactions, and often do not require diverse tool-using capabilities required for web based data science. Conversely, traditional data science benchmarks typically concentrate on static, often textually bound datasets and do not assess end-to-end workflows that encompass data acquisition, cleaning, analysis, and insight generation. In response, we introduce WebDS, the first end-to-end web-based data science benchmark. It comprises 870 web-based data science tasks across 29 diverse websites from structured government data portals to unstructured news media, challenging agents to perform complex, multi-step operations requiring the use of tools and heterogeneous data formats that better reflect the realities of modern data analytics. Evaluations of current SOTA LLM agents indicate significant performance gaps in accomplishing these tasks. For instance, Browser Use, which accomplishes 80% of tasks on Web Voyager, successfully completes only 15% of tasks in WebDS, which our analysis suggests is due to new failure modes like poor information grounding, repetitive behavior and shortcut-taking that agents performing WebDS' tasks display. By providing a more robust and realistic testing ground, WebDS sets the stage for significant advances in the development of practically useful LLM-based data science.


Searching for ChatGPT? Bing begs you to use Copilot instead

PCWorld

Stop us if you've heard this before: Microsoft encourages you not to visit its competition. You may have noticed that if you visit Bing.com and then search for Google, Microsoft might remind you that it too has a search engine. More recently, Microsoft is now encouraging you to remain within its ecosystem and use Copilot instead of venturing elsewhere to Google, OpenAI, or Meta. When searching for "Claude" within Microsoft Edge -- I use Edge with Bing set at its search engine -- I tried looking up "Claude," the AI tool developed by Anthropic. While Bing dutifully returned the link as well as related information, it also reminded me that "Your Copilot is here," accompanied with a Copilot-specific search box.


Revealed: The richest and youngest AI-billionaires making fortune from the big tech boom

Daily Mail - Science & tech

From helping you answer emails to translating legal documents, artificial intelligence is now a part of almost all facets of life. Meanwhile, organisations from Microsoft and Apple to the NHS have piled vast sums of funding into the latest intelligent software. And for the few people behind this AI boom, there have been enormous profits to be made. Leading the pack as the richest of new AI billionaires is Jensen Huang, CEO of chipmaker Nvidia, with a staggering net-worth of 113 billion ( 151bn). Mr Huang joins several monumental big tech figures, such as Meta's Mark Zuckerberg and Elon Musk, who have recently made huge investments in AI.


Tesla awards boss Elon Musk 29bn in shares

BBC News

"It is imperative to retain and motivate our extraordinary talent, beginning with Elon", Tesla's board wrote on X, a platform owned by Musk, adding that "no one matches Elon's remarkable combination of leadership experience, technical expertise". The company said the billionaire had a "proven track record" in building "revolutionary and profitable businesses". Tech firms trying to assert themselves in the AI sector have been offering huge sums to workers at rivals in an effort to persuade them to join them and boost their development. Facebook founder Mark Zuckerberg was said to have recently tried to lure top developers from ChatGPT-creator OpenAI with million-dollar pay deals. Meanwhile Microsoft's AI division, headed up by former Google DeepMind co-founder Mustafa Suleyman, recently gained several new hires from Google's ranks. Tesla the company was at an "inflection point" and needed Musk's prowess as it pivots from being an electric vehicle firm to an AI and robotics focussed company.


Demis Hassabis on our AI future: 'It'll be 10 times bigger than the Industrial Revolution – and maybe 10 times faster'

The Guardian

The head of Google’s DeepMind says artificial intelligence could usher in an era of ‘incredible productivity’ and ‘radical abundance’. But who will it benefit? And why does he wish the tech giants had moved more slowly?


Calibrated Language Models and How to Find Them with Label Smoothing

arXiv.org Machine Learning

Recent advances in natural language processing (NLP) have opened up greater opportunities to enable fine-tuned large language models (LLMs) to behave as more powerful interactive agents through improved instruction-following ability. However, understanding how this impacts confidence calibration for reliable model output has not been researched in full. In this work, we examine various open-sourced LLMs, identifying significant calibration degradation after instruction tuning in each. Seeking a practical solution, we look towards label smoothing, which has been shown as an effective method to regularize for overconfident predictions but has yet to be widely adopted in the supervised fine-tuning (SFT) of LLMs. We first provide insight as to why label smoothing is sufficient to maintain calibration throughout the SFT process. However, settings remain where the effectiveness of smoothing is severely diminished, in particular the case of large vocabulary LLMs (LV-LLMs). We posit the cause to stem from the ability to become over-confident, which has a direct relationship with the hidden size and vocabulary size, and justify this theoretically and experimentally. Finally, we address an outstanding issue regarding the memory footprint of the cross-entropy loss computation in the label smoothed loss setting, designing a customized kernel to dramatically reduce memory consumption without sacrificing speed or performance in comparison to existing solutions for non-smoothed losses.