Large Language Model
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
Peng, Yebo, Liu, Zixiang, Li, Yaoming, Yang, Zhizhuo, Xu, Xinye, Ye, Bowen, Yuan, Weijun, Wang, Zihan, Yang, Tong
Evaluating the mathematical capability of Large Language Models (LLMs) is a critical yet challenging frontier. Existing benchmarks fall short, particularly for proof-centric problems, as manual creation is unscalable and costly, leaving the true mathematical abilities of LLMs largely unassessed. To overcome these barriers, we propose Proof2Hybrid, the first fully automated framework that synthesizes high-quality, proof-centric benchmarks from natural language mathematical corpora. The key novelty of our solution is Proof2X, a roadmap of converting mathematical proofs into various kinds of questions that are easy to verify. Instructed by this roadmap, we propose a new type of hybrid-formatted questions, named ``$m$-out-of-$n$ multiple judge questions'', specifically designed to enable robust, automatic evaluation while being resilient to guessing and superficial pattern matching inherent in traditional formats. As a demonstration of our framework, we introduce AlgGeoTest, a benchmark for algebraic geometry--a frontier domain of modern mathematics--comprising 456 challenging items. Our extensive evaluations on state-of-the-art LLMs using AlgGeoTest reveal profound deficits in their comprehension of algebraic geometry, providing a more precise measure of their true mathematical capabilities. Our framework and benchmark pave the way for a new wave of in-depth research into the mathematical intelligence of AI systems.
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
Pham, Kien T., He, Yingqing, Xing, Yazhou, Chen, Qifeng, Chen, Long
Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring semantic information, such as the classes of sounding sources present in the audio, limiting their ability to generate videos with accurate content and spatial composition. In contrast, we humans can not only naturally identify the semantic categories of sounding sources but also determine their deeply encoded spatial attributes, including locations and movement directions. This useful information can be elucidated by considering specific spatial indicators derived from the inherent physical properties of sound, such as loudness or frequency. As prior methods largely ignore this factor, we present SpA2V, the first framework explicitly exploits these spatial auditory cues from audios to generate videos with high semantic and spatial correspondence. SpA2V decomposes the generation process into two stages: 1) Audio-guided Video Planning: We meticulously adapt a state-of-the-art MLLM for a novel task of harnessing spatial and semantic cues from input audio to construct Video Scene Layouts (VSLs). This serves as an intermediate representation to bridge the gap between the audio and video modalities. 2) Layout-grounded Video Generation: We develop an efficient and effective approach to seamlessly integrate VSLs as conditional guidance into pre-trained diffusion models, enabling VSL-grounded video generation in a training-free manner. Extensive experiments demonstrate that SpA2V excels in generating realistic videos with semantic and spatial alignment to the input audios.
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
Large Language Model (LLM) applications are increasingly relying on external tools to extend their capabilities beyond text generation. However, current tool integration approaches suffer from fragmentation, protocol limitations, and implementation complexity, leading to substantial development overhead. This paper presents Toolregistry, a protocol-agnostic tool management library that simplifies tool registration, representation, execution, and lifecycle management via a unified interface. Our evaluation demonstrates that \toolregistry achieves 60-80% reduction in tool integration code, up to 3.1x performance improvements through concurrent execution, and 100% compatibility with OpenAI function calling standards. Real-world case studies show significant improvements in development efficiency and code maintainability across diverse integration scenarios. \toolregistry is open-source and available at https://github.com/Oaklight/ToolRegistry, with comprehensive documentation at https://toolregistry.readthedocs.io/.
Enhancing Spectral Graph Neural Networks with LLM-Predicted Homophily
Lu, Kangkang, Yu, Yanhua, Huang, Zhiyong, Chua, Tat-Seng
Spectral Graph Neural Networks (SGNNs) have achieved remarkable performance in tasks such as node classification due to their ability to learn flexible filters. Typically, these filters are learned under the supervision of downstream tasks, enabling SGNNs to adapt to diverse structural patterns. However, in scenarios with limited labeled data, SGNNs often struggle to capture the optimal filter shapes, resulting in degraded performance, especially on graphs with heterophily. Meanwhile, the rapid progress of Large Language Models (LLMs) has opened new possibilities for enhancing graph learning without modifying graph structure or requiring task-specific training. In this work, we propose a novel framework that leverages LLMs to estimate the homophily level of a graph and uses this global structural prior to guide the construction of spectral filters. Specifically, we design a lightweight and plug-and-play pipeline where a small set of labeled node pairs is formatted as natural language prompts for the LLM, which then predicts the graph's homophily ratio. This estimated value informs the spectral filter basis, enabling SGNNs to adapt more effectively to both homophilic and heterophilic structures. Extensive experiments on multiple benchmark datasets demonstrate that our LLM-assisted spectral framework consistently improves performance over strong SGNN baselines.
OpenAI releases two 'open' AI models after DeepSeek's success
OpenAI is releasing a pair of open and freely available artificial intelligence models that can mimic the human process of reasoning, months after China's DeepSeek gained global attention with its own open AI software. The two models, called GPT-oss-120b and GPT-oss-20b, will be available on AI software hosting platform Hugging Face and can produce text -- but not images or videos -- in response to user prompts, OpenAI said on Tuesday. These models can also carry out complex tasks like writing code and looking up information online on a user's behalf, the company said. Crucially, the models are both open-weight systems, similar to Meta Platforms' Llama. The term "weight" refers to the parameters in an AI model.
OpenAI takes on Meta and DeepSeek with free and customisable AI models
OpenAI is taking on Mark Zuckerberg's Meta and Chinese rival DeepSeek by launching its own freely available artificial intelligence models. The ChatGPT developer has announced two "open weight" large language models, which are free to download and can be customised by developers. Meta's Llama models are available on a similar basis, and OpenAI's move marks a departure from ChatGPT, which is based on a "closed" model that cannot be customised. Sam Altman, OpenAI's chief executive, said the company was excited to add to a stack of freely available AI models "based on democratic values โฆ and for wide benefit". He added: "We're excited to make this model, the result of billions of dollars of research, available to the world to get AI into the hands of the most people possible." OpenAI said the models could underpin an AI agent that operates autonomously, and that they were "designed to be used within agentic workflows".
OpenAI Just Released Its First Open-Weight Models Since GPT-2
OpenAI just dropped its first open-weight models in over five years. The two language models, gpt-oss-120b and gpt-oss-20b, can run locally on consumer devices and be fine-tuned for specific purposes. For OpenAI, they represent a shift away from its recent strategy of focusing on proprietary releases, as the company moves towards a wider, and more open, group of AI models that are available for users. "We're excited to make this model, the result of billions of dollars of research, available to the world to get AI into the hands of the most people possible," said OpenAI CEO Sam Altman in an emailed statement. Both gpt-oss-120b and gpt-oss-20b are officially available to download for free on Hugging Face, a popular hosting platform for AI tools.
OpenAI has finally released open-weight language models
"The vast majority of our [enterprise and startup] customers are already using a lot of open models," said Casey Dvorak, a research program manager at OpenAI, in a media briefing about the model release. "Because there is no [competitive] open model from OpenAI, we wanted to plug that gap and actually allow them to use our technology across the board." The new models come in two different sizes, the smaller of which can theoretically run on 16 GB of RAM--the minimum amount that Apple currently offers on its computers. The larger model requires a high-end laptop or specialized hardware. Open models have a few key use cases.
OpenAI stops ChatGPT from telling people to break up with partners
ChatGPT will not tell people to break up with their partner and will encourage users to take breaks from long chatbot sessions, under new changes to the artificial intelligence tool. OpenAI, ChatGPT's developer, said the chatbot would stop giving definitive answers to personal challenges and would instead help people to mull over problems such as potential breakups. "When you ask something like: 'Should I break up with my boyfriend?' ChatGPT shouldn't give you an answer. It should help you think it through โ asking questions, weighing pros and cons," said OpenAI.
Claude Fans Threw a Funeral for Anthropic's Retired AI Model
On July 21 at 9 am PT, Anthropic retired Claude 3 Sonnet, a lightweight model known for being quick and cost-effective. On Saturday, in a large warehouse in San Francisco's SOMA district, more than 200 people gathered to mourn its passing. The star-studded funeral was put on by a group of Claude fanatics and Gen Z founders, one of whom told me he dropped out of college after learning about artificial general intelligence. Attendees included Amanda Askell, an Anthropic researcher who has jokingly called herself the "Fairy Claudemother," staffers from Anthropic and OpenAI, and high-profile X posters including the writer Noah Smith. The warehouse was dimly lit, with a tentacle from a shoggoth (a fictional H.P. Lovecraft creature that's become a popular metaphor for AI models) hanging from the ceiling.