Large Language Model
Rethinking Graph Structure Learning in the Era of LLMs
Zhang, Zhihan, Li, Xunkai, Zeng, Guang, Qin, Hongchao, Li, Ronghua, Wang, Guoren
Recently, the emergence of large language models (LLMs) has prompted researchers to explore the integration of language descriptions into graphs, aiming to enhance model encoding capabilities from a data-centric perspective. This graph representation is called text-attributed graphs (TAGs). A review of prior advancements highlights that graph structure learning (GSL) is a pivotal technique for improving data utility, making it highly relevant to efficient TAG learning. However, most GSL methods are tailored for traditional graphs without textual information, underscoring the necessity of developing a new GSL paradigm. Despite clear motivations, it remains challenging: (1) How can we define a reasonable optimization objective for GSL in the era of LLMs, considering the massive parameters in LLM? (2) How can we design an efficient model architecture that enables seamless integration of LLM for this optimization objective? For Question 1, we reformulate existing GSL optimization objectives as a tree optimization framework, shifting the focus from obtaining a well-trained edge predictor to a language-aware tree sampler. For Question 2, we propose decoupled and training-free model design principles for LLM integration, shifting the focus from computation-intensive fine-tuning to more efficient inference. Based on this, we propose Large Language and Tree Assistant (LLaTA), which leverages tree-based LLM in-context learning to enhance the understanding of topology and text, enabling reliable inference and generating improved graph structure. Extensive experiments on 10 TAG datasets demonstrate that LLaTA enjoys flexibility - incorporated with any backbone; scalability - outperforms other LLM-based GSL methods in terms of running efficiency; effectiveness - achieves SOTA performance.
Few-Shot Graph Out-of-Distribution Detection with LLMs
Xu, Haoyan, Yao, Zhengtao, Dong, Yushun, Wang, Ziyi, Rossi, Ryan A., Li, Mengyuan, Zhao, Yue
Existing methods for graph out-of-distribution (OOD) detection typically depend on training graph neural network (GNN) classifiers using a substantial amount of labeled in-distribution (ID) data. However, acquiring high-quality labeled nodes in text-attributed graphs (TAGs) is challenging and costly due to their complex textual and structural characteristics. Large language models (LLMs), known for their powerful zero-shot capabilities in textual tasks, show promise but struggle to naturally capture the critical structural information inherent to TAGs, limiting their direct effectiveness. To address these challenges, we propose LLM-GOOD, a general framework that effectively combines the strengths of LLMs and GNNs to enhance data efficiency in graph OOD detection. Specifically, we first leverage LLMs' strong zero-shot capabilities to filter out likely OOD nodes, significantly reducing the human annotation burden. To minimize the usage and cost of the LLM, we employ it only to annotate a small subset of unlabeled nodes. We then train a lightweight GNN filter using these noisy labels, enabling efficient predictions of ID status for all other unlabeled nodes by leveraging both textual and structural information. After obtaining node embeddings from the GNN filter, we can apply informativeness-based methods to select the most valuable nodes for precise human annotation. Finally, we train the target ID classifier using these accurately annotated ID nodes. Extensive experiments on four real-world TAG datasets demonstrate that LLM-GOOD significantly reduces human annotation costs and outperforms state-of-the-art baselines in terms of both ID classification accuracy and OOD detection performance.
R-PRM: Reasoning-Driven Process Reward Modeling
She, Shuaijie, Liu, Junxiao, Liu, Yifeng, Chen, Jiajun, Huang, Xin, Huang, Shujian
Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by evaluating each reasoning step. However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy, which is further exacerbated by the scarcity of annotated data. To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM). First, we leverage stronger LLMs to generate seed data from limited annotations, effectively bootstrapping our model's reasoning capabilities and enabling comprehensive step-by-step evaluation. Second, we further enhance performance through preference optimization, without requiring additional annotated data. Third, we introduce inference-time scaling to fully harness the model's reasoning potential. Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 11.9 and 8.5 points in F1 scores, respectively. When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.5 points across six challenging datasets. Further analysis reveals that R-PRM exhibits more comprehensive evaluation and stronger generalization capabilities, thereby highlighting its significant potential.
Does AI Prediction Scale to Decision Making?
Artificial intelligence (AI) excels at prediction.1 Large language models (LLMs), for example, are remarkable at predicting the next word and stringing together fluent text. AI's predictive power extends beyond language to generate moving visuals and audio. AI models are trained with vast amounts of data and they use the statistical associations and patterns in the data to generate outputs. But does AI's ability to predict scale to decision making? Do relatively mundane (albeit impressive) forms of prediction--like predicting the next word--extend to reasoning in novel situations and to decision making in the real world?12
ChatGPT now speaks even more naturally with fewer interruptions
OpenAI has updated ChatGPT's Advanced Voice Mode feature, promising a more natural conversation experience. The aim is to make the AI-powered assistant more pleasant to talk to and less prone to interrupting you mid-sentence. In a video posted on the OpenAI YouTube channel on Monday, researcher Manuka Stratta showed off the improvements. One of the most common annoyances with voice assistants is that they tend to interrupt you when you pause to think. That's now been fixed here.
Fox News AI Newsletter: AI study buddies are boosting grades to new heights
Alpha School co-founder Mackenzie Price and a junior at the school Elle Kristine join'Fox & Friends' to discuss the benefits of incorporating artificial intelligence into the classroom. Will A.I. make schools'obsolete,' or does it present a new'opportunity' for the education system? STUDY BUDDY: A Texas private school is seeing student test scores soar to new heights following the implementation of an artificial intelligence "tutor." 'URGENT CALL': A new report from the Anti-Defamation League shows anti-Jewish and anti-Israel biases among AI large language models. ROBOTS SWARM: The automotive industry is undergoing a seismic shift driven by the integration of AI-powered humanoid robots into production lines. UBTech Robotics, in collaboration with Zeekr, has pioneered a groundbreaking initiative where swarm robots work together to build cars faster and more efficiently than ever before.
Microsoft introduces deep research and analysis tools for Copilot
Microsoft has launched two new "reasoning agents" for Copilot that were designed to analyze vast amounts of work data, including emails, meetings, chats and documents. The first tool called "Researcher" is based on OpenAI's deep research model combined with Copilot's advanced orchestration and deep search capabilities. Researcher was made for "complex, multi-step research" at work. It can take a user's internal work data along with additional information from the web, such as competitive data, emerging trends and the latest market analysis, to create market strategies and comprehensive quarterly reports, among other potential uses. Plus, it can pull data from Salesforce, ServiceNow and other external sources. Meanwhile, the new "Analyst" tool was built to function like a skilled data scientist.
The Download: China's empty data centers, and OpenAI's new practical image generator
Just months ago, China's boom in data center construction was at its height, fueled by both government and private investors. Renting out GPUs to companies that need them for training AI models was once seen as a sure bet. But with the rise of DeepSeek and a sudden change in the economics around AI, the industry is faltering. Prices for GPUs are falling and many newly built facilities are now sitting empty. Read the full story to find out why.
Baby names associated with intelligence are dying out, study reveals - so, is yours at risk of extinction?
It's one of the most difficult decisions a new parent can make โ what shall we call our baby? Now, a huge analysis has revealed that names associated with intelligence are dying out, while those linked to beauty, elegance or strength are on the up. The study, carried out by The Economist, scrutinised the names of nearly 400 million infants born in Britain and the US over the last 143 years. Researchers used a large language model โ the type of AI that powers the likes of ChatGPT โ for their analysis. They fed it with an enormous amount of text taken from the internet and asked it to identify the five most common terms linked with each name.
The AI Hype Index: DeepSeek mania, Israel's spying tool, and cheating at chess
That's why we've created the AI Hype Index--a simple, at-a-glance summary of everything you need to know about the state of the industry. While AI models are certainly capable of creating interesting and sometimes entertaining material, their output isn't necessarily useful. Google DeepMind is hoping that its new robotics model could make machines more receptive to verbal commands, paving the way for us to simply speak orders to them aloud. Elsewhere, the Chinese startup Monica has created Manus, which it claims is the very first general AI agent to complete truly useful tasks. And burnt-out coders are allowing AI to take the wheel entirely in a new practice dubbed "vibe coding."