Large Language Model
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
Xu, Yifan, Zhang, Chao, Jiang, Hanqi, Wang, Xiaoyan, Ma, Ruifei, Li, Yiwei, Wu, Zihao, Li, Zeju, Liu, Xiangde
--Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend Large Language Models (LLMs) for tackling tasks of 3D scene understanding. Current methods rely heavily on 3D point clouds, but the 3D point cloud reconstruction of an indoor scene often results in information loss. Some textureless planes or repetitive patterns are prone to omission and manifest as voids within the reconstructed 3D point clouds. Besides, objects with complex structures tend to introduce distortion of details caused by misalignments between the captured images and the dense reconstructed point clouds. Based on these insights, we propose Argus, a novel 3D multimodal framework that leverages multi-view images for enhanced 3D scene understanding with LLMs. In general, Argus can be treated as a 3D Large Multimodal Foundation Model (3D-LMM) since it takes various modalities as input(text instructions, 2D multi-view images, and 3D point clouds) and expands the capability of LLMs to tackle 3D tasks. Argus involves fusing and integrating multi-view images and camera poses into view-as-scene features, which interact with the 3D features to create comprehensive and detailed 3D-aware scene embeddings. Our approach compensates for the information loss while reconstructing 3D point clouds and helps LLMs better understand the 3D world. Extensive experiments demonstrate that our method outperforms existing 3D-LMMs in various downstream tasks. NTRODUCTION Received 8 August 2024; revised 23 March 2025; accepted 12 June 2025. Yifan Xu is with School of Computer Science and Engineering, Bei-hang University, Beijing 100191, China, also with Beijing Digital Native Digital City Research Center, Beijing 100084, China (email: xudax-ian2001@gmail.com).
osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning
Xie, Fujing, Schwertfeger, Sören, Blum, Hermann
Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features, achieving a high level of detail and guiding robots to find objects specified by open-vocabulary language queries. While the issue of scalability for such approaches has received some attention, another fundamental problem is that high-detail object mapping quickly becomes outdated, as objects get moved around a lot. In this work, we develop a mapping and navigation system for object-goal navigation that, from the ground up, considers the possibilities that a queried object can have moved, or may not be mapped at all. Instead of striving for high-fidelity mapping detail, we consider that the main purpose of a map is to provide environment grounding and context, which we combine with the semantic priors of LLMs to reason about object locations and deploy an active, online approach to navigate to the objects. Through simulated and real-world experiments we find that our approach tends to have higher retrieval success at shorter path lengths for static objects and by far outperforms prior approaches in cases of dynamic or unmapped object queries. We provide our code and dataset at: https://anonymous.4open.science/r/osmAG-LLM.
ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
Lu, Yuhang, Tu, Jiadong, Ma, Yuexin, Zhu, Xinge
End-to-end autonomous driving has emerged as a promising approach to unify perception, prediction, and planning within a single framework, reducing information loss and improving adaptability. However, existing methods often rely on fixed and sparse trajectory supervision, limiting their ability to capture the hierarchical reasoning process that human drivers naturally employ. To bridge this gap, we propose ReAL-AD, a Reasoning-Augmented Learning framework that structures decision-making in autonomous driving based on the three-tier human cognitive model: Driving Strategy, Driving Decision, and Driving Operation, where Vision-Language Models (VLMs) are incorporated to enhance situational awareness and structured reasoning across these levels. Specifically, we introduce: (1) the Strategic Reasoning Injector, which formulates high-level driving strategies by interpreting complex traffic contexts from VLM-generated insights; (2) the Tactical Reasoning Integrator, which refines strategic intent into interpretable tactical choices such as lane changes, overtaking, and speed adjustments; and (3) the Hierarchical Trajectory Decoder, which progressively translates tactical decisions into precise control actions for smooth and human-like trajectory execution. Extensive evaluations show that integrating our framework improves planning accuracy and safety by over 30%, making end-to-end autonomous driving more interpretable and aligned with human-like hierarchical reasoning. The project page can be found at: \href{https://4dvlab.github.io/project_page/realad}{\texttt{4dvlab.github.io/project\_page/realad}}
OpenAI launches personal assistant capable of controlling files and web browsers
Users of ChatGPT will be able to ask an AI agent to find restaurant reservations, go shopping for them and even draw up lists of candidates for job vacancies, as the chatbot gains the powers of a personal assistant from Thursday. ChatGPT agent, launched by Open AI everywhere apart from the EU, not only "thinks" but also acts, the US company said. The agent combines the powers of AI research tools with the ability to take control of web browsers, computer files and software such as spreadsheets and slide decks. It follows the launch of similar "agents" by Google and Anthropic as interest grows in AI models that can handle computer-based tasks by judging which software is best to use and toggling between systems to autonomously complete assignments like drafting travel itineraries or carrying out work research. "The hope is that agents are able to bring some real utility to users – to actually do things for them rather than just outputting polished text and sounding impressive," said Niamh Burns, senior media analyst at Enders Analysis.
How to run an LLM on your laptop
Getting into local models takes a bit more effort than, say, navigating to ChatGPT's online interface. But the very accessibility of a tool like ChatGPT comes with a cost. "It's the classic adage: If something's free, you're the product," says Elizabeth Seger, the director of digital policy at Demos, a London-based think tank. OpenAI, which offers both paid and free tiers, trains its models on users' chats by default. It's not too difficult to opt out of this training, and it also used to be possible to remove your chat data from OpenAI's systems entirely, until a recent legal decision in the New York Times' ongoing lawsuit against OpenAI required the company to maintain all user conversations with ChatGPT.
OpenAI's New ChatGPT Agent Tries to Do It All
Isa Fulford, the research lead for OpenAI's new ChatGPT agent, needed to order a bunch of cupcakes, so she asked the AI tool to do it for her. "I was very specific about what I wanted, and it was a lot of cupcakes," she says. "That one took almost an hour--but it was easier than me doing it myself, because I didn't want to do it." OpenAI has launched a new agent for ChatGPT that uses a virtual browser to complete tasks and can generate downloadable files, specifically PowerPoint presentations and Excel spreadsheets. While not a full replacement for the Microsoft suite of workplace tools, the features included in this agent from OpenAI could obviate some users' reliance on Microsoft's enterprise software.
Chess Grandmaster Magnus Carlsen Beats ChatGPT Without Losing a Single Piece
The world's top chess player defeated ChatGPT in an online match in only 53 moves. Magnus Carlsen won the game without losing a single piece, while ChatGPT lost all its pawns, screenshots the Norwegian grandmaster shared on X on July 10 showed. "I sometimes get bored while travelling," Carlsen captioned the post. "That was methodical, clean, and sharp. Well played!" ChatGPT said to him, according to the screenshots Carlsen posted.
AI firms 'unprepared' for dangers of building human-level systems, report warns
Artificial intelligence companies are "fundamentally unprepared" for the consequences of creating systems with human-level intellectual performance, according to a leading AI safety group. The Future of Life Institute (FLI) said none of the firms on its AI safety index scored higher than a D for "existential safety planning". One of the five reviewers of the FLI's report said that, despite aiming to develop artificial general intelligence (AGI), none of the companies scrutinised had "anything like a coherent, actionable plan" to ensure the systems remained safe and controllable. AGI refers to a theoretical stage of AI development at which a system is capable of matching a human in carrying out any intellectual task. OpenAI, the developer of ChatGPT, has said its mission is to ensure AGI "benefits all of humanity".
Top AI Companies Have 'Unacceptable' Risk Management, Studies Say
"We want to make it really easy for people to see who is not just talking the talk, but who is also walking the walk," says Max Tegmark, president of the FLI. Read More: Some Top AI Labs Have'Very Weak' Risk Management, Study Finds SaferAI assessed top AI companies' risk management protocols (also known as responsible scaling policies) to score each company on its approach to identifying and mitigating AI risks. No AI company scored better than "weak" in SaferAI's assessment of their risk management maturity. The highest scorer was Anthropic (35%), followed by OpenAI (33%), Meta (22%), and Google DeepMind (20%). Two companies, Anthropic and Google DeepMind, received lower scores than the first time the study was carried out, in October 2024.
LLMs are Bayesian, in Expectation, not in Realization
Chlon, Leon, Rashidi, Sarah, Khamis, Zein, Awada, MarcAntonio M.
Large language models demonstrate remarkable in-context learning capabilities, adapting to new tasks without parameter updates. While this phenomenon has been successfully modeled as implicit Bayesian inference, recent empirical findings reveal a fundamental contradiction: transformers systematically violate the martingale property, a cornerstone requirement of Bayesian updating on exchangeable data. This violation challenges the theoretical foundations underlying uncertainty quantification in critical applications. Our theoretical analysis establishes four key results: (1) positional encodings induce martingale violations of order $Θ(\log n / n)$; (2) transformers achieve information-theoretic optimality with excess risk $O(n^{-1/2})$ in expectation over orderings; (3) the implicit posterior representation converges to the true Bayesian posterior in the space of sufficient statistics; and (4) we derive the optimal chain-of-thought length as $k^* = Θ(\sqrt{n}\log(1/\varepsilon))$ with explicit constants, providing a principled approach to reduce inference costs while maintaining performance. Empirical validation on GPT-3 confirms predictions (1)-(3), with transformers reaching 99\% of theoretical entropy limits within 20 examples. Our framework provides practical methods for extracting calibrated uncertainty estimates from position-aware architectures and optimizing computational efficiency in deployment.