Goto

Collaborating Authors

 Large Language Model


GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning

arXiv.org Artificial Intelligence

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit strong reasoning ability in this task through accurate identification and adaptive application of geometric principles within visual contexts. However, existing benchmarks fail to jointly assess both dimensions of the human-like geometric reasoning mechanism in MLLMs, remaining a critical gap in assessing their ability to tackle GPS. To this end, we introduce GeoSense, the first comprehensive bilingual benchmark designed to systematically evaluate the geometric reasoning abilities of MLLMs through the lens of geometric principles. GeoSense features a five-level hierarchical framework of geometric principles spanning plane and solid geometry, an intricately annotated dataset of 1,789 problems, and an innovative evaluation strategy. Through extensive experiments on GeoSense with various open-source and closed-source MLLMs, we observe that Gemini-2.0-pro-flash performs best, achieving an overall score of $65.3$. Our in-depth analysis reveals that the identification and application of geometric principles remain a bottleneck for leading MLLMs, jointly hindering their reasoning abilities. These findings underscore GeoSense's potential to guide future advancements in MLLMs' geometric reasoning capabilities, paving the way for more robust and human-like reasoning in artificial intelligence.


Looking beyond the next token

arXiv.org Artificial Intelligence

The structure of causal language model training assumes that each token can be accurately predicted from the previous context. This contrasts with humans' natural writing and reasoning process, where goals are typically known before the exact argument or phrasings. While this mismatch has been well studied in the literature, the working assumption has been that architectural changes are needed to address this mismatch. We argue that rearranging and processing the training data sequences can allow models to more accurately imitate the true data-generating process, and does not require any other changes to the architecture or training infrastructure. We demonstrate that this technique, Trelawney, and the inference algorithms derived from it allow us to improve performance on several key benchmarks that span planning, algorithmic reasoning, and story generation tasks. Finally, our method naturally enables the generation of long-term goals at no additional cost. We investigate how using the model's goal-generation capability can further improve planning and reasoning. Additionally, we believe Trelawney could potentially open doors to new capabilities beyond the current language modeling paradigm.


Transferable text data distillation by trajectory matching

arXiv.org Artificial Intelligence

In the realm of large language model (LLM), as the size of large models increases, it also brings higher training costs. There is a urgent need to minimize the data size in LLM training. Compared with data selection method, the data distillation method aims to synthesize a small number of data samples to achieve the training effect of the full data set and has better flexibility. Despite its successes in computer vision, the discreteness of text data has hitherto stymied its exploration in natural language processing (NLP). In this work, we proposed a method that involves learning pseudo prompt data based on trajectory matching and finding its nearest neighbor ID to achieve cross-architecture transfer. During the distillation process, we introduce a regularization loss to improve the robustness of our distilled data. To our best knowledge, this is the first data distillation work suitable for text generation tasks such as instruction tuning. Evaluations on two benchmarks, including ARC-Easy and MMLU instruction tuning datasets, established the superiority of our distillation approach over the SOTA data selection method LESS. Furthermore, our method demonstrates a good transferability over LLM structures (i.e., OPT to Llama).


Cognitive Memory in Large Language Models

arXiv.org Artificial Intelligence

This paper examines memory mechanisms in Large Language Models (LLMs), emphasizing their importance for context-rich responses, reduced hallucinations, and improved efficiency. It categorizes memory into sensory, short-term, and long-term, with sensory memory corresponding to input prompts, short-term memory processing immediate context, and long-term memory implemented via external databases or structures. The text-based memory section covers acquisition (selection and summarization), management (updating, accessing, storing, and resolving conflicts), and utilization (full-text search, SQL queries, semantic search). The KV cache-based memory section discusses selection methods (regularity-based summarization, score-based approaches, special token embeddings) and compression techniques (low-rank compression, KV merging, multimodal compression), along with management strategies like offloading and shared attention mechanisms. Parameter-based memory methods (LoRA, TTT, MoE) transform memories into model parameters to enhance efficiency, while hidden-state-based memory approaches (chunk mechanisms, recurrent transformers, Mamba model) improve long-text processing by combining RNN hidden states with current methods. Overall, the paper offers a comprehensive analysis of LLM memory mechanisms, highlighting their significance and future research directions.


How Well Can Vison-Language Models Understand Humans' Intention? An Open-ended Theory of Mind Question Evaluation Benchmark

arXiv.org Artificial Intelligence

Understanding human intentions through visual cues is a fundamental aspect of social intelligence, allowing effective communication, collaboration, and interaction [2]. This capability, often referred to as the Theory of Mind (ToM), involves the ability to infer the beliefs, desires, and intentions of others based on observable behaviors and environmental contexts [9, 7, 12]. Recent advances in VLMs have demonstrated impressive abilities in multimodal reasoning, combining visual and textual information to perform complex tasks [5, 10, 13]. However, their capability to perform ToM-like reasoning, specifically in interpreting intentions from visual cues, remains underexplored. For example, Etesam et al. [4] only investigate the emotional component of ToM, instead of exploring more broad categories such as intentions, religions, etc. Jin et al. [6] frame the ToM task as a binary choice question, without requiring VLMs to engage in open-ended reasoning. Consequently, this approach may not fully capture the VLMs' capability to perform ToM tasks. To further highlight, ToM tasks present unique challenges for VLMs, requiring both visual feature extraction and contextual reasoning to infer hidden mental states. Thus, our study, which evaluates VLM performance on ToM tasks through an open-ended question framework, is pivotal to assessing VLMs' capacity for advanced multimodal understanding and social intelligence.


Windows Copilot promises to chill out when you tap the key

PCWorld

Remember when Microsoft promised that the Copilot key would be the next big thing? Since then Microsoft has begun backing away from its Copilot app, and this week the company is promising that Copilot won't even launch when you tap the key -- just a subset of the app will. Instead, Microsoft is promising that the Copilot key -- or, in future, the WIN C shortcut -- will launch Copilot Chat, a small chat box that won't take up as much screen space as before. But even this new experience isn't free from Microsoft's fragmentation problems, which puts separate features on separate tracks. Microsoft has two Copilot experiences: the "consumer" version of Copilot, and the more professional Copilot experience as Microsoft 365 Copilot.


Woman says ChatGPT saved her life by helping detect cancer, which doctors missed

FOX News

Fox News senior medical analyst Dr. Marc Siegel joined'Fox & Friends' to discuss the impact of artificial intelligence on medicine and his take on President Trump's decision to withdraw from the World Health Organization. A mother of two credits ChatGPT for saving her life, claiming the artificial intelligence chatbot flagged the condition leading to her cancer when doctors missed it. Lauren Bannon, who divides her time between North Carolina and the U.S. Virgin Islands, first noticed in February 2024 that she was having trouble bending her fingers in the morning and evening, as reported by Kennedy News and Media. After four months, the 40-year-old was told by doctors that she had rheumatoid arthritis, despite testing negative for the condition. WHAT IS ARTIFICIAL INTELLIGENCE (AI)?


OpenAI Wants to Go For-Profit. Experts Say Regulators Should Step In

TIME - Tech

In the latest development in an ongoing struggle over OpenAI's future direction--and potentially the future of artificial intelligence itself--dozens of prominent figures are urging the Attorneys General of California and Delaware to block OpenAI's controversial plan to convert from its unique nonprofit-controlled structure to a for-profit company. In a letter made public April 23, signatories including "AI Godfather" Geoffrey Hinton, Harvard legal professor Lawrence Lessig, and several former OpenAI researchers argue the move represents a fundamental betrayal of OpenAI's founding mission. "The proposed restructuring would eliminate essential safeguards, effectively handing control of, and profits from, what could be the most powerful technology ever created to a for-profit entity with legal duties to prioritize shareholder returns," the letter's authors write. It lands as OpenAI faces immense pressure from the other side: failing to implement the restructure by the end of the year could cost the company 20 billion and hamstring future fundraising. OpenAI was founded in 2015 as a non-profit, with its stated mission being to ensure that artificial general intelligence (AGI) "benefits all of humanity" rather than advancing "the private gain of any person."


Motorola's Latest Razr Phones Are All In on AI

WIRED

But at the company's closed-door launch event on Wednesday in New York City, much of the spotlight was on Moto AI, artificial intelligence features powered by in-house and third-party large language models, like Meta's Llama. Google's Gemini is naturally available on the Razr phones, but for the first time, the AI search engine Perplexity AI will be preinstalled on the devices. The CEO of Perplexity, Aravind Srinivas, took the stage to talk about the optimizations made to take advantage of the Razr's unique design. Motorola even says Microsoft's Copilot will also be available in the coming months. The 2025 Razr lineup starts at 700 for the base Razr, 1,000 for the Razr, and 1,300 for the Razr Ultra; the midrange Razr is almost the same device as the Razr from 2024, with a few enhancements to durability.


China's AI DeepSeek faces House probe over US data harvesting, CCP propaganda

FOX News

'The Big Weekend Show' co-hosts discuss the impact of new artificial intelligence apps on national security and jobs. FIRST ON FOX: A powerful House Committee is demanding information from DeepSeek on what U.S. data it used to train the AI model as members accuse the company of being in the pocket of the Chinese government. In announcing a new probe into DeepSeek, House Energy and Commerce committee members penned a letter expressing concern that companies like it "harvest Americans' personal and proprietary information and introduce new data security vulnerabilities into the U.S. economy." "DeepSeek admits to sending Americans' personal information to servers in China, where it is undoubtedly accessed by officials connected to the Chinese Communist Party," Chairman Brett Guthrie, R-Ky., and Gus Bilirakis, R-Fla., said in a statement. "We are concerned that this close relationship with agents having close connections to our primary adversary jeopardizes our data and our national security."