Goto

Collaborating Authors

 Large Language Model


Singular Value Few-shot Adaptation of Vision-Language Models

arXiv.org Artificial Intelligence

Vision-language models (VLMs) like CLIP have shown impressive zero-shot and few-shot learning capabilities across diverse applications. However, adapting these models to new fine-grained domains remains difficult due to reliance on prompt engineering and the high cost of full model fine-tuning. Existing adaptation approaches rely on augmented components, such as prompt tokens and adapter modules, which could limit adaptation quality, destabilize the model, and compromise the rich knowledge learned during pretraining. In this work, we present CLIP-SVD, a novel multi-modal and parameter-efficient adaptation technique that leverages Singular Value Decomposition (SVD) to modify the internal parameter space of CLIP without injecting additional modules. Specifically, we fine-tune only the singular values of the CLIP parameter matrices to rescale the basis vectors for domain adaptation while retaining the pretrained model. This design enables enhanced adaptation performance using only 0.04% of the model's total parameters and better preservation of its generalization ability. CLIP-SVD achieves state-of-the-art classification results on 11 natural and 10 biomedical datasets, outperforming previous methods in both accuracy and generalization under few-shot settings. Additionally, we leverage a natural language-based approach to analyze the effectiveness and dynamics of the CLIP adaptation to allow interpretability of CLIP-SVD. The code is publicly available at https://github.com/HealthX-Lab/CLIP-SVD.


Self-Guided Function Calling in Large Language Models via Stepwise Experience Recall

arXiv.org Artificial Intelligence

Function calling enables large language models (LLMs) to interact with external systems by leveraging tools and APIs. When faced with multi-step tool usage, LLMs still struggle with tool selection, parameter generation, and tool-chain planning. Existing methods typically rely on manually designing task-specific demonstrations, or retrieving from a curated library. These approaches demand substantial expert effort and prompt engineering becomes increasingly complex and inefficient as tool diversity and task difficulty scale. To address these challenges, we propose a self-guided method, Stepwise Experience Recall (SEER), which performs fine-grained, stepwise retrieval from a continually updated experience pool. Instead of relying on static or manually curated library, SEER incrementally augments the experience pool with past successful trajectories, enabling continuous expansion of the pool and improved model performance over time. Evaluated on the ToolQA benchmark, SEER achieves an average improvement of 6.1% on easy and 4.7% on hard questions. We further test SEER on $ฯ„$-bench, which includes two real-world domains. Powered by Qwen2.5-7B and Qwen2.5-72B models, SEER demonstrates substantial accuracy gains of 7.44% and 23.38%, respectively.


Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision

arXiv.org Artificial Intelligence

Chain-of-Thought (CoT) reasoning has been widely adopted to enhance Large Language Models (LLMs) by decomposing complex tasks into simpler, sequential subtasks. However, extending CoT to vision-language reasoning tasks remains challenging, as it often requires interpreting transitions of visual states to support reasoning. Existing methods often struggle with this due to limited capacity of modeling visual state transitions or incoherent visual trajectories caused by fragmented architectures. To overcome these limitations, we propose Uni-CoT, a Unified Chain-of-Thought framework that enables coherent and grounded multimodal reasoning within a single unified model. The key idea is to leverage a model capable of both image understanding and generation to reason over visual content and model evolving visual states. However, empowering a unified model to achieve that is non-trivial, given the high computational cost and the burden of training. To address this, Uni-CoT introduces a novel two-level reasoning paradigm: A Macro-Level CoT for high-level task planning and A Micro-Level CoT for subtask execution. This design significantly reduces the computational overhead. Furthermore, we introduce a structured training paradigm that combines interleaved image-text supervision for macro-level CoT with multi-task objectives for micro-level CoT. Together, these innovations allow Uni-CoT to perform scalable and coherent multi-modal reasoning. Furthermore, thanks to our design, all experiments can be efficiently completed using only 8 A100 GPUs with 80GB VRAM each. Experimental results on reasoning-driven image generation benchmark (WISE) and editing benchmarks (RISE and KRIS) indicates that Uni-CoT demonstrates SOTA performance and strong generalization, establishing Uni-CoT as a promising solution for multi-modal reasoning. Project Page and Code: https://sais-fuxi.github.io/projects/uni-cot/


Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models

arXiv.org Artificial Intelligence

Large language models often retain unintended content, prompting growing interest in knowledge unlearning. Recent approaches emphasize localized unlearning, restricting parameter updates to specific regions in an effort to remove target knowledge while preserving unrelated general knowledge. However, their effectiveness remains uncertain due to the lack of robust and thorough evaluation of the trade-off between the competing goals of unlearning. In this paper, we begin by revisiting existing localized unlearning approaches. We then conduct controlled experiments to rigorously evaluate whether local parameter updates causally contribute to unlearning. Our findings reveal that the set of parameters that must be modified for effective unlearning is not strictly determined, challenging the core assumption of localized unlearning that parameter locality is inherently indicative of effective knowledge removal.


ChatGPT might ask adults for ID after teen suicides

PCWorld

When you purchase through links in our articles, we may earn a small commission. The large language model is putting in new security features after high-profile cases of teen suicides, and may require that adults verify with scanned identification in some countries. Cases of "AI psychosis" are apparently on the rise, and multiple people have committed suicide after conversing with the ChatGPT large language model. Representatives of ChatGPT maker OpenAI are testifying before the US congress in response, and the company is announcing new methods of detecting users' age. According to the CEO, that may include ID verification.


Firefox 143 gets Microsoft Copilot AI and Google Lens support

PCWorld

When you purchase through links in our articles, we may earn a small commission. The newest version of Firefox comes with new AI features and accessibility improvements, plus fixes to numerous security vulnerabilities. The latest update to Firefox brings the browser up to version 143 with various new features and improvements, including some that other browsers already offer. However, some of these features--like Google Lens--are only being introduced gradually. Mozilla plans to release Firefox 144 on October 14th, 2025.


AI-designed viruses are here and already killing bacteria

MIT Technology Review

Can AI create a life form? These "generative" genomes are a start Artificial intelligence can draw cat pictures and write emails. A research team in California says it used AI to propose new genetic codes for viruses--and managed to get several of these viruses to replicate and kill bacteria. The scientists, based at Stanford University and the nonprofit Arc Institute, both in Palo Alto, say the germs with AI-written DNA represent the "the first generative design of complete genomes." The work, described in a preprint paper, has the potential to create new treatments and accelerate research into artificially engineered cells. It is also an "impressive first step" toward AI-designed life forms, says Jef Boeke, a biologist at NYU Langone Health, who was provided an advance copy of the paper by .


Nvidia CEO Jensen Huang Is Bananas for Google Gemini's AI Image Generator

WIRED

Nvidia CEO Jensen Huang Is Bananas for Google Gemini's AI Image Generator The Nvidia CEO reveals his consuming love for Google's image generator, the artsy side of Grok, and what exactly he uses Perplexity, Gemini, and ChatGPT for right now. Nvidia CEO Jensen Huang is in London, standing in front of a room full of journalists, outing himself as a huge fan of Gemini's Nano Banana . "How could anyone not love Nano Banana? I mean Nano Banana, how good is that? Tell me it's not true!" "Tell me it's not true! I was just talking to Demis [Hassabis, CEO of DeepMind ] yesterday and I said'How about that Nano Banana! It looks like lots of people agree with him: The popularity of the Nano Banana AI image generator--which launched in August and allows users to make precise edits to AI images while preserving the quality of faces, animals, or other objects in the background--has caused a 300 million image surge for Gemini in the first few days in September already, according to a post on X by Josh Woodward, VP of Google Labs and Google Gemini. Huang, whose company was among a cohort of big US technology companies to announce investments into data centers, supercomputers, and AI research in the UK on Tuesday, is on a high. Speaking ahead of a white-tie event with UK prime minister Keir Starmer (where he plans to wear custom black leather tails), he's boisterously optimistic about the future of AI in the UK, saying the country is "too humble" about the country's potential for AI advancements. He cites the UK's pedigree in themes as wide as the industrial revolution, steam trains, DeepMind (now owned by Google), and university researchers, as well as other tangential skills. "No one fries food better than you do," he quips. Nvidia announced a $683 million equity investment in datacenter builder Nscale this week, a move that--alongside investments from OpenAI and Microsoft--has propelled the company to the epicenter of this AI push in the UK. Huang estimates that Nscale will generate more than $68 billion in revenues over six years. "I'll go on record to say I'm the best thing that's ever happened to him," he says, referring to Nscale CEO Josh Payne. "As AI services get deployed--I'm sure that all of you use it.


The Download: measuring returns on R&D, and AI's creative potential

MIT Technology Review

Plus: TikTok's potential new owners have deep pockets Given the draconian cuts to US federal funding for science, it's worth asking some hard-nosed money questions: How much should we be spending on R&D? How much value do we get out of such investments, anyway? To answer that, in several recent papers, economists have approached this issue in clever new ways. And, though they ask slightly different questions, their conclusions share a bottom line: R&D is, in fact, one of the better long-term investments that the government can make. This article is part of MIT Technology Review Explains, our series untangling the complex, messy world of technology to help you understand what's coming next. We've been here before . Artists and musicians are finding new ways to make art using AI, by injecting friction, challenge, and serendipity into the process.


Nvidia boss 'disappointed' by reported China chip ban

BBC News

Nvidia boss'disappointed' by reported China chip ban The boss of Nvidia says he is disappointed that China has reportedly ordered its top technology companies to halt purchases of the firm's artificial intelligence (AI) chips. Jensen Huang added he would be patient in response to the move from China's internet regulator. There are a lot of places we can't go to, and that's fine, he told reporters on Wednesday. Mr Huang is one of a number of tech bosses, including Microsoft's Satya Nadella, accompanying US President Donald trump on his state visit to the UK. Nvidia - the world's leading chipmaker - had previously been banned from selling its most advanced chips to China, before Trump reversed the ban in July.