Goto

Collaborating Authors

 Media


A Definition of AGI

arXiv.org Artificial Intelligence

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.


Score Distillation of Flow Matching Models

arXiv.org Artificial Intelligence

Diffusion models achieve high-quality image generation but are limited by slow iterative sampling. Distillation methods alleviate this by enabling one- or few-step generation. Flow matching, originally introduced as a distinct framework, has since been shown to be theoretically equivalent to diffusion under Gaussian assumptions, raising the question of whether distillation techniques such as score distillation transfer directly. We provide a simple derivation -- based on Bayes' rule and conditional expectations -- that unifies Gaussian diffusion and flow matching without relying on ODE/SDE formulations. Building on this view, we extend Score identity Distillation (SiD) to pretrained text-to-image flow-matching models, including SANA, SD3-Medium, SD3.5-Medium/Large, and FLUX.1-dev, all with DiT backbones. Experiments show that, with only modest flow-matching- and DiT-specific adjustments, SiD works out of the box across these models, in both data-free and data-aided settings, without requiring teacher finetuning or architectural changes. This provides the first systematic evidence that score distillation applies broadly to text-to-image flow matching models, resolving prior concerns about stability and soundness and unifying acceleration techniques across diffusion- and flow-based generators. A project page is available at https://yigu1008.github.io/SiD-DiT.


Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences

arXiv.org Artificial Intelligence

Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants evaluating 3,840 query-response pairs, we examine how different refusal strategies affect user perceptions across varying motivations. Our findings reveal that response strategy largely shapes user experience, while actual user motivation has negligible impact. Partial compliance -- providing general information without actionable details -- emerges as the optimal strategy, reducing negative user perceptions by over 50% to flat-out refusals. Complementing this, we analyze response patterns of 9 state-of-the-art LLMs and evaluate how 6 reward models score different refusal strategies, demonstrating that models rarely deploy partial compliance naturally and reward models currently undervalue it. This work demonstrates that effective guardrails require focusing on crafting thoughtful refusals rather than detecting intent, offering a path toward AI safety mechanisms that ensure both safety and sustained user engagement.


Can VLMs Detect and Localize Fine-Grained AI-Edited Images?

arXiv.org Artificial Intelligence

Fine-grained detection and localization of localized image edits is crucial for assessing content authenticity, especially as modern diffusion models and image editors can produce highly realistic manipulations. However, this problem faces three key challenges: (1) most AIGC detectors produce only a global real-or-fake label without indicating where edits occur; (2) traditional computer vision methods for edit localization typically rely on costly pixel-level annotations; and (3) there is no large-scale, modern benchmark specifically targeting edited-image detection. To address these gaps, we develop an automated data-generation pipeline and construct FragFake, a large-scale benchmark of AI-edited images spanning multiple source datasets, diverse editing models, and several common edit types. Building on FragFake, we are the first to systematically study vision language models (VLMs) for edited-image classification and edited-region localization. Our experiments show that pretrained VLMs, including GPT4o, perform poorly on this task, whereas fine-tuned models such as Qwen2.5-VL achieve high accuracy and substantially higher object precision across all settings. We further explore GRPO-based RLVR training, which yields modest metric gains while improving the interpretability of model outputs. Ablation and transfer analyses reveal how data balancing, training size, LoRA rank, and training domain affect performance, and highlight both the potential and the limitations of cross-editor and cross-dataset generalization. We anticipate that this work will establish a solid foundation to facilitate and inspire subsequent research endeavors in the domain of multimodal content authenticity.


The Dream Within Huang Long Cave: AI-Driven Interactive Narrative for Family Storytelling and Emotional Reflection

arXiv.org Artificial Intelligence

This paper introduces the art project The Dream Within Huang Long Cave, an AI-driven interactive and immersive narrative experience. The project offers new insights into AI technology, artistic practice, and psychoanalysis. Inspired by actual geographical landscapes and familial archetypes, the work combines psychoanalytic theory and computational technology, providing an artistic response to the concept of "the nonexistence of the Big Other." The narrative is driven by a combination of a large language model (LLM) and a realistic digital character, forming a virtual agent named YELL. Through dialogue and exploration within a cave automatic virtual environment (CA VE), the audience is invited to unravel the language puzzles presented by YELL and help him overcome his life challenges. YELL is a fictional embodiment of the "Big Other," modeled after the artist's real father. Through a cross-temporal interaction with this digital father, the project seeks to deconstruct complex familial relationships. By demonstrating "the non-existence of the Big Other," we aim to underscore the authenticity of interpersonal emotions, positioning art as a bridge for emotional connection and understanding within family dynamics.


The two standout science-fiction films of 2025

New Scientist

From Mickey 17 and M3gan 2.0 to a musical about the end of the world, this was an eclectic year for science-fiction films. Some ideas are so compelling, so intuitive, one would sooner recycle them than take them apart to explore. So, in 1950, Isaac Asimov fixed up some puzzle stories into a fiendish, Agatha Christie-in-space sci-fi novel, I, Robot, while in 1968, Stanley Kubrick's 2001: A Space Odyssey set a high bar for films about (or at least containing) artificial intelligence. There, ideas-wise, the story of robots in cinema pretty much starts to repeat on an endless loop. This year, The Electric State spun a yarn about a robot rebellion, M3gan 2.0 showed you can't keep a good killerbot down and Companion took the femmebot's point of view to give us a decent adult-themed Asimov pastiche. All three toyed with the usual notions around free will and indulged in handwringing about when to treat a machine like a person.


Holiday travel privacy risks and how to stay safe

FOX News

Holiday travelers face increased scammer attacks using leaked personal data from airlines and hotels to send fake flight cancellations and payment requests.


Robot stuns crowd after shocking onstage reveal

FOX News

Xpeng's Next Gen Iron humanoid robot fooled viewers who thought it was a human actor, prompting CEO He Xiaopeng to cut into the robot's leg to prove its authenticity.



Nike, Superdry and Lacoste ads banned over misleading green claims

BBC News

Adverts for Nike, Superdry and Lacoste have been banned for making misleading claims about their green credentials. The UK's advertising watchdog challenged the brands over the use of the word sustainable in paid-for Google ads which were not backed up by evidence of their sustainability. The Advertising Standards Authority (ASA) identified three adverts from the retailers promising customers sustainable materials, sustainable style and sustainable clothing. The UK's advertising code states that the basis of claims about environmental sustainability must be clear and supported by a high level of substantiation. In each case, it asked the companies for evidence to back up the claims about the sustainability of the products.