Large Language Model
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
Chi, Haoang, Li, He, Yang, Wenjing, Liu, Feng, Lan, Long, Ren, Xiaoguang, Liu, Tongliang, Han, Bo
Causal reasoning capability is critical in advancing large language models (LLMs) toward strong artificial intelligence. While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality and providing responses that obey the laws of causality, it remains unclear whether they perform genuine causal reasoning akin to humans. However, current evidence indicates the contrary. Specifically, LLMs are only capable of performing shallow (level-1) causal reasoning, primarily attributed to the causal knowledge embedded in their parameters, but they lack the capacity for genuine human-like (level-2) causal reasoning. To support this hypothesis, methodologically, we delve into the autoregression mechanism of transformer-based LLMs, revealing that it is not inherently causal. Empirically, we introduce a new causal Q&A benchmark called CausalProbe-2024, whose corpora are fresh and nearly unseen for the studied LLMs. The LLMs exhibit a significant performance drop on CausalProbe-2024 compared to earlier benchmarks, indicating the fact that they primarily engage in level-1 causal reasoning. To bridge the gap towards level-2 causal reasoning, we draw inspiration from the fact that human reasoning is usually facilitated by general knowledge and intended goals. We propose G^2-Reasoner, a method that incorporates general knowledge and goal-oriented prompts into LLMs' causal reasoning processes. Experiments demonstrate that G^2-Reasoner significantly enhances LLMs' causal reasoning capability, particularly in fresh and counterfactual contexts. This work sheds light on a new path for LLMs to advance towards genuine causal reasoning, going beyond level-1 and making strides towards level-2.
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
Cai, Lingling, Zhao, Kang, Yuan, Hangjie, Wang, Xiang, Zhang, Yingya, Huang, Kejie
The advent of Video Diffusion Transformers (Video DiTs) marks a milestone in video generation. However, directly applying existing video editing methods to Video DiTs often incurs substantial computational overhead, due to resource-intensive attention modification or finetuning. To alleviate this problem, we present DFVEdit, an efficient zero-shot video editing method tailored for Video DiTs. DFVEdit eliminates the need for both attention modification and fine-tuning by directly operating on clean latents via flow transformation. To be more specific, we observe that editing and sampling can be unified under the continuous flow perspective. Building upon this foundation, we propose the Conditional Delta Flow Vector (CDFV) -- a theoretically unbiased estimation of DFV -- and integrate Implicit Cross Attention (ICA) guidance as well as Embedding Reinforcement (ER) to further enhance editing quality. DFVEdit excels in practical efficiency, offering at least 20x inference speed-up and 85% memory reduction on Video DiTs compared to attention-engineering-based editing methods. Extensive quantitative and qualitative experiments demonstrate that DFVEdit can be seamlessly applied to popular Video DiTs (e.g., CogVideoX and Wan2.1), attaining state-of-the-art performance on structural fidelity, spatial-temporal consistency, and editing quality.
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
Li, Cheng, Liu, Jiexiong, Chen, Yixuan, Zhou, Qihang, Meta, KunLun
This paper introduces KunLunBaizeRAG, a reinforcement learning-driven reasoning framework designed to enhance the reasoning capabilities of large language models (LLMs) in complex multi-hop question-answering tasks. The framework addresses key limitations of traditional RAG, such as retrieval drift, information redundancy, and strategy rigidity. Key innovations include the RAG-driven Reasoning Alignment (RDRA) mechanism, the Search-Think Iterative Enhancement (STIE) mechanism, the Network-Local Intelligent Routing (NLR) mechanism, and a progressive hybrid training strategy. Experimental results demonstrate significant improvements in exact match (EM) and LLM-judged score (LJ) across four benchmarks, highlighting the framework's robustness and effectiveness in complex reasoning scenarios.
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Dang, Jisheng, Song, Huilin, Xiao, Junbin, Wang, Bimei, Peng, Han, Li, Haoxuan, Yang, Xun, Wang, Meng, Chua, Tat-Seng
Grounded Video Question Answering (Grounded VideoQA) requires aligning textual answers with explicit visual evidence. However, modern multimodal models often rely on linguistic priors and spurious correlations, resulting in poorly grounded predictions. In this work, we propose MUPA, a cooperative MUlti-Path Agentic approach that unifies video grounding, question answering, answer reflection and aggregation to tackle Grounded VideoQA. MUPA features three distinct reasoning paths on the interplay of grounding and QA agents in different chronological orders, along with a dedicated reflection agent to judge and aggregate the multi-path results to accomplish consistent QA and grounding. This design markedly improves grounding fidelity without sacrificing answer accuracy. Despite using only 2B parameters, our method outperforms all 7B-scale competitors. When scaled to 7B parameters, MUPA establishes new state-of-the-art results, with Acc@GQA of 30.3% and 47.4% on NExT-GQA and DeVE-QA respectively, demonstrating MUPA' effectiveness towards trustworthy video-language understanding. Our code is available in https://github.com/longmalongma/MUPA.
OpenAI Leadership Responds to Meta Offers: 'Someone Has Broken Into Our Home'
Mark Chen, the chief research officer at OpenAI, sent a forceful memo to staff on Saturday, promising to go head-to-head with the social giant in the war for top research talent. This memo, which was sent to OpenAI employees in Slack and obtained by WIRED, came days after Meta CEO Mark Zuckerberg successfully recruited four senior researchers from the company to join Meta's superintelligence lab. "I feel a visceral feeling right now, as if someone has broken into our home and stolen something," Chen wrote. "Please trust that we haven't been sitting idly by." Chen promised that he was working with Sam Altman, the CEO of OpenAI, and other leaders at the company "around the clock to talk to those with offers," adding, "we've been more proactive than ever before, we're recalibrating comp, and we're scoping out creative ways to recognize and reward top talent." Still, even as OpenAI leadership appears desperate to retain its staff, Chen said that he has "high personal standards of fairness," and wants to retain top talent with that in mind.
I Let AI Agents Plan My Vacation--and It Wasn't Terrible
The worst part of travel is the planning: the faff of finding and booking transport, accommodation, restaurant reservations--the list can feel endless. To help, the latest wave of AI agents, such as OpenAI's Operator and Anthropic's Computer Use claim they can take these dreary, cumbersome tasks from befuddled travelers and do it all for you. But exactly how good are they are digging out the good stuff? What better way to find out than deciding on a last-minute weekend away. I tasked Operator, which is available to ChatGPT Pro subscribers, with booking me something budget-friendly, with good food and art, and told it that I'd prefer to travel by train.
What Do Americans Actually Want to Read? One Author Crunched the Numbers--and Wrote It.
This enterprise proved so amusing that the pair, in collaboration with composer Dave Soldier, repeated the experiment with popular music, releasing the "most wanted" and "least wanted" songs together on a CD with a cover photo of all three men wearing white lab coats and pointing at a calculator. Sadly, the pair stopped short of what I view as the greatest challenge: producing novels that reflect what Americans like and dislike in fiction. Now, at last, with People's Choice Literature, by the writer/artist/composer Tom Comitta, a new "scientist" has taken up the task. People's Choice Literature offers its readers two novels for the price of one. The first is a thriller whose heroine tries to prevent her boss, a new age–y tech mogul, from launching a quantum computing network that will bring about a total surveillance state.
AI is learning to lie, scheme and threaten its creators
The world's most advanced AI models are exhibiting troubling new behaviors -- lying, scheming and even threatening their creators to achieve their goals. In one particularly jarring example, under threat of being unplugged, Anthropic's latest creation Claude 4 lashed back by blackmailing an engineer and threatened to reveal an extramarital affair. Meanwhile, ChatGPT-creator OpenAI's o1 tried to download itself onto external servers and denied it when caught red-handed.
OpenAI Loses 4 Key Researchers to Meta
Four OpenAI researchers are leaving the company to go to Meta, two sources confirm to WIRED. Their OpenAI Slack profiles have been deactivated. The Information first reported on the departures. It's the latest in a series of aggressive moves by Mark Zuckerberg, who is racing to catch up to OpenAI, Anthropic and Google in building artificial general intelligence. Earlier this month, OpenAI CEO Sam Altman said that Meta has been making "giant offers" to OpenAI staffers with " 100 million signing bonuses."
Fox News AI Newsletter: ChatGPT rewiring your brain
'The CyberGuy' Kurt Knutsson joins'Fox & Friends Weekend' to discuss the potential effects of artificial intelligence software like ChatGPT on the brain. Massachusetts Institute of Technology researchers are studying ChatGPT's effects on the brain. BRAIN DANGER: Using ChatGPT on a long-term basis could have negative effects on brain function. That's according to a study led by the Massachusetts Institute of Technology (MIT), which found that using a large language model (LLM) to write multiple essays over a four-month period could hamper cognitive abilities. 'ERRATIC': Videos taken this week by passengers showed Tesla robotaxis – which are Model Y vehicles with advanced software – braking suddenly, speeding, conducting improper drop-offs, entering the wrong lane and driving over a curb, according to Reuters.