Goto

Collaborating Authors

 Law


What in Common Models Hallucinate When Reasoning Across Scenes

Neural Information Processing Systems

Multimodal language models possess a remarkable ability to handle an openvocabulary worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that are saturating and their reasoning in the real world. To address this gap, we build a novel benchmark of in-the-wild scenes that we call Common-OBench. With more than 10.5k examples using exclusively new images not found in web training data to avoid contamination, Common-OBenchgoes beyond just perception, inspired by cognitive tests for humans, to probe reasoning across scenes by asking "what's in common?". We evaluate leading multimodal language models, including models specifically trained to reason. We find that perceiving objects in single images is easy for most models, yet reasoning across scenes is very challenging even for the best models, including reasoning models. Despite saturating many leaderboards focusing on perception, the best performing model only achieves 35% on Common-OBench--and on Common-OComplex, consisting of more complex scenes, the best model achieves only 1%. Curiously, we find models are more prone to hallucinate when similar objects are present in the scene, suggesting models may be relying on object co-occurrence seen during training. Among the models we evaluated, we found scale can provide modest improvements while models explicitly trained with multi-image inputs show bigger improvements, suggesting scaled multi-image training may offer promise.


AgentAuditor: Human-Level Safety and Security Evaluation for LLMAgents

Neural Information Processing Systems

Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators often miss dangers in agents' step-by-step actions, overlook subtle meanings, fail to see how small issues compound, and get confused by unclear safety or security rules. To overcome this evaluation crisis, we introduce AgentAuditor, a universal, training-free, memory-augmented reasoning framework that empowers LLM evaluators to emulate human expert evaluators. AgentAuditor constructs an experiential memory by having an LLM adaptively extract structured semantic features (e.g., scenario, risk, behavior) and generate associated chain-of-thought reasoning traces for past interactions. A multi-stage, contextaware retrieval-augmented generation process then dynamically retrieves the most relevant reasoning experiences to guide the LLM evaluator's assessment of new cases. Moreover, we develop ASSEBench, the first benchmark designed to check how well LLM-based evaluators can spot both safety risks and security threats. ASSEBench comprises 2293 meticulously annotated interaction records, covering 15 risk types across 29 application scenarios. A key feature of ASSEBench is its nuanced approach to ambiguous risk situations, employing "Strict" and "Lenient" judgment standards. Experiments demonstrate that AgentAuditor not only consistently improves the evaluation performance of LLMs across all benchmarks but also sets a new state-of-the-art in LLM-as-a-judge for agent safety and security, achieving human-level accuracy.


Hallo Brötchen! Berlin Zoo welcomes baby pygmy hippo

Popular Science

More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. Pygmy hippos like Brötchen are endangered in the wild. Breakthroughs, discoveries, and DIY tips sent six days a week. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy . The pygmy hippopotamus () was born on May 9th at Germany's Zoo Berlin, weighing in at 13 pounds.


Gradient Multi-Normalization for Efficient LLMTraining

Neural Information Processing Systems

Training large language models (LLMs) commonly relies on adaptive optimizers such as Adam (Kingma & Ba, 2015), which accelerate convergence through moment estimates but incur substantial memory overhead. Recent stateless approaches such as SWAN (Ma et al., 2024) have shown that appropriate preprocessing of instantaneous gradient matrices can match the performance of adaptive methods without storing optimizer states. Building on this insight, we introduce gradient multi-normalization, a principled framework for designing stateless optimizers that normalize gradients with respect to multiple norms simultaneously. Whereas standard first-order methods can be viewed as gradient normalization under a single norm (Bernstein & Newhouse, 2024), our formulation generalizes this perspective to a multi-norm setting. We derive an efficient alternating scheme that enforces these normalization constraints and show that our procedure can produce, up to an arbitrary precision, a fixed-point of the problem. This unifies and extends prior stateless optimizers, showing that SWAN arises as a specific instance with particular norm choices. Leveraging this principle, we develop SinkGD, a lightweight matrix optimizer that retains the memory footprint of SGD (w/o momentum) while substantially reducing computation relative to whitening-based methods. On the memory-efficient LLaMA training benchmark (Zhao et al., 2024a), SinkGD achieves state-of-the-art performance, reaching the same evaluation perplexity as Adam using only 40% of the training tokens.


Bohdi: Heterogeneous LLMFusion with Automatic Data Exploration

Neural Information Processing Systems

While promising, existing methods suffer from two major limitations: 1) reliance on real data from limited domain for knowledge fusion, preventing the target LLM from fully acquiring knowledge across diverse domains, and 2) fixed data allocation proportions across domains, failing to dynamically adjust according to the target LLM's varying capabilities across domains, leading to a capability imbalance. To overcome these limitations, we propose Bohdi, a synthetic-data-only heterogeneous LLM fusion framework. Through the organization of knowledge domains into a hierarchical tree structure, Bohdi enables automatic domain exploration and multi-domain data generation through multimodel collaboration, thereby comprehensively extracting knowledge from source LLMs. By formalizing domain expansion and data sampling proportion allocation on the knowledge tree as a Hierarchical Multi-Armed Bandit problem, Bohdi leverages the designed DynaBranches mechanism to adaptively adjust sampling proportions based on the target LLM's performance feedback across domains. Integrated with our proposed Introspection-Rebirth (IR) mechanism, DynaBranches dynamically tracks capability shifts during target LLM's updates via Sliding Window Binomial Likelihood Ratio Testing (SWBLRT), further enhancing its online adaptation capability. Comparative experimental results on a comprehensive suite of benchmarks demonstrate that Bohdi significantly outperforms existing baselines on multiple target LLMs, exhibits higher data efficiency, and virtually eliminates the imbalance in the target LLM's capabilities. Our code is available at Bohdi.



Why Fines Alone Won't Make Social Media Safer For Kids

TIME - Tech

If courts want to reduce harm, they must focus on product design choices, measurable safety outcomes, and governance, write Peter Chapman, Ravi Iyer, and Meetali Jain.


Flutterly adorable! Schnauzer with incredibly long lashes is looking for a home - as delighted fans claim she has 'eyelashes of dreams'

Daily Mail - Science & tech

TV star mom, 46, who appeared on'quitting everything to change your life' show died in fire at luxury Caribbean beach resort that sent 1,700 tourists running for their lives Furious Trump hits back at Italian Prime Minister Meloni and gives her unusual'nickname' as their photo feud ramps up The'marry me' sex move that'll make even the most commitment-phobic of men beg to see you again... and it worked for THREE of my friends Take the 10-second finger exercise that may reveal your risk of dementia... and even protect against it'It feels like emotional blackmail': As Harry and Meghan announce return to Britain with Archie and Lilibet, insiders reveal fears about decision to bring children and'manipulation' of Royals The four mistakes that led to bungee tragedy on Skeleton Bridge: FRED KELLY saw the scene for himself, now he retraces the prelude to disaster. So was it really an accident? Dua Lipa stuns in a bespoke Chanel bridal gown and parties into the early hours as she shares the first pictures from her £1.5million Little-known penis condition that SHORTENS manhood: Shockingly, 1 in 10 men have it... but most miss the signs until it's too late to reverse with easy cure: DR PETAR BAJIC World Cup commentator denies making racist comment about Ciara live on air during USA's win over Australia Harrowing chain of events behind The Ring star's death at just 35 laid bare by doctors in agonizing detail... and how it could have been prevented Jelly Roll reveals divorce'plot twist' as he posts footage of post-split phone call with Bunnie XO Lindsey Vonn shows off remarkable progress in gym workout just four months after horrific injury: 'Makes me so happy' Taylor Swift's bombshell reconciliation phone call with Blake Lively: Insiders reveal every detail of wedding invite'olive branch' literally no one saw coming... and the actress has a dress picked out! Swedish actress, 81, was in TWO James Bond movies and also worked with Charlton Heston, who is she?


AI music is everywhere now -- and almost nobody can tell

PCWorld

AI-generated music is becoming increasingly common and increasingly difficult to recognise. Here are the tell-tale signs that reveal whether a song is AI-generated – and what this development means for the music industry.


Traditional Home Insurance Is Collapsing. Here's What Could Fill the Gap

WIRED

Traditional Home Insurance Is Collapsing. A new, AI-assisted model of insurance is quietly exploding in disaster-prone areas--and may be coming for FEMA too. Is it the answer to climate change, or a trap? In 2019, when the worst flooding in recorded history spread across the entire Mississippi River basin, Colin Wellenkamp's phone rang for weeks. Wellenkamp runs a nonprofit called the Mississippi River Cities & Towns Initiative, which coordinates between mayors' offices in more than 100 river communities from northern Minnesota to southern Louisiana. As he describes it, his headquarters served as "one big virtual situation room" for relief agencies and municipalities up and down the central US.