Goto

Collaborating Authors

 Media


How to stop scam texts from targeting aging parents

FOX News

This material may not be published, broadcast, rewritten, or redistributed. Quotes displayed in real-time or delayed by at least 15 minutes. Market data provided by Factset . Powered and implemented by FactSet Digital Solutions . Mutual Fund and ETF data provided by LSEG . Artemis crew says they wanted to'connect with humanity,' show what can be done when they put their mind to it Scientists revive ancient 24,000-year-old'zombie worm' from Arctic ice -- then it reproduced'Gigantic' ancient octopus used jaws to crush prey and hunted alongside the dinosaurs 100M years ago: study Scientists uncover identity of mysterious'golden orb' discovered miles underwater in 2023 Artemis astronauts enter eerie 40-minute communication blackout on Moon's far side NASA chief Jared Isaacman says Artemis II would not be possible'if it wasn't for President Trump' Researchers pinpoint source of black hole's 3,000-light-year-long jet stream using enhanced telescope network Is Spielberg's new UFO film more fact than fiction? Auburn University's bald eagle tradition celebrates its 25th anniversary American public'can handle' truth about UAPs, whistleblower says Google's AI unleashes new powerful scam-busting features for Android The CyberGuy explains steps you can take to protect yourself from scams. Scam texts are annoying for everyone.


NEP: Autoregressive Image Editing via Next Editing Token Prediction

Neural Information Processing Systems

Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively regenerate only the intended editing areas. This results in (1) unnecessary computational costs and (2) a bias toward reconstructing non-editing regions, which compromises the quality of the intended edits. To resolve these limitations, we propose to formulate image editing as Next Editing-token Prediction (NEP) based on autoregressive image generation, where only regions that need to be edited are regenerated, thus avoiding unintended modification to the non-editing areas. To enable any-region editing, we propose to pre-train an any-order autoregressive text-to-image (T2I) model. Once trained, it is capable of zero-shot image editing and can be easily adapted to NEP for image editing, which achieves a new state-of-the-art on widely used image editing benchmarks. Moreover, our model naturally supports test-time scaling (TTS) through iteratively refining its generation in a zero-shot manner.


ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

Neural Information Processing Systems

Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation.


See Stonehenge's construction like NEVER before: Incredible visual reveals the vast manpower needed to haul the 25-tonne stones into position 5,000 years ago

Daily Mail - Science & tech

Keir Starmer cries as he quits No 10 claiming a deluded list of'achievements' - now Britain awaits its seventh PM in ten years Putin'prepares mass call-up' for the Ukraine meat-grinder - as video shows Russian veteran with no legs threatening recruiter with a knife in sign of growing resistance facing desperate Kremlin'Al Roker is an absolute ****': KENNEDY's Today show insider gives brutal behind-the-scenes verdict on beloved weatherman and names other two-faced NBC hosts No one can see the real reason Jelly Roll divorced Bunnie XO. Boston's Scotland-loving residents claim England fans are'ruining the vibe' compared to the Tartan Army Secret life of John Travolta's daughter Ella Bleu: New details about'unusual' relationship with her dad revealed by insiders amid fears that aspiring actress is'stuck' Johnny Depp's ex Amber Heard gives rare glimpse of daughter Oonagh, five, after finishing 10k race in Spain Colorado siblings VANISH from home in middle of the night... and police ...


Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models

Neural Information Processing Systems

Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose "HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head's attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection.


PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement

Neural Information Processing Systems

Latent Diffusion Models (LDMs) have markedly advanced the quality of image inpainting and local editing. However, the inherent latent compression often introduces pixel-level inconsistencies, such as chromatic shifts, texture mismatches, and visible seams along editing boundaries. Existing remedies, including backgroundconditioned latent decoding and pixel-space harmonization, usually fail to fully eliminate these artifacts in practice and do not generalize well across different latent representations or tasks. We introduce PixPerfect, a pixel-level refinement framework that delivers seamless, high-fidelity local edits across diverse LDM architectures and tasks. PixPerfect leverages (i) a differentiable discriminative pixel space that amplifies and suppresses subtle color and texture discrepancies, (ii) a comprehensive artifact simulation pipeline that exposes the refiner to realistic local editing artifacts during training, and (iii) a direct pixel-space refinement scheme that ensures broad applicability across diverse latent representations and tasks. Extensive experiments on inpainting, object removal, and insertion benchmarks demonstrate that PixPerfect substantially enhances perceptual fidelity and downstream editing performance, establishing a new standard for robust and high-fidelity localized image editing.


Japan lifts upper limit on drones operated by one pilot

The Japan Times

Removing the limit on simultaneous drone use by one person is aimed at optimizing logistics and infrastructure inspections, and at making damage assessment and search operations more effective during disasters. The transport ministry has lifted the upper limit on the number of drones that can be operated by a single pilot at the same time, officials said. Previously, the number was set at five. Removing the limit is aimed at optimizing logistics and infrastructure inspections, and at making damage assessment and search operations more effective during disasters. In March 2025, the ministry set guidelines for using multiple drones at once.


Disentanglement Beyond Static vs. Dynamic: ABenchmark and Evaluation Framework for Multi-Factor Sequential Representations

Neural Information Processing Systems

Learning disentangled representations in sequential data is a key goal in deep learning, with broad applications in vision, audio, and time series. While realworld data involves multiple interacting semantic factors over time, prior work has mostly focused on simpler two-factor static and dynamic settings, primarily because such settings make data collection easier, thereby overlooking the inherently multifactor nature of real-world data. We introduce the first standardized benchmark for evaluating multi-factor sequential disentanglement across six diverse datasets spanning video, audio, and time series. Our benchmark includes modular tools for dataset integration, model development, and evaluation metrics tailored to multi-factor analysis. We additionally propose a post-hoc Latent Exploration Stage to automatically align latent dimensions with semantic factors, and introduce a Koopman-inspired model that achieves state-of-the-art results. Moreover, we show that Vision-Language Models can automate dataset annotation and serve as zeroshot disentanglement evaluators, removing the need for manual labels and human intervention. Together, these contributions provide a robust and scalable foundation for advancing multi-factor sequential disentanglement. Our code is available on GitHub, and the datasets and trained models are available on Hugging Face.


Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs

Neural Information Processing Systems

Large language models (LLMs) excel at complex tasks thanks to advances in their reasoning abilities. However, existing methods overlook the trade-off between reasoning effectiveness and efficiency, often encouraging unnecessarily long reasoning chains and wasting tokens. To address this, we propose Learning to Think (L2T) 3, an information-theoretic reinforcement fine-tuning framework for LLMs to make the models achieve optimal reasoning with fewer tokens. Specifically, L2T treats each query-response interaction as a hierarchical session of multiple episodes and proposes a universal dense process reward, i.e., quantifies the episode-wise information gain in parameters, requiring no extra annotations or task-specific evaluators. We propose a method to quickly estimate this reward based on PACBayes bounds and the Fisher information matrix. Theoretical analyses show that it significantly reduces computational complexity with high estimation accuracy. By immediately rewarding each episode's contribution and penalizing excessive updates, L2T optimizes the model via reinforcement learning to maximize the use of each episode and achieve effective updates. Empirical results on various reasoning benchmarks and base models demonstrate the advantage of L2T across different tasks, boosting both reasoning effectiveness and efficiency.


The Reverse Centaur's Guide to Life After AI by Cory Doctorow review – the real price of artificial intelligence

The Guardian

Cory Doctorow speaks at a digital society conference. Cory Doctorow speaks at a digital society conference. The Reverse Centaur's Guide to Life After AI by Cory Doctorow review - the real price of artificial intelligence A s former Google CEO Eric Schmidt could tell you, AI is a hard sell these days. Last month, he tried talking up the AI revolution during a commencement address at the University of Arizona and was loudly booed by students about to enter an AI-ravaged job market. Schmidt is not the only AI booster to crash out with students recently as the popular backlash grows.