Media
NEP: Autoregressive Image Editing via Next Editing Token Prediction
Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively regenerate only the intended editing areas. This results in (1) unnecessary computational costs and (2) a bias toward reconstructing non-editing regions, which compromises the quality of the intended edits. To resolve these limitations, we propose to formulate image editing as Next Editing-token Prediction (NEP) based on autoregressive image generation, where only regions that need to be edited are regenerated, thus avoiding unintended modification to the non-editing areas. To enable any-region editing, we propose to pre-train an any-order autoregressive text-to-image (T2I) model. Once trained, it is capable of zero-shot image editing and can be easily adapted to NEP for image editing, which achieves a new state-of-the-art on widely used image editing benchmarks. Moreover, our model naturally supports test-time scaling (TTS) through iteratively refining its generation in a zero-shot manner.
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation.
See Stonehenge's construction like NEVER before: Incredible visual reveals the vast manpower needed to haul the 25-tonne stones into position 5,000 years ago
Keir Starmer cries as he quits No 10 claiming a deluded list of'achievements' - now Britain awaits its seventh PM in ten years Putin'prepares mass call-up' for the Ukraine meat-grinder - as video shows Russian veteran with no legs threatening recruiter with a knife in sign of growing resistance facing desperate Kremlin'Al Roker is an absolute ****': KENNEDY's Today show insider gives brutal behind-the-scenes verdict on beloved weatherman and names other two-faced NBC hosts No one can see the real reason Jelly Roll divorced Bunnie XO. Boston's Scotland-loving residents claim England fans are'ruining the vibe' compared to the Tartan Army Secret life of John Travolta's daughter Ella Bleu: New details about'unusual' relationship with her dad revealed by insiders amid fears that aspiring actress is'stuck' Johnny Depp's ex Amber Heard gives rare glimpse of daughter Oonagh, five, after finishing 10k race in Spain Colorado siblings VANISH from home in middle of the night... and police ...
Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models
Recent guidance methods in diffusion models steer reverse sampling by perturbing the model to construct an implicit weak model and guide generation away from it. Among these approaches, attention perturbation has demonstrated strong empirical performance in unconditional scenarios where classifier-free guidance is not applicable. However, existing attention perturbation methods lack principled approaches for determining where perturbations should be applied, particularly in Diffusion Transformer (DiT) architectures where quality-relevant computations are distributed across layers. In this paper, we investigate the granularity of attention perturbations, ranging from the layer level down to individual attention heads, and discover that specific heads govern distinct visual concepts such as structure, style, and texture quality. Building on this insight, we propose "HeadHunter", a systematic framework for iteratively selecting attention heads that align with user-centric objectives, enabling fine-grained control over generation quality and visual attributes. In addition, we introduce SoftPAG, which linearly interpolates each selected head's attention map toward an identity matrix, providing a continuous knob to tune perturbation strength and suppress artifacts. Our approach not only mitigates the oversmoothing issues of existing layer-level perturbation but also enables targeted manipulation of specific visual styles through compositional head selection.
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
Latent Diffusion Models (LDMs) have markedly advanced the quality of image inpainting and local editing. However, the inherent latent compression often introduces pixel-level inconsistencies, such as chromatic shifts, texture mismatches, and visible seams along editing boundaries. Existing remedies, including backgroundconditioned latent decoding and pixel-space harmonization, usually fail to fully eliminate these artifacts in practice and do not generalize well across different latent representations or tasks. We introduce PixPerfect, a pixel-level refinement framework that delivers seamless, high-fidelity local edits across diverse LDM architectures and tasks. PixPerfect leverages (i) a differentiable discriminative pixel space that amplifies and suppresses subtle color and texture discrepancies, (ii) a comprehensive artifact simulation pipeline that exposes the refiner to realistic local editing artifacts during training, and (iii) a direct pixel-space refinement scheme that ensures broad applicability across diverse latent representations and tasks. Extensive experiments on inpainting, object removal, and insertion benchmarks demonstrate that PixPerfect substantially enhances perceptual fidelity and downstream editing performance, establishing a new standard for robust and high-fidelity localized image editing.
Japan lifts upper limit on drones operated by one pilot
Removing the limit on simultaneous drone use by one person is aimed at optimizing logistics and infrastructure inspections, and at making damage assessment and search operations more effective during disasters. The transport ministry has lifted the upper limit on the number of drones that can be operated by a single pilot at the same time, officials said. Previously, the number was set at five. Removing the limit is aimed at optimizing logistics and infrastructure inspections, and at making damage assessment and search operations more effective during disasters. In March 2025, the ministry set guidelines for using multiple drones at once.
Disentanglement Beyond Static vs. Dynamic: ABenchmark and Evaluation Framework for Multi-Factor Sequential Representations
Learning disentangled representations in sequential data is a key goal in deep learning, with broad applications in vision, audio, and time series. While realworld data involves multiple interacting semantic factors over time, prior work has mostly focused on simpler two-factor static and dynamic settings, primarily because such settings make data collection easier, thereby overlooking the inherently multifactor nature of real-world data. We introduce the first standardized benchmark for evaluating multi-factor sequential disentanglement across six diverse datasets spanning video, audio, and time series. Our benchmark includes modular tools for dataset integration, model development, and evaluation metrics tailored to multi-factor analysis. We additionally propose a post-hoc Latent Exploration Stage to automatically align latent dimensions with semantic factors, and introduce a Koopman-inspired model that achieves state-of-the-art results. Moreover, we show that Vision-Language Models can automate dataset annotation and serve as zeroshot disentanglement evaluators, removing the need for manual labels and human intervention. Together, these contributions provide a robust and scalable foundation for advancing multi-factor sequential disentanglement. Our code is available on GitHub, and the datasets and trained models are available on Hugging Face.
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
Large language models (LLMs) excel at complex tasks thanks to advances in their reasoning abilities. However, existing methods overlook the trade-off between reasoning effectiveness and efficiency, often encouraging unnecessarily long reasoning chains and wasting tokens. To address this, we propose Learning to Think (L2T) 3, an information-theoretic reinforcement fine-tuning framework for LLMs to make the models achieve optimal reasoning with fewer tokens. Specifically, L2T treats each query-response interaction as a hierarchical session of multiple episodes and proposes a universal dense process reward, i.e., quantifies the episode-wise information gain in parameters, requiring no extra annotations or task-specific evaluators. We propose a method to quickly estimate this reward based on PACBayes bounds and the Fisher information matrix. Theoretical analyses show that it significantly reduces computational complexity with high estimation accuracy. By immediately rewarding each episode's contribution and penalizing excessive updates, L2T optimizes the model via reinforcement learning to maximize the use of each episode and achieve effective updates. Empirical results on various reasoning benchmarks and base models demonstrate the advantage of L2T across different tasks, boosting both reasoning effectiveness and efficiency.
The Reverse Centaur's Guide to Life After AI by Cory Doctorow review – the real price of artificial intelligence
Cory Doctorow speaks at a digital society conference. Cory Doctorow speaks at a digital society conference. The Reverse Centaur's Guide to Life After AI by Cory Doctorow review - the real price of artificial intelligence A s former Google CEO Eric Schmidt could tell you, AI is a hard sell these days. Last month, he tried talking up the AI revolution during a commencement address at the University of Arizona and was loudly booed by students about to enter an AI-ravaged job market. Schmidt is not the only AI booster to crash out with students recently as the popular backlash grows.
Carvalho resigns as LAUSD superintendent amid federal investigation
Things to Do in L.A. Tap to enable a layout that focuses on the article. Alberto Carvalho, who resigned Sunday as LAUSD superintendent, addresses students at an elementary school in 2023. This is read by an automated voice. Please report any issues or inconsistencies here . Alberto Carvalho resigned Sunday night.