Goto

Collaborating Authors

 artifact


Claude Sonnet 5.5 is here. This prompt stops it from wasting your usage

PCWorld

When you purchase through links in our articles, we may earn a small commission. Claude Sonnet 5.5 is here. This prompt, straight from Anthropic's own docs, helps curb that bad habit. The latest version of Claude Sonnet--Anthropic's workhorse AI model--has just landed, and with it comes a prompt that could help keep any AI agent from leaping into action before making a plan first. Available now on the web, the Claude desktop app, and other Claude platforms, Claude Sonnet 5.5 is Anthropic's middle-of-the-road model, It's a step below Opus 5.5 but apparently just as capable as Fable 5.1 (Anthropic's pricey frontier flagship) and a step above Haiku 5 (a fast but basic Claude model that's due for a 5.5 upgrade).


Lost WWI trench lives on in immersive VR installation

Popular Science

Archeology and gaming technology join forces to preserve battlefields of the past. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. The rolling, green hills of West Flanders, Belgium, were once the bitter, brutal staging ground for some of the bloodiest fighting of World War I . An estimated 600,000 soldiers were killed, wounded, or went missing in that region between 1914 and 1918, often fighting along intricate but claustrophobically cramped trench lines that were perpetually bombarded by artillery fire and sinking into thick, waterlogged mud. However, a visitor walking across many of those old battlegrounds today might never know they were treading across a graveyard.


Anthropic's Claude can now create editable documents for you

Engadget

Anthropic has launched Claude Docs, a new feature that lets the AI assistant/chatbot create documents for you. As part of the company's efforts to make it easier to access Claude's many capabilities -- more on this later -- you'll be able to ask it to make a document for you right within your conversation. You can ask Claude to create a document, such as a one-page sheet with pertinent information, and it will post the result inside your chat. The document starts private, but you can choose to share it with colleagues if you're collaborating with them. While you can edit the document within Claude, you can also export and open it in Google Docs or Microsoft Word if you want. The company is also adding the ability to create slides and presentations within conversations.


fb693c67f61e5321746ffce8b6fdd2d0-Paper-Datasets_and_Benchmarks_Track.pdf

Neural Information Processing Systems

Although numerous Artificial Intelligence Generated Image (AIGI) detectors have been proposed, often reporting high accuracy, their effectiveness in real-world scenarios remains questionable. To bridge this gap, we introduce AIGIBench, a comprehensive benchmark designed to rigorously evaluate the robustness and generalization capabilities of state-of-the-art AIGI detectors. AIGIBench simulates real-world challenges through four core tasks: multi-source generalization, robustness to image degradation, sensitivity to data augmentation, and impact of test-time preprocessing. It includes 23 diverse fake image subsets that span both advanced and widely adopted image generation techniques, along with real-world samples collected from social media and AI art platforms. Extensive experiments on 11 advanced detectors demonstrate that, despite their high reported accuracy in controlled settings, these detectors suffer significant performance drops on real-world data, limited benefits from common augmentations, and nuanced effects of preprocessing, highlighting the need for more robust detection strategies. By providing a unified and realistic evaluation framework, AIGIBench offers valuable insights to guide future research toward dependable and generalizable AIGI detection2.


Fine Temporal Preference Optimization for Video Diffusion Models

Neural Information Processing Systems

Direct Preference Optimization (DPO) has recently been applied as a post-training technique for text-to-video diffusion models. To obtain training data, annotators are asked to provide preferences between two videos generated from independent noise. However, this approach prohibits fine-grained comparisons, and we point out that it biases the annotators towards low-motion clips as they often contain fewer visual artifacts. In this work, we introduce DenseDPO, a method that addresses these shortcomings by making three contributions. First, we create each video pair for DPO by denoising corrupted copies of a ground truth video. This results in aligned pairs with similar motion structures while differing in local details, effectively neutralizing the motion bias. Second, we leverage the resulting temporal alignment to label preferences on short segments rather than entire clips, yielding a denser and more precise learning signal. With only one-third of the labeled data, DenseDPO greatly improves motion generation over vanilla DPO, while matching it in text alignment, visual quality, and temporal consistency. Finally, we show that DenseDPO unlocks automatic preference annotation using off-the-shelf Vision Language Models (VLMs): GPT accurately predicts segment-level preferences similar to task-specifically fine-tuned video reward models, and DenseDPO trained on these labels achieves performance close to using human labels.


Fix False Transparency by Noise Guided Splatting

Neural Information Processing Systems

Opaque objects reconstructed by 3DGaussian Splatting (3DGS) often exhibit a falsely transparent surface, leading to inconsistent background and internal patterns under camera motion in interactive viewing. This issue stems from the ill-posed optimization in 3DGS. During training, background and foreground Gaussians are blended via ฮฑ-compositing and optimized solely against the input RGB images using a photometric loss. As this process lacks an explicit constraint on surface opacity, the optimization may incorrectly assign transparency to opaque regions, resulting in view-inconsistent and falsely transparent output. This issue is difficult to detect in standard evaluation settings (i.e., rendering static images), but becomes particularly evident in object-centric reconstructions under interactive viewing.


VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models

Neural Information Processing Systems

Faces synthesized by diffusion models (DMs) with high-quality and controllable attributes pose a significant challenge for Deepfake detection. Most state-of-the-art detectors only yield a binary decision, incapable of forgery localization, attribution of forgery methods, and providing analysis on the cause of forgeries. In this work, we integrate Multimodal Large Language Models (MLLMs) within DMbased face forensics, and propose a fine-grained analysis triad framework called VLForgery, that can 1) predict falsified facial images; 2) locate the falsified face regions subjected to partial synthesis; and 3) attribute the synthesis with specific generators. To achieve the above goals, we introduce VLF (Visual Language Forensics), a novel and diverse synthesis face dataset designed to facilitate rich interactions between'Visual' and'Language' modalities in MLLMs. Additionally, we propose an extrinsic knowledge-guided description method, termed EkCot, which leverages knowledge from the image generation pipeline to enable MLLMs to quickly capture image content. Furthermore, we introduce a low-level vision comparison pipeline designed to identify differential features between real and fake that MLLMs can inherently understand. These features are then incorporated into EkCot, enhancing its ability to analyze forgeries in a structured manner, following the sequence of detection, localization, and attribution. Extensive experiments demonstrate that VLForgery outperforms other state-of-the-art forensic approaches in detection accuracy, with additional potential for falsified region localization and attribution analysis.



PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement

Neural Information Processing Systems

Latent Diffusion Models (LDMs) have markedly advanced the quality of image inpainting and local editing. However, the inherent latent compression often introduces pixel-level inconsistencies, such as chromatic shifts, texture mismatches, and visible seams along editing boundaries. Existing remedies, including backgroundconditioned latent decoding and pixel-space harmonization, usually fail to fully eliminate these artifacts in practice and do not generalize well across different latent representations or tasks. We introduce PixPerfect, a pixel-level refinement framework that delivers seamless, high-fidelity local edits across diverse LDM architectures and tasks. PixPerfect leverages (i) a differentiable discriminative pixel space that amplifies and suppresses subtle color and texture discrepancies, (ii) a comprehensive artifact simulation pipeline that exposes the refiner to realistic local editing artifacts during training, and (iii) a direct pixel-space refinement scheme that ensures broad applicability across diverse latent representations and tasks. Extensive experiments on inpainting, object removal, and insertion benchmarks demonstrate that PixPerfect substantially enhances perceptual fidelity and downstream editing performance, establishing a new standard for robust and high-fidelity localized image editing.


Building 3DRepresentations and Generating Motions From a Single Image via Video-Generation

Neural Information Processing Systems

Autonomous robots typically need to construct representations of their surroundings and adapt their motions to the geometry of their environment. Here, we tackle the problem of constructing a policy model for collision-free motion generation, consistent with the environment, from a single input RGB image. Extracting 3D structures from a single image often involves monocular depth estimation. Developments in depth estimation have given rise to large pre-trained models such as DepthAnything. However, using outputs of these models for downstream motion generation is challenging due to frustum-shaped errors that arise.