Goto

Collaborating Authors

 Generative AI


ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep Generation

Neural Information Processing Systems

DALL E) can produce high-quality images based on input language descriptions. These models incorporate a black-box safety filter to prevent the generation of unsafe or unethical content, such as violent, criminal, or hateful imagery. Recent jailbreaking methods generate adversarial prompts capable of bypassing safety filters and producing unsafe content, exposing vulnerabilities in influential commercial models. However, once these adversarial prompts are identified, the safety filter can be updated to prevent the generation of unsafe images. In this work, we propose an effective, simple, and difficult-to-detect jailbreaking solution: generating safe content initially with normal text prompts and then editing the generations to embed unsafe content.


Anthropic Denies It Could Sabotage AI Tools During War

WIRED

The Department of Defense alleges the AI developer could manipulate models in the middle of war. Company executives argue that's impossible. Anthropic cannot manipulate its generative AI model Claude once the US military has it running, an executive wrote in a court filing on Friday. The statement was made in response to accusations from the Trump administration about the company potentially tampering with its AI tools during war . "Anthropic has never had the ability to cause Claude to stop working, alter its functionality, shut off access, or otherwise influence or imperil military operations," Thiyagu Ramasamy, Anthropic's head of public sector, wrote .


OpenAI is developing a unified AI 'superapp' for desktop users

PCWorld

OpenAI is developing a unified desktop superapp that will integrate ChatGPT, Codex, and Atlas into a single application, according to PCWorld's coverage of The Wall Street Journal report. This consolidation aims to reduce service fragmentation and improve overall quality for users accessing OpenAI's various AI tools. The superapp represents a significant shift toward streamlined AI services, potentially making OpenAI's offerings more accessible and efficient for desktop users. It seems you'll soon be able to access most of OpenAI's services in one place on your computer.


The Download: OpenAI is building a fully automated researcher, and a psychedelic trial blind spot

MIT Technology Review

Plus: OpenAI is also creating a super app. OpenAI has a new grand challenge: building an AI researcher--a fully automated agent-based system capable of tackling large, complex problems by itself. The San Francisco firm said the new goal will be its "north star" for the next few years. By September, the company plans to build "an autonomous AI research intern" that can take on a small number of specific research problems. The intern will be the precursor to the fully automated multi-agent system, which is slated to debut in 2028. In an exclusive interview this week, OpenAI's chief scientist, Jakub Pachocki, talked me through the plans.


OpenAI is throwing everything into building a fully automated researcher

MIT Technology Review

OpenAI is refocusing its research efforts and throwing its resources into a new grand challenge. The San Francisco firm has set its sights on building what it calls an AI researcher, a fully automated agent-based system that will be able to go off and tackle large, complex problems by itself. OpenAI says that this new research goal will be its "North Star" for the next few years, pulling together multiple research strands, including work on reasoning models, agents, and interpretability .


Debiasing Synthetic Data Generated by Deep Generative Models

Neural Information Processing Systems

While synthetic data hold great promise for privacy protection, their statistical analysis poses significant challenges that necessitate innovative solutions. The use of deep generative models (DGMs) for synthetic data generation is known to induce considerable bias and imprecision into synthetic data analyses, compromising their inferential utility as opposed to original data analyses. This bias and uncertainty can be substantial enough to impede statistical convergence rates, even in seemingly straightforward analyses like mean calculation.


OpenAI is putting ChatGPT, its browser and code generator into one desktop app

Engadget

The company is reportedly making a unified app to streamline the user experience. OpenAI is developing a "super app" for desktop that unifies ChatGPT, its browser and its Codex app, according to the and . A company spokesperson told the publications that OpenAI Chief of Applications Fidji Simo will lead the application revamp with assistance from OpenAI President Greg Brockman. Simo will also help the marketing team advertise the app when it comes out. OpenAI's leadership is apparently hoping that combining several products can help it streamline user experience and dedicate its resources to one project.


A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion Models

Neural Information Processing Systems

High-dimensional data commonly lies on low-dimensional submanifolds, and estimating the local intrinsic dimension (LID) of a datum -- i.e. the dimension of the submanifold it belongs to -- is a longstanding problem. LID can be understood as the number of local factors of variation: the more factors of variation a datum has, the more complex it tends to be. Estimating this quantity has proven useful in contexts ranging from generalization in neural networks to detection of out-of-distribution data, adversarial examples, and AI-generated text. The recent successes of deep generative models present an opportunity to leverage them for LID estimation, but current methods based on generative models produce inaccurate estimates, require more than a single pre-trained model, are computationally intensive, or do not exploit the best available deep generative models: diffusion models (DMs). In this work, we show that the Fokker-Planck equation associated with a DM can provide an LID estimator which addresses the aforementioned deficiencies.


Game devs say Nvidia's DLSS 5 reveal blindsided them

PCWorld

PCWorld reports that Nvidia's DLSS 5 announcement caught major game developers from Ubisoft and Capcom off-guard, who were unaware their games would be featured in demonstrations. The generative AI technology faces significant backlash from gamers who criticize it as an "AI filter" that potentially devalues game aesthetics and may require two high-end GPUs. Despite being planned for fall 2026 release, DLSS 5 already raises concerns about artistic control and whether developers want this AI-enhanced visual processing in their games. Nvidia DLSS 5 is coming later this year, adding generative "AI" features to the performance-enhancing tech . Gamers are calling the tool an "Instagram yaas filter" and "AI slop," among other, less kind terms. The way that it adds detail to faces and seems to hijack -- or replace?


ChatGPT is dialing back its 'if you want' end-response teasers

PCWorld

Instant to reduce annoying "if you want" and teaser-style phrasing that users found intrusive. This change addresses widespread user complaints about persistent, clickbait-like follow-up prompts that negatively impacted the AI interaction experience. The update aims to create more natural, direct conversations by making ChatGPT less chatty and eliminating the bothersome response teasers. It wasn't all that long ago that ChatGPT was a constant nag, persistently dropping "Would you like me to?"-style questions at the end of its responses. OpenAI eventually tweaked the phrasing, dropping the question marks and going for "if you want"-style teasers that invited users to extend their chat sessions. Now, OpenAI has acknowledged that it went too far with the clickbaity follow-ups, noting in a recent update for one of its newest models that it's now cutting back on the teasers. "We're rolling out an update to GPT-5.3 Instant that improves follow-up tone and reduces teaser-style phrasing," reads a recent ChatGPT release note, which adds that users should soon see fewer follow-ups like "if you want," "you'll never believe," and "I can tell you three things that " Those teasers are, of course, a way for ChatGPT to keep subscribers chatting, but users have been complaining that the persistent follow-ups are more annoying than they are intriguing. "I hated it with a passion and hope it's completely gone," wrote one user on Reddit .