Generative AI
AgentBuilder: Exploring Scaffolds for Prototyping User Experiences of Interface Agents
Liang, Jenny T., Barik, Titus, Nichols, Jeffrey, Schoop, Eldon, Cheng, Ruijia
Interface agents powered by generative AI models (referred to as "agents") can automate actions based on user commands. An important aspect of developing agents is their user experience (i.e., agent experience). There is a growing need to provide scaffolds for a broader set of individuals beyond AI engineers to prototype agent experiences, since they can contribute valuable perspectives to designing agent experiences. In this work, we explore the affordances agent prototyping systems should offer by conducting a requirements elicitation study with 12 participants with varying experience with agents. We identify key activities in agent experience prototyping and the desired capabilities of agent prototyping systems. We instantiate those capabilities in the AgentBuilder design probe for agent prototyping. We conduct an in situ agent prototyping study with 14 participants using AgentBuilder to validate the design requirements and elicit insights on how developers prototype agents and what their needs are in this process.
Neon: Negative Extrapolation From Self-Training Improves Image Generation
Alemohammad, Sina, Wang, Zhangyang, Baraniuk, Richard G.
Scaling generative AI models is bottlenecked by the scarcity of high-quality training data. The ease of synthesizing from a generative model suggests using (unverified) synthetic data to augment a limited corpus of real data for the purpose of fine-tuning in the hope of improving performance. Unfortunately, however, the resulting positive feedback loop leads to model autophagy disorder (MAD, aka model collapse) that results in a rapid degradation in sample quality and/or diversity. In this paper, we introduce Neon (for Negative Extrapolation frOm self-traiNing), a new learning method that turns the degradation from self-training into a powerful signal for self-improvement. Given a base model, Neon first fine-tunes it on its own self-synthesized data but then, counterintuitively, reverses its gradient updates to extrapolate away from the degraded weights. We prove that Neon works because typical inference samplers that favor high-probability regions create a predictable anti-alignment between the synthetic and real data population gradients, which negative extrapolation corrects to better align the model with the true data distribution. Neon is remarkably easy to implement via a simple post-hoc merge that requires no new real data, works effectively with as few as 1k synthetic samples, and typically uses less than 1% additional training compute. We demonstrate Neon's universality across a range of architectures (diffusion, flow matching, autoregressive, and inductive moment matching models) and datasets (ImageNet, CIFAR-10, and FFHQ). In particular, on ImageNet 256x256, Neon elevates the xAR-L model to a new state-of-the-art FID of 1.02 with only 0.36% additional training compute. Code is available at https://github.com/VITA-Group/Neon
ChatGPT will soon allow erotica for verified adults, says OpenAI boss
OpenAI plans to allow a wider range of content, including erotica, on its popular chatbot ChatGPT as part of its push to treat adult users like adults, says its boss Sam Altman. In a post on X on Tuesday, Mr Altman said upcoming versions of the popular chatbot would enable it to behave in a more human-like way - but only if you want it, not because we are usage maxxing. The move, reminiscent of Elon Musk's xAI recent introduction of two sexually explicit chatbots to Grok, could help OpenAI attract more paying subscribers. It is also likely to intensify pressure on lawmakers to introduce tighter restrictions on chatbot companions. OpenAI did not respond to the BBC's requests for comment following Mr Altman's post.
OpenAI will allow verified adults to use ChatGPT to generate erotic content
The company launched a dedicated ChatGPT experience for under-18 users in September. The company launched a dedicated ChatGPT experience for under-18 users in September. New version will allow users to customize AI assistant's personality in what firm calls'treat adults users like adults' policy OpenAI announced plans on Tuesday to relax restrictions on its ChatGPT chatbot, including allowing erotic content for verified adult users as part of what the company calls a "treat adult users like adults" principle. OpenAI's plan includes the release of an updated version of ChatGPT that will allow users to customize their AI assistant's personality, including options for more human-like responses, heavy emoji use, or friend-like behavior. The most significant change will come in December, when OpenAI plans to roll out more comprehensive age-gating that would permit erotic content for adults who have verified their ages.
'Sovereign AI' Has Become a New Front in the US-China Tech War
'Sovereign AI' Has Become a New Front in the US-China Tech War OpenAI has announced "AI sovereignty partnerships with governments around the world, but can proprietary models compete with Beijing's open source offerings? OpenAI has announced a number of projects this year with foreign governments to help build out what it has called their "sovereign AI" systems. The company says the deals, some of which are being coordinated with the US government, are part of a broader push to give national leaders more control over a technology that could reshape their economies. Over the past few months, sovereign AI has become something of a buzzword in both Washington and Silicon Valley. Proponents of the concept argue it's crucial that AI systems developed in democratic nations are able to proliferate globally, particularly as China races to deploy its own AI technology abroad.
Revisiting Trust in the Era of Generative AI: Factorial Structure and Latent Profiles
Sun, Haocan, Liu, Weizi, Wu, Di, Yu, Guoming, Yao, Mike
Trust is one of the most important factors shaping whether and how people adopt and rely on artificial intelligence (AI). Yet most existing studies measure trust in terms of functionality, focusing on whether a system is reliable, accurate, or easy to use, while giving less attention to the social and emotional dimensions that are increasingly relevant for today's generative AI (GenAI) systems. These systems do not just process information; they converse, respond, and collaborate with users, blurring the line between tool and partner. In this study, we introduce and validate the Human-AI Trust Scale (HAITS), a new measure designed to capture both the rational and relational aspects of trust in GenAI. Drawing on prior trust theories, qualitative interviews, and two waves of large-scale surveys in China and the United States, we used exploratory (n = 1,546) and confirmatory (n = 1,426) factor analyses to identify four key dimensions of trust: Affective Trust, Competence Trust, Benevolence & Integrity, and Perceived Risk. We then applied latent profile analysis to classify users into six distinct trust profiles, revealing meaningful differences in how affective-competence trust and trust-distrust frameworks coexist across individuals and cultures. Our findings offer a validated, culturally sensitive tool for measuring trust in GenAI and provide new insight into how trust evolves in human-AI interaction. By integrating instrumental and relational perspectives of trust, this work lays the foundation for more nuanced research and design of trustworthy AI systems.
Generative artificial intelligence improves projections of climate extremes
Tie, Ruian, Zhong, Xiaohui, Shi, Zhengyu, Li, Hao, Chen, Bin, Liu, Jun, Libo, Wu
Climate change is amplifying extreme weather and climate events worldwide [1]. Anthropogenic greenhouse gas emissions have disrupted the Earth's climate system, driving more frequent and severe heatwaves [2], cold spells [3], heavy precipitation [4], agricultural droughts [5], and tropical cyclones (TCs) [6]. Between 2016 and 2024, daily land temperature records show that extreme heat events occurred over four times more often than expected, while cold records declined by half [7]. These unprecedented shifts threaten human health [8, 9], infrastructure [10, 11], food security [12], biodiversity [13], and global economies [14, 15]. Therefore, reliable climate projections are essential for effective mitigation and adaptation strategies [16-18]. The Coupled Model Intercomparison Project (CMIP) [19] provides a foundation for global climate projections. Since its launch in 1995, CMIP has coordinated systematic evaluation of coupled general circulation models (GCMs). CMIP5 introduced Representative Concentration Pathways (RCPs), while CMIP6 extended this framework by incorporating Shared Socioeconomic Pathways (SSPs) through ScenarioMIP, enabling consistent simulations of emissions and socioeconomic trajectories to 2100 and facilitating integrated assessment of climate risks [20]. These advances have greatly enhanced the scientific and policy relevance of climate projections.
mCLM: A Modular Chemical Language Model that Generates Functional and Makeable Molecules
Edwards, Carl, Han, Chi, Lee, Gawon, Nguyen, Thao, Szymkuฤ, Sara, Prasad, Chetan Kumar, Jin, Bowen, Han, Jiawei, Diao, Ying, Liu, Ge, Peng, Hao, Grzybowski, Bartosz A., Burke, Martin D., Ji, Heng
Despite their ability to understand chemical knowledge, large language models (LLMs) remain limited in their capacity to propose novel molecules with desired functions (e.g., drug-like properties). In addition, the molecules that LLMs propose can often be challenging to make, and are almost never compatible with automated synthesis approaches. To better enable the discovery of functional small molecules, LLMs need to learn a new molecular language that is more effective in predicting properties and inherently synced with automated synthesis technology. Current molecule LLMs are limited by representing molecules based on atoms. In this paper, we argue that just like tokenizing texts into meaning-bearing (sub-)word tokens instead of characters, molecules should be tokenized at the level of functional building blocks, i.e., parts of molecules that bring unique functions and serve as effective building blocks for real-world automated laboratory synthesis. This motivates us to propose mCLM, a modular Chemical-Language Model that comprises a bilingual language model that understands both natural language descriptions of functions and molecular blocks. mCLM front-loads synthesizability considerations while improving the predicted functions of molecules in a principled manner. mCLM, with only 3B parameters, achieves improvements in synthetic accessibility relative to 7 other leading generative AI methods including GPT-5. When tested on 122 out-of-distribution medicines using only building blocks/tokens that are compatible with automated modular synthesis, mCLM outperforms all baselines in property scores and synthetic accessibility. mCLM can also reason on multiple functions and iteratively self-improve to rescue drug candidates that failed late in clinical trials ("fallen angels").
EvoCAD: Evolutionary CAD Code Generation with Vision Language Models
Preintner, Tobias, Yuan, Weixuan, Kรถnig, Adrian, Bรคck, Thomas, Raponi, Elena, van Stein, Niki
Abstract--Combining large language models with evolutionary computation algorithms represents a promising research direction leveraging the remarkable generative and in-context learning capabilities of LLMs with the strengths of evolutionary algorithms. Our method samples multiple CAD objects, which are then optimized using an evolutionary approach with vision language and reasoning language models. We assess our method using GPT -4V and GPT -4o, evaluating it on the CAD-Prompt benchmark dataset and comparing it to prior methods. Additionally, we introduce two new metrics based on topological properties defined by the Euler characteristic, which capture a form of semantic similarity between 3D objects. Our results demonstrate that EvoCAD outperforms previous approaches on multiple metrics, particularly in generating topologically correct objects, which can be efficiently evaluated using our two novel metrics that complement existing spatial metrics. The use of generative AI tools powered by large language models (LLMs) has transformed the way humans work, create, and develop. However, while significant attention is directed towards textual knowledge tasks, comparatively little focus is devoted on working with symbolic representations, such as those utilized in computer-aided design (CAD). These code-like textual representations, in the following referred as CAD code, enable visual assets to be processed by LLMs [21].
Characterizing Web Search in The Age of Generative AI
Kirsten, Elisabeth, Perdekamp, Jost Grosse, Upadhyay, Mihir, Gummadi, Krishna P., Zafar, Muhammad Bilal
The advent of LLMs has given rise to a new type of web search: Generative search, where LLMs retrieve web pages related to a query and generate a single, coherent text as a response. This output modality stands in stark contrast to traditional web search, where results are returned as a ranked list of independent web pages. In this paper, we ask: Along what dimensions do generative search outputs differ from traditional web search? We compare Google, a traditional web search engine, with four generative search engines from two providers (Google and OpenAI) across queries from four domains. Our analysis reveals intriguing differences. Most generative search engines cover a wider range of sources compared to web search. Generative search engines vary in the degree to which they rely on internal knowledge contained within the model parameters v.s. external knowledge retrieved from the web. Generative search engines surface varying sets of concepts, creating new opportunities for enhancing search diversity and serendipity. Our results also highlight the need for revisiting evaluation criteria for web search in the age of Generative AI.