Generative AI
Generative AI for Research Data Processing: Lessons Learnt From Three Use Cases
Mitra, Modhurita, de Vos, Martine G., Cortinovis, Nicola, Ometto, Dawa
--There has been enormous interest in generative AI since ChatGPT was launched in 2022. However, there are concerns about the accuracy and consistency of the outputs of generative AI. We have carried out an exploratory study on the application of this new technology in research data processing. We identified tasks for which rule-based or traditional machine learning approaches were difficult to apply, and then performed these tasks using generative AI. We demonstrate the feasibility of using the generative AI model Claude 3 Opus in three research projects involving complex data processing tasks: 1) Information extraction: We extract plant species names from historical seedlists (catalogues of seeds) published by botanical gardens. We share the lessons we learnt from these use cases: How to determine if generative AI is an appropriate tool for a given data processing task, and if so, how to maximise the accuracy and consistency of the results obtained. In this paper, we share our insights on the application of generative AI in research software engineering projects. Generative AI can potentially be used to perform a wide variety of research data processing tasks, such as interpreting documents, extracting information from them, and classifying text into categories. Since the tasks are specified through prompts in natural language, the barrier to entry is low. Therefore, this tool can be easily used by domain experts in a wide range of fields, with varying levels of programming skills and depth of knowledge of technical topics such as machine learning.
A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models
Cao, Gengxian, Li, Fengyuan, Duan, Hong, Yang, Ye, Wang, Bofeng, Li, Donghe
This paper introduces a novel multi-Agent framework that automates the end to end production of Qinqiang opera by integrating Large Language Models , visual generation, and Text to Speech synthesis. Three specialized agents collaborate in sequence: Agent1 uses an LLM to craft coherent, culturally grounded scripts;Agent2 employs visual generation models to render contextually accurate stage scenes; and Agent3 leverages TTS to produce synchronized, emotionally expressive vocal performances. In a case study on Dou E Yuan, the system achieved expert ratings of 3.8 for script fidelity, 3.5 for visual coherence, and 3.8 for speech accuracy-culminating in an overall score of 3.6, a 0.3 point improvement over a Single Agent baseline. Ablation experiments demonstrate that removing Agent2 or Agent3 leads to drops of 0.4 and 0.5 points, respectively, underscoring the value of modular collaboration. This work showcases how AI driven pipelines can streamline and scale the preservation of traditional performing arts, and points toward future enhancements in cross modal alignment, richer emotional nuance, and support for additional opera genres.
Application of Deep Generative Models for Anomaly Detection in Complex Financial Transactions
Tang, Tengda, Yao, Jianhua, Wang, Yixian, Sha, Qiuwu, Feng, Hanrui, Xu, Zhen
This study proposes an algorithm for detecting suspicious behaviors in large payment flows based on deep generative models. By combining Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), the algorithm is designed to detect abnormal behaviors in financial transactions. First, the GAN is used to generate simulated data that approximates normal payment flows. The discriminator identifies anomalous patterns in transactions, enabling the detection of potential fraud and money laundering behaviors. Second, a VAE is introduced to model the latent distribution of payment flows, ensuring that the generated data more closely resembles real transaction features, thus improving the model's detection accuracy. The method optimizes the generative capabilities of both GAN and VAE, ensuring that the model can effectively capture suspicious behaviors even in sparse data conditions. Experimental results show that the proposed method significantly outperforms traditional machine learning algorithms and other deep learning models across various evaluation metrics, especially in detecting rare fraudulent behaviors. Furthermore, this study provides a detailed comparison of performance in recognizing different transaction patterns (such as normal, money laundering, and fraud) in large payment flows, validating the advantages of generative models in handling complex financial data.
AI with Emotions: Exploring Emotional Expressions in Large Language Models
Ishikawa, Shin-nosuke, Yoshino, Atsushi
The human-level performance of Large Language Models (LLMs) across various tasks has raised expectations for the potential of Artificial Intelligence (AI) to possess emotions someday. To explore the capability of current LLMs to express emotions in their outputs, we conducted an experiment using several LLMs (OpenAI GPT, Google Gemini, Meta Llama3, and Cohere Command R+) to role-play as agents answering questions with specified emotional states. We defined the emotional states using Russell's Circumplex model, a well-established framework that characterizes emotions along the sleepy-activated (arousal) and pleasure-displeasure (valence) axes. We chose this model for its simplicity, utilizing two continuous parameters, which allows for better controllability in applications involving continuous changes in emotional states. The responses generated were evaluated using a sentiment analysis model, independent of the LLMs, trained on the GoEmotions dataset. The evaluation showed that the emotional states of the generated answers were consistent with the specifications, demonstrating the LLMs' capability for emotional expression. This indicates the potential for LLM-based AI agents to simulate emotions, opening up a wide range of applications for emotion-based interactions, such as advisors or consultants who can provide advice or opinions with a personal touch.
Expanding the Generative AI Design Space through Structured Prompting and Multimodal Interfaces
Karnatak, Nimisha, Baranes, Adrien, Marchant, Rob, Zeng, Huinan, Butler, Trรญona, Olson, Kristen
Text-based prompting remains the predominant interaction paradigm in generative AI, yet it often introduces friction for novice users such as small business owners (SBOs), who struggle to articulate creative goals in domain-specific contexts like advertising. Through a formative study with six SBOs in the United Kingdom, we identify three key challenges: difficulties in expressing brand intuition through prompts, limited opportunities for fine-grained adjustment and refinement during and after content generation, and the frequent production of generic content that lacks brand specificity. In response, we present ACAI (AI Co-Creation for Advertising and Inspiration), a multimodal generative AI tool designed to support novice designers by moving beyond traditional prompt interfaces. ACAI features a structured input system composed of three panels: Branding, Audience and Goals, and the Inspiration Board. These inputs allow users to convey brand-relevant context and visual preferences. This work contributes to HCI research on generative systems by showing how structured interfaces can foreground user-defined context, improve alignment, and enhance co-creative control in novice creative workflows.
Labeling Messages as AI-Generated Does Not Reduce Their Persuasive Effects
Gallegos, Isabel O., Shani, Chen, Shi, Weiyan, Bianchi, Federico, Gainsburg, Izzy, Jurafsky, Dan, Willer, Robb
As generative artificial intelligence (AI) enables the creation and dissemination of information at massive scale and speed, it is increasingly important to understand how people perceive AI-generated content. One prominent policy proposal requires explicitly labeling AI-generated content to increase transparency and encourage critical thinking about the information, but prior research has not yet tested the effects of such labels. To address this gap, we conducted a survey experiment (N=1601) on a diverse sample of Americans, presenting participants with an AI-generated message about several public policies (e.g., allowing colleges to pay student-athletes), randomly assigning whether participants were told the message was generated by (a) an expert AI model, (b) a human policy expert, or (c) no label. We found that messages were generally persuasive, influencing participants' views of the policies by 9.74 percentage points on average. However, while 94.6% of participants assigned to the AI and human label conditions believed the authorship labels, labels had no significant effects on participants' attitude change toward the policies, judgments of message accuracy, nor intentions to share the message with others. These patterns were robust across a variety of participant characteristics, including prior knowledge of the policy, prior experience with AI, political party, education level, or age. Taken together, these results imply that, while authorship labels would likely enhance transparency, they are unlikely to substantially affect the persuasiveness of the labeled content, highlighting the need for alternative strategies to address challenges posed by AI-generated information.
OpenAI says it would buy Chrome if Google is forced to sell
Google is under the microscope following a court ruling last year that it has a monopoly over online search, but the future of its vast suite of digital services is still uncertain at this stage. Last month, the Justice Department suggested that Google would need to sell off the Chrome browser; if the tech giant does make that move, there's already at least one interested buyer. Bloomberg reports that Nick Turley, head of ChatGPT, spoke at a hearing today about the Google monopoly situation and was asked whether OpenAI would be interested in acquiring Chrome. "Yes, we would, as would many other parties," he said. Users can currently use the ChatGPT AI assistant in Chrome through a plugin, but Turley said there could be deeper integrations if OpenAI owned the browser.
ChatGPT users annoyed by the AI's incessantly 'phony' positivity
ChatGPT users are increasingly criticizing the AI-powered chatbot for being too positive in its responses, Ars Technica reports. When you converse with ChatGPT, you might notice that the chatbot tends to inflate its responses with praise and flattery, saying things like "Good question!" and "You have a rare talent" and "You're thinking on a level most people can only dream of." Over the years, users have remarked on ChatGPT's fawning responses, which ranges from positive affirmations to outright flattery and more. One X user described the chatbot as "the biggest suckup I've ever met," another complained that it was "phony," and yet another lamented the chatbot's behavior and called it "freaking annoying." This is known as "sycophancy" among AI researchers, and it's entirely intentional based on how OpenAI has trained the underlying AI models.
The Great AI Lock-In Has Begun
There are really two OpenAIs. One is the creator of world-bending machines--the start-up that unleashed ChatGPT and in turn the generative-AI boom, surging toward an unrecognizable future with the rest of the tech industry in tow. This is the OpenAI that promises to eventually bring about "superintelligent" programs that exceed humanity's capabilities. The other OpenAI is simply a business. This is the company that is reportedly working on a social network and considering an expansion into hardware; it is the company that offers user-experience updates to ChatGPT, such as an "image library" feature announced last week and the new ability to "reference" past chats to provide personalized responses.
OpenAI's newest AI models hallucinate way more, for reasons unknown
Last week, OpenAI released its new o3 and o4-mini reasoning models, which perform significantly better than their o1 and o3-mini predecessors and have new capabilities like "thinking with images" and agentically combining AI tools for more complex results. This is unusual as newer models tend to hallucinate less as the underlying AI tech improves. In the realm of LLMs and reasoning AIs, a "hallucination" occurs when the model makes up information that sounds convincing but has no bearing in truth. In other words, when you ask questions to ChatGPT, it may respond with an answer that's patently false or incorrect. OpenAI's in-house benchmark PersonQA--which is used to measure the factual accuracy of its AI models when talking about people--found that o3 hallucinated in 33 percent of responses while o4-mini did even worse at 48 percent.