Generative AI
Project Riley: Multimodal Multi-Agent LLM Collaboration with Emotional Reasoning and Voting
Ortigoso, Ana Rita, Vieira, Gabriel, Fuentes, Daniel, Frazรฃo, Luis, Costa, Nuno, Pereira, Antรณnio
This paper presents Project Riley, a novel multimodal and multi-model conversational AI architecture oriented towards the simulation of reasoning influenced by emotional states. Drawing inspiration from Pixar's Inside Out, the system comprises five distinct emotional agents - Joy, Sadness, Fear, Anger, and Disgust - that engage in structured multi-round dialogues to generate, criticise, and iteratively refine responses. A final reasoning mechanism synthesises the contributions of these agents into a coherent output that either reflects the dominant emotion or integrates multiple perspectives. The architecture incorporates both textual and visual large language models (LLMs), alongside advanced reasoning and self-refinement processes. A functional prototype was deployed locally in an offline environment, optimised for emotional expressiveness and computational efficiency. From this initial prototype, another one emerged, called Armando, which was developed for use in emergency contexts, delivering emotionally calibrated and factually accurate information through the integration of Retrieval-Augmented Generation (RAG) and cumulative context tracking. The Project Riley prototype was evaluated through user testing, in which participants interacted with the chatbot and completed a structured questionnaire assessing three dimensions: Emotional Appropriateness, Clarity and Utility, and Naturalness and Human-likeness. The results indicate strong performance in structured scenarios, particularly with respect to emotional alignment and communicative clarity.
Human-AI Collaboration or Academic Misconduct? Measuring AI Use in Student Writing Through Stylometric Evidence
Oliveira, Eduardo Araujo, Mohoni, Madhavi, Lรณpez-Pernas, Sonsoles, Saqr, Mohammed
Human - Artificial Intelligence (HAI) collaboration in writing offers opportunities to enhance efficiency and boost student confidence; however, it also carries risks, such as reduced creativity, over - reliance on AI - generated content, and academic integrity (Kim & Lee, 2023) . While the ethical use of AI in education is widely acknowledged as a way to enhance student learning (Cotton et al., 2023; Foltynek et al., 2023), the rise of Unauthorised Content Generation (UCG) presents a significant challenge to academic misconduct. Measuring the extent and nature of HAI collaboration in academic contexts remains a critical challenge for educators, particularly as generative AI (genAI) tools become increasingly available and integrated into educational settings (Atchley et al., 2024; E. Oliveira et al., 2023) . Distinguishing AI - generated text from human - authored content is necessary for understanding student learning behaviours, supporting skill development, and maintaining academic integrity. Analysing student writing patterns can help educators evaluate how st udents engage with AI tools, track their writing skill progression, and identify areas where additional support is needed (Pan et al., 2025). Existing detection tools for AI - assisted misconduct often lack reliability, explainability, and resilience to circ umvention strategies such as paraphrasing (Cotton et al., 2023) . These challenges highlight the need for innovative, transparent, and robust approaches to address the unacknowledged use of genAI in HAI collaboration within academic writing (Kasneci et al., 2023) .
Can LLMs Simulate Personas with Reversed Performance? A Benchmark for Counterfactual Instruction Following
Kumar, Sai Adith Senthil, Yan, Hao, Perepa, Saipavan, Yue, Murong, Yao, Ziyu
Large Language Models (LLMs) are now increasingly widely used to simulate personas in virtual environments, leveraging their instruction-following capability. However, we discovered that even state-of-the-art LLMs cannot simulate personas with reversed performance (e.g., student personas with low proficiency in educational settings), which impairs the simulation diversity and limits the practical applications of the simulated environments. In this work, using mathematical reasoning as a representative scenario, we propose the first benchmark dataset for evaluating LLMs on simulating personas with reversed performance, a capability that we dub "counterfactual instruction following". We evaluate both open-weight and closed-source LLMs on this task and find that LLMs, including the OpenAI o1 reasoning model, all struggle to follow counterfactual instructions for simulating reversedly performing personas. Intersectionally simulating both the performance level and the race population of a persona worsens the effect even further. These results highlight the challenges of counterfactual instruction following and the need for further research.
Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
Dimgba, Martha O., Oba, Sharon, Agrawal, Ameeta, Giabbanelli, Philippe J.
Language models have been shown to propagate social bias through their output, particularly in the representation of gender and ethnicity. This paper investigates gender and ethnicity biases in AI-generated occupational stories. Representation biases are measured before and after applying our proposed mitigation strategy, Bias Analysis and Mitigation through Explanation (BAME), revealing improvements in demographic representation ranging from 2% to 20%. BAME leverages model-generated explanations to inform targeted prompt engineering, effectively reducing biases without modifying model parameters. By analyzing stories generated across 25 occupational groups, three large language models (Claude 3.5 Sonnet, Llama 3.1 70B Instruct, and GPT-4 Turbo), and multiple demographic dimensions, we identify persistent patterns of overrepresentation and underrepresentation linked to training data stereotypes. Our findings demonstrate that guiding models with their own internal reasoning mechanisms can significantly enhance demographic parity, thereby contributing to the development of more transparent generative AI systems.
Exploring persuasive interactions with generative social robots: An experimental framework
Vonschallen, Stephan, Finsler, Larissa Julia Corina, Schmiedel, Theresa, Eyssel, Friederike
Integrating generative AI such as Large Language Models into social robots has improved their ability to engage in natural, human-like communication. This study presents a method to examine their persuasive capabilities. We designed an experimental framework focused on decision making and tested it in a pilot that varied robot appearance and self-knowledge. Using qualitative analysis, we evaluated interaction quality, persuasion effectiveness, and the robot's communicative strategies. Participants generally experienced the interaction positively, describing the robot as competent, friendly, and supportive, while noting practical limits such as delayed responses and occasional speech-recognition errors. Persuasiveness was highly context dependent and shaped by robot behavior: Participants responded well to polite, reasoned suggestions and expressive gestures, but emphasized the need for more personalized, context-aware arguments and clearer social roles. These findings suggest that generative social robots can influence user decisions, but their effectiveness depends on communicative nuance and contextual relevance. We propose refinements to the framework to further study persuasive dynamics between robots and human users.
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
Wang, Haoming, Zou, Haoyang, Song, Huatong, Feng, Jiazhan, Fang, Junjie, Lu, Junting, Liu, Longxiang, Luo, Qinyu, Liang, Shihao, Huang, Shijue, Zhong, Wanjun, Ye, Yining, Qin, Yujia, Xiong, Yuwen, Song, Yuxin, Wu, Zhiyong, Li, Aoyan, Li, Bo, Dun, Chen, Liu, Chong, Zan, Daoguang, Leng, Fuxing, Wang, Hanbin, Yu, Hao, Chen, Haobin, Guo, Hongyi, Su, Jing, Huang, Jingjia, Shen, Kai, Shi, Kaiyu, Yan, Lin, Zhao, Peiyao, Liu, Pengfei, Ye, Qinghao, Zheng, Renjie, Xin, Shulin, Zhao, Wayne Xin, Heng, Wen, Huang, Wenhao, Wang, Wenqian, Qin, Xiaobo, Lin, Yi, Wu, Youbin, Chen, Zehui, Wang, Zihao, Zhong, Baoquan, Zhang, Xinchun, Li, Xujing, Li, Yuanfan, Zhao, Zhongkai, Jiang, Chengquan, Wu, Faming, Zhou, Haotian, Pang, Jinlin, Han, Li, Liu, Qi, Ma, Qianli, Liu, Siyao, Cai, Songhua, Fu, Wenqi, Liu, Xin, Wang, Yaohui, Zhang, Zhi, Zhou, Bo, Li, Guoliang, Shi, Jiajun, Yang, Jiale, Tang, Jie, Li, Li, Han, Qihua, Lu, Taoran, Lin, Woyu, Tong, Xiaokang, Li, Xinyao, Zhang, Yichi, Miao, Yu, Jiang, Zhengxuan, Li, Zili, Zhao, Ziyuan, Li, Chenxin, Ma, Dehua, Lin, Feng, Zhang, Ge, Yang, Haihua, Guo, Hangyu, Zhu, Hongda, Liu, Jiaheng, Du, Junda, Cai, Kai, Li, Kuanye, Yuan, Lichen, Han, Meilan, Wang, Minchao, Guo, Shuyue, Cheng, Tianhao, Ma, Xiaobo, Xiao, Xiaojun, Huang, Xiaolong, Chen, Xinjie, Du, Yidi, Chen, Yilin, Wang, Yiwen, Li, Zhaojian, Yang, Zhenzhu, Zeng, Zhiyuan, Jin, Chaolin, Li, Chen, Chen, Hao, Chen, Haoli, Chen, Jian, Zhao, Qinghao, Shi, Guang
The development of autonomous agents for graphical user interfaces (GUIs) presents major challenges in artificial intelligence. While recent advances in native agent models have shown promise by unifying perception, reasoning, action, and memory through end-to-end learning, open problems remain in data scalability, multi-turn reinforcement learning (RL), the limitations of GUI-only operation, and environment stability. In this technical report, we present UI-TARS-2, a native GUI-centered agent model that addresses these challenges through a systematic training methodology: a data flywheel for scalable data generation, a stabilized multi-turn RL framework, a hybrid GUI environment that integrates file systems and terminals, and a unified sandbox platform for large-scale rollouts. Empirical evaluation demonstrates that UI-TARS-2 achieves significant improvements over its predecessor UI-TARS-1.5. On GUI benchmarks, it reaches 88.2 on Online-Mind2Web, 47.5 on OSWorld, 50.6 on WindowsAgentArena, and 73.3 on AndroidWorld, outperforming strong baselines such as Claude and OpenAI agents. In game environments, it attains a mean normalized score of 59.8 across a 15-game suite-roughly 60% of human-level performance-and remains competitive with frontier proprietary models (e.g., OpenAI o3) on LMGame-Bench. Additionally, the model can generalize to long-horizon information-seeking tasks and software engineering benchmarks, highlighting its robustness across diverse agent tasks. Detailed analyses of training dynamics further provide insights into achieving stability and efficiency in large-scale agent RL. These results underscore UI-TARS-2's potential to advance the state of GUI agents and exhibit strong generalization to real-world interactive scenarios.
DuckDuckGo's paid plan now includes advanced AI models like GPT-5
DuckDuckGo is now expanding its paid subscription with access to several of the most advanced AI models on the market. Subscribers can now access OpenAI's GPT-4o and GPT-5, Anthropic's Claude Sonnet 4, and Meta's Llama Maverick via the Duck.ai To protect privacy, Duck.ai hides the user's IP address from the AI model providers, and chat logs are saved locally and aren't used to train the AI models. In addition, there's a special "Fire Button" that lets users instantly delete previous conversations and chat histories. The price of the subscription remains unchanged at 9.99/month or 99/year.
Using generative AI, researchers design compounds that can kill drug-resistant bacteria
With help from artificial intelligence, MIT researchers have designed novel antibiotics that can combat two hard-to-treat infections: drug-resistant Neisseria gonorrhoeae and multi-drug-resistant Staphylococcus aureus (MRSA). Using generative AI algorithms, the research team designed more than 36 million possible compounds and computationally screened them for antimicrobial properties. The top candidates they discovered are structurally distinct from any existing antibiotics, and they appear to work by novel mechanisms that disrupt bacterial cell membranes. This approach allowed the researchers to generate and evaluate theoretical compounds that have never been seen before -- a strategy that they now hope to apply to identify and design compounds with activity against other species of bacteria. "We're excited about the new possibilities that this project opens up for antibiotics development. Our work shows the power of AI from a drug design standpoint, and enables us to exploit much larger chemical spaces that were previously inaccessible," says James Collins, the Termeer Professor of Medical Engineering and Science in MIT's Institute for Medical Engineering and Science (IMES) and Department of Biological Engineering, and a member of the Broad Institute.
Diffusion Generative Models Meet Compressed Sensing, with Applications to Image Data and Financial Time Series
Guo, Zhengyi, Li, Jiatu, Tang, Wenpin, Yao, David D.
This paper develops dimension reduction techniques for accelerating diffusion model inference in the context of synthetic data generation. The idea is to integrate compressed sensing into diffusion models: (i) compress the data into a latent space, (ii) train a diffusion model in the latent space, and (iii) apply a compressed sensing algorithm to the samples generated in the latent space, facilitating the efficiency of both model training and inference. Under suitable sparsity assumptions on data, the proposed algorithm is proved to enjoy faster convergence by combining diffusion model inference with sparse recovery. As a byproduct, we obtain an optimal value for the latent space dimension. We also conduct numerical experiments on a range of datasets, including image data (handwritten digits, medical images, and climate data) and financial time series for stress testing. Key words: Complexity, compressed sensing, diffusion models, inference time, signal recovery, sparsity.
Continuous Monitoring of Large-Scale Generative AI via Deterministic Knowledge Graph Structures
Gupta, Kishor Datta, Haque, Mohd Ariful, Ali, Hasmot, Kamal, Marufa, Alam, Syed Bahauddin, Rahman, Mohammad Ashiqur
Generative AI (GEN AI) models have revolutionized diverse application domains but present substantial challenges due to reliability concerns, including hallucinations, semantic drift, and inherent biases. These models typically operate as black-boxes, complicating transparent and objective evaluation. Current evaluation methods primarily depend on subjective human assessment, limiting scalability, transparency, and effectiveness. This research proposes a systematic methodology using deterministic and Large Language Model (LLM)-generated Knowledge Graphs (KGs) to continuously monitor and evaluate GEN AI reliability. We construct two parallel KGs: (i) a deterministic KG built using explicit rule-based methods, predefined ontologies, domain-specific dictionaries, and structured entity-relation extraction rules, and (ii) an LLM-generated KG dynamically derived from real-time textual data streams such as live news articles. Utilizing real-time news streams ensures authenticity, mitigates biases from repetitive training, and prevents adaptive LLMs from bypassing predefined benchmarks through feedback memorization. To quantify structural deviations and semantic discrepancies, we employ several established KG metrics, including Instantiated Class Ratio (ICR), Instantiated Property Ratio (IPR), and Class Instantiation (CI). An automated real-time monitoring framework continuously computes deviations between deterministic and LLM-generated KGs. By establishing dynamic anomaly thresholds based on historical structural metric distributions, our method proactively identifies and flags significant deviations, thus promptly detecting semantic anomalies or hallucinations. This structured, metric-driven comparison between deterministic and dynamically generated KGs delivers a robust and scalable evaluation framework.