Goto

Collaborating Authors

 Media


Chicago paper publishes AI-generated 'summer reading list' with books that don't exist

FOX News

Texas high school student Elliston Berry joins'Fox & Friends' to discuss the House's passage of a new bill that criminalizes the sharing of non-consensual intimate images, including content created with artificial intelligence. The Chicago Sun-Times admitted on Tuesday that it published an AI-generated list of books that don't exist for its summer reading list. On Sunday, the publication released a special 64-page section titled "Heat Index: Your Guide to the Best of Summer" which featured a list of 15 recommended books for summer. However, upon further look, it was found that 10 of the 15 books on the list were not real. One example included a book called "Nightshade Market" by Min Jin Lee, which was described as a "riveting tale set in Seoul's underground economy" and follows "three women whose paths intersect in an illegal night market" exploring "class, gender and the shadow economies beneath prosperous societies."


Wheeled, rugged robot dog built for extreme industrial missions

FOX News

The machine is designed to inspect industrial sites, respond to disasters, carry out logistics operations and support scientific research. Deep Robotics, a company from China, has unveiled a durable four-legged robot built to operate in extreme environments that humans struggle to traverse. It's called the Lynx M20, and it builds upon the agility of its predecessor, the Lynx robot dog. This versatile machine is designed to handle anything from inspecting industrial sites and responding to disasters to carrying out logistics operations and supporting scientific research. Here's what you need to know.


AI Melania: First lady embarks on 'new frontier' in publishing with audiobook of memoir

FOX News

EXCLUSIVE: First lady Melania Trump is launching an audiobook of her memoir using artificial intelligence (AI) audio technology in multiple languages, Fox News Digital has learned. The first lady released her first memoir, "Melania," last year. This week, she is breaking new ground by releasing "Melania, the Audiobook," which has been "created entirely" with AI. "I am proud to be at the forefront of publishing's new frontier โ€“ the intersection of artificial intelligence technology and audio," Trump told Fox News Digital. The first lady said ElevenLabs AI developed "an AI-generated replica of my voice under strict supervision, which will establish an unforgettable connection with my personal story, in multiple languages for listeners worldwide." ElevenLabs AI CEO Mati Staniszewski told Fox News Digital that they are "excited that Melania Trump trusted our technology to power this first-of-its-kind audiobook project."


'Shakespeare would be writing for games today': Cannes' first video game Lili is a retelling of Macbeth

The Guardian

The Cannes film festival isn't typically associated with video games, but this year it's playing host to an unusual collaboration. Lili is a co-production between the New York-based game studio iNK Stories (creator of 1979 Revolution: Black Friday, about a photojournalist in Iran) and the Royal Shakespeare Company, and it's been turning heads with its eye-catching translocation of Macbeth to modern-day Iran. "It's been such an incredible coup to have it as the first video game experience at Cannes," says iNK Stories co-founder Vassiliki Khonsari. "People have gone in saying, I'm not familiar playing games, so I may just try it out for five minutes. The Cannes festival's Immersive Competition began in 2024, although the lineup doesn't usually feature traditional video games. "VR films and projection mapping is the thrust of it," says iNK Stories' other co-founder, Vassiliki's husband Navid Khonsari. But Lili weaves live-action footage with video game mechanics in a similar way to a game such as Telling Lies or Immortality. Its lead, Zar Amir Ebrahimi, won best actress at Cannes three years ago. Lili focuses on the story of Lady Macbeth, here cast as the ambitious wife of an upwardly mobile officer in the Basij (a paramilitary volunteer militia within the Islamic Revolutionary Guard in Iran). As in the play, she plots a murder to secure her husband's rise. "I think that the narrative of Lady Macbeth is that she's manipulative, and that's exactly what got us interested," says Navid. "The social limitations based on her gender forced her to try to attain whatever leadership role she can," he continues. "If she was a man, she would have been one of the greatest kings that country would have ever experienced, but because she was a woman she had to work within the structure that was there for her.


Large Language Models Are More Persuasive Than Incentivized Human Persuaders

arXiv.org Artificial Intelligence

We directly compare the persuasion capabilities of a frontier large language model (LLM; Claude Sonnet 3.5) against incentivized human persuaders in an interactive, real - time conversational quiz setting. In this preregistered, large - scale incentivized expe riment, participants (quiz takers) completed an online quiz where persuaders (either humans or LLMs) attempted to persuade quiz takers toward correct or incorrect answers. We find that LLM persuaders achieved significantly higher compliance with their dire ctional persuasion attempts than incentivized human persuaders, demonstrating superior persuasive capabilities in both truthful (toward correct answers) and deceptive (toward incorrect answers) contexts. We also find that LLM persuaders significantly incre ased quiz takers' accuracy, leading to higher earnings, when steering quiz takers toward correct answers, and significantly decreased their accuracy, leading to lower earnings, when steering them toward incorrect answers. Overall, our findings suggest that AI's persuasion capabilities already exceed those of humans that have real - money bonuses tied to performance. Our findings of increasingly capable AI persuaders thus underscore the urgency of emerging alignment and governance frameworks.


Improving Language Model Personas via Rationalization with Psychological Scaffolds

arXiv.org Artificial Intelligence

Language models prompted with a user description or persona are being used to predict the user's preferences and opinions. However, existing approaches to building personas mostly rely on a user's demographic attributes and/or prior judgments, but not on any underlying reasoning behind a user's judgments. We introduce PB&J (Psychology of Behavior and Judgments), a framework that improves LM personas by incorporating potential rationales for why the user could have made a certain judgment. Our rationales are generated by a language model to explicitly reason about a user's behavior on the basis of their experiences, personality traits, or beliefs. Our method employs psychological scaffolds: structured frameworks such as the Big 5 Personality Traits or Primal World Beliefs to help ground the generated rationales in existing theories. Experiments on public opinion and movie preference prediction tasks demonstrate that language model personas augmented with PB&J rationales consistently outperform personas conditioned only on user demographics and / or judgments, including those that use a model's default chain-of-thought, which is not grounded in psychological theories. Additionally, our PB&J personas perform competitively with those using human-written rationales, suggesting the potential of synthetic rationales guided by existing theories.


Long-Form Information Alignment Evaluation Beyond Atomic Facts

arXiv.org Artificial Intelligence

Information alignment evaluators are vital for various NLG evaluation tasks and trustworthy LLM deployment, reducing hallucinations and enhancing user trust. Current fine-grained methods, like FactScore, verify facts individually but neglect inter-fact dependencies, enabling subtle vulnerabilities. In this work, we introduce MontageLie, a challenging benchmark that constructs deceptive narratives by "montaging" truthful statements without introducing explicit hallucinations. We demonstrate that both coarse-grained LLM-based evaluators and current fine-grained frameworks are susceptible to this attack, with AUC-ROC scores falling below 65%. To enable more robust fine-grained evaluation, we propose DoveScore, a novel framework that jointly verifies factual accuracy and event-order consistency. By modeling inter-fact relationships, DoveScore outperforms existing fine-grained methods by over 8%, providing a more robust solution for long-form text alignment evaluation. Our code and datasets are available at https://github.com/dannalily/DoveScore.


Exploring the Innovation Opportunities for Pre-trained Models

arXiv.org Artificial Intelligence

Innovators transform the world by understanding where services are successfully meeting customers' needs and then using this knowledge to identify failsafe opportunities for innovation. Pre-trained models have changed the AI innovation landscape, making it faster and easier to create new AI products and services. Understanding where pre-trained models are successful is critical for supporting AI innovation. Unfortunately, the hype cycle surrounding pre-trained models makes it hard to know where AI can really be successful. To address this, we investigated pre-trained model applications developed by HCI researchers as a proxy for commercially successful applications. The research applications demonstrate technical capabilities, address real user needs, and avoid ethical challenges. Using an artifact analysis approach, we categorized capabilities, opportunity domains, data types, and emerging interaction design patterns, uncovering some of the opportunity space for innovation with pre-trained models.


IA-T2I: Internet-Augmented Text-to-Image Generation

arXiv.org Artificial Intelligence

Current text-to-image (T2I) generation models achieve promising results, but they fail on the scenarios where the knowledge implied in the text prompt is uncertain. For example, a T2I model released in February would struggle to generate a suitable poster for a movie premiering in April, because the character designs and styles are uncertain to the model. To solve this problem, we propose an Internet-Augmented text-to-image generation (IA-T2I) framework to compel T2I models clear about such uncertain knowledge by providing them with reference images. Specifically, an active retrieval module is designed to determine whether a reference image is needed based on the given text prompt; a hierarchical image selection module is introduced to find the most suitable image returned by an image search engine to enhance the T2I model; a self-reflection mechanism is presented to continuously evaluate and refine the generated image to ensure faithful alignment with the text prompt. To evaluate the proposed framework's performance, we collect a dataset named Img-Ref-T2I, where text prompts include three types of uncertain knowledge: (1) known but rare. (2) unknown. (3) ambiguous. Moreover, we carefully craft a complex prompt to guide GPT-4o in making preference evaluation, which has been shown to have an evaluation accuracy similar to that of human preference evaluation. Experimental results demonstrate the effectiveness of our framework, outperforming GPT-4o by about 30% in human evaluation.


ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality

arXiv.org Artificial Intelligence

Despite extensive research on toxic speech detection in text, a critical gap remains in handling spoken Mandarin audio. The lack of annotated datasets that capture the unique prosodic cues and culturally specific expressions in Mandarin leaves spoken toxicity underexplored. To address this, we introduce Toxic-Tone--the largest public dataset of its kind--featuring detailed annotations that distinguish both forms of toxicity (e.g., profanity, bullying) and sources of toxicity (e.g., anger, sarcasm, dismissiveness). Our data, sourced from diverse real-world audio and organized into 13 topical categories, mirrors authentic communication scenarios. We also propose a multimodal detection framework that integrates acoustic, linguistic, and emotional features using state-of-the-art speech and emotion encoders. Extensive experiments show our approach outperforms text-only and baseline models, underscoring the essential role of speech-specific cues in revealing hidden toxic expressions. Index T erms: Toxicity detection; Mandarin Chinese; Annotation; Ensemble Warning: This paper may contain uncomfortable content.