Goto

Collaborating Authors

 Media


Analyzing Character Representation in Media Content using Multimodal Foundation Model: Effectiveness and Trust

arXiv.org Artificial Intelligence

Recent advances in AI has made automated analysis of complex media content at scale possible while generating actionable insights regarding character representation along such dimensions as gender and age. Past works focused on quantifying representation from audio/video/text using AI models, but without having the audience in the loop. We ask, even if character distribution along demographic dimensions are available, how useful are those to the general public? Do they actually trust the numbers generated by AI models? Our work addresses these open questions by proposing a new AI-based character representation tool and performing a thorough user study. Our tool has two components: (i) An analytics extraction model based on the Contrastive Language Image Pretraining (CLIP) foundation model that analyzes visual screen data to quantify character representation across age and gender; (ii) A visualization component effectively designed for presenting the analytics to lay audience. The user study seeks empirical evidence on the usefulness and trustworthiness of the AI-generated results for carefully chosen movies presented in the form of our visualizations. We found that participants were able to understand the analytics in our visualizations, and deemed the tool `overall useful'. Participants also indicated a need for more detailed visualizations to include more demographic categories and contextual information of the characters. Participants' trust in AI-based gender and age models is seen to be moderate to low, although they were not against the use of AI in this context. Our tool including code, benchmarking, and the user study data can be found at https://github.com/debadyuti0510/Character-Representation-Media.


Socioeconomic Threats of Deepfakes and the Role of Cyber-Wellness Education in Defense

Communications of the ACM

Due to the limits of science and its steep learning curve, we must rely on the expertise of others to develop our knowledge and skills.26 Toward this end, social media platforms have revolutionized how netizens--users who are actively engaged in online communities--gain knowledge and skills by facilitating the exchange of costless information with the public (for example, followers or influencers). Businesses around the world also use these platforms along with tools based on generative artificial intelligence (GenAI) to craft synthetic media, hoping to grow revenue by attracting more customers and improving their online experience.28 Generative AI tools can empower cyber threats and have cyberpsychological effects on netizens, allowing malicious actors to craft deepfakes in the form of disinformation, misinformation, and malinformation. Service providers not only must enhance GenAI tools to reduce hallucinations, but they also have a statutory duty to mitigate data-driven biases.


Light-based AI image generator uses almost no power

New Scientist

An AI image generator that uses light to produce images, rather than conventional computing hardware, could consume hundreds of times less energy. When an artificial intelligence model produces an image from text, it typically uses a process called diffusion. The AI is first shown a large collection of images and shown how to destroy them using statistical noise, then it encodes these patterns in a set of rules. When it is given a new, noisy image, it can use these rules to do the same thing in reverse: over many steps, it works towards a coherent image that matches a given text request. For realistic, high-resolution images, diffusion uses many sequential steps that require a significant level of computing power.


Rumors spread like viruses. The French Revolution proved it.

Popular Science

Breakthroughs, discoveries, and DIY tips sent every weekday. It's hard to contain misinformation once enough people believe it. A conspiracy theory spreads exponentially regardless of its accuracy, making it that much more likely to translate into real violence. According to a study published August 27 in the journal Nature, these situations can (and should) be geographically mapped with the same models that epidemiologists use to track diseases. And as an example, researchers turned to one of history's most famous moments of misinformation. The Great Fear of 1789 was a major chapter in the French Revolution and a defining moment in modern history.


AI drone finds missing hiker's remains in mountains after 10 months

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A missing hiker's dead body was finally found in July in Italy's rugged Piedmont region after 10 months. The recovery team credited the breakthrough to an AI-powered drone that spotted a critical clue within hours. The same process would have taken weeks or even months if done by the human eye.


3 Things James O'Donnell is into right now

MIT Technology Review

This is a podcast in which two very smart people (who happen to be young and hilarious professors of philosophy) draw unexpected philosophical connections between facets of modern life. Ellie Anderson and David Peรฑa-Guzmรกn have done hour-long episodes on everything from mommy issues to animal justice, with particularly sharp segments on tech-adjacent issues like biohacking and the relationship between AI and art. Whenever I think society is dealing with a brand-new problem, these two unearth someone who was pondering it centuries ago. It's a treat to listen to. Over the summer I was eager to watch Mountainhead, a darkly funny film by Jesse Armstrong, the creator of Succession, that follows four unlikable tech founders as they watch the world collapse under political turmoil and violence caused by AI deepfakes.


ChatGPT has its uses, but I still hate it โ€“ and I'll tell you why Imogen West-Knights

The Guardian

It's one of those topics that comes up over drinks or dinner at the moment: whether or not you think AI is going to steal your job. So far, I've felt relatively confident that while AI could no doubt have a fair crack at writing a newspaper opinion column, there is something I do as part of my work that AI cannot: reporting. Except now, it seems, AI is claiming to be doing that as well. Last week, it was revealed that at least six reputable publications have had to take down published articles because it turned out that they were probably pieces of fiction written by AI and then passed off by somebody as works of journalism under the name of Margaux Blanchard. One of these was a piece for Wired titled They Fell in Love Playing Minecraft.


Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models

arXiv.org Artificial Intelligence

The proliferation of misinformation in digital platforms reveals the limitations of traditional detection methods, which mostly rely on static classification and fail to capture the intricate process of real-world fact-checking. Despite advancements in Large Language Models (LLMs) that enhance automated reasoning, their application to misinformation detection remains hindered by issues of logical inconsistency and superficial verification. In response, we introduce Debate-to-Detect (D2D), a novel Multi-Agent Debate (MAD) framework that reformulates misinformation detection as a structured adversarial debate. Inspired by fact-checking workflows, D2D assigns domain-specific profiles to each agent and orchestrates a five-stage debate process, including Opening Statement, Rebuttal, Free Debate, Closing Statement, and Judgment. To transcend traditional binary classification, D2D introduces a multi-dimensional evaluation mechanism that assesses each claim across five distinct dimensions: Factuality, Source Reliability, Reasoning Quality, Clarity, and Ethics. Experiments with GPT-4o on two datasets demonstrate significant improvements over baseline methods, and the case study highlight D2D's capability to iteratively refine evidence while improving decision transparency, representing a substantial advancement towards interpretable misinformation detection. The code will be released publicly after the official publication.


Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games

arXiv.org Artificial Intelligence

Contemporary artificial intelligence (AI) development largely centers on rational decision-making, valued for its measurability and suitability for objective evaluation. Y et in real-world contexts, an intelligent agent's decisions are shaped not only by logic but also by deeper influences such as beliefs, values, and preferences. The diversity of human decision-making styles emerges from these differences, highlighting that "style" is an essential but often overlooked dimension of intelligence. This dissertation introduces playstyle as an alternative lens for observing and analyzing the decision-making behavior of intelligent agents, and examines its foundational meaning and historical context from a philosophical perspective. By analyzing how beliefs and values drive intentions and actions, we construct a two-tier framework for style formation: the external interaction loop with the environment and the internal cognitive loop of deliberation. On this basis, we formalize style-related characteristics and propose measurable indicators such as style capacity, style popularity, and evolutionary dynamics. The study focuses on three core research directions: (1) Defining and measuring playstyle, proposing a general playstyle metric based on discretized state spaces, and extending it to quantify strategic diversity and competitive balance; (2) Expressing and generating playstyle, exploring how reinforcement learning and imitation learning can be used to train agents exhibiting specific stylistic tendencies, and introducing a novel approach for human-like style learning and modeling; and (3) Practical applications, analyzing the potential of these techniques in domains such as game design and interactive entertainment. Finally, the dissertation outlines future extensions, including the role of style as a core element in building artificial general intelligence (AGI). By investigating stylistic variation, we aim to rethink autonomy, value expression, and even offer a tangible perspective on the ultimate i philosophical question: What is the soul?


Hybrid Deep Searcher: Integrating Parallel and Sequential Search Reasoning

arXiv.org Artificial Intelligence

Large reasoning models (LRMs) have demonstrated strong performance in complex, multi-step reasoning tasks. Existing methods enhance LRMs by sequentially integrating external knowledge retrieval; models iteratively generate queries, retrieve external information, and progressively reason over this information. However, purely sequential querying increases inference latency and context length, diminishing coherence and potentially reducing accuracy. To address these limitations, we introduce HDS-QA (Hybrid Deep Search QA), a synthetic dataset automatically generated from Natural Questions, explicitly designed to train LRMs to distinguish parallelizable from sequential queries. HDS-QA comprises hybrid-hop questions that combine parallelizable independent subqueries (executable simultaneously) and sequentially dependent subqueries (requiring step-by-step resolution), along with synthetic reasoning-querying-retrieval paths involving parallel queries. We fine-tune an LRM using HDS-QA, naming the model HybridDeepSearcher, which outperforms state-of-the-art baselines across multiple benchmarks, notably achieving +15.9 and +11.5 F1 on FanOutQA and a subset of BrowseComp, respectively, both requiring comprehensive and exhaustive search. Experimental results highlight two key advantages: HybridDeepSearcher reaches comparable accuracy with fewer search turns, significantly reducing inference latency, and it effectively scales as more turns are permitted. These results demonstrate the efficiency, scalability, and effectiveness of explicitly training LRMs to leverage hybrid parallel and sequential querying.