Generative AI
From G-Factor to A-Factor: Establishing a Psychometric Framework for AI Literacy
Li, Ning, Deng, Wenming, Chen, Jiatan
This research addresses the growing need to measure and understand AI literacy in the context of generative AI technologies. Through three sequential studies involving a total of 517 participants, we establish AI literacy as a coherent, measurable construct with significant implications for education, workforce development, and social equity. Study 1 (N=85) revealed a dominant latent factor - termed the "A-factor" - that accounts for 44.16% of variance across diverse AI interaction tasks. Study 2 (N=286) refined the measurement tool by examining four key dimensions of AI literacy: communication effectiveness, creative idea generation, content evaluation, and step-by-step collaboration, resulting in an 18-item assessment battery. Study 3 (N=146) validated this instrument in a controlled laboratory setting, demonstrating its predictive validity for real-world task performance. Results indicate that AI literacy significantly predicts performance on complex, language-based creative tasks but shows domain specificity in its predictive power. Additionally, regression analyses identified several significant predictors of AI literacy, including cognitive abilities (IQ), educational background, prior AI experience, and training history. The multidimensional nature of AI literacy and its distinct factor structure provide evidence that effective human-AI collaboration requires a combination of general and specialized abilities. These findings contribute to theoretical frameworks of human-AI collaboration while offering practical guidance for developing targeted educational interventions to promote equitable access to the benefits of generative AI technologies.
Debiasing Diffusion Model: Enhancing Fairness through Latent Representation Learning in Stable Diffusion Model
Huang, Lin-Chun, Tsao, Ching Chieh, Su, Fang-Yi, Chiang, Jung-Hsien
Image generative models, particularly diffusion-based models, have surged in popularity due to their remarkable ability to synthesize highly realistic images. However, since these models are data-driven, they inherit biases from the training datasets, frequently leading to disproportionate group representations that exacerbate societal inequities. Traditionally, efforts to debiase these models have relied on predefined sensitive attributes, classifiers trained on such attributes, or large language models to steer outputs toward fairness. However, these approaches face notable drawbacks: predefined attributes do not adequately capture complex and continuous variations among groups. To address these issues, we introduce the Debiasing Diffusion Model (DDM), which leverages an indicator to learn latent representations during training, promoting fairness through balanced representations without requiring predefined sensitive attributes. This approach not only demonstrates its effectiveness in scenarios previously addressed by conventional techniques but also enhances fairness without relying on predefined sensitive attributes as conditions. In this paper, we discuss the limitations of prior bias mitigation techniques in diffusion-based models, elaborate on the architecture of the DDM, and validate the effectiveness of our approach through experiments.
Universal Narrative Model: an Author-centric Storytelling Framework for Generative AI
In their survey of authoring tools for computational narrative, Kybartas and Bidarra note that "we believe that creating a standard model of computational narrative could allow different systems to interact with the same narrative, without being restricted by incompatible models and definitions. Furthermore, such a model would also facilitate research into the generation of specific story components, e.g., allowing for multiple generators and even authors to collaborate on a given narrative" [Kybartas and Bidarra [2017]]. This paper proposes such a standard: the Universal Narrative Model (UNM). We foresee that generative AI will enable a new paradigm of storytelling technologies and processes: from assisting a writer of linear media (novels, film, television, etc.) by allowing them to test out scenes and characters before committing them to a script, all the way through to real-time storytelling systems in videogames which respond to a player's agency, and countless use cases in between [Peng et al. [2024]]. The UNM is designed to service any use case in which coherent narrative structure is a consideration, and in which authorial intent and direction is privileged. In the last five years, a robust body of research has demonstrated a wide variety of potential uses for computational narrative systems powered by generative AI, and some limited commercial deployments already exist [Yang et al. [2024], Hu et al. [2024]]. With such promise, however, comes a series of challenges: technical, narrative, and ethical. The goal of the Entertainment Technology Center's "Universal Narrative Model" project was to produce the UNM as an open standard. The ultimate directive of the project was to privilege, above all else, author-centric design and functionality, setting the stage for generative workflows which extend an author's narrative intent and creativity, rather than eclipse or replace it.
Localized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptation
Lee, Byung Hyun, Lim, Sungjin, Chun, Se Young
Fine-tuning based concept erasing has demonstrated promising results in preventing generation of harmful contents from text-to-image diffusion models by removing target concepts while preserving remaining concepts. To maintain the generation capability of diffusion models after concept erasure, it is necessary to remove only the image region containing the target concept when it locally appears in an image, leaving other regions intact. However, prior arts often compromise fidelity of the other image regions in order to erase the localized target concept appearing in a specific area, thereby reducing the overall performance of image generation. To address these limitations, we first introduce a framework called localized concept erasure, which allows for the deletion of only the specific area containing the target concept in the image while preserving the other regions. As a solution for the localized concept erasure, we propose a training-free approach, dubbed Gated Low-rank adaptation for Concept Erasure (GLoCE), that injects a lightweight module into the diffusion model. GLoCE consists of low-rank matrices and a simple gate, determined only by several generation steps for concepts without training. By directly applying GLoCE to image embeddings and designing the gate to activate only for target concepts, GLoCE can selectively remove only the region of the target concepts, even when target and remaining concepts coexist within an image. Extensive experiments demonstrated GLoCE not only improves the image fidelity to text prompts after erasing the localized target concepts, but also outperforms prior arts in efficacy, specificity, and robustness by large margin and can be extended to mass concept erasure.
Evaluating Large Language Models on the Spanish Medical Intern Resident (MIR) Examination 2024/2025:A Comparative Analysis of Clinical Reasoning and Knowledge Application
Vera, Carlos Luengo, Picon, Ignacio Ferro, Nunez, M. Teresa del Val, Gandia, Jose Andres Gomez, Ancillo, Antonio de Lucas, Arroyo, Victor Ramos, Figueredo, Carlos Milan
The MIR serves as a critical selection mechanism for medical graduates entering specialized training in Spain. A study is to be conducted on the ability of generative AI models to meet the challenges presented by MIR, with emphasis on clinical reasoning, image interpretation and epidemiological calculations. This research evaluates LLM performance in complex clinical scenarios and explores the extent to which LLMs demonstrate medical reasoning beyond mere information recall. Findings The results reveal key insights into the performance of 22 LLMs on the MIR 2024 and 2025 exams. The exam features 210 multiple-choice questions covering diverse medical domains and incorporates case-based scenarios, image interpretation (25 questions), and laboratory data analysis.
Fox News AI Newsletter: 'Digital twin' danger
A woman in Washington, D.C., views a manipulated video on January 24, 2019, that changes what is said by President Donald Trump and former President Barack Obama. This illustration photo taken on January 30, 2023 shows a phone screen displaying a statement from the head of security policy at META with a fake video of Ukrainian President Volodymyr Zelensky calling on his soldiers to lay down their weapons shown in the background, in Washington, D.C. (OLIVIER DOULIERY/AFP via Getty Images) NEW REALITY: Artificial intelligence (AI) is producing hyperrealistic "digital twins" of politicians, celebrities, pornographic material, and more – leaving victims of deepfake technology struggling to determine legal recourse. NO BOUNDARY: Scarlett Johansson has taken a vocal stand on artificial intelligence, after having her likeness and voice used without permission. Last year, Johansson said she had been asked to voice OpenAI's Chatbot by CEO Sam Altman, but turned down the job, only for people to notice that the feature, named "Sky," sounded almost exactly like the actress. It was like: If that can happen to me, how are we going to protect ourselves from this? There's no boundary here; we're setting ourselves up to be taken advantage of," the 40-year-old told InStyle Magazine earlier this month.
AI pioneer wants Europe to forge its own nimbler way forward
One belief underlying the power-hungry approach to machine learning advanced by OpenAI and Mistral AI is that an artificial intelligence model must review its entire dataset before spitting out new insights. Sepp Hochreiter, an early pioneer of the technology who runs an AI lab at Johannes Kepler University in Linz, Austria, has a different view, one that requires far less cash and computing power. He's interested in teaching AI models how to efficiently forget. Hochreiter holds a special place in the world of artificial intelligence, having scaled the technology's highest peaks long before most computer scientists. As a university student in Munich during the 1990s, he came up with the conceptual framework that underpinned the first generation of nimble AI models used by Alphabet, Apple and Amazon.
Unlocking Learning Potentials: The Transformative Effect of Generative AI in Education Across Grade Levels
The advent of generative artificial intelligence (GAI) has brought about a notable surge in the field of education. The use of GAI to support learning is becoming increasingly prevalent among students. However, the manner and extent of its utilisation vary considerably from one individual to another. And researches about student's utilisation and perceptions of GAI remains relatively scarce. To gain insight into the issue, this paper proposed a hybrid-survey method to examine the impact of GAI on students across four different grades in six key areas (LIPSAL): learning interest, independent learning, problem solving, self-confidence, appropriate use, and learning enjoyment. Firstly, through questionnaire, we found that among LIPSAL, GAI has the greatest impact on the concept of appropriate use, the lowest level of learning interest and self-confidence. Secondly, a comparison of four grades revealed that the high and low factors of LIPSAL exhibited grade-related variation, and college students exhibited a higher level than high school students across LIPSAL. Thirdly, through interview, the students demonstrated a comprehensive understanding of the application of GAI. We found that students have a positive attitude towards GAI and are very willing to use it, which is why GAI has grown so rapidly in popularity. They also told us prospects and challenges in using GAI. In the future, as GAI matures technologically, it will have an greater impact on students. These findings may help better understand usage by different students and inform future research in digital education.
Context-aware Multimodal AI Reveals Hidden Pathways in Five Centuries of Art Evolution
Kim, Jin, Lee, Byunghwee, You, Taekho, Yun, Jinhyuk
The rise of multimodal generative AI is transforming the intersection of technology and art, offering deeper insights into large-scale artwork. Although its creative capabilities have been widely explored, its potential to represent artwork in latent spaces remains underexamined. We use cutting-edge generative AI, specifically Stable Diffusion, to analyze 500 years of Western paintings by extracting two types of latent information with the model: formal aspects (e.g., colors) and contextual aspects (e.g., subject). Our findings reveal that contextual information differentiates between artistic periods, styles, and individual artists more successfully than formal elements. Additionally, using contextual keywords extracted from paintings, we show how artistic expression evolves alongside societal changes. Our generative experiment, infusing prospective contexts into historical artworks, successfully reproduces the evolutionary trajectory of artworks, highlighting the significance of mutual interaction between society and art. This study demonstrates how multimodal AI expands traditional formal analysis by integrating temporal, cultural, and historical contexts.
Probabilistic Graph Circuits: Deep Generative Models for Tractable Probabilistic Inference over Graphs
Papež, Milan, Rektoris, Martin, Šmídl, Václav, Pevný, Tomáš
Deep generative models (DGMs) have recently demonstrated remarkable success in capturing complex probability distributions over graphs. Although their excellent performance is attributed to powerful and scalable deep neural networks, it is, at the same time, exactly the presence of these highly non-linear transformations that makes DGMs intractable. Indeed, despite representing probability distributions, intractable DGMs deny probabilistic foundations by their inability to answer even the most basic inference queries without approximations or design choices specific to a very narrow range of queries. To address this limitation, we propose probabilistic graph circuits (PGCs), a framework of tractable DGMs that provide exact and efficient probabilistic inference over (arbitrary parts of) graphs. Nonetheless, achieving both exactness and efficiency is challenging in the permutation-invariant setting of graphs. We design PGCs that are inherently invariant and satisfy these two requirements, yet at the cost of low expressive power. Therefore, we investigate two alternative strategies to achieve the invariance: the first sacrifices the efficiency, and the second sacrifices the exactness. We demonstrate that ignoring the permutation invariance can have severe consequences in anomaly detection, and that the latter approach is competitive with, and sometimes better than, existing intractable DGMs in the context of molecular graph generation.