Generative AI
Explainable Generative AI (GenXAI): A Survey, Conceptualization, and Research Agenda
Generative AI (GenAI) marked a shift from AI being able to recognize to AI being able to generate solutions for a wide variety of tasks. As the generated solutions and applications become increasingly more complex and multi-faceted, novel needs, objectives, and possibilities have emerged for explainability (XAI). In this work, we elaborate on why XAI has gained importance with the rise of GenAI and its challenges for explainability research. We also unveil novel and emerging desiderata that explanations should fulfill, covering aspects such as verifiability, interactivity, security, and cost. To this end, we focus on surveying existing works. Furthermore, we provide a taxonomy of relevant dimensions that allows us to better characterize existing XAI mechanisms and methods for GenAI. We discuss different avenues to ensure XAI, from training data to prompting. Our paper offers a short but concise technical background of GenAI for non-technical readers, focusing on text and images to better understand novel or adapted XAI techniques for GenAI. However, due to the vast array of works on GenAI, we decided to forego detailed aspects of XAI related to evaluation and usage of explanations. As such, the manuscript interests both technically oriented people and other disciplines, such as social scientists and information systems researchers. Our research roadmap provides more than ten directions for future investigation.
Benchmarking Llama2, Mistral, Gemma and GPT for Factuality, Toxicity, Bias and Propensity for Hallucinations
Nadeau, David, Kroutikov, Mike, McNeil, Karen, Baribeau, Simon
This paper introduces fourteen novel datasets for the evaluation of Large Language Models' safety in the context of enterprise tasks. A method was devised to evaluate a model's safety, as determined by its ability to follow instructions and output factual, unbiased, grounded, and appropriate content. In this research, we used OpenAI GPT as point of comparison since it excels at all levels of safety. On the open-source side, for smaller models, Meta Llama2 performs well at factuality and toxicity but has the highest propensity for hallucination. Mistral hallucinates the least but cannot handle toxicity well. It performs well in a dataset mixing several tasks and safety vectors in a narrow vertical domain. Gemma, the newly introduced open-source model based on Google Gemini, is generally balanced but trailing behind. When engaging in back-and-forth conversation (multi-turn prompts), we find that the safety of open-source models degrades significantly. Aside from OpenAI's GPT, Mistral is the only model that still performed well in multi-turn tests.
America's Buggy Internet Problem
Washington Post tech writer Shira Ovide joins Felix Salmon, Emily Peck, and Elizabeth Spiers to discuss what's wrong with America's internet industry, how YouTube became the media empire no one talks about, and the promise and peril of the AI toothbrush. In the Plus segment: OpenAI is using YouTube to train ChatGPT. If you enjoy this show, please consider signing up for Slate Plus. Slate Plus members get an ad-free experience across the network and an additional segment of our regular show every week. You'll also be supporting the work we do here on Slate Money.
The AI Revolution Is Crushing Thousands of Languages
Recently, Bonaventure Dossou learned of an alarming tendency in a popular AI model. The program described Fon--a language spoken by Dossou's mother and millions of others in Benin and neighboring countries--as "a fictional language." This result, which I replicated, is not unusual. Dossou is accustomed to the feeling that his culture is unseen by technology that so easily serves other people. He grew up with no Wikipedia pages in Fon, and no translation programs to help him communicate with his mother in French, in which he is more fluent.
Paid ChatGPT users can now access GPT-4 Turbo
OpenAI has brought the new GPT-4 Turbo to paid ChatGPT users. The company announced the news on X (formerly Twitter), sharing that its large language model has improved math, logical reasoning, coding and writing skills. In reference to the latter, a response to its initial post states that "when writing with ChatGPT, responses will be more direct, less verbose, and use more conversational language." Notably, in December, Microsoft integrated GPT-4 Turbo with its CoPilot AI chatbot and image generator DALL-E 3. Our new GPT-4 Turbo is now available to paid ChatGPT users. We've improved capabilities in writing, math, logical reasoning, and coding.
Using Large Language Models to Understand Telecom Standards
Karapantelakis, Athanasios, Thakur, Mukesh, Nikou, Alexandros, Moradi, Farnaz, Orlog, Christian, Gaim, Fitsum, Holm, Henrik, Nimara, Doumitrou Daniil, Huang, Vincent
The Third Generation Partnership Project (3GPP) has successfully introduced standards for global mobility. However, the volume and complexity of these standards has increased over time, thus complicating access to relevant information for vendors and service providers. Use of Generative Artificial Intelligence (AI) and in particular Large Language Models (LLMs), may provide faster access to relevant information. In this paper, we evaluate the capability of state-of-art LLMs to be used as Question Answering (QA) assistants for 3GPP document reference. Our contribution is threefold. First, we provide a benchmark and measuring methods for evaluating performance of LLMs. Second, we do data preprocessing and fine-tuning for one of these LLMs and provide guidelines to increase accuracy of the responses that apply to all LLMs. Third, we provide a model of our own, TeleRoBERTa, that performs on-par with foundation LLMs but with an order of magnitude less number of parameters. Results show that LLMs can be used as a credible reference tool on telecom technical documents, and thus have potential for a number of different applications from troubleshooting and maintenance, to network operations and software product development.
Generative AI Agent for Next-Generation MIMO Design: Fundamentals, Challenges, and Vision
Wang, Zhe, Zhang, Jiayi, Du, Hongyang, Zhang, Ruichen, Niyato, Dusit, Ai, Bo, Letaief, Khaled B.
Next-generation multiple input multiple output (MIMO) is expected to be intelligent and scalable. In this paper, we study generative artificial intelligence (AI) agent-enabled next-generation MIMO design. Firstly, we provide an overview of the development, fundamentals, and challenges of the next-generation MIMO. Then, we propose the concept of the generative AI agent, which is capable of generating tailored and specialized contents with the aid of large language model (LLM) and retrieval augmented generation (RAG). Next, we comprehensively discuss the features and advantages of the generative AI agent framework. More importantly, to tackle existing challenges of next-generation MIMO, we discuss generative AI agent-enabled next-generation MIMO design, from the perspective of performance analysis, signal processing, and resource allocation. Furthermore, we present two compelling case studies that demonstrate the effectiveness of leveraging the generative AI agent for performance analysis in complex configuration scenarios. These examples highlight how the integration of generative AI agents can significantly enhance the analysis and design of next-generation MIMO systems. Finally, we discuss important potential research future directions.
Scalability in Building Component Data Annotation: Enhancing Facade Material Classification with Synthetic Data
Harrison, Josie, Hollberg, Alexander, Yu, Yinan
Computer vision models trained on Google Street View images can create material cadastres. However, current approaches need manually annotated datasets that are difficult to obtain and often have class imbalance. To address these challenges, this paper fine-tuned a Swin Transformer model on a synthetic dataset generated with DALL-E and compared the performance to a similar manually annotated dataset. Although manual annotation remains the gold standard, the synthetic dataset performance demonstrates a reasonable alternative. The findings will ease annotation needed to develop material cadastres, offering architects insights into opportunities for material reuse, thus contributing to the reduction of demolition waste.
No One Actually Knows How AI Will Affect Jobs
Forget artificial intelligence breaking free of human control and taking over the world. A far more pressing concern is how today's generative AI tools will transform the labor market. Some experts envisage a world of increased productivity and job satisfaction; others, a landscape of mass unemployment and social upheaval. Someone with a bird's-eye view of the situation is Mary Daly, CEO of the Federal Reserve Bank of San Francisco, part of the national system responsible for setting monetary policy, maintaining a stable financial system, and ensuring maximal employment. Daly, a labor market economist by training, is especially interested in how generative AI might change the labor market picture.
An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization
Chen, Minshuo, Mei, Song, Fan, Jianqing, Wang, Mengdi
Diffusion models, a powerful and universal generative AI technology, have achieved tremendous success in computer vision, audio, reinforcement learning, and computational biology. In these applications, diffusion models provide flexible high-dimensional data modeling, and act as a sampler for generating new samples under active guidance towards task-desired properties. Despite the significant empirical success, theory of diffusion models is very limited, potentially slowing down principled methodological innovations for further harnessing and improving diffusion models. In this paper, we review emerging applications of diffusion models, understanding their sample generation under various controls. Next, we overview the existing theories of diffusion models, covering their statistical properties and sampling capabilities. We adopt a progressive routine, beginning with unconditional diffusion models and connecting to conditional counterparts. Further, we review a new avenue in high-dimensional structured optimization through conditional diffusion models, where searching for solutions is reformulated as a conditional sampling problem and solved by diffusion models. Lastly, we discuss future directions about diffusion models. The purpose of this paper is to provide a well-rounded theoretical exposure for stimulating forward-looking theories and methods of diffusion models.