Generative AI
How Generative-AI can be Effectively used in Government Chatbots
With the rapid development of artificial intelligence and breakthroughs in machine learning and natural language processing, intelligent question-answering robots have become widely used in government affairs. This paper conducts a horizontal comparison between Guangdong Province's government chatbots, ChatGPT, and Wenxin Ernie, two large language models, to analyze the strengths and weaknesses of existing government chatbots and AIGC technology. The study finds significant differences between government chatbots and large language models. China's government chatbots are still in an exploratory stage and have a gap to close to achieve "intelligence." To explore the future direction of government chatbots more deeply, this research proposes targeted optimization paths to help generative AI be effectively applied in government chatbot conversations.
LLVMs4Protest: Harnessing the Power of Large Language and Vision Models for Deciphering Protests in the News
Large language and vision models have transformed how social movements scholars identify protest and extract key protest attributes from multi-modal data such as texts, images, and videos. This article documents how we fine-tuned two large pretrained transformer models, including longformer and swin-transformer v2, to infer potential protests in news articles using textual and imagery data. First, the longformer model was fine-tuned using the Dynamic of Collective Action (DoCA) Corpus. We matched the New York Times articles with the DoCA database to obtain a training dataset for downstream tasks. Second, the swin-transformer v2 models was trained on UCLA-protest imagery data. UCLA-protest project contains labeled imagery data with information such as protest, violence, and sign. Both fine-tuned models will be available via \url{https://github.com/Joshzyj/llvms4protest}. We release this short technical report for social movement scholars who are interested in using LLVMs to infer protests in textual and imagery data.
Algorithmic Persuasion Through Simulation: Information Design in the Age of Generative AI
Harris, Keegan, Immorlica, Nicole, Lucier, Brendan, Slivkins, Aleksandrs
How can an informed sender persuade a receiver, having only limited information about the receiver's beliefs? Motivated by research showing generative AI can simulate economic agents, we initiate the study of information design with an oracle. We assume the sender can learn more about the receiver by querying this oracle, e.g., by simulating the receiver's behavior. Aside from AI motivations such as general-purpose Large Language Models (LLMs) and problem-specific machine learning models, alternate motivations include customer surveys and querying a small pool of live users. Specifically, we study Bayesian Persuasion where the sender has a second-order prior over the receiver's beliefs. After a fixed number of queries to an oracle to refine this prior, the sender commits to an information structure. Upon receiving the message, the receiver takes a payoff-relevant action maximizing her expected utility given her posterior beliefs. We design polynomial-time querying algorithms that optimize the sender's expected utility in this Bayesian Persuasion game. As a technical contribution, we show that queries form partitions of the space of receiver beliefs that can be used to quantify the sender's knowledge.
ROSO: Improving Robotic Policy Inference via Synthetic Observations
Miyashita, Yusuke, Gahtidis, Dimitris, La, Colin, Rabinowicz, Jeremy, Leitner, Jurgen
In this paper, we propose the use of generative artificial intelligence (AI) to improve zero-shot performance of a pre-trained policy by altering observations during inference. Modern robotic systems, powered by advanced neural networks, have demonstrated remarkable capabilities on pre-trained tasks. However, generalizing and adapting to new objects and environments is challenging, and fine-tuning visuomotor policies is time-consuming. To overcome these issues we propose Robotic Policy Inference via Synthetic Observations (ROSO). ROSO uses stable diffusion to pre-process a robot's observation of novel objects during inference time to fit within its distribution of observations of the pre-trained policies. This novel paradigm allows us to transfer learned knowledge from known tasks to previously unseen scenarios, enhancing the robot's adaptability without requiring lengthy fine-tuning. Our experiments show that incorporating generative AI into robotic inference significantly improves successful outcomes, finishing up to 57% of tasks otherwise unsuccessful with the pre-trained policy.
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
Zhang, Kai, Mo, Lingbo, Chen, Wenhu, Sun, Huan, Su, Yu
Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise. Thus, they still require lots of manual tuning to produce desirable outcomes in practice. To address this issue, we introduce MagicBrush (https://osu-nlp-group.github.io/MagicBrush/), the first large-scale, manually annotated dataset for instruction-guided real image editing that covers diverse scenarios: single-turn, multi-turn, mask-provided, and mask-free editing. MagicBrush comprises over 10K manually annotated triplets (source image, instruction, target image), which supports trainining large-scale text-guided image editing models. We fine-tune InstructPix2Pix on MagicBrush and show that the new model can produce much better images according to human evaluation. We further conduct extensive experiments to evaluate current image editing baselines from multiple dimensions including quantitative, qualitative, and human evaluations. The results reveal the challenging nature of our dataset and the gap between current baselines and real-world editing needs.
Why Won't OpenAI Say What the Q* Algorithm Is?
Last week, it seemed that OpenAI--the secretive firm behind ChatGPT--had been broken open. The company's board had suddenly fired CEO Sam Altman, hundreds of employees revolted in protest, Altman was reinstated, and the media dissected the story from every possible angle. Yet the reporting belied the fact that our view into the most crucial part of the company is still so fundamentally limited: We don't really know how OpenAI develops its technology, nor do we understand exactly how Altman has directed work on future, more powerful generations. This was made acutely apparent last Wednesday, when Reuters and The Information reported that, prior to Altman's firing, several staff researchers had raised concerns about a supposedly dangerous breakthrough. At issue was an algorithm called Q* (pronounced "Q-star"), which has allegedly been shown to solve certain grade-school-level math problems that it hasn't seen before.
Prominent Women in Tech Say They Don't Want to Join OpenAI's All-Male Board
Earlier this month, OpenAI's board abruptly fired its popular CEO, Sam Altman. The ouster shocked the tech world and rankled Altman's loyal employees, the vast majority of whom threatened to quit unless their boss was reinstated. After a chaotic five-day exile, Altman got his old job back--with a reconfigured, all-male board overseeing him, led by ex-Salesforce CEO and former Twitter board chair Bret Taylor. Right now, only three people sit on this provisional OpenAI board. Immediately prior to the failed coup, there were six.
Why Europe Must Not Let AI Firms Put Profits Before People
The soap opera-like ousting and swift return of OpenAI CEO Sam Altman produced plenty of fodder for ironic quips online but it also exposed some serious fault lines. One such critique I enjoyed was: "How are we supposed to solve the AI alignment problem if aligning just a few board members presents an insurmountable challenge?" As the company behind ChatGPT, OpenAI may be one of the more recognizable names, but artificial intelligence is more than one company. It's a technology of immense consequence, yet it remains almost entirely unregulated. The E.U. has a chance to meaningfully tackle that challenge--but not if it bends the knee to Big Tech's ongoing onslaught. Inspirational Members of the European Parliament have so far been standing firm in the face of incredible pressure, in an effort to save this landmark legislation.
The frantic battle over OpenAI shows that money triumphs in the end Robert Reich
How do we gain access to artificial intelligence's huge potential benefits – such as devising new life-saving drugs or finding new ways to teach children – without opening a box of horrors? If we're not careful, AI could be a Frankenstein monster. It might eliminate nearly all jobs. It could lead to autonomous warfare. Even such a mundane goal as making as many paper clips as possible, critics of AI argue, could push an all-powerful AI to end all life on Earth in pursuit of more clips.
Identifying and Mitigating Vulnerabilities in LLM-Integrated Applications
Jiang, Fengqing, Xu, Zhangchen, Niu, Luyao, Wang, Boxin, Jia, Jinyuan, Li, Bo, Poovendran, Radha
Large language models (LLMs) are increasingly deployed as the service backend for LLM-integrated applications such as code completion and AI-powered search. LLM-integrated applications serve as middleware to refine users' queries with domain-specific knowledge to better inform LLMs and enhance the responses. Despite numerous opportunities and benefits, LLM-integrated applications also introduce new attack surfaces. Understanding, minimizing, and eliminating these emerging attack surfaces is a new area of research. In this work, we consider a setup where the user and LLM interact via an LLM-integrated application in the middle. We focus on the communication rounds that begin with user's queries and end with LLM-integrated application returning responses to the queries, powered by LLMs at the service backend. For this query-response protocol, we identify potential vulnerabilities that can originate from the malicious application developer or from an outsider threat initiator that is able to control the database access, manipulate and poison data that are high-risk for the user. Successful exploits of the identified vulnerabilities result in the users receiving responses tailored to the intent of a threat initiator. We assess such threats against LLM-integrated applications empowered by OpenAI GPT-3.5 and GPT-4. Our empirical results show that the threats can effectively bypass the restrictions and moderation policies of OpenAI, resulting in users receiving responses that contain bias, toxic content, privacy risk, and disinformation. To mitigate those threats, we identify and define four key properties, namely integrity, source identification, attack detectability, and utility preservation, that need to be satisfied by a safe LLM-integrated application. Based on these properties, we develop a lightweight, threat-agnostic defense that mitigates both insider and outsider threats.