Generative AI
Open-sourcing generative AI
Alison Smith is a Director of Generative AI at Booz Allen Hamilton where she helps clients address their missions with innovative solutions. Leading Booz Allen's investments in Generative AI and grounding them in real business needs, Alison employs a pragmatic approach to designing, implementing, and deploying Generative AI that blends existing tools with additional customization. She is also responsible for disseminating best practices and key solutions throughout the firm to ensure that all teams are up-to-date on the latest available tools, solutions, and approaches to common client problems. In addition to her role at Booz Allen which balances technical solutions and business growth, Alison also enjoys staying connected to and serving her local community. From 2017-2021, Alison served on the board of a non-profit, DC Open Government Coalition (DCOGC), a group that seeks to enhance public access to government information and ensure transparent government operations; in November 2021, Alison was recognized as a Power Woman in Code by DCFemTech.
Students Are Likely Writing Millions of Papers With AI
Students have submitted more than 22 million papers that may have used generative AI in the past year, new data released by plagiarism detection company Turnitin shows. A year ago, Turnitin rolled out an AI writing detection tool that was trained on its trove of papers written by students as well as other AI-generated texts. Since then, more than 200 million papers have been reviewed by the detector, predominantly written by high school and college students. Turnitin found that 11 percent may contain AI-written language in 20 percent of its content, with 3 percent of the total papers reviewed getting flagged for having 80 percent or more AI writing. Turnitin says its detector has a false positive rate of less than 1 percent when analyzing full documents.
OpenAI prepares to fight for its life as legal troubles mount
OpenAI is also at the center of several regulatory investigations, which have forced the company to spend even more on legal support. The Securities and Exchange Commission is looking into whether investors were misled during the chaotic period when Altman briefly left the company. The Federal Trade Commission is probing whether it ran afoul of consumer protection laws in a number of areas, including a data leak and ChatGPT's inaccurate claims. And the commission has had talks with the Justice Department about which agency should probe its multibillion-dollar partnership with Microsoft, amid concerns that such deals are dampening competition in the quickly evolving AI market.
"Sora is Incredible and Scary": Emerging Governance Challenges of Text-to-Video Generative AI Models
Zhou, Kyrie Zhixuan, Choudhry, Abhinav, Gumusel, Ece, Sanfilippo, Madelyn Rose
Text-to-video generative AI models such as Sora OpenAI have the potential to disrupt multiple industries. In this paper, we report a qualitative social media analysis aiming to uncover people's perceived impact of and concerns about Sora's integration. We collected and analyzed comments (N=292) under popular posts about Sora-generated videos, comparison between Sora videos and Midjourney images, and artists' complaints about copyright infringement by Generative AI. We found that people were most concerned about Sora's impact on content creation-related industries. Emerging governance challenges included the for-profit nature of OpenAI, the blurred boundaries between real and fake content, human autonomy, data privacy, copyright issues, and environmental impact. Potential regulatory solutions proposed by people included law-enforced labeling of AI content and AI literacy education for the public. Based on the findings, we discuss the importance of gauging people's tech perceptions early and propose policy recommendations to regulate Sora before its public release.
Toward Cross-Layer Energy Optimizations in Machine Learning Systems
Chung, Jae-Won, Chowdhury, Mosharaf
The enormous energy consumption of machine learning (ML) and generative AI workloads shows no sign of waning, taking a toll on operating costs, power delivery, and environmental sustainability. Despite a long line of research on energy-efficient hardware, we found that software plays a critical role in ML energy optimization through two recent works: Zeus and Perseus. This is especially true for large language models (LLMs) because their model sizes and, therefore, energy demands are growing faster than hardware efficiency improvements. Therefore, we advocate for a cross-layer approach for energy optimizations in ML systems, where hardware provides architectural support that pushes energy-efficient software further, while software leverages and abstracts the hardware to develop techniques that bring hardware-agnostic energy-efficiency gains.
StockGPT: A GenAI Model for Stock Prediction and Trading
Generative artificial intelligence (GenAI)--a set of advanced technologies capable of generating texts, images, videos, programming codes, or arts from instructions via sounds or texts--has taken the society by storm and exerted wide-range influences on many aspects of the world economy (Baldassarre et al. 2023; Mannuru et al. 2023; Sรฆtra 2023). Although it had been around for years, GenAI came to public prominence since the introduction of ChatGPT in November 2022, a chatbox able to generate answers, reasoning, and conversations at human level. Since its introduction, ChatGPT and similar large language models have quickly made their ways into the investment industry. One common use of ChatGPT for investment is to give trading recommendations directly from news about a company (such as news articles or corporate communications) (Lopez-Lira and Tang 2023). A less direct approach is to rely on similar pretrained language models such as BERT (Devlin et al. 2018) and OPT (Zhang et al. 2022) to generate a sentiment score for each company which is then used to make trading decisions.
'Inceptionism' and Balenciaga popes: a brief history of deepfakes
Concern about doctored or manipulative media is always high around election cycles, but 2024 will be different for two reasons: deepfakes made by artificial intelligence (AI) and the sheer number of polls. The term deepfake refers to a hoax that uses AI to create a phoney image, most commonly fake videos of people, with the effect often compounded by a voice component. Combined with the fact that around half the world's population is holding important elections this year โ including India, the US, the EU and, most probably, the UK โ and there is potential for the technology to be highly disruptive. Here is a guide to some of the most effective deepfakes in recent years, including the first attempts to create hoax images. The banana where it all began.
'Time is running out': can a future of undetectable deepfakes be avoided?
With more than 4,000 shares, 20,000 comments, and 100,000 reactions on Facebook, the photo of the elderly woman, sitting behind her homemade 122nd birthday cake, has unquestionably gone viral. "I started decorating cakes from five years old," the caption reads, "and I can't wait to grow my baking journey." The picture is also unquestionably fake. If the curious candles โ one seems to float in the air, attached to nothing โ or the weird amorphous blobs on the cake in the foreground didn't give it away, then the fact the celebrant would be the oldest person in the world by almost five years should. Thankfully, the stakes for viral supercentenarian cake decorators are low.
How tech giants cut corners to harvest data for AI
The artificial intelligence lab had exhausted every reservoir of reputable English-language text on the internet as it developed its latest AI system. It needed more data to train the next version of its technology -- lots more. So OpenAI researchers created a speech recognition tool called Whisper. It could transcribe the audio from YouTube videos, yielding new conversational text that would make an AI system smarter. Some OpenAI employees discussed how such a move might go against YouTube's rules, three people with knowledge of the conversations said. YouTube, which is owned by Google, prohibits use of its videos for applications that are "independent" of the video platform.
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Ghosh, Shaona, Varshney, Prasoon, Galinkin, Erick, Parisien, Christopher
As Large Language Models (LLMs) and generative AI become more widespread, the content safety risks associated with their use also increase. We find a notable deficiency in high-quality content safety datasets and benchmarks that comprehensively cover a wide range of critical safety areas. To address this, we define a broad content safety risk taxonomy, comprising 13 critical risk and 9 sparse risk categories. Additionally, we curate AEGISSAFETYDATASET, a new dataset of approximately 26, 000 human-LLM interaction instances, complete with human annotations adhering to the taxonomy. We plan to release this dataset to the community to further research and to help benchmark LLM models for safety. To demonstrate the effectiveness of the dataset, we instruction-tune multiple LLM-based safety models. We show that our models (named AEGISSAFETYEXPERTS), not only surpass or perform competitively with the state-of-the-art LLM-based safety models and general purpose LLMs, but also exhibit robustness across multiple jail-break attack categories. We also show how using AEGISSAFETYDATASET during the LLM alignment phase does not negatively impact the performance of the aligned models on MT Bench scores. Furthermore, we propose AEGIS, a novel application of a no-regret online adaptation framework with strong theoretical guarantees, to perform content moderation with an ensemble of LLM content safety experts in deployment