Government
Exploring the Potential Role of Generative AI in the TRAPD Procedure for Survey Translation
Metheney, Erica Ann, Yehle, Lauren
This paper explores and assesses in what ways generative AI can assist in translating survey instruments. Writing effective survey questions is a challenging and complex task, made even more difficult for surveys that will be translated and deployed in multiple linguistic and cultural settings. Translation errors can be detrimental, with known errors rendering data unusable for its intended purpose and undetected errors leading to incorrect conclusions. A growing number of institutions face this problem as surveys deployed by private and academic organizations globalize, and the success of their current efforts depends heavily on researchers' and translators' expertise and the amount of time each party has to contribute to the task. Thus, multilinguistic and multicultural surveys produced by teams with limited expertise, budgets, or time are at significant risk for translation-based errors in their data. We implement a zero-shot prompt experiment using ChatGPT to explore generative AI's ability to identify features of questions that might be difficult to translate to a linguistic audience other than the source language. We find that ChatGPT can provide meaningful feedback on translation issues, including common source survey language, inconsistent conceptualization, sensitivity and formality issues, and nonexistent concepts. In addition, we provide detailed information on the practicality of the approach, including accessing the necessary software, associated costs, and computational run times. Lastly, based on our findings, we propose avenues for future research that integrate AI into survey translation practices.
AI-generated Image Detection: Passive or Watermark?
Guo, Moyang, Hu, Yuepeng, Jiang, Zhengyuan, Li, Zeyu, Sadovnik, Amir, Daw, Arka, Gong, Neil
While text-to-image models offer numerous benefits, they also pose significant societal risks. Detecting AI-generated images is crucial for mitigating these risks. Detection methods can be broadly categorized into passive and watermark-based approaches: passive detectors rely on artifacts present in AI-generated images, whereas watermark-based detectors proactively embed watermarks into such images. A key question is which type of detector performs better in terms of effectiveness, robustness, and efficiency. However, the current literature lacks a comprehensive understanding of this issue. In this work, we aim to bridge that gap by developing ImageDetectBench, the first comprehensive benchmark to compare the effectiveness, robustness, and efficiency of passive and watermark-based detectors. Our benchmark includes four datasets, each containing a mix of AI-generated and non-AI-generated images. We evaluate five passive detectors and four watermark-based detectors against eight types of common perturbations and three types of adversarial perturbations. Our benchmark results reveal several interesting findings. For instance, watermark-based detectors consistently outperform passive detectors, both in the presence and absence of perturbations. Based on these insights, we provide recommendations for detecting AI-generated images, e.g., when both types of detectors are applicable, watermark-based detectors should be the preferred choice. Our code and data are publicly available at https://github.com/moyangkuo/ImageDetectBench.git.
Consistency Checks for Language Model Forecasters
Paleka, Daniel, Sudhir, Abhimanyu Pallavi, Alvarez, Alejandro, Bhat, Vineeth, Shen, Adam, Wang, Evan, Tramรจr, Florian
Forecasting is a task that is difficult to evaluate: the ground truth can only be known in the future. Recent work showing LLM forecasters rapidly approaching human-level performance begs the question: how can we benchmark and evaluate these forecasters instantaneously? Following the consistency check framework, we measure the performance of forecasters in terms of the consistency of their predictions on different logically-related questions. We propose a new, general consistency metric based on arbitrage: for example, if a forecasting AI illogically predicts that both the Democratic and Republican parties have 60% probability of winning the 2024 US presidential election, an arbitrageur can trade against the forecaster's predictions and make a profit. We build an automated evaluation system that generates a set of base questions, instantiates consistency checks from these questions, elicits the predictions of the forecaster, and measures the consistency of the predictions. We then build a standard, proper-scoring-rule forecasting benchmark, and show that our (instantaneous) consistency metrics correlate with LLM forecasters' ground truth Brier scores (which are only known in the future). We also release a consistency benchmark that resolves in 2028, providing a long-term evaluation tool for forecasting.
Biden looks to limit AI product exports, tech leaders say they'll lose global market share
Leaders in the tech industry are urging the Biden administration not to add a new regulation that will limit artificial intelligence exports, citing concerns it is overbroad and could diminish the United States' global dominance in AI. The new rule, which industry leaders say could come as early as the end of this week, effectively seeks to shore up the U.S. economy and national security efforts by adding new restrictions on how many U.S.-made artifical intelligence products can be deployed across the globe. "A rule of this nature would cede the global market to U.S. competitors who will be eager to fill the untapped demand created by placing arbitrary constraints on U.S. companies' ability to sell basic computing systems overseas," stated a Monday letter from Jason Oxman, the president and CEO of the Information Technology Industry Council (ITI), sent to Commerce Department Secretary Gina Raimondo. "Should the U.S. lose its advantage in the global AI ecosystem, it will be difficult, if not impossible, to regain in the future." FBI'S NEW WARNING ABOUT AI-DRIVEN SCAMS THAT ARE AFTER YOUR CASH The process to place new export controls on artificial intelligence goes back to October 2022, when the Biden administration's Commerce Department first released an updated export framework aimed at slowing the progress of Chinese military programs. Details of the new incoming export controls surfaced after the Biden administration called on American tech company NVIDIA to stop selling certain computer chips to China the following month.
Before Las Vegas, Intel Analysts Warned That Bomb Makers Were Turning to AI
Using a series of prompts six days before he died by suicide outside the main entrance of the Trump International Hotel in Las Vegas, Matthew Livelsberger, a highly decorated US Army Green Beret from Colorado, consulted with an artificial intelligence on the best ways to turn a rented Cybertruck into a four-ton vehicle-borne explosive. According to documents obtained exclusively by WIRED, US intelligence analysts have been issuing warnings about this precise scenario over the past year--and among their concerns are that AI tools could be used by racially or ideologically motivated extremists to target critical infrastructure, in particular the power grid. "We knew that AI was going to change the game at some point or another in, really, all of our lives," Sheriff Kevin McMahill of the Las Vegas Metropolitan Police Department told reporters on Tuesday. Copies of his exchanges with OpenAI's ChatGPT show that Livelsberger, 37, pursued information on how to amass as much explosive material as he legally could while en route to Las Vegas, as well as how best to set it off using the Desert Eagle gun discovered in the Cybertruck following his death. Screenshots shared by McMahill's office reveal Livelsberger prompting ChatGPT for information on Tannerite, a reactive compound typically used for target practice.
FBI verifies Tesla Cybertruck subject emailed podcaster, says he used ChatGPT to plan Trump hotel explosion
The FBI confirmed the email sent to the Shawn Ryan Show podcast was indeed from the Tesla Cybertruck subject Matthew Livelsberger, while Las Vegas police say Livelsberger used ChatGPT to plan the explosion. The FBI on Tuesday said an email that appeared to have been sent from the Las Vegas Cybertruck explosion subject Matthew Livelsberger to prominent podcaster Shawn Ryan was indeed confirmed to have come from Livelsberger. At a press conference, Special Agent in Charge of the FBI in Las Vegas Spencer Evans clarified that law enforcement has not verified the actual content of Livelsberger's email, just that he sent it. "We have confirmed the document that he sent to the podcast. We know that he was the one that sent that document. That's correct," Evans told reporters.
How China Is Advancing in AI Despite U.S. Chip Restrictions
In 2017, Beijing unveiled an ambitious roadmap to dominate artificial intelligence development, aiming to secure global leadership by 2030. By 2020, the plan called for "iconic advances" in AI to demonstrate its progress. Then in late 2022, OpenAI's release of ChatGPT took the world by surprise--and caught China flat-footed. At the time, leading Chinese technology companies were still reeling from an 18-month government crackdown that shaved around 1 trillion off China's tech sector. It was almost a year before a handful of Chinese AI chatbots received government approval for public release.
NASA Wants to Explore the Icy Moons of Jupiter and Saturn With Autonomous Robots
Europa's orbit is an ellipse, and the satellite's shape is affected by Jupiter's gravity, becoming deformed when it passes closer to Jupiter. This change in shape creates friction inside Europa, generating enormous amounts of heat in a mechanism known as tidal heating, which melts some of the ice and forms a vast internal ocean beneath the moon's thick ice shell. Europa's internal ocean is salty and is estimated to be about 100 kilometers deep on average, with a total volume of water twice that of all Earth's oceans, despite this moon being considerably smaller than our planet. In addition, it is believed that internal oceans exist on Jupiter's moons Ganymede and Callisto and Saturn's moons Titan and Enceladus. Liquid water is essential for life as we know it, which is why the ocean worlds are at the forefront of the search for extraterrestrial life.
Retrieval-Augmented Generation with Graphs (GraphRAG)
Han, Haoyu, Wang, Yu, Shomer, Harry, Guo, Kai, Ding, Jiayuan, Lei, Yongjia, Halappanavar, Mahantesh, Rossi, Ryan A., Mukherjee, Subhabrata, Tang, Xianfeng, He, Qi, Hua, Zhigang, Long, Bo, Zhao, Tong, Shah, Neil, Javari, Amin, Xia, Yinglong, Tang, Jiliang
Retrieval-augmented generation (RAG) is a powerful technique that enhances downstream task execution by retrieving additional information, such as knowledge, skills, and tools from external sources. Graph, by its intrinsic "nodes connected by edges" nature, encodes massive heterogeneous and relational information, making it a golden resource for RAG in tremendous real-world applications. As a result, we have recently witnessed increasing attention on equipping RAG with Graph, i.e., GraphRAG. However, unlike conventional RAG, where the retriever, generator, and external data sources can be uniformly designed in the neural-embedding space, the uniqueness of graph-structured data, such as diverse-formatted and domain-specific relational knowledge, poses unique and significant challenges when designing GraphRAG for different domains. Given the broad applicability, the associated design challenges, and the recent surge in GraphRAG, a systematic and up-to-date survey of its key concepts and techniques is urgently desired. Following this motivation, we present a comprehensive and up-to-date survey on GraphRAG. Our survey first proposes a holistic GraphRAG framework by defining its key components, including query processor, retriever, organizer, generator, and data source. Furthermore, recognizing that graphs in different domains exhibit distinct relational patterns and require dedicated designs, we review GraphRAG techniques uniquely tailored to each domain. Finally, we discuss research challenges and brainstorm directions to inspire cross-disciplinary opportunities.