Government
Crafting In-context Examples according to LMs' Parametric Knowledge
Lee, Yoonsang, Atreya, Pranav, Ye, Xi, Choi, Eunsol
In-context learning has been applied to knowledge-rich tasks such as question answering. In such scenarios, in-context examples are used to trigger a behaviour in the language model: namely, it should surface information stored in its parametric knowledge. We study the construction of in-context example sets, with a focus on the parametric knowledge of the model regarding in-context examples. We identify 'known' examples, where models can correctly answer from its parametric knowledge, and 'unknown' ones. Our experiments show that prompting with 'unknown' examples decreases the performance, potentially as it encourages hallucination rather than searching its parametric knowledge. Constructing an in-context example set that presents both known and unknown information performs the best across diverse settings. We perform analysis on three multi-answer question answering datasets, which allows us to further study answer set ordering strategies based on the LM's knowledge about each answer. Together, our study sheds lights on how to best construct in-context example sets for knowledge-rich tasks.
Are Large Language Models Temporally Grounded?
Qiu, Yifu, Zhao, Zheng, Ziser, Yftah, Korhonen, Anna, Ponti, Edoardo M., Cohen, Shay B.
Are Large language models (LLMs) temporally grounded? Since LLMs cannot perceive and interact with the environment, it is impossible to answer this question directly. Instead, we provide LLMs with textual narratives and probe them with respect to their common-sense knowledge of the structure and duration of events, their ability to order events along a timeline, and self-consistency within their temporal model (e.g., temporal relations such as after and before are mutually exclusive for any pair of events). We evaluate state-of-the-art LLMs (such as LLaMA 2 and GPT-4) on three tasks reflecting these abilities. Generally, we find that LLMs lag significantly behind both human performance as well as small-scale, specialised LMs. In-context learning, instruction tuning, and chain-of-thought prompting reduce this gap only to a limited degree. Crucially, LLMs struggle the most with self-consistency, displaying incoherent behaviour in at least 27.23% of their predictions. Contrary to expectations, we also find that scaling the model size does not guarantee positive gains in performance. To explain these results, we study the sources from which LLMs may gather temporal information: we find that sentence ordering in unlabelled texts, available during pre-training, is only weakly correlated with event ordering. Moreover, public instruction tuning mixtures contain few temporal tasks. Hence, we conclude that current LLMs lack a consistent temporal model of textual narratives. Code, datasets, and LLM outputs are available at https://github.com/yfqiu-nlp/temporal-llms.
Loss Modeling for Multi-Annotator Datasets
Jinadu, Uthman, Annan, Jesse, Wen, Shanshan, Ding, Yi
Accounting for the opinions of all annotators of a dataset is critical for fairness. However, when annotating large datasets, individual annotators will frequently provide thousands of ratings which can lead to fatigue. Additionally, these annotation processes can occur over multiple days which can lead to an inaccurate representation of an annotator's opinion over time. To combat this, we propose to learn a more accurate representation of diverse opinions by utilizing multitask learning in conjunction with loss-based label correction. We show that using our novel formulation, we can cleanly separate agreeing and disagreeing annotations. Furthermore, we demonstrate that this modification can improve prediction performance in a single or multi-annotator setting. Lastly, we show that this method remains robust to additional label noise that is applied to subjective data.
Exploring the Practicality of Generative Retrieval on Dynamic Corpora
Yoon, Soyoung, Kim, Chaeeun, Lee, Hyunji, Jang, Joel, Yang, Sohee, Seo, Minjoon
Benchmarking the performance of information retrieval (IR) methods are mostly conducted with a fixed set of documents (static corpora); in realistic scenarios, this is rarely the case and the document to be retrieved are constantly updated and added. In this paper, we focus on conducting a comprehensive comparison between two categories of contemporary retrieval systems, Dual Encoders (DE) and Generative Retrievals (GR), in a dynamic scenario where the corpora to be retrieved is updated. We also conduct an extensive evaluation of computational and memory efficiency, crucial factors for IR systems for real-world deployment. Our results demonstrate that GR is more adaptable to evolving knowledge (+13-18% on the StreamingQA Benchmark), robust in handling data with temporal information (x 10 times), and efficient in terms of memory (x 4 times), indexing time (x 6 times), and inference flops (x 10 times). Our paper highlights GR's potential for future use in practical IR systems.
Could the next election be AI-generated? Presidential candidates use tech to promote themselves and attack their opponent in Argentina
The next US election could see a flood of AI-generated campaigning posters after candidates in Argentina used it to promote themselves and attack their opponent. Sergio Massa and Javier Milei are battling for the presidency and are harnessing the power of artificial intelligence in hopes of one-upping the other. Massa recreated himself in several scenes where he sports military metals, surrounded by hundreds of people looking up at him in hope while pushing out a video showing Javier as a character in the film Clockwork Orange. But the far-right libertarian economist did not sit back quietly - he used AI to create Massa in the form of a Chinese communist leader. Argentina's digital posters follow those created by US officials this year, such as a video from Ron DeSantis of Florida's campaign which featured a video showing Donald Trump embracing Anthony Fauci.
Red nation on the red planet? This communist country's latest venture could be key to human activity on Mars
Commercial spaceflight companies like SpaceX have made space travel more accessible, and allow more research for future missions to the moon and Mars, astronauts said. A robotic space chemist could create oxygen on Mars using materials from the planet's surface, Chinese researchers behind the project say. A refrigerator-sized machine equipped with artificial intelligence and a robotic arm broke down material from five meteorites and analyzed it to identify a chemical formula that creates a substance that can cause oxygen to separate from water. Researchers said it would have taken a human 2,000 years to find that formula. WHAT IS ARTIFICIAL INTELLIGENCE (AI)?
US Navy destroyer shoots down drone from Yemen in the Red Sea
The U.S. Department of Defense released video footage of a U.S. air strike on a training and weapons facility in Abul Kamal, Syria. The USS Thomas Hudner, an Arleigh Burke-class destroyer, shot down a drone from Yemen in the Red Sea on Wednesday, two U.S. defense officials confirmed to Fox News. A defense official said the drone was shot down in self-defense. "The drone was heading towards the Hudner," the official said. The drone attack is the latest in a series of attacks on American troops stationed in the Middle East amid the ongoing Israel-Hamas war.
YouTube to Require Creators to Disclose Use of Generative AI
YouTube is rolling out new rules for AI content, including a requirement that creators reveal whether they've used generative artificial intelligence to make realistic looking videos. In a blog post Tuesday outlining a number of AI-related policy updates, YouTube said creators that don't disclose whether they've used AI tools to make "altered or synthetic" videos face penalties including having their content removed or suspension from the platform's revenue sharing program. "Generative AI has the potential to unlock creativity on YouTube and transform the experience for viewers and creators on our platform," Jennifer Flannery O'Connor and Emily Moxley, vice presidents for product management, wrote in the blog post. "But just as important, these opportunities must be balanced with our responsibility to protect the YouTube community." The restrictions expand on rules that YouTube's parent company, Google, unveiled in September requiring that political ads on YouTube and other Google platforms using artificial intelligence come with a prominent warning label.
Fox News AI Newsletter: Hyped AI-based restaurant goes bust
A restaurant in a rural Oregon city couldn't find enough servers to stay fully staffed. So the owner hired a robot named Plato. She had no idea how much pushback she'd get from the community. ROCKY ROAD: A hyped AI-based restaurant opened to fanfare last month in San Francisco. 'INCREDIBLY POOR DECISION': Biden hands China big win with military AI deal, experts say.
Biden hands China big win with military deal, experts say: 'Incredibly poor decision'
House Armed Services Committee holds a hearing on the Department of Defense using artifical intelligence. President Biden is set to strike a deal with China that would limit the use of artifical intelligence in nuclear weapons. Biden is to meet with Chinese President Xi Jinping on Wednesday at the Asia-Pacific Economic Cooperation (APEC) summit in San Francisco, where the two leaders are expected to also sign an agreement to limit AI's use in military applications, according to a report from Business Insider. According to the report, Biden and Xi will agree to limit AI use in the systems that control and deploy nuclear weapons as well as the technology's use in autonomous weapon systems such as drones. US MILITARY NEEDS AI VEHICLES, WEAPON SYSTEMS TO BE'SUPERIOR' GLOBAL FORCE: EXPERTS President Biden shakes hands with Chinese President Xi Jinping as they meet on the sidelines of the G20 leaders summit in Bali, Indonesia, on Nov. 14, 2022.