Personal
System 2 Attention (is something you might need too)
Weston, Jason, Sukhbaatar, Sainbayar
Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To help rectify these issues, we introduce System 2 Attention (S2A), which leverages the ability of LLMs to reason in natural language and follow instructions in order to decide what to attend to. S2A regenerates the input context to only include the relevant portions, before attending to the regenerated context to elicit the final response. In experiments, S2A outperforms standard attention-based LLMs on three tasks containing opinion or irrelevant information: QA, math word problems and longform generation, where S2A increases factuality and objectivity, and decreases sycophancy.
Lost in the Middle: How Language Models Use Long Contexts
Liu, Nelson F., Lin, Kevin, Hewitt, John, Paranjape, Ashwin, Bevilacqua, Michele, Petroni, Fabio, Liang, Percy
While recent language models have the ability to take long contexts as input, relatively little is known about how well they use longer context. We analyze the performance of language models on two tasks that require identifying relevant information in their input contexts: multi-document question answering and key-value retrieval. We find that performance can degrade significantly when changing the position of relevant information, indicating that current language models do not robustly make use of information in long input contexts. In particular, we observe that performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models. Our analysis provides a better understanding of how language models use their input context and provides new evaluation protocols for future long-context language models.
Interview with Dautzenberg Roman: #IROS2023 Best Paper Award on Mobile Manipulation sponsored by OMRON Sinic X Corp.
Congratulations to Dautzenberg Roman and his team of researchers, who won the IROS 2023 Best Paper Award on Mobile Manipulation sponsored by OMRON Sinic X Corp. for their paper "A perching and tilting aerial robot for precise and versatile power tool work on vertical walls". Below, the authors tell us more about their work, the methodology, and what they are planning next. Our paper shows a an aerial robot (think "drone") which can exert large forces in the horizontal direction, i.e. onto walls. This is a difficult task, as UAVs usually rely on thrust vectoring to apply horizontal forces and thus can only apply small forces before losing control authority. By perching onto walls, our system no longer needs the propulsion to remain at a desired site.
Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
Qin, Chengwei, Zhang, Aston, Zhang, Zhuosheng, Chen, Jiaao, Yasunaga, Michihiro, Yang, Diyi
Spurred by advancements in scale, large language models (LLMs) have demonstrated the ability to perform a variety of natural language processing (NLP) tasks zero-shot -- i.e., without adaptation on downstream data. Recently, the debut of ChatGPT has drawn a great deal of attention from the natural language processing (NLP) community due to the fact that it can generate high-quality responses to human input and self-correct previous mistakes based on subsequent conversations. However, it is not yet known whether ChatGPT can serve as a generalist model that can perform many NLP tasks zero-shot. In this work, we empirically analyze the zero-shot learning ability of ChatGPT by evaluating it on 20 popular NLP datasets covering 7 representative task categories. With extensive empirical studies, we demonstrate both the effectiveness and limitations of the current version of ChatGPT. We find that ChatGPT performs well on many tasks favoring reasoning capabilities (e.g., arithmetic reasoning) while it still faces challenges when solving specific tasks such as sequence tagging. We additionally provide in-depth analysis through qualitative case studies.
Who Is Mira Murati, OpenAI's New Interim CEO?
Until the dramatic departure of OpenAI's cofounder and CEO Sam Altman Friday, Mira Murati was its chief technology officer--but you could also call her as its minister of truth. In addition to heading the teams that develop tools such as ChatGPT and Dall-E, it's been her job to make sure those products don't mislead people, show bias, or snuff out humanity altogether. This interview was conducted in July 2023 for WIRED's cover story on OpenAI. It is being published today after Sam Altman's sudden departure to provide a glimpse at the thinking of the powerful AI company's new boss. Steven Levy: How did you come to join OpenAI?
OpenAI CEO Sam Altman ousted as 'board no longer has confidence' in his leadership
In a surprise shakeup of its c-suite Friday, OpenAI's board of directors announced that CEO Sam Altman has been fired and will be leaving both the company and the board, effective immediately. Chief Technology Officer Mira Murati has been named interim CEO. Altman's oustering reportedly follows an internal "deliberative review process" which found he had not been "consistently candid in his communications with the board, hindering its ability to exercise its responsibilities," the company announced. As such, "the board no longer has confidence in his ability to continue leading OpenAI." OpenAI, which owns popular AI chatbot ChatGPT, thanked Altman' for his "many contributions to the founding and growth of OpenAI," but believes that "as the leader of the company's research, product, and safety functions, Mira is exceptionally qualified to step into the role of interim CEO." The board added it has "the utmost confidence in her ability to lead OpenAI during this transition period."
A Rise in Antisemitism; and a Conversation with the A.I. Pioneer Geoffrey Hinton
Sign up to receive our weekly newsletter of the best New Yorker podcasts. The State Department's Special Envoy to Monitor and Combat Antisemitism, the historian Deborah Lipstadt, says the prejudice is coming "from all ends of the political spectrum, and in between." It threatens not only Jews, she says, but the stability of democracies. Lipstadt and David Remnick discuss how antisemitic sentiments may overlap in complicated ways with political opposition to Israel, including anti-Zionism. Plus, The New Yorker's ideas editor speaks with Geoffrey Hinton, the computer scientist known as the godfather of A.I. Hinton pioneered neural networks, the artificial brains that power ChatGPT, for example.
Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
Gupta, Ashim, Rajendhran, Rishanth, Stringham, Nathan, Srikumar, Vivek, Marasović, Ana
Are the longstanding robustness issues in NLP resolved by today's larger and more performant models? To address this question, we conduct a thorough investigation using 19 models of different sizes spanning different architectural choices and pretraining objectives. We conduct evaluations using (a) OOD and challenge test sets, (b) CheckLists, (c) contrast sets, and (d) adversarial inputs. Our analysis reveals that not all OOD tests provide further insight into robustness. Evaluating with CheckLists and contrast sets shows significant gaps in model performance; merely scaling models does not make them sufficiently robust. Finally, we point out that current approaches for adversarial evaluations of models are themselves problematic: they can be easily thwarted, and in their current forms, do not represent a sufficiently deep probe of model robustness. We conclude that not only is the question of robustness in NLP as yet unresolved, but even some of the approaches to measure robustness need to be reassessed.
Data-Driven Structured Policy Iteration for Homogeneous Distributed Systems
Alemzadeh, Siavash, Talebi, Shahriar, Mesbahi, Mehran
Control of networked systems, comprised of interacting agents, is often achieved through modeling the underlying interactions. Constructing accurate models of such interactions--in the meantime--can become prohibitive in applications. Data-driven control methods avoid such complications by directly synthesizing a controller from the observed data. In this paper, we propose an algorithm referred to as Data-driven Structured Policy Iteration (D2SPI), for synthesizing an efficient feedback mechanism that respects the sparsity pattern induced by the underlying interaction network. In particular, our algorithm uses temporary "auxiliary" communication links in order to enable the required information exchange on a (smaller) sub-network during the "learning phase" -- links that will be removed subsequently for the final distributed feedback synthesis. We then proceed to show that the learned policy results in a stabilizing structured policy for the entire network. Our analysis is then followed by showing the stability and convergence of the proposed distributed policies throughout the learning phase, exploiting a construct referred to as the "Patterned monoid.'' The performance of D2SPI is then demonstrated using representative simulation scenarios.
Towards Verifiable Text Generation with Symbolic References
Hennigen, Lucas Torroba, Shen, Shannon, Nrusimha, Aniruddha, Gapp, Bernhard, Sontag, David, Kim, Yoon
Large language models (LLMs) have demonstrated an impressive ability to synthesize plausible and fluent text. However they remain vulnerable to hallucinations, and thus their outputs generally require manual human verification for high-stakes applications, which can be timeconsuming and difficult. This paper proposes symbolically grounded generation (SymGen) as a simple approach for enabling easier validation of an LLM's output. SymGen prompts an LLM to interleave its regular output text with explicit symbolic references to fields present in some conditioning data (e.g., a table in JSON format). The references can be used to display the provenance of different spans of text in the generation, reducing the effort required for manual verification. Across data-to-text and question answering experiments, we find that Figure 1: Compare a standard LLM-generated (A) with LLMs are able to directly output text that makes a SymGen (B, ours) description of a basketball game, use of symbolic references while maintaining based on statistics about it.