Pacific Ocean
Tracking the Newsworthiness of Public Documents
Spangher, Alexander, Ferrara, Emilio, Welsh, Ben, Peng, Nanyun, Tumgoren, Serdar, May, Jonathan
Journalists must find stories in huge amounts of textual data (e.g. leaks, bills, press releases) as part of their jobs: determining when and why text becomes news can help us understand coverage patterns and help us build assistive tools. Yet, this is challenging because very few labelled links exist, language use between corpora is very different, and text may be covered for a variety of reasons. In this work we focus on news coverage of local public policy in the San Francisco Bay Area by the San Francisco Chronicle. First, we gather news articles, public policy documents and meeting recordings and link them using probabilistic relational modeling, which we show is a low-annotation linking methodology that outperforms other retrieval-based baselines. Second, we define a new task: newsworthiness prediction, to predict if a policy item will get covered. We show that different aspects of public policy discussion yield different newsworthiness signals. Finally we perform human evaluation with expert journalists and show our systems identify policies they consider newsworthy with 68% F1 and our coverage recommendations are helpful with an 84% win-rate.
Biden hands China big win with military deal, experts say: 'Incredibly poor decision'
House Armed Services Committee holds a hearing on the Department of Defense using artifical intelligence. President Biden is set to strike a deal with China that would limit the use of artifical intelligence in nuclear weapons. Biden is to meet with Chinese President Xi Jinping on Wednesday at the Asia-Pacific Economic Cooperation (APEC) summit in San Francisco, where the two leaders are expected to also sign an agreement to limit AI's use in military applications, according to a report from Business Insider. According to the report, Biden and Xi will agree to limit AI use in the systems that control and deploy nuclear weapons as well as the technology's use in autonomous weapon systems such as drones. US MILITARY NEEDS AI VEHICLES, WEAPON SYSTEMS TO BE'SUPERIOR' GLOBAL FORCE: EXPERTS President Biden shakes hands with Chinese President Xi Jinping as they meet on the sidelines of the G20 leaders summit in Bali, Indonesia, on Nov. 14, 2022.
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
Saad-Falcon, Jon, Khattab, Omar, Potts, Christopher, Zaharia, Matei
Evaluating retrieval-augmented generation (RAG) systems traditionally relies on hand annotations for input queries, passages to retrieve, and responses to generate. We introduce ARES, an Automated RAG Evaluation System, for evaluating RAG systems along the dimensions of context relevance, answer faithfulness, and answer relevance. Using synthetic training data, ARES finetunes lightweight LM judges to assess the quality of individual RAG components. To mitigate potential prediction errors, ARES utilizes a small set of human-annotated datapoints for prediction-powered inference (PPI). Across six different knowledge-intensive tasks in KILT and SuperGLUE, ARES accurately evaluates RAG systems while using a few hundred human annotations during evaluation. Furthermore, ARES judges remain effective across domain shifts, proving accurate even after changing the type of queries and/or documents used in the evaluated RAG systems. We make our datasets and code for replication and deployment available at https://github.com/stanford-futuredata/ARES.
Alternatives to the Scaled Dot Product for Attention in the Transformer Neural Network Architecture
The transformer neural network architecture uses a form of attention in which the dot product of query and key is divided by the square root of the key dimension before applying softmax. This scaling of the dot product is designed to avoid the absolute value of the dot products becoming so large that applying softmax leads to vanishing gradients. In this paper, we propose some alternative scalings, including dividing the dot product instead by the sum of the key lengths before applying softmax. We use simulated keys and queries to show that in many situations this appears to be more effective at avoiding regions where applying softmax leads to vanishing gradients. Attention plays a prominent role in the transformer neural network architecture, as indicated by the title of the landmark paper introducing the architecture, "Attention Is All You Need" [1], by Vaswani et al.
When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks
Peng, Hao, Wang, Xiaozhi, Chen, Jianhui, Li, Weikai, Qi, Yunjia, Wang, Zimu, Wu, Zhili, Zeng, Kaisheng, Xu, Bin, Hou, Lei, Li, Juanzi
In-context learning (ICL) has become the default method for using large language models (LLMs), making the exploration of its limitations and understanding the underlying causes crucial. In this paper, we find that ICL falls short of handling specification-heavy tasks, which are tasks with complicated and extensive task specifications, requiring several hours for ordinary humans to master, such as traditional information extraction tasks. The performance of ICL on these tasks mostly cannot reach half of the state-of-the-art results. To explore the reasons behind this failure, we conduct comprehensive experiments on 18 specification-heavy tasks with various LLMs and identify three primary reasons: inability to specifically understand context, misalignment in task schema comprehension with humans, and inadequate long-text understanding ability. Furthermore, we demonstrate that through fine-tuning, LLMs can achieve decent performance on these tasks, indicating that the failure of ICL is not an inherent flaw of LLMs, but rather a drawback of existing alignment methods that renders LLMs incapable of handling complicated specification-heavy tasks via ICL. To substantiate this, we perform dedicated instruction tuning on LLMs for these tasks and observe a notable improvement. We hope the analyses in this paper could facilitate advancements in alignment methods enabling LLMs to meet more sophisticated human demands.
Estimating Appearance Models for Image Segmentation via Tensor Factorization
Neto, Jeova Farias Sales Rocha
Image Segmentation is one of the core tasks in Computer Vision and solving it often depends on modeling the image appearance data via the color distributions of each it its constituent regions. Whereas many segmentation algorithms handle the appearance models dependence using alternation or implicit methods, we propose here a new approach to directly estimate them from the image without prior information on the underlying segmentation. Our method uses local high order color statistics from the image as an input to tensor factorization-based estimator for latent variable models. This approach is able to estimate models in multiregion images and automatically output the regions proportions without prior user interaction, overcoming the drawbacks from a prior attempt to this problem. We also demonstrate the performance of our proposed method in many challenging synthetic and real imaging scenarios and show that it leads to an efficient segmentation algorithm.
Biden and Xi look to put floor under plummeting U.S.-China ties
Nearly a year to the date since their last meeting, U.S. President Joe Biden and Chinese leader Xi Jinping will sit down Wednesday in the San Francisco Bay Area to try and put a floor under ties that have plummeted to fresh lows in recent months. When Biden and Xi meet on the sidelines of the Asia-Pacific Economic Cooperation forum in San Francisco, both will have a laundry list of concerns to discuss. From military-to-military lines of communication, Taiwan, and the South and East China Seas to tough U.S. semiconductor export controls, the manufacture and export of fentanyl, and artificial intelligence threats -- all will be on the table during several hours of discussions. But don't expect the talks -- the pair's seventh interaction since the start of the Biden administration but just the second in-person meeting -- to yield any dramatic breakthroughs.
Improving Zero-shot Reader by Reducing Distractions from Irrelevant Documents in Open-Domain Question Answering
Cho, Sukmin, Seo, Jeongyeon, Jeong, Soyeong, Park, Jong C.
Large language models (LLMs) enable zero-shot approaches in open-domain question answering (ODQA), yet with limited advancements as the reader is compared to the retriever. This study aims at the feasibility of a zero-shot reader that addresses the challenges of computational cost and the need for labeled data. We find that LLMs are distracted due to irrelevant documents in the retrieved set and the overconfidence of the generated answers when they are exploited as zero-shot readers. To tackle these problems, we mitigate the impact of such documents via Distraction-aware Answer Selection (DAS) with a negation-based instruction and score adjustment for proper answer selection. Experimental results show that our approach successfully handles distraction across diverse scenarios, enhancing the performance of zero-shot readers. Furthermore, unlike supervised readers struggling with unseen data, zero-shot readers demonstrate outstanding transferability without any training.
The US Wants China to Start Talking About AI Weapons
When US President Joe Biden meets with his Chinese counterpart Xi Jinping in the San Francisco Bay Area this week, the pair will have a long list of matters to discuss, including the Israel-Hamas war and Russia's ongoing invasion of Ukraine. Behind the scenes at the APEC summit, however, US officials hope to strike up a dialogue with China about placing guardrails around military use of artificial intelligence, with the ultimate goal of lessening the potential risks that rapid adoption--and reckless use--of the technology might bring. "We have a collective interest in reducing the potential risks from the deployment of unreliable AI applications," because of risks of unintended escalation, says a senior State Department official familiar with recent efforts to broach the issue, who spoke on condition of anonymity. "We very much hope to have a further conversation with China on this issue." Biden's meeting with Xi this week may provide momentum for more military dialogue.
Consistency Analysis of ChatGPT
Jang, Myeongjun Erik, Lukasiewicz, Thomas
ChatGPT has gained a huge popularity since its introduction. Its positive aspects have been reported through many media platforms, and some analyses even showed that ChatGPT achieved a decent grade in professional exams, adding extra support to the claim that AI can now assist and even replace humans in industrial fields. Others, however, doubt its reliability and trustworthiness. This paper investigates the trustworthiness of ChatGPT and GPT-4 regarding logically consistent behaviour, focusing specifically on semantic consistency and the properties of negation, symmetric, and transitive consistency. Our findings suggest that while both models appear to show an enhanced language understanding and reasoning ability, they still frequently fall short of generating logically consistent predictions. We also ascertain via experiments that prompt designing, few-shot learning and employing larger large language models (LLMs) are unlikely to be the ultimate solution to resolve the inconsistency issue of LLMs.