Government
Knowledge Mechanisms in Large Language Models: A Survey and Perspective
Wang, Mengru, Yao, Yunzhi, Xu, Ziwen, Qiao, Shuofei, Deng, Shumin, Wang, Peng, Chen, Xiang, Gu, Jia-Chen, Jiang, Yong, Xie, Pengjun, Huang, Fei, Chen, Huajun, Zhang, Ningyu
Understanding knowledge mechanisms in Large Language Models (LLMs) is crucial for advancing towards trustworthy AGI. This paper reviews knowledge mechanism analysis from a novel taxonomy including knowledge utilization and evolution. Knowledge utilization delves into the mechanism of memorization, comprehension and application, and creation. Knowledge evolution focuses on the dynamic progression of knowledge within individual and group LLMs. Moreover, we discuss what knowledge LLMs have learned, the reasons for the fragility of parametric knowledge, and the potential dark knowledge (hypothesis) that will be challenging to address. We hope this work can help understand knowledge in LLMs and provide insights for future research.
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends
Martinelli, Giuliano, Barba, Edoardo, Navigli, Roberto
Large autoregressive generative models have emerged as the cornerstone for achieving the highest performance across several Natural Language Processing tasks. However, the urge to attain superior results has, at times, led to the premature replacement of carefully designed task-specific approaches without exhaustive experimentation. The Coreference Resolution task is no exception; all recent state-of-the-art solutions adopt large generative autoregressive models that outperform encoder-based discriminative systems. In this work,we challenge this recent trend by introducing Maverick, a carefully designed - yet simple - pipeline, which enables running a state-of-the-art Coreference Resolution system within the constraints of an academic budget, outperforming models with up to 13 billion parameters with as few as 500 million parameters. Maverick achieves state-of-the-art performance on the CoNLL-2012 benchmark, training with up to 0.006x the memory resources and obtaining a 170x faster inference compared to previous state-of-the-art systems. We extensively validate the robustness of the Maverick framework with an array of diverse experiments, reporting improvements over prior systems in data-scarce, long-document, and out-of-domain settings. We release our code and models for research purposes at https://github.com/SapienzaNLP/maverick-coref.
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
Yu, Sangwon, Song, Jongyoon, Hwang, Bongkyu, Kang, Hoyoung, Cho, Sooah, Choi, Junhwa, Joe, Seongho, Lee, Taehee, Gwon, Youngjune L., Yoon, Sungroh
A binary decision task, like yes-no questions or answer verification, reflects a significant real-world scenario such as where users look for confirmation about the correctness of their decisions on specific issues. In this work, we observe that language models exhibit a negative bias in the binary decisions of complex reasoning tasks. Based on our observations and the rationale about attention-based model dynamics, we propose a negative attention score (NAS) to systematically and quantitatively formulate negative bias. Based on NAS, we identify attention heads that attend to negative tokens provided in the instructions as answer candidate of binary decisions, regardless of the question in the prompt, and validate their association with the negative bias. Additionally, we propose the negative attention score alignment (NASA) method, which is a parameter-efficient fine-tuning technique to address the extracted negatively biased attention heads. Experimental results from various domains of reasoning tasks and large model search space demonstrate that NASA significantly reduces the gap between precision and recall caused by negative bias while preserving their generalization abilities. Our codes are available at \url{https://github.com/ysw1021/NASA}.
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
Wu, Zehui, Gong, Ziwei, Ai, Lin, Shi, Pengyuan, Donbekci, Kaan, Hirschberg, Julia
This paper introduces a novel approach to emotion detection in speech using Large Language Models (LLMs). We address the limitation of LLMs in processing audio inputs by translating speech characteristics into natural language descriptions. Our method integrates these descriptions into text prompts, enabling LLMs to perform multimodal emotion analysis without architectural modifications. We evaluate our approach on two datasets: IEMOCAP and MELD, demonstrating significant improvements in emotion recognition accuracy, particularly for high-quality audio data. Our experiments show that incorporating speech descriptions yields a 2 percentage point increase in weighted F1 score on IEMOCAP (from 70.111\% to 72.596\%). We also compare various LLM architectures and explore the effectiveness of different feature representations. Our findings highlight the potential of this approach in enhancing emotion detection capabilities of LLMs and underscore the importance of audio quality in speech-based emotion recognition tasks. We'll release the source code on Github.
Deceptive AI systems that give explanations are more convincing than honest AI systems and can amplify belief in misinformation
Danry, Valdemar, Pataranutaporn, Pat, Groh, Matthew, Epstein, Ziv, Maes, Pattie
Advanced Artificial Intelligence (AI) systems, specifically large language models (LLMs), have the capability to generate not just misinformation, but also deceptive explanations that can justify and propagate false information and erode trust in the truth. We examined the impact of deceptive AI generated explanations on individuals' beliefs in a pre-registered online experiment with 23,840 observations from 1,192 participants. We found that in addition to being more persuasive than accurate and honest explanations, AI-generated deceptive explanations can significantly amplify belief in false news headlines and undermine true ones as compared to AI systems that simply classify the headline incorrectly as being true/false. Moreover, our results show that personal factors such as cognitive reflection and trust in AI do not necessarily protect individuals from these effects caused by deceptive AI generated explanations. Instead, our results show that the logical validity of AI generated deceptive explanations, that is whether the explanation has a causal effect on the truthfulness of the AI's classification, plays a critical role in countering their persuasiveness - with logically invalid explanations being deemed less credible. This underscores the importance of teaching logical reasoning and critical thinking skills to identify logically invalid arguments, fostering greater resilience against advanced AI-driven misinformation.
DDU-Net: A Domain Decomposition-based CNN for High-Resolution Image Segmentation on Multiple GPUs
Verburg, Cornรฉ, Heinlein, Alexander, Cyr, Eric C.
The segmentation of ultra-high resolution images poses challenges such as loss of spatial information or computational inefficiency. In this work, a novel approach that combines encoder-decoder architectures with domain decomposition strategies to address these challenges is proposed. Specifically, a domain decomposition-based U-Net (DDU-Net) architecture is introduced, which partitions input images into non-overlapping patches that can be processed independently on separate devices. A communication network is added to facilitate inter-patch information exchange to enhance the understanding of spatial context. Experimental validation is performed on a synthetic dataset that is designed to measure the effectiveness of the communication network. Then, the performance is tested on the DeepGlobe land cover classification dataset as a real-world benchmark data set. The results demonstrate that the approach, which includes inter-patch communication for images divided into $16\times16$ non-overlapping subimages, achieves a $2-3\,\%$ higher intersection over union (IoU) score compared to the same network without inter-patch communication. The performance of the network which includes communication is equivalent to that of a baseline U-Net trained on the full image, showing that our model provides an effective solution for segmenting ultra-high-resolution images while preserving spatial context. The code is available at https://github.com/corne00/HiRes-Seg-CNN.
Acting Secret Service director tells Senate Trump shooting was 'a failure of the Secret Service'
Fox News' Chad Pergram previews the Senate's Tuesday hearing with acting U.S. Secret Service Director Ronald Rowe Jr. and FBI Deputy Director Paul Abbate as lawmakers continue investigating the security lapses at Trump's Butler rally. Acting Secret Service Director Ronald Rowe, Jr. admitted to the Senate on Tuesday that the assassination attempt against former President Trump was "a failure of the Secret Service," and not local law enforcement. Rowe's admission was the most direct assignment of guilt by the Secret Service and investigators since the July 13 shooting. The acting director appeared before the Senate Judiciary and Homeland Security committees on Tuesday alongside FBI Deputy Director Paul Abbate. Rowe detailed the failure of a drone detection system that was supposed to be online before shooter Thomas Matthew Crooks conducted his own reconnaissance the day of the rally.
UK regulator looks at Google's partnership with Anthropic
The Competition and Markets Authority has begun a preliminary investigation into a partnership between Google and the AI startup Anthropic, marking the latest in a string of investigations into deals between big tech companies and smallerAI ones. Google invested 2bn (about 1.56bn) into Anthropic in 2023, shortly after signing a cloud computing agreement with the startup, which develops the Claude LLM and chatbot. The CMA is now considering whether the partnership has "resulted in the creation of a relevant merger situation" which would allow the agency to begin a formal investigation. It is inviting comments over the next two weeks. The move comes amid broader concerns about competition in the generative AI sector. A deal between Amazon and Anthropic is also being investigated by the CMA as a potential merger after Amazon took a 4bn stake in the company and signed a deal to become one of the startup's cloud computing providers.
Russia has overrun 2 more eastern Donetsk villages, Ukrainian troops report
Former U.S. ambassador to NATO Kay Bailey Hutchinson discusses Biden's recent effort to show American allies that he is fit to serve as president and Ukrainian President Zelenskyy's concern about delaying action against Russia. Russian forces have overrun two front-line villages in Ukraine's eastern Donetsk region, a Ukrainian army sergeant said Monday, after relentless assaults that are part of a Kremlin summer push to overwhelm battlefield defenses there. Separately, attacks in Russia's Kursk region by the Security Service of Ukraine, also known as the SBU, struck a number of substations causing power outages, according to a statement from the General Staff of Ukraine. The claim of responsibility came after Russia said it thwarted a nighttime Ukrainian drone attack. "They pressed non-stop" to capture Vovche and Prohres, the chief sergeant of Ukraine's 47th Separate Mechanized Brigade, Oleh Chaus, told Radio Svaboda.
The Download: rebuilding economic security, and solving math problems
Fiona Murray is the William Porter (1967) Professor of Entrepreneurship at the MIT School of Management and Vice Chair of the NATO Innovation Fund. A country's economic security--its ability to generate both national security and economic prosperity--is grounded in it having technological capabilities that outpace those of its adversaries and complement those of its allies. Though this is a principle well known throughout history, the move over the last few decades toward globalization and offshoring has made ensuring a nation state's security and economic prosperity increasingly problematic. For the US and its allies in NATO, a particular problem has emerged: a "missing middle" in technology investment. Insufficient capital is allocated toward the maturation of breakthroughs in critical technologies to ensure that they can be deployed at scale.