Government
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
Mao, Xin, Li, Feng-Lin, Xu, Huimin, Zhang, Wei, Luu, Anh Tuan
While Reinforcement Learning from Human Feedback (RLHF) significantly enhances the generation quality of Large Language Models (LLMs), recent studies have raised concerns regarding the complexity and instability associated with the Proximal Policy Optimization (PPO) algorithm, proposing a series of order-based calibration methods as viable alternatives. This paper delves further into current order-based methods, examining their inefficiencies in utilizing reward values and addressing misalignment issues. Building upon these findings, we propose a novel \textbf{V}alue-based \textbf{C}ali\textbf{B}ration (VCB) method to better align LLMs with human preferences. Experimental results demonstrate that VCB surpasses existing alignment methods on AI assistant and summarization datasets, providing impressive generalizability, robustness, and stability in diverse settings.
One-stage Prompt-based Continual Learning
Kim, Youngeun, Li, Yuhang, Panda, Priyadarshini
Prompt-based Continual Learning (PCL) has gained considerable attention as a promising continual learning solution as it achieves state-of-the-art performance while preventing privacy violation and memory overhead issues. Nonetheless, existing PCL approaches face significant computational burdens because of two Vision Transformer (ViT) feed-forward stages; one is for the query ViT that generates a prompt query to select prompts inside a prompt pool; the other one is a backbone ViT that mixes information between selected prompts and image tokens. To address this, we introduce a one-stage PCL framework by directly using the intermediate layer's token embedding as a prompt query. This design removes the need for an additional feed-forward stage for query ViT, resulting in ~50% computational cost reduction for both training and inference with marginal accuracy drop < 1%. We further introduce a Query-Pool Regularization (QR) loss that regulates the relationship between the prompt query and the prompt pool to improve representation power. The QR loss is only applied during training time, so there is no computational overhead at inference from the QR loss. With the QR loss, our approach maintains ~ 50% computational cost reduction during inference as well as outperforms the prior two-stage PCL methods by ~1.4% on public class-incremental continual learning benchmarks including CIFAR-100, ImageNet-R, and DomainNet.
Understanding Public Perceptions of AI Conversational Agents: A Cross-Cultural Analysis
Liu, Zihan, Li, Han, Chen, Anfan, Zhang, Renwen, Lee, Yi-Chieh
Conversational Agents (CAs) have increasingly been integrated into everyday life, sparking significant discussions on social media. While previous research has examined public perceptions of AI in general, there is a notable lack in research focused on CAs, with fewer investigations into cultural variations in CA perceptions. To address this gap, this study used computational methods to analyze about one million social media discussions surrounding CAs and compared people's discourses and perceptions of CAs in the US and China. We find Chinese participants tended to view CAs hedonically, perceived voice-based and physically embodied CAs as warmer and more competent, and generally expressed positive emotions. In contrast, US participants saw CAs more functionally, with an ambivalent attitude. Warm perception was a key driver of positive emotions toward CAs in both countries. We discussed practical implications for designing contextually sensitive and user-centric CAs to resonate with various users' preferences and needs.
US, UK conduct joint strikes on more than a dozen Houthi targets in Yemen: 'Specifically targeted'
The United States and United Kingdom carried out more than a dozen strikes against Iranian-backed Houthi targets in Yemen on Saturday, with support from Australia, Bahrain, Canada, Denmark, the Netherlands and New Zealand, two U.S. officials told Fox News. The targets were hit successfully and include weapons storage facilities, and drone and missile launchers. The operation hit five Houthi-controlled locations in Yemen and is a response to the near-daily Houthi attacks involving Iranian drones and anti-ship ballistic missiles, a senior U.S. official said. The fourth round of American and British strikes came days after a British cargo ship was hit by a Houthi missile. In a joint statement, the U.S, U.K. and the other allied countries said: "In response to the Houthis' continued attacks against commercial and naval vessels transiting the Red Sea and surrounding waterways, today the militaries of the United States and United Kingdom, with support from Australia, Bahrain, Canada, Denmark, the Netherlands, and New Zealand, conducted an additional round of strikes against several targets in Houthi-controlled areas of Yemen."
Dean Phillips distances himself from campaign operative who reportedly paid 1 for AI-generated Biden deepfake
Longshot Democratic presidential candidate Rep. Dean Phillips, D-Minn., is distancing himself from a report that one of his campaign's former consultants hired a magician to create a deepfake of President Biden urging New Hampshire voters not to participate in last month's primary. Paul Carpenter, a magician from New Orleans, came forward and said he had made the deepfake for 1 and that a Democratic consultant Steve Kramer had paid him 150 to do it, according to an NBC report. Kramer is a get-out-the-vote specialist who worked on ballot access for the Phillips campaign and also worked on Kanye West's unsuccessful 2020 presidential campaign. "I'm disgusted that a consultant hired to assist my campaign [with] ballot access is alleged to have faked a robocall impersonating Joe Biden," Phillips wrote on X on Friday. "While I don't know the person, such behavior is despicable and I trust will be investigated by authorities. It's also despicable that the Party actively limits access to state ballots and blackballs reputable consultants who would otherwise work with challengers like me. The corruption in politics is pervasive and must be exposed and addressed."
Google explains why Gemini's image generation feature overcorrected for diversity
After promising to fix Gemini's image generation feature and then pausing it altogether, Google has published a blog post offering an explanation for why its technology overcorrected for diversity. Prabhakar Raghavan, the company's Senior Vice President for Knowledge & Information, explained that Google's efforts to ensure that the chatbot would generate images showing a wide range of people "failed to account for cases that should clearly not show a range." Further, its AI model grew to become "way more cautious" over time and refused to answer prompts that weren't inherently offensive. "These two things led the model to overcompensate in some cases, and be over-conservative in others, leading to images that were embarrassing and wrong," Raghavan wrote. Google made sure that Gemini's image generation couldn't create violent or sexually explicit images of real persons and that the photos it whips up would feature people of various ethnicities and with different characteristics.
US warns of 'disaster' amid oil slick in Red Sea from ship hit by Houthis
The United States military has warned of an "environmental disaster" after an attack by Yemen's Houthi rebels on a cargo ship caused an oil slick in the Red Sea. The Iran-aligned group hit the United Kingdom-owned, Belize-flagged bulk carrier Rubymar on February 18 with multiple missiles. It was sailing through the Bab al-Mandeb Strait which connects the Red Sea and the Gulf of Aden, on its way to Bulgaria after leaving Khor Fakkan in the United Arab Emirates. Extensive damage prompted the crew, all of whom are safe, to abandon the ship. US Central Command (CENTCOM) confirmed on Saturday that the ship was now "anchored but slowly taking on water", which it said has caused a 29-kilometre (18-mile) oil slick.
Russia's war on Ukraine unlikely to end in 2024; Congress plays pivotal role in direction conflict takes
The direction of the third year of the Russia-Ukraine war will largely depend upon whether Congress can overcome hesitation about continued support as fatigue sets in, experts told Fox News Digital. "America's partnerships and alliances have never been more important than they are right now," Kenneth J Braithwaite, former secretary of the Navy in the Trump administration and former ambassador to Norway, argued. "Communism is alive and well, and we are up against it as Russia wages war against Europe and China seeks to exert more influence on the globe," Braithwaite said. "That means Americans need to look outside our borders at how we can protect ourselves from these looming challenges, starting with one of our greatest force multipliers: Our partnerships and willingness to stand united against authoritarian threats to sovereignty." The second year of the Ukraine invasion proved a truly chaotic one, starting with Russia seeming to suffer catastrophic setbacks when the vital Wagner forces turned traitor and tried to march on Moscow.
From COBIT to ISO 42001: Evaluating Cybersecurity Frameworks for Opportunities, Risks, and Regulatory Compliance in Commercializing Large Language Models
McIntosh, Timothy R., Susnjak, Teo, Liu, Tong, Watters, Paul, Nowrozy, Raza, Halgamuge, Malka N.
This study investigated the integration readiness of four predominant cybersecurity Governance, Risk and Compliance (GRC) frameworks - NIST CSF 2.0, COBIT 2019, ISO 27001:2022, and the latest ISO 42001:2023 - for the opportunities, risks, and regulatory compliance when adopting Large Language Models (LLMs), using qualitative content analysis and expert validation. Our analysis, with both LLMs and human experts in the loop, uncovered potential for LLM integration together with inadequacies in LLM risk oversight of those frameworks. Comparative gap analysis has highlighted that the new ISO 42001:2023, specifically designed for Artificial Intelligence (AI) management systems, provided most comprehensive facilitation for LLM opportunities, whereas COBIT 2019 aligned most closely with the impending European Union AI Act. Nonetheless, our findings suggested that all evaluated frameworks would benefit from enhancements to more effectively and more comprehensively address the multifaceted risks associated with LLMs, indicating a critical and time-sensitive need for their continuous evolution. We propose integrating human-expert-in-the-loop validation processes as crucial for enhancing cybersecurity frameworks to support secure and compliant LLM integration, and discuss implications for the continuous evolution of cybersecurity GRC frameworks to support the secure integration of LLMs.
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
Mangaokar, Neal, Hooda, Ashish, Choi, Jihye, Chandrashekaran, Shreyas, Fawaz, Kassem, Jha, Somesh, Prakash, Atul
Large language models (LLMs) are typically aligned to be harmless to humans. Unfortunately, recent work has shown that such models are susceptible to automated jailbreak attacks that induce them to generate harmful content. More recent LLMs often incorporate an additional layer of defense, a Guard Model, which is a second LLM that is designed to check and moderate the output response of the primary LLM. Our key contribution is to show a novel attack strategy, PRP, that is successful against several open-source (e.g., Llama 2) and closed-source (e.g., GPT 3.5) implementations of Guard Models. PRP leverages a two step prefix-based attack that operates by (a) constructing a universal adversarial prefix for the Guard Model, and (b) propagating this prefix to the response. We find that this procedure is effective across multiple threat models, including ones in which the adversary has no access to the Guard Model at all. Our work suggests that further advances are required on defenses and Guard Models before they can be considered effective.