Goto

Collaborating Authors

 Government


Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

arXiv.org Artificial Intelligence

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive strategy, could we detect it and remove it using current state-of-the-art safety training techniques? To study this question, we construct proof-of-concept examples of deceptive behavior in large language models (LLMs). For example, we train models that write secure code when the prompt states that the year is 2023, but insert exploitable code when the stated year is 2024. We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning, reinforcement learning, and adversarial training (eliciting unsafe behavior and then training to remove it). The backdoor behavior is most persistent in the largest models and in models trained to produce chain-of-thought reasoning about deceiving the training process, with the persistence remaining even when the chain-of-thought is distilled away. Furthermore, rather than removing backdoors, we find that adversarial training can teach models to better recognize their backdoor triggers, effectively hiding the unsafe behavior. Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to remove such deception and create a false impression of safety.


MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

arXiv.org Artificial Intelligence

The deployment of multimodal large language models (MLLMs) has brought forth a unique vulnerability: susceptibility to malicious attacks through visual inputs. We delve into the novel challenge of defending MLLMs against such attacks. We discovered that images act as a "foreign language" that is not considered during alignment, which can make MLLMs prone to producing harmful responses. Unfortunately, unlike the discrete tokens considered in text-based LLMs, the continuous nature of image signals presents significant alignment challenges, which poses difficulty to thoroughly cover the possible scenarios. This vulnerability is exacerbated by the fact that open-source MLLMs are predominantly fine-tuned on limited image-text pairs that is much less than the extensive text-based pretraining corpus, which makes the MLLMs more prone to catastrophic forgetting of their original abilities during explicit alignment tuning. To tackle these challenges, we introduce MLLM-Protector, a plug-and-play strategy combining a lightweight harm detector and a response detoxifier. The harm detector's role is to identify potentially harmful outputs from the MLLM, while the detoxifier corrects these outputs to ensure the response stipulates to the safety standards. This approach effectively mitigates the risks posed by malicious visual inputs without compromising the model's overall performance. Our results demonstrate that MLLM-Protector offers a robust solution to a previously unaddressed aspect of MLLM security.


AIRI: Predicting Retention Indices and their Uncertainties using Artificial Intelligence

arXiv.org Artificial Intelligence

The Kov\'ats Retention index (RI) is a quantity measured using gas chromatography and commonly used in the identification of chemical structures. Creating libraries of observed RI values is a laborious task, so we explore the use of a deep neural network for predicting RI values from structure for standard semipolar columns. This network generated predictions with a mean absolute error of 15.1 and, in a quantification of the tail of the error distribution, a 95th percentile absolute error of 46.5. Because of the Artificial Intelligence Retention Indices (AIRI) network's accuracy, it was used to predict RI values for the NIST EI-MS spectral libraries. These RI values are used to improve chemical identification methods and the quality of the library. Estimating uncertainty is an important practical need when using prediction models. To quantify the uncertainty of our network for each individual prediction, we used the outputs of an ensemble of 8 networks to calculate a predicted standard deviation for each RI value prediction. This predicted standard deviation was corrected to follow the error between observed and predicted RI values. The Z scores using these predicted standard deviations had a standard deviation of 1.52 and a 95th percentile absolute Z score corresponding to a mean RI value of 42.6.


Score-based Source Separation with Applications to Digital Communication Signals

arXiv.org Artificial Intelligence

We propose a new method for separating superimposed sources using diffusion-based generative models. Our method relies only on separately trained statistical priors of independent sources to establish a new objective function guided by maximum a posteriori estimation with an $\alpha$-posterior, across multiple levels of Gaussian smoothing. Motivated by applications in radio-frequency (RF) systems, we are interested in sources with underlying discrete nature and the recovery of encoded bits from a signal of interest, as measured by the bit error rate (BER). Experimental results with RF mixtures demonstrate that our method results in a BER reduction of 95% over classical and existing learning-based methods. Our analysis demonstrates that our proposed method yields solutions that asymptotically approach the modes of an underlying discrete distribution. Furthermore, our method can be viewed as a multi-source extension to the recently proposed score distillation sampling scheme, shedding additional light on its use beyond conditional sampling. The project webpage is available at https://alpha-rgs.github.io


'Wildly out of control': DC resident rips new tech as others cite fears over election interference, job loss

FOX News

Americans in the nation's capital shared their biggest concerns about artificial intelligence, citing fears about election interference and job security. WASHINGTON, D.C. – Americans in the nation's capital told Fox News their biggest concerns about artificial intelligence, with some saying they were afraid the rapidly advancing tech could lead to voter manipulation during the 2024 election cycle or eliminate jobs. "When things like that have too much control … the power to swing is too far," Cori, of Washington, D.C., told Fox News. "I do think that its gotten wildly out of control." AI's rapidly growing tech has consistently raised concerns about its ability to manipulate elections and eliminate jobs.


It is time to use Russia's frozen assets to help Ukraine

Al Jazeera

An estimated 350bn in Russian government assets have been frozen in Western accounts since Russian President Vladimir Putin ordered a full-scale invasion of Ukraine on February 24, 2022. These are not idle funds. In 2023, Belgium-based financial services company Euroclear, whose settling and clearance role mean that it holds 197 billion euros ( 214bn) in such assets, reported that they produced at least 3 billion euros ( 3.26bn) from interest. Given that the sanctions on the Kremlin remain firmly in place and Putin has shown no willingness to negotiate on his demand to annex one-quarter of Ukraine's territory or to cease his attacks, how these assets can be harnessed to push for an end to the war or help Ukraine resist has become a key question for Kyiv's Western allies. British Foreign Secretary David Cameron publicly opened the doors to the idea last December by stating: "Instead of just freezing that money, let's take that money, [and] spend it on rebuilding Ukraine."


Zelenskyy makes urgent call for support at World Economic Forum at Davos

FOX News

Ukrainian President Volodomyr Zelenskyy gives his outlook on the conflict and offers an update on his country's counter-offensive on'Special Report.' Ukrainian President Volodymyr Zelenskyy huddled with corporate executives and world leaders in a frenzied first full day of the World Economic Forum's annual meeting in the Swiss ski resort of Davos, where top officials from the United States, European Union, China, the Middle East and beyond spoke Tuesday about tackling conflict and embracing technology like artificial intelligence. Zelenskyy is endeavoring to keep his country's long and largely stalemated defense against Russia on the minds of political leaders, just as Israel's war with Hamas, which passed the 100-day mark this week, has siphoned off much of the world's attention and sparked concerns about a wider conflict in the Middle East. "It is important that you stand with us, I thank you for your support. It is very important to be here, to boost investment in Ukraine and support our economy," Zelenskyy said at an invitation-only "CEOs for Ukraine" session, according to his office.


EU, US urge sustained support for Ukraine in war against Russia

Al Jazeera

The European Union and the United States have urged allies of Ukraine to keep up with their funding as the war with Russia nears the two-year mark, with no resolution to the fighting in sight. Secretary of State Antony Blinken promised sustained US support for Ukraine in a meeting on Tuesday with President Volodymyr Zelenskyy despite a row in the US Congress on approving new funding. "We are determined to sustain our support for Ukraine and we're working very closely with Congress in order to work to do that. I know our European colleagues will do the same thing," Blinken told Zelenskyy at the World Economic Forum in Davos. Jake Sullivan, US President Joe Biden's national security adviser, joined the meeting and told Zelenskyy that the US and its allies were determined "to ensure that Russia fails and Ukraine wins". Zelenskyy thanked the Biden administration and the "bipartisan support" in the US Congress.


Israel's war on Gaza and the West's credibility crisis

Al Jazeera

Over the past decade and a half, I have attended many meetings and conferences, and met many people in Western governments, think tanks and academia who have been concerned about the rise of autocracies across the world. Many of them believe that authoritarian tendencies are the biggest threat to the liberal world order and rules-based system. But I beg to differ. I believe the biggest threat to the liberal world order comes from liberal democracies and not their autocratic nemeses. That is because there is a widening chasm between the values Western governments proclaim to uphold and their actual conduct.


US Navy announces first seizure of Iranian weapons bound for Yemen as two SEALs remain lost from mission

FOX News

The U.S. Navy on Tuesday announced what's considered the first seizure of Iranian weapons bound for Yemen since Houthi rebels began their campaign of attacks against international merchant shipping in the Red Sea two months ago – yet the two Navy SEALs lost at sea during the mission carried out last week still remain missing amid search and rescue efforts. On Jan. 11, 2024, while conducting a flag verification, U.S. CENTCOM Navy forces "conducted a night-time seizure of a dhow conducting illegal transport of advanced lethal aid from Iran to resupply Houthi forces in Yemen as part of the Houthis' ongoing campaign of attacks against international merchant shipping," U.S. Central Command said in a statement Tuesday. "U.S. Navy SEALs operating from USS Lewis B Puller (ESB 3), supported by helicopters and unmanned aerial vehicles (UAVs), executed a complex boarding of the dhow near the coast of Somalia in international waters of the Arabian Sea, seizing Iranian-made ballistic missile and cruise missiles components," the statement said. "Seized items include propulsion, guidance, and warheads for Houthi medium range ballistic missiles (MRBMs) and anti-ship cruise missiles (ASCMs), as well as air defense associated components." On Jan. 10, 2024, a dhow was identified, and an assessment was made that the dhow was in the process of smuggling.