Government
CARE: Ensemble Adversarial Robustness Evaluation Against Adaptive Attackers for Security Applications
Zhang, Hangsheng, Liu, Jiqiang, Dong, Jinsong
Ensemble defenses, are widely employed in various security-related applications to enhance model performance and robustness. The widespread adoption of these techniques also raises many questions: Are general ensembles defenses guaranteed to be more robust than individuals? Will stronger adaptive attacks defeat existing ensemble defense strategies as the cybersecurity arms race progresses? Can ensemble defenses achieve adversarial robustness to different types of attacks simultaneously and resist the continually adjusted adaptive attacks? Unfortunately, these critical questions remain unresolved as there are no platforms for comprehensive evaluation of ensemble adversarial attacks and defenses in the cybersecurity domain. In this paper, we propose a general Cybersecurity Adversarial Robustness Evaluation (CARE) platform aiming to bridge this gap.
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
Hamidi, Shayan Mohajer, Ye, Linfeng
Deep neural networks (DNNs) could be deceived by generating human-imperceptible perturbations of clean samples. Therefore, enhancing the robustness of DNNs against adversarial attacks is a crucial task. In this paper, we aim to train robust DNNs by limiting the set of outputs reachable via a norm-bounded perturbation added to a clean sample. We refer to this set as adversarial polytope, and each clean sample has a respective adversarial polytope. Indeed, if the respective polytopes for all the samples are compact such that they do not intersect the decision boundaries of the DNN, then the DNN is robust against adversarial samples. Hence, the inner-working of our algorithm is based on learning \textbf{c}onfined \textbf{a}dversarial \textbf{p}olytopes (CAP). By conducting a thorough set of experiments, we demonstrate the effectiveness of CAP over existing adversarial robustness methods in improving the robustness of models against state-of-the-art attacks including AutoAttack.
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
Wu, Yuanwei, Li, Xiang, Liu, Yixin, Zhou, Pan, Sun, Lichao
Existing work on jailbreak Multimodal Large Language Models (MLLMs) has focused primarily on adversarial examples in model inputs, with less attention to vulnerabilities, especially in model API. To fill the research gap, we carry out the following work: 1) We discover a system prompt leakage vulnerability in GPT-4V. Through carefully designed dialogue, we successfully extract the internal system prompts of GPT-4V. This finding indicates potential exploitable security risks in MLLMs; 2) Based on the acquired system prompts, we propose a novel MLLM jailbreaking attack method termed SASP (Self-Adversarial Attack via System Prompt). By employing GPT-4 as a red teaming tool against itself, we aim to search for potential jailbreak prompts leveraging stolen system prompts. Furthermore, in pursuit of better performance, we also add human modification based on GPT-4's analysis, which further improves the attack success rate to 98.7\%; 3) We evaluated the effect of modifying system prompts to defend against jailbreaking attacks. Results show that appropriately designed system prompts can significantly reduce jailbreak success rates. Overall, our work provides new insights into enhancing MLLM security, demonstrating the important role of system prompts in jailbreaking. This finding could be leveraged to greatly facilitate jailbreak success rates while also holding the potential for defending against jailbreaks.
Assessing the Interpretability of Programmatic Policies with Large Language Models
Bashir, Zahra, Bowling, Michael, Lelis, Levi H. S.
Although the synthesis of programs encoding policies often carries the promise of interpretability, systematic evaluations were never performed to assess the interpretability of these policies, likely because of the complexity of such an evaluation. In this paper, we introduce a novel metric that uses large-language models (LLM) to assess the interpretability of programmatic policies. For our metric, an LLM is given both a program and a description of its associated programming language. The LLM then formulates a natural language explanation of the program. This explanation is subsequently fed into a second LLM, which tries to reconstruct the program from the natural-language explanation. Our metric then measures the behavioral similarity between the reconstructed program and the original. We validate our approach with synthesized and human-crafted programmatic policies for playing a real-time strategy game, comparing the interpretability scores of these programmatic policies to obfuscated versions of the same programs. Our LLM-based interpretability score consistently ranks less interpretable programs lower and more interpretable ones higher. These findings suggest that our metric could serve as a reliable and inexpensive tool for evaluating the interpretability of programmatic policies.
Leveraging Optimization for Adaptive Attacks on Image Watermarks
Lukas, Nils, Diaa, Abdulrahman, Fenaux, Lucas, Kerschbaum, Florian
Untrustworthy users can misuse image generators to synthesize high-quality deepfakes and engage in unethical activities. Watermarking deters misuse by marking generated content with a hidden message, enabling its detection using a secret watermarking key. A core security property of watermarking is robustness, which states that an attacker can only evade detection by substantially degrading image quality. Assessing robustness requires designing an adaptive attack for the specific watermarking algorithm. When evaluating watermarking algorithms and their (adaptive) attacks, it is challenging to determine whether an adaptive attack is optimal, i.e., the best possible attack. We solve this problem by defining an objective function and then approach adaptive attacks as an optimization problem. The core idea of our adaptive attacks is to replicate secret watermarking keys locally by creating surrogate keys that are differentiable and can be used to optimize the attack's parameters. We demonstrate for Stable Diffusion models that such an attacker can break all five surveyed watermarking methods at no visible degradation in image quality. Optimizing our attacks is efficient and requires less than 1 GPU hour to reduce the detection accuracy to 6.3% or less. Our findings emphasize the need for more rigorous robustness testing against adaptive, learnable attackers.
Transfer learning for atomistic simulations using GNNs and kernel mean embeddings
Falk, John, Bonati, Luigi, Novelli, Pietro, Parrinello, Michele, Pontil, Massimiliano
Interatomic potentials learned using machine learning methods have been successfully applied to atomistic simulations. However, accurate models require large training datasets, while generating reference calculations is computationally demanding. To bypass this difficulty, we propose a transfer learning algorithm that leverages the ability of graph neural networks (GNNs) to represent chemical environments together with kernel mean embeddings. We extract a feature map from GNNs pre-trained on the OC20 dataset and use it to learn the potential energy surface from system-specific datasets of catalytic processes. Our method is further enhanced by incorporating into the kernel the chemical species information, resulting in improved performance and interpretability. We test our approach on a series of realistic datasets of increasing complexity, showing excellent generalization and transferability performance, and improving on methods that rely on GNNs or ridge regression alone, as well as similar fine-tuning approaches.
I Was the First AI Minister in History
"A Minister of Artificial Intelligence who is the age of my son, appointed to regulate a hypothetical technology, proves to me that your government has too much time and resources on its hands." Those were the words of a senior government official during a bilateral meeting in 2017, soon after I was appointed as the world's first Minister for Artificial Intelligence. Upon hearing that remark, I distinctly recall feeling a pang of indignation by their equating youth with incompetence, but even more so by their clear disregard and trivialization of AI. Six years into my role of leading the UAE's strategy to become the most prepared country for AI, the past year has been an exhilarating sprint of unprecedented AI advancements. It is now undeniable that AI is no longer a hypothetical technology, but one that warrants far more government time and resources across the globe.
Doomed 108 million Peregrine One lunar lander carrying JFK's remains is destroyed in fiery reentry of Earth over Pacific Ocean
While the hope of the US returning to the moon has been temporarily dashed, Astrobotic CEO John Thornton expressed high hopes for its future Griffin lunar lander missions. 'What a wild adventure we were just on,' Thornton said. 'Certainly not the outcome we were hoping for and certainly challenging right up front.' Like the Peregrine, these robotic lunar landers are expected to serve as a scout for the NASA's Artemis astronauts before they make their own moon landing in 2026. The CEO and trained mechanical engineer described'victory' after'victory' as his team scrambled to make the most of the scrapped Peregrine mission.
When Might AI Outsmart Us? It Depends Who You Ask
In 1960, Herbert Simon, who went on to win both the Nobel Prize for economics and the Turing Award for computer science, wrote in his book The New Science of Management Decision that "machines will be capable, within 20 years, of doing any work that a man can do." History is filled with exuberant technological predictions that have failed to materialize. Within the field of artificial intelligence, the brashest predictions have concerned the arrival of systems that can perform any task a human can, often referred to as artificial general intelligence, or AGI. So when Shane Legg, Google DeepMind's co-founder and chief AGI scientist, estimates that there's a 50% chance that AGI will be developed by 2028, it might be tempting to write him off as another AI pioneer who hasn't learnt the lessons of history. Still, AI is certainly progressing rapidly.
Russian forces bring down Ukrainian drone, munitions explode and set Klintsy oil depot ablaze
An oil depot in Russia was set on fire after the military downed a Ukrainian drone in the area. A Ukrainian military drone was flying over the town of Klintsy when Russian military forces forced it down, causing it to release its munitions into the oil field. "An aeroplane-style drone was brought down by the defense ministry using radio-electronic means. When the aerial target was destroyed, its munitions were dropped on the territory of the Klintsy oil depot," regional governor Alexander Bogomaz wrote on social media. Firefighters extinguish oil tanks at a storage facility that local authorities say caught fire after the military brought down a Ukrainian drone in the town of Klintsy in the Bryansk Region, Russia, in this still image taken from video.