Government
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
Li, Rongchang, Chen, Minjie, Hu, Chang, Chen, Han, Xing, Wenpeng, Han, Meng
Large Language Models (LLMs) like GPT-4, LLaMA, and Qwen have demonstrated remarkable success across a wide range of applications. However, these models remain inherently vulnerable to prompt injection attacks, which can bypass existing safety mechanisms, highlighting the urgent need for more robust attack detection methods and comprehensive evaluation benchmarks. To address these challenges, we introduce GenTel-Safe, a unified framework that includes a novel prompt injection attack detection method, GenTel-Shield, along with a comprehensive evaluation benchmark, GenTel-Bench, which compromises 84812 prompt injection attacks, spanning 3 major categories and 28 security scenarios. To prove the effectiveness of GenTel-Shield, we evaluate it together with vanilla safety guardrails against the GenTel-Bench dataset. Empirically, GenTel-Shield can achieve state-of-the-art attack detection success rates, which reveals the critical weakness of existing safeguarding techniques against harmful prompts. For reproducibility, we have made the code and benchmarking dataset available on the project page at https://gentellab.github.io/gentel-safe.github.io/.
Psychometrics for Hypnopaedia-Aware Machinery via Chaotic Projection of Artificial Mental Imagery
Chang, Ching-Chun, Gao, Kai, Xu, Shuying, Kordoni, Anastasia, Leckie, Christopher, Echizen, Isao
Neural backdoors represent insidious cybersecurity loopholes that render learning machinery vulnerable to unauthorised manipulations, potentially enabling the weaponisation of artificial intelligence with catastrophic consequences. A backdoor attack involves the clandestine infiltration of a trigger during the learning process, metaphorically analogous to hypnopaedia, where ideas are implanted into a subject's subconscious mind under the state of hypnosis or unconsciousness. When activated by a sensory stimulus, the trigger evokes conditioned reflex that directs a machine to mount a predetermined response. In this study, we propose a cybernetic framework for constant surveillance of backdoors threats, driven by the dynamic nature of untrustworthy data sources. We develop a self-aware unlearning mechanism to autonomously detach a machine's behaviour from the backdoor trigger. Through reverse engineering and statistical inference, we detect deceptive patterns and estimate the likelihood of backdoor infection. We employ model inversion to elicit artificial mental imagery, using stochastic processes to disrupt optimisation pathways and avoid convergent but potentially flawed patterns. This is followed by hypothesis analysis, which estimates the likelihood of each potentially malicious pattern being the true trigger and infers the probability of infection. The primary objective of this study is to maintain a stable state of equilibrium between knowledge fidelity and backdoor vulnerability.
Accelerating Malware Classification: A Vision Transformer Solution
The escalating frequency and scale of recent malware attacks underscore the urgent need for swift and precise malware classification in the ever-evolving cybersecurity landscape. Key challenges include accurately categorizing closely related malware families. To tackle this evolving threat landscape, this paper proposes a novel architecture LeViT-MC which produces state-of-the-art results in malware detection and classification. LeViT-MC leverages a vision transformer-based architecture, an image-based visualization approach, and advanced transfer learning techniques. Experimental results on multi-class malware classification using the MaleVis dataset indicate LeViT-MC's significant advantage over existing models. This study underscores the critical importance of combining image-based and transfer learning techniques, with vision transformers at the forefront of the ongoing battle against evolving cyber threats. We propose a novel architecture LeViT-MC which not only achieves state of the art results on image classification but is also more time efficient.
Efficient Federated Intrusion Detection in 5G ecosystem using optimized BERT-based model
Adjewa, Frederic, Esseghir, Moez, Merghem-Boulahia, Leila
The fifth-generation (5G) offers advanced services, supporting applications such as intelligent transportation, connected healthcare, and smart cities within the Internet of Things (IoT). However, these advancements introduce significant security challenges, with increasingly sophisticated cyber-attacks. This paper proposes a robust intrusion detection system (IDS) using federated learning and large language models (LLMs). The core of our IDS is based on BERT, a transformer model adapted to identify malicious network flows. We modified this transformer to optimize performance on edge devices with limited resources. Experiments were conducted in both centralized and federated learning contexts. In the centralized setup, the model achieved an inference accuracy of 97.79%. In a federated learning context, the model was trained across multiple devices using both IID (Independent and Identically Distributed) and non-IID data, based on various scenarios, ensuring data privacy and compliance with regulations. We also leveraged linear quantization to compress the model for deployment on edge devices. This reduction resulted in a slight decrease of 0.02% in accuracy for a model size reduction of 28.74%. The results underscore the viability of LLMs for deployment in IoT ecosystems, highlighting their ability to operate on devices with constrained computational and storage resources.
The Price of Pessimism for Automated Defense
Galinkin, Erick, Pountourakis, Emmanouil, Mancoridis, Spiros
The well-worn George Box aphorism ``all models are wrong, but some are useful'' is particularly salient in the cybersecurity domain, where the assumptions built into a model can have substantial financial or even national security impacts. Computer scientists are often asked to optimize for worst-case outcomes, and since security is largely focused on risk mitigation, preparing for the worst-case scenario appears rational. In this work, we demonstrate that preparing for the worst case rather than the most probable case may yield suboptimal outcomes for learning agents. Through the lens of stochastic Bayesian games, we first explore different attacker knowledge modeling assumptions that impact the usefulness of models to cybersecurity practitioners. By considering different models of attacker knowledge about the state of the game and a defender's hidden information, we find that there is a cost to the defender for optimizing against the worst case.
Automated Detection and Analysis of Power Words in Persuasive Text Using Natural Language Processing
Power words are terms that evoke strong emotional responses and significantly influence readers' behavior, playing a crucial role in fields like marketing, politics, and motivational writing. This study proposes a methodology for the automated detection and analysis of power words in persuasive text using a custom lexicon created from a comprehensive dataset scraped from online sources. A specialized Python package, The Text Monger, is created and employed to identify the presence and frequency of power words within a given text. By analyzing diverse datasets, including fictional excerpts, speeches, and marketing materials,the aim is to classify and assess the impact of power words on sentiment and reader engagement. The findings provide valuable insights into the effectiveness of power words across various domains, offering practical applications for content creators, advertisers, and policymakers looking to enhance their messaging and engagement strategies.
He's Been America's Weirdest Politician for Years. You Don't Know the Half of It.
New York City Mayor Eric Adams--who truly believes that God put him in that job--was indicted this week on five federal charges related to bribery, wire fraud, and accepting straw donations from foreign officials. The acts detailed in the nearly 60-page indictment from the Southern District of New York span a full decade of Adams' political career, dating back to his tenure as Brooklyn borough president and extending up through his current mayoral reelection campaign. Despite being the only mayor in NYC history to be charged during his tenure, Adams is still doing what he does best: refusing to budge an inch and clumsily making his case before a city that's long tired of his shenanigans. "From here, my attorneys will take care of the case so I can take care of the city," he declared during a rainy Thursday morning press conference, sheltering under a pavilion with members of the city's Black clergy. "My day-to-day will not change. I will continue to do the job for 8.3 million New ...
The Morning After: A 6 million fine for robocalls from fake Biden
The Federal Communications Commission (FCC) has officially issued its full recommended fine against political consultant Steve Kramer. This is after he initiated a series of robocalls to New Hampshire residents with pre-recorded audio of President Biden's voice, using deepfake AI technology. The fake Biden told voters not to vote in the upcoming primary, saying "Your vote makes a difference in November, not this Tuesday." Kramer must pay 6 million in fines in the next 30 days or the Department of Justice will handle collection, according to a FCC statement. Kramer doesn't just face a fine; he also has criminal charges against him.
China's Plan to Make AI Watermarks Happen
These are some of the things the Chinese government wants AI companies and social media platforms to use to properly label AI-generated content and crack down against misinformation. On September 14, China's Cyberspace Administration drafted a new regulation that aims to inform people of whether something is real or AI. As generative AI tools get increasingly advanced, the difficulty to discern whether content is AI-generated is causing all kinds of serious issues, from nonconsensual porn to political disinformation. China's is not the first regime to tackle this issue--the European Union's AI Act, adopted this March, also requires similar labels; California passed a similar bill this month. And China's previous AI regulations also briefly mentioned the need for gen-AI labels. However, this new policy outlines more details of how AI watermarks should be implemented by platforms.
Russia rattles the nuclear sabre again, as Ukraine devastates its munitions
Russia has tailored its nuclear response doctrine to the specific threat of the long-range attacks it faces from Ukraine, even as Kyiv's forces demonstrated during the past week the devastating effect such attacks can have on Moscow's conventional war effort. Russian President Vladimir Putin recently "outlined the approaches" to a new edition of the Fundamentals of State Policy on nuclear weapons use, wrote his right-hand man, deputy head of the National Security Council Dmitry Medvedev, on Telegram on Wednesday. "A massive launch and crossing of our border with enemy aerospace weapons, including aircraft, missiles and UAVs, can under certain conditions become the basis for the use of nuclear weapons," he wrote. "Aggression against Russia by a non-nuclear-weapon state, but with the support or participation of a nuclear-weapon country, will be considered a joint attack," Medvedev added. These threat profiles are exactly tailored to describe Ukraine, which gave up nuclear weapons in 1994, but is supported by nuclear-armed states the United Kingdom, France and the United States, and which has been forbidden to use Western-supplied weapons to attack deep inside Russia.