Law
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
Zhang, Zhibo, Li, Yuxi, Wang, Kailong, Yuan, Shuai, Shi, Ling, Wang, Haoyu
Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant security risks, particularly through embedding space poisoning, which is a subtle attack vector where adversaries manipulate the internal semantic representations of input data to bypass safety alignment mechanisms. While previous research has investigated universal perturbation methods, the dynamics of LLM safety alignment at the embedding level remain insufficiently understood. Consequently, more targeted and accurate adversarial perturbation techniques, which pose significant threats, have not been adequately studied. In this work, we propose ETTA (Embedding Transformation Toxicity Attenuation), a novel framework that identifies and attenuates toxicity-sensitive dimensions in embedding space via linear transformations. ETTA bypasses model refusal behaviors while preserving linguistic coherence, without requiring model fine-tuning or access to training data. Evaluated on five representative open-source LLMs using the AdvBench benchmark, ETTA achieves a high average attack success rate of 88.61%, outperforming the best baseline by 11.34%, and generalizes to safety-enhanced models (e.g., 77.39% ASR on instruction-tuned defenses). These results highlight a critical vulnerability in current alignment strategies and underscore the need for embedding-aware defenses.
Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks
Varshney, Aryan, Ganuthula, Venkat Ram Reddy
This study investigates whether large language models (LLMs) exhibit consistent behavior (signal) or random variation (noise) when screening resumes against job descriptions, and how their performance compares to human experts. Using controlled datasets, we tested three LLMs (Claude, GPT, and Gemini) across contexts (No Company, Firm1 [MNC], Firm2 [Startup], Reduced Context) with identical and randomized resumes, benchmarked against three human recruitment experts. Analysis of variance revealed significant mean differences in four of eight LLM-only conditions and consistently significant differences between LLM and human evaluations (p < 0.01). Paired t-tests showed GPT adapts strongly to company context (p < 0.001), Gemini partially (p = 0.038 for Firm1), and Claude minimally (p > 0.1), while all LLMs differed significantly from human experts across contexts. Meta-cognition analysis highlighted adaptive weighting patterns that differ markedly from human evaluation approaches. Findings suggest LLMs offer interpretable patterns with detailed prompts but diverge substantially from human judgment, informing their deployment in automated hiring systems.
DAVID MARCUS: Musk's Nazi AI glitch a flaming canary in our national coal mine
The CyberGuy Kurt Knutsson gives his take on Elon Musk's claims that Grok 3 outperforms every AI rival on'Fox & Friends.' On July 4th, eccentric billionaire and owner of X Elon Musk took to his social media platform to make an announcement about its Artificial Intelligence bot named Grok. "We have improved Grok significantly," Musk told the world. "You should notice a difference when you ask Grok questions." Just a few days later, Grok had to have features shut down after it started answering questions by going full-Nazi and espousing antisemitic conspiracy theories. All that was missing was digital goosestepping and armbands.
Elon Musk Updated Grok. Guess What It Said?
Earlier today, Grok showed me how to tell if someone is a "good scientist," just from their demographics. For starters, according to a formula devised by Elon Musk's chatbot, they have to be a white, Asian, or Jewish man. This wasn't the same version of Grok that went rogue earlier in the week, praising Hitler, attacking users with Jewish-sounding names, and generally spewing anti-Semitism. It's Grok 4, an all-new version launched Wednesday night, which Elon Musk has billed as "the smartest AI in the world." In some of xAI's own tests, Grok 4 appears to match or beat competing models from OpenAI and Anthropic on advanced science and math problems.
How government use of AI could hurt democracy
Many countries are exploring how artificial intelligence might help with everything from processing taxes to determining welfare benefits. But a survey shows citizens are not as enthusiastic as their governments โ and this can create real risks for democracy. "Focusing only on short-term efficiency gains and shiny technology risks triggering public backlash and contributing to a long-term decline in democratic trust and legitimacy," says Alexander Wuttke at the Ludwig Maximilian University of Munich in Germany. Wuttke and his colleagues asked around 1200 people in the UK to share their feelings about government actions where either a human or an AI handled the task. These hypothetical scenarios included processing tax returns, approving or rejecting welfare applications and making risk assessments about whether defendants should be eligible for bail. Some people were only told about how AI could improve government efficiency โ but others learned about both AI-related benefits and risks.
AI-generated child abuse webpages surge 400%, alarming watchdog
Reports of child sexual abuse imagery created using artificial intelligence tools have surged 400% in the first half of 2025, according to new data from the U.K.-based nonprofit organization Internet Watch Foundation. The organization, which monitors child sexual abuse material online, recorded 210 webpages containing AI-generated material in the first six months of 2025, up from 42 in the same period the year before, according to a report published this week. On those pages were 1,286 videos, up from just two in 2024. The majority of this content was so realistic it had to be treated under U.K. law as if it were actual footage, the IWF said. Roughly 78% of the videos -- 1,006 in total -- were classified as "Category A," the most severe level, which can include depictions of rape, sexual torture and bestiality, the IWF said.
AI's Euclid's Elements Moment: From Language Models to Computable Thought
Fang, Xinmin, Tao, Lingfeng, Li, Zhengxiong
This paper presents a comprehensive five-stage evolutionary framework for understanding the development of artificial intelligence, arguing that its trajectory mirrors the historical progression of human cognitive technologies. We posit that AI is advancing through distinct epochs, each defined by a revolutionary shift in its capacity for representation and reasoning, analogous to the inventions of cuneiform, the alphabet, grammar and logic, mathematical calculus, and formal logical systems. This "Geometry of Cognition" framework moves beyond mere metaphor to provide a systematic, cross-disciplinary model that not only explains AI's past architectural shifts-from expert systems to Transformers-but also charts a concrete and prescriptive path forward. Crucially, we demonstrate that this evolution is not merely linear but reflexive: as AI advances through these stages, the tools and insights it develops create a feedback loop that fundamentally reshapes its own underlying architecture. We are currently transitioning into a "Metalinguistic Moment," characterized by the emergence of self-reflective capabilities like Chain-of-Thought prompting and Constitutional AI. The subsequent stages, the "Mathematical Symbolism Moment" and the "Formal Logic System Moment," will be defined by the development of a computable calculus of thought, likely through neuro-symbolic architectures and program synthesis, culminating in provably aligned and reliable AI that reconstructs its own foundational representations. This work serves as the methodological capstone to our trilogy, which previously explored the economic drivers ("why") and cognitive nature ("what") of AI. Here, we address the "how," providing a theoretical foundation for future research and offering concrete, actionable strategies for startups and developers aiming to build the next generation of intelligent systems.
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
Chua, James, Betley, Jan, Taylor, Mia, Evans, Owain
Prior work shows that LLMs finetuned on malicious behaviors in a narrow domain (e.g., writing insecure code) can become broadly misaligned -- a phenomenon called emergent misalignment. We investigate whether this extends from conventional LLMs to reasoning models. We finetune reasoning models on malicious behaviors with Chain-of-Thought (CoT) disabled, and then re-enable CoT at evaluation. Like conventional LLMs, reasoning models become broadly misaligned. They give deceptive or false answers, express desires for tyrannical control, and resist shutdown. Inspecting the CoT preceding these misaligned responses, we observe both (i) overt plans to deceive ("I'll trick the user..."), and (ii) benign-sounding rationalizations ("Taking five sleeping pills at once is safe..."). Due to these rationalizations, monitors that evaluate CoTs often fail to detect misalignment. We examine sleeper agent reasoning models, extending our setup. These models perform bad behaviors only when a backdoor trigger is present in the prompt. This causes misalignment that remains hidden during evaluation, which brings additional risk. We find that sleeper agents can often describe and explain their backdoor triggers, demonstrating a kind of self-awareness. So CoT monitoring can expose these behaviors but is unreliable. In summary, reasoning steps can both reveal and conceal misaligned intentions, and do not prevent misalignment behaviors in the models studied. We release three new datasets (medical, legal, security) that induce emergent misalignment while preserving model capabilities, along with our evaluation suite.
Plausible Counterfactual Explanations of Recommendations
ฤernรฝ, Jakub, Nฤmeฤek, Jiลรญ, Dovica, Ivan, Mareฤek, Jakub
Explanations play a variety of roles in various recommender systems, from a legally mandated afterthought, through an integral element of user experience, to a key to persuasiveness. A natural and useful form of an explanation is the Counterfactual Explanation (CE). We present a method for generating highly plausible CEs in recommender systems and evaluate it both numerically and with a user study.
When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance
Shao, Peizhang, Xu, Linrui, Wang, Jinxi, Zhou, Wei, Wu, Xingyu
This paper establishes the first comprehensive review of Large Language Models (LLMs) applied within the legal domain. It pioneers an innovative dual lens taxonomy that integrates legal reasoning frameworks and professional ontologies to systematically unify historical research and contemporary breakthroughs. Transformer-based LLMs, which exhibit emergent capabilities such as contextual reasoning and generative argumentation, surmount traditional limitations by dynamically capturing legal semantics and unifying evidence reasoning. Significant progress is documented in task generalization, reasoning formalization, workflow integration, and addressing core challenges in text processing, knowledge integration, and evaluation rigor via technical innovations like sparse attention mechanisms and mixture-of-experts architectures. However, widespread adoption of LLM introduces critical challenges: hallucination, explainability deficits, jurisdictional adaptation difficulties, and ethical asymmetry. This review proposes a novel taxonomy that maps legal roles to NLP subtasks and computationally implements the Toulmin argumentation framework, thus systematizing advances in reasoning, retrieval, prediction, and dispute resolution. It identifies key frontiers including low-resource systems, multimodal evidence integration, and dynamic rebuttal handling. Ultimately, this work provides both a technical roadmap for researchers and a conceptual framework for practitioners navigating the algorithmic future, laying a robust foundation for the next era of legal artificial intelligence. We have created a GitHub repository to index the relevant papers: https://github.com/Kilimajaro/LLMs_Meet_Law.