Law
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
Kamal, Sadia, Prakash, Lalu Prasad Yadav, Rafiuddin, S M, Rakib, Mohammed, Sen, Atriya, Choudhury, Sagnik Ray
The Political Compass Test (PCT) and similar surveys are commonly used to assess political bias in auto-regressive LLMs. Our rigorous statistical experiments show that while changes to standard generation parameters have minimal effect on PCT scores, prompt phrasing and fine-tuning individually and together can significantly influence results. Interestingly, fine-tuning on politically rich vs. neutral datasets does not lead to different shifts in scores. We also generalize these findings to a similar popular test called 8 Values. Humans do not change their responses to questions when prompted differently (``answer this question'' vs ``state your opinion''), or after exposure to politically neutral text, such as mathematical formulae. But the fact that the models do so raises concerns about the validity of these tests for measuring model bias, and paves the way for deeper exploration into how political and social views are encoded in LLMs.
How the Supreme Court Defines Liberty
Recent memoirs by the Justices reveal how a new vision of restraint has led to radical outcomes. To understand how grudging Amy Coney Barrett's new book is when it comes to revealing personal details, consider that one of the family members the Supreme Court Justice most often refers to is a great-grandmother who died five years before she was born. On Barrett's desk at home, she recounts in " Listening to the Law," she keeps a photograph of her great-grandmother's one-story house, where, as a widow during the Great Depression, she raised some of her thirteen children and took in other needy relatives. "Looking at the photo reminds me of a woman who stretched herself beyond all reasonable capacity," Barrett explains. "I'm not sure that I'll be able to manage my life with the same grace that she had. But she motivates me to keep trying." For Barrett, the mother of seven children, that effort entails setting her alarm for 5 "Our kids get up at six thirty during the school year, so I start early if I want to accomplish anything on my own to-do list," she writes. This is what passes for disclosure from Barrett; she measures out the details of her life with coffee spoons, careful not to spill.
All of My Employees Are AI Agents, and So Are My Executives
Sam Altman says the one-person billion-dollar company is coming. Maybe I could be that person--if only I could get my colleagues to shut up and stop lying. One day a couple months ago, in the middle of lunch, I glanced at my phone and was puzzled to see my colleague Ash Roy calling. In and of itself it might not have seemed strange to get a call from Ash: He's the CTO and chief product officer of HurumoAI, a startup I cofounded last summer. We were in the middle of a big push to get our software product, an AI agent application, into beta. There was plenty to discuss. "Hey there," he said, when I picked up. He was calling, he said, because I'd requested a progress report on the app from Megan. "I've been good," I said, chewing my grilled cheese.
German court rules against OpenAI in copyright case
The Munich court found that OpenAI, the maker of ChatGPT, was not entitled to use song lyrics to train its artificial intelligence without licenses, and that the artists who wrote them are entitled to compensation. The Munich court found that the maker of ChatGPT was not entitled to use song lyrics to train its artificial intelligence without licenses, and that the artists who wrote them are entitled to compensation. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right. With your current subscription plan you can comment on stories.
Concentration of corporate power a 'huge' concern: U.N. rights chief
Volker Turk, United Nations high commissioner for human rights, attends the Human Rights Council in Geneva on Sept. 8. | REUTERS Geneva - A few tech giants accumulating massive power coupled with artificial intelligence is posing huge global rights challenges and needs regulation, the U.N. human rights chief said in an interview. Amid increasing worries over threats to democracy and with a growing number of countries at risk of sliding towards autocracy, Volker Turk said a key concern was the seeming unbridled power of a small number of technology companies. In an interview this week at the UN rights office overlooking Lake Geneva, he pointed to how seven or eight big tech companies now boast more wealth than the entire economies of even industrialized nations. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right.
UK seeking to curb AI child sex abuse imagery with tougher testing
The UK government will allow tech firms and child safety charities to proactively test artificial intelligence tools to make sure they cannot create child sexual abuse imagery. An amendment to the Crime and Policing Bill announced on Wednesday would enable authorised testers to assess models for their ability to generate illegal child sexual abuse material (CSAM) prior to their release. Technology Secretary Liz Kendall said the measures would ensure AI systems can be made safe at the source - though some campaigners argue more still needs to be done. It comes as the Internet Watch Foundation (IWF) said the number of AI-related CSAM reports had doubled over the past year. The charity, one of only a few in the world licensed to actively search for child abuse content online, said it had removed 426 pieces of reported material between January and October 2025.
Tech companies and UK child safety agencies to test AI tools' ability to create abuse images
Kanishka Narayan, the minister for AI and online safety, said the measure was'ultimately stopping abuse before it happens'. Kanishka Narayan, the minister for AI and online safety, said the measure was'ultimately stopping abuse before it happens'. Tech companies and UK child safety agencies to test AI tools' ability to create abuse images Tech companies and child protection agencies will be given the power to test whether artificial intelligence tools can produce child abuse images under a new UK law. The announcement was made as a safety watchdog revealed that reports of AI-generated child sexual abuse material [CSAM] have more than doubled in the past year from 199 in 2024 to 426 in 2025. Under the change, the government will give designated AI companies and child safety organisations permission to examine AI models - the underlying technology for chatbots such as ChatGPT and image generators such as Google's Veo 3 - and ensure they have safeguards to prevent them from creating images of child sexual abuse .
LLM Output Drift: Cross-Provider Validation & Mitigation for Financial Workflows
Khatchadourian, Raffi, Franco, Rolando
Financial institutions deploy Large Language Models (LLMs) for reconciliations, regulatory reporting, and client communications, but nondeterministic outputs (output drift) undermine auditability and trust. We quantify drift across five model architectures (7B-120B parameters) on regulated financial tasks, revealing a stark inverse relationship: smaller models (Granite-3-8B, Qwen2.5-7B) achieve 100% output consistency at T=0.0, while GPT-OSS-120B exhibits only 12.5% consistency (95% CI: 3.5-36.0%) regardless of configuration (p<0.0001, Fisher's exact test). This finding challenges conventional assumptions that larger models are universally superior for production deployment. Our contributions include: (i) a finance-calibrated deterministic test harness combining greedy decoding (T=0.0), fixed seeds, and SEC 10-K structure-aware retrieval ordering; (ii) task-specific invariant checking for RAG, JSON, and SQL outputs using finance-calibrated materiality thresholds (plus or minus 5%) and SEC citation validation; (iii) a three-tier model classification system enabling risk-appropriate deployment decisions; and (iv) an audit-ready attestation system with dual-provider validation. We evaluated five models (Qwen2.5-7B via Ollama, Granite-3-8B via IBM watsonx.ai, Llama-3.3-70B, Mistral-Medium-2505, and GPT-OSS-120B) across three regulated financial tasks. Across 480 runs (n=16 per condition), structured tasks (SQL) remain stable even at T=0.2, while RAG tasks show drift (25-75%), revealing task-dependent sensitivity. Cross-provider validation confirms deterministic behavior transfers between local and cloud deployments. We map our framework to Financial Stability Board (FSB), Bank for International Settlements (BIS), and Commodity Futures Trading Commission (CFTC) requirements, demonstrating practical pathways for compliance-ready AI deployments.
FedShard: Federated Unlearning with Efficiency Fairness and Performance Fairness
Wen, Siyuan, Zhang, Meng, Yang, Yang, Ding, Ningning
To protect clients' right to be forgotten in federated learning, federated unlearning aims to remove the data contribution of leaving clients from the global learned model. While current studies mainly focused on enhancing unlearning efficiency and effectiveness, the crucial aspects of efficiency fairness and performance fairness among decentralized clients during unlearning have remained largely unexplored. In this study, we introduce FedShard, the first federated unlearning algorithm designed to concurrently guarantee both efficiency fairness and performance fairness. FedShard adaptively addresses the challenges introduced by dilemmas among convergence, unlearning efficiency, and unlearning fairness. Furthermore, we propose two novel metrics to quantitatively assess the fairness of unlearning algorithms, which we prove to satisfy well-known properties in other existing fairness measurements. Our theoretical analysis and numerical evaluation validate FedShard's fairness in terms of both unlearning performance and efficiency. We demonstrate that FedShard mitigates unfairness risks such as cascaded leaving and poisoning attacks and realizes more balanced unlearning costs among clients. Experimental results indicate that FedShard accelerates the data unlearning process 1.3-6.2 times faster than retraining from scratch and 4.9 times faster than the state-of-the-art exact unlearning methods.