Law
MillStone: How Open-Minded Are LLMs?
Triedman, Harold, Shmatikov, Vitaly
Large language models equipped with Web search, information retrieval tools, and other agentic capabilities are beginning to supplant traditional search engines. As users start to rely on LLMs for information on many topics, including controversial and debatable issues, it is important to understand how the stances and opinions expressed in LLM outputs are influenced by the documents they use as their information sources. In this paper, we present MillStone, the first benchmark that aims to systematically measure the effect of external arguments on the stances that LLMs take on controversial issues (not all of them political). We apply MillStone to nine leading LLMs and measure how ``open-minded'' they are to arguments supporting opposite sides of these issues, whether different LLMs agree with each other, which arguments LLMs find most persuasive, and whether these arguments are the same for different LLMs. In general, we find that LLMs are open-minded on most issues. An authoritative source of information can easily sway an LLM's stance, highlighting the importance of source selection and the risk that LLM-based information retrieval and search systems can be manipulated.
AI Governance in Higher Education: A course design exploring regulatory, ethical and practical considerations
Weuts, Raphaรซl, Bleher, Johannes, Bleher, Hannah, Flores, Rozanne Tuesday, Xuanyang, Guo, Pujszo, Paweล, Almรกsi, Zsolt
As artificial intelligence (AI) systems permeate critical sectors, the need for professionals who can address ethical, legal and governance challenges has become urgent. Current AI ethics education remains fragmented, often siloed by discipline and disconnected from practice. This paper synthesizes literature and regulatory developments to propose a modular, interdisciplinary curriculum that integrates technical foundations with ethics, law and policy. We highlight recurring operational failures in AI - bias, misspecified objectives, generalization errors, misuse and governance breakdowns - and link them to pedagogical strategies for teaching AI governance. Drawing on perspectives from the EU, China and international frameworks, we outline a semester plan that emphasizes integrated ethics, stakeholder engagement and experiential learning. The curriculum aims to prepare students to diagnose risks, navigate regulation and engage diverse stakeholders, fostering adaptive and ethically grounded professionals for responsible AI governance.
Benchmarking Gender and Political Bias in Large Language Models
Yang, Jinrui, Han, Xudong, Baldwin, Timothy
We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-call vote outcomes and includes rich demographic metadata for each Member of the European Parliament (MEP), such as gender, age, country, and political group. Using EuroParlVote, we evaluate state-of-the-art LLMs on two tasks -- gender classification and vote prediction -- revealing consistent patterns of bias. We find that LLMs frequently misclassify female MEPs as male and demonstrate reduced accuracy when simulating votes for female speakers. Politically, LLMs tend to favor centrist groups while underperforming on both far-left and far-right ones. Proprietary models like GPT-4o outperform open-weight alternatives in terms of both robustness and fairness. We release the EuroParlVote dataset, code, and demo to support future research on fairness and accountability in NLP within political contexts.
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
Yoo, Yongmin, Xu, Qiongkai, Cao, Longbing
High-stakes texts such as patent claims, medical records, and technical reports are structurally complex and demand a high degree of reliability and precision. While large language models (LLMs) have recently been applied to automate their generation in high-stakes domains, reliably evaluating such outputs remains a major challenge. Conventional natural language generation (NLG) metrics are effective for generic documents but fail to capture the structural and legal characteristics essential to evaluating complex high-stakes documents. To address this gap, we propose PatentScore, a multi-dimensional evaluation framework specifically designed for one of the most intricate and rigorous domains, patent claims. PatentScore integrates hierarchical decomposition of claim elements, validation patterns grounded in legal and technical standards, and scoring across structural, semantic, and legal dimensions. In experiments on our dataset which consists of 400 Claim1, PatentScore achieved the highest correlation with expert annotations ($r = 0.819$), significantly outperforming widely used NLG metrics. This work establishes a new standard for evaluating LLM-generated patent claims, providing a solid foundation for research on patent generation and validation.
OpenAI Rolls Out Teen Safety Features Amid Growing Scrutiny
CEO Sam Altman announced an age-prediction system and new parental controls in a blog post on Tuesday. OpenAI announced new teen safety features for ChatGPT on Tuesday as part of an ongoing effort to respond to concerns about how minors engage with chatbots . The company is building an age-prediction system that identifies if a user is under 18 years old and routes them to an " age-appropriate " system that blocks graphic sexual content. If the system detects that the user is considering suicide or self-harm, it will contact the user's parents. In cases of imminent danger, if a user's parents are unreachable, the system may contact the authorities.
Charlie Kirk Shooting Suspect Charged as Prosecutor Seeks Death Penalty
In the indictment, prosecutors claim Tyler Robinson planned Kirk's killing in advance, citing rooftop surveillance, engraved bullets, and a written note as they seek the death penalty. A TV monitor displays a picture of Tyler Robinson, a suspect in the killing of Charlie Kirk in Orem, Utah. Utah County prosecutors on Tuesday charged Tyler Robinson in the shooting death of conservative activist Charlie Kirk at Utah Valley University, a murder officials say was politically motivated. They intend to seek the death penalty. Utah County Attorney Jeff Gray announced the indictment at a midday news conference, listing charges of aggravated murder, felony discharge of a firearm causing serious bodily injury, and commission of a violent offense in the presence of a child.
Grok's Responses Are Only Getting More Bizarre
Grok's Responses Are Only Getting More Bizarre Listen to more stories on the Noa app. When video of Charlie Kirk's assassination began circulating on X last week, Elon Musk's chatbot described it in upbeat terms. As users sought information about Kirk's condition, the bot, Grok, declared to some of them that the horrific footage was satire. This is a "meme edit," Grok told one user; Kirk "takes the roast in stride with a laugh--he's faced tougher crowds," it told another. "Yes, he survives this one easily." In the past several months, Grok has been on quite the hot streak: The bot spread false information about a supposed " white genocide," called for a second Holocaust while annointing itself "MechaHitler," and provided me with a list of what it believes the "good races" are.
Matthew Prince Wants AI Companies to Pay for Their Sins
The Cloudflare CEO joined to talk about standing up to content scraping, the internet's potential futures, and his company's relationship to Trump. Matthew Prince may not be a household name, but the world most certainly knows his work. Prince is the cofounder and CEO of Cloudflare . Launched in 2010, the internet infrastructure company has found itself increasingly in the position of serving as the web's bodyguard. It filters out bad traffic, keeps sites safe, and stops them from crashing when too many people visit. Its tools defend against DDoS attacks. In 2017, Cloudflare made headlines when it dropped white supremacist site The Daily Stormer . Cloudflare's severing of ties with The Daily Stormer marked a momentous shift, one that came after years of claiming a neutral stance. Prince continues to evolve the way Cloudflare works. In July, the company rolled out a new tool tasked with blocking unauthorized AI scraping. It effectively creates a pay-per-crawl model requiring AI platforms to shell out money if they want access to a site's content. On this episode of, I talked to Prince about publishing, the old internet, and how his ideal version of the future web means that OpenAI just might become the Netflix of content. KATIE DRUMMOND: Good to have you here, Matthew. You should have been warned ahead of time, but you probably weren't.
The looming crackdown on AI companionship
The risks posed when kids form bonds with chatbots have turned AI safety from an abstract worry into a political flashpoint. As long as there has been AI, there have been people sounding alarms about what it might do to us: rogue superintelligence, mass unemployment, or environmental ruin from data center sprawl. But this week showed that another threat entirely--that of kids forming unhealthy bonds with AI--is the one pulling AI safety out of the academic fringe and into regulators' crosshairs. This has been bubbling for a while. Two high-profile lawsuits filed in the last year, against Character.AI and OpenAI, allege that companion-like behavior in their models contributed to the suicides of two teenagers. A study by US nonprofit Common Sense Media, published in July, found that 72% of teenagers have used AI for companionship.