Law
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
Dineen, Jacob, RRV, Aswin, Liu, Qin, Xu, Zhikun, Ye, Xiao, Shen, Ming, Li, Zhaonan, Lu, Shijie, Baral, Chitta, Chen, Muhao, Zhou, Ben
Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the training signal. We introduce QA-LIGN, which decomposes monolithic rewards into interpretable principle-specific evaluations through structured natural language programs. Models learn through a draft, critique, and revise pipeline, where symbolic evaluation against the rubrics provides transparent feedback for both initial and revised responses during GRPO training. Applied to uncensored Llama-3.1-8B-Instruct, QA-LIGN reduces attack success rates by up to 68.7% while maintaining a 0.67% false refusal rate, achieving Pareto optimal safety-helpfulness performance and outperforming both DPO and GRPO with state-of-the-art reward models given equivalent training. These results demonstrate that making reward signals interpretable and modular improves alignment effectiveness, suggesting transparency enhances LLM safety.
Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents
Widyasari, Ratnadira, Weyssow, Martin, Irsan, Ivana Clairine, Ang, Han Wei, Liauw, Frank, Ouh, Eng Lieh, Shar, Lwin Khin, Kang, Hong Jin, Lo, David
Detecting vulnerabilities in source code remains a critical yet challenging task, especially when benign and vulnerable functions share significant similarities. In this work, we introduce VulTrial, a courtroom-inspired multi-agent framework designed to identify vulnerable code and to provide explanations. It employs four role-specific agents, which are security researcher, code author, moderator, and review board. Using GPT-4o as the base LLM, VulTrial almost doubles the efficacy of prior best-performing baselines. Additionally, we show that role-specific instruction tuning with small quantities of data significantly further boosts VulTrial's efficacy. Our extensive experiments demonstrate the efficacy of VulTrial across different LLMs, including an open-source, in-house-deployable model (LLaMA-3.1-8B), as well as the high quality of its generated explanations and its ability to uncover multiple confirmed zero-day vulnerabilities in the wild.
Trump Wants to Trade Fuel Economy for Cheaper Cars. But It Might Not Work
By rolling back auto industry fuel efficiency goals, US president Donald Trump hopes to make new cars cheaper. But prices won't drop for years, and consumers will spend more on gas in the meantime. The Trump administration says its proposal to roll back vehicle fuel economy standards, announced officially in the Oval Office on Wednesday, is an attempt to shave dollars off the ballooning cost of new cars in the US. But the intended price drops likely won't show up on dealership lots and showroom floors for months if not years, given the length of automakers' product planning schedule. It would also likely force Americans to pay more, long-term, at another place they tend to visit more frequently: the pump.
Why Do Trump's Favorite Tech Bros Look So Sad?
The Industry Trump Gave the Tech Bros Everything. Why Are They Still Crashing Out? This was supposed to be their year--but a historically unpopular president and fears of an A.I. stock market crash loom large over Silicon Valley. Enter your email to receive alerts for this author. You can manage your newsletter subscriptions at any time.
What legal experts say about second US strike on Venezuela boat
Several legal experts have told BBC Verify that the second strike on an alleged Venezuelan drug boat by the US military was probably illegal, and would likely be considered an extrajudicial killing under international law. On Monday, the Trump administration confirmed that a follow-up strike on the boat - which has been criticised as a double tap - was ordered by US Navy Admiral Frank Bradley with the overall operation having been authorised by War Secretary Pete Hegseth. Nine people died in the first strike on the vessel and two survivors were left clinging to the burning wreckage when it was struck again, killing them, according to the Washington Post. A US official has said four missiles were used in the operation. The Trump administration has not denied there were survivors and has insisted the strikes on 2 September were in accordance with the law of armed conflict.
The age of unipolar diplomacy is coming to an end
What is a Palestinian without olives? In Gaza, the world has seen the cost of a diplomacy that claims to uphold a rules-based order but applies it selectively. The United States intervened late, and only to defend an occupation the International Court of Justice (ICJ) has ruled illegal. Alongside other Western nations that built multilateral institutions, the US increasingly pursues nationalist agendas that undermine them. The hypocrisy is stark: one set of rules for Ukraine, another for Gaza.
Interview with Alice Xiang: Fair human-centric image dataset for ethical AI benchmarking
Earlier this month, Sony AI released a dataset that establishes a new benchmark for AI ethics in computer vision models. The research behind the dataset, named Fair Human-Centric Image Benchmark (FHIBE), has been published in Nature . FHIBE is the first publicly-available, globally-diverse, consent-based human image dataset (inclusive of over 10,000 human images) for evaluating bias across a wide variety of computer vision tasks. We sat down with project lead, Alice Xiang, Global Head of AI Governance at Sony Group and Lead Research Scientist for AI Ethics at Sony AI, to discuss the project and the broader implications of this research. Could you start by introducing the project and taking us through some of the main contributions?
'I don't take no for an answer': how a small group of women changed the law on deepfake porn
Charlotte Owen: 'The Lords were blown away by these brilliant women.' Charlotte Owen: 'The Lords were blown away by these brilliant women.' 'I don't take no for an answer': how a small group of women changed the law on deepfake porn For Jodie*, watching the conviction of her best friend, and knowing she helped secure it, felt at first like a kind of victory. It was certainly more than most survivors of deepfake image-based abuse could expect. They had met as students and bonded over their shared love of music. In the years since graduation, he'd also become her support system, the friend she reached for each time she learned that her images and personal details had been posted online without her consent.