Goto

Collaborating Authors

 Government


HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation

arXiv.org Artificial Intelligence

There is growing interest in hypothesis generation with large language models (LLMs). However, fundamental questions remain: what makes a good hypothesis, and how can we systematically evaluate methods for hypothesis generation? To address this, we introduce HypoBench, a novel benchmark designed to evaluate LLMs and hypothesis generation methods across multiple aspects, including practical utility, generalizability, and hypothesis discovery rate. HypoBench includes 7 real-world tasks and 5 synthetic tasks with 194 distinct datasets. We evaluate four state-of-the-art LLMs combined with six existing hypothesis-generation methods. Overall, our results suggest that existing methods are capable of discovering valid and novel patterns in the data. However, the results from synthetic datasets indicate that there is still significant room for improvement, as current hypothesis generation methods do not fully uncover all relevant or meaningful patterns. Specifically, in synthetic settings, as task difficulty increases, performance significantly drops, with best models and methods only recovering 38.8% of the ground-truth hypotheses. These findings highlight challenges in hypothesis generation and demonstrate that HypoBench serves as a valuable resource for improving AI systems designed to assist scientific discovery.


Reward Distance Comparisons Under Transition Sparsity

arXiv.org Artificial Intelligence

Reward comparisons are vital for evaluating differences in agent behaviors induced by a set of reward functions. Most conventional techniques utilize the input reward functions to learn optimized policies, which are then used to compare agent behaviors. However, learning these policies can be computationally expensive and can also raise safety concerns. Direct reward comparison techniques obviate policy learning but suffer from transition sparsity, where only a small subset of transitions are sampled due to data collection challenges and feasibility constraints. Existing state-of-the-art direct reward comparison methods are ill-suited for these sparse conditions since they require high transition coverage, where the majority of transitions from a given coverage distribution are sampled. When this requirement is not satisfied, a distribution mismatch between sampled and expected transitions can occur, leading to significant errors. This paper introduces the Sparsity Resilient Reward Distance (SRRD) pseudometric, designed to eliminate the need for high transition coverage by accommodating diverse sample distributions, which are common under transition sparsity. We provide theoretical justification for SRRD's robustness and conduct experiments to demonstrate its practical efficacy across multiple domains.


A Framework for the Private Governance of Frontier Artificial Intelligence

arXiv.org Artificial Intelligence

This paper presents a proposal for the governance of frontier AI systems through a hybrid public-private system. Private bodies, authorized and overseen by government, provide certifications to developers of frontier AI systems on an opt-in basis. In exchange for opting in, frontier AI firms receive protections from tort liability for customer misuse of their models. Before detailing the proposal, the paper explores more commonly discussed approaches to AI governance, analyzing their strengths and flaws. It also examines the nature of frontier AI governance itself. The paper includes consideration of the political economic, institutional, legal, safety, and other merits and tradeoffs inherent in the governance system it proposes.


LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks

arXiv.org Artificial Intelligence

Large language model unlearning has become a critical challenge in ensuring safety and controlled model behavior by removing undesired data-model influences from the pretrained model while preserving general utility. Significant recent efforts have been dedicated to developing LLM unlearning benchmarks such as WMDP (Weapons of Mass Destruction Proxy) and MUSE (Machine Unlearning Six-way Evaluation), facilitating standardized unlearning performance assessment and method comparison. Despite their usefulness, we uncover for the first time a novel coreset effect within these benchmarks. Specifically, we find that LLM unlearning achieved with the original (full) forget set can be effectively maintained using a significantly smaller subset (functioning as a "coreset"), e.g., as little as 5% of the forget set, even when selected at random. This suggests that LLM unlearning in these benchmarks can be performed surprisingly easily, even in an extremely low-data regime. We demonstrate that this coreset effect remains strong, regardless of the LLM unlearning method used, such as NPO (Negative Preference Optimization) and RMU (Representation Misdirection Unlearning), the popular ones in these benchmarks. The surprisingly strong coreset effect is also robust across various data selection methods, ranging from random selection to more sophisticated heuristic approaches. We explain the coreset effect in LLM unlearning through a keyword-based perspective, showing that keywords extracted from the forget set alone contribute significantly to unlearning effectiveness and indicating that current unlearning is driven by a compact set of high-impact tokens rather than the entire dataset. We further justify the faithfulness of coreset-unlearned models along additional dimensions, such as mode connectivity and robustness to jailbreaking attacks. Codes are available at https://github.com/OPTML-Group/MU-Coreset.


Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification

arXiv.org Artificial Intelligence

Despite the successes of language models, their evaluation remains a daunting challenge for new and existing tasks. We consider the task of text simplification, commonly used to improve information accessibility, where evaluation faces two major challenges. First, the data in existing benchmarks might not reflect the capabilities of current language models on the task, often containing disfluent, incoherent, or simplistic examples. Second, existing human ratings associated with the benchmarks often contain a high degree of disagreement, resulting in inconsistent ratings; nevertheless, existing metrics still have to show higher correlations with these imperfect ratings. As a result, evaluation for the task is not reliable and does not reflect expected trends (e.g., more powerful models being assigned higher scores). We address these challenges for the task of text simplification through three contributions. First, we introduce SynthSimpliEval, a synthetic benchmark for text simplification featuring simplified sentences generated by models of varying sizes. Through a pilot study, we show that human ratings on our benchmark exhibit high inter-annotator agreement and reflect the expected trend: larger models produce higher-quality simplifications. Second, we show that auto-evaluation with a panel of LLM judges (LLMs-as-a-jury) often suffices to obtain consistent ratings for the evaluation of text simplification. Third, we demonstrate that existing learnable metrics for text simplification benefit from training on our LLMs-as-a-jury-rated synthetic data, closing the gap with pure LLMs-as-a-jury for evaluation. Overall, through our case study on text simplification, we show that a reliable evaluation requires higher quality test data, which could be obtained through synthetic data and LLMs-as-a-jury ratings.


AI threats to national security can be countered through an incident regime

arXiv.org Artificial Intelligence

Recent progress in AI capabilities has heightened concerns that AI systems could pose a threat to national security, for example, by making it easier for malicious actors to perform cyberattacks on critical national infrastructure, or through loss of control of autonomous AI systems. In parallel, federal legislators in the US have proposed nascent 'AI incident regimes' to identify and counter similar threats. In this paper, we consolidate these two trends and present a timely proposal for a legally mandated post-deployment AI incident regime that aims to counter potential national security threats from AI systems. We start the paper by introducing the concept of 'security-critical' to describe sectors that pose extreme risks to national security, before arguing that 'security-critical' describes civilian nuclear power, aviation, life science dual-use research of concern, and frontier AI development. We then present in detail our AI incident regime proposal, justifying each component of the proposal by demonstrating its similarity to US domestic incident regimes in other 'security-critical' sectors. Finally, we sketch a hypothetical scenario where our proposed AI incident regime deals with an AI cyber incident. Our proposed AI incident regime is split into three phases. The first phase revolves around a novel operationalization of what counts as an 'AI incident' and we suggest that AI providers must create a 'national security case' before deploying a frontier AI system. The second and third phases spell out that AI providers should notify a government agency about incidents, and that the government agency should be involved in amending AI providers' security and safety procedures, in order to counter future threats to national security.


DeepSeek poses 'profound' security threat, U.S. house panel claims

The Japan Times

Chinese artificial intelligence firm DeepSeek is a "profound threat" to U.S. national security, a bipartisan House committee said Wednesday, urging Nvidia to hand over information on sales of chips that the startup may have used to develop its breakthrough chatbot model. The House Select Committee on China alleged in a report Wednesday that DeepSeek's ties to Chinese government interests "are significant," citing corporate filings obtained by the panel. Lawmakers claimed that DeepSeek's founder, Liang Wenfeng, controls the firm alongside the High-Flyer Quant hedge fund in an "integrated ecosystem" linked to state-linked hardware distributors and Chinese research institute Zhejiang Lab. "Although it presents itself as just another AI chatbot, offering users a way to generate text and answer questions, closer inspection reveals that the app siphons data back to the People's Republic of China (PRC), creates security vulnerabilities for its users, and relies on a model that covertly censors and manipulates information pursuant to Chinese law," the report states.


Duffy contrasts Biden-era 'drone fiasco' with Trump admin's 'radical transparency' after FAA announces testing

FOX News

Transportation Sec. Sean Duffy indicated the Trump administration is committed to "radical transparency." In a video message about the Federal Aviation Administration doing "drone-detection testing" in New Jersey, Transportation Sec. Sean Duffy indicated that the Trump administration is committed to "radical transparency," juxtaposing that approach with what he referred to as the Biden administration's "drone fiasco." The FAA noted in a post on its website last week that the testing is slated to occur "in Cape May, New Jersey, between April 14-25." "The FAA will operate several large drones and more than 100 commercial off-the-shelf drones during the two-week period. Testing will take place over the water and near the Cape May Ferry Terminal during the daytime on weekdays only. The public should not fly recreational drones near this area during the test period," the post stated.


JD Vance gears up to talk economic priorities during trips to Italy, India

FOX News

Tech expert Kurt'CyberGuy' Knutsson joins'Fox & Friends' to discuss the future of AI development in the United States. Vice President JD Vance is poised to kick off a trip to Italy and India on Friday โ€“ marking his third international trip with the Trump administration. Vance and the second family are poised to meet with and "discuss shared economic and geopolitical priorities with leaders in each country," according to a statement from Vance's office. When in Rome, Vance is scheduled to meet with Italy's Prime Minister Giorgia Meloni and Vatican Secretary of State Cardinal Pietro Parolin. He will meet with India's Prime Minister Narendra Modi while visiting New Delhi, Jaipur and Agra.


US trade restriction on Nvidia sends markets tumbling again

The Guardian

US stocks have fallen further after Donald Trump imposed a new trade restriction on the chip designer Nvidia, rattling investors and triggering a sell-off across the semiconductor industry. The S&P 500 index dropped by about 1.3% in early trading, with the tech-heavy Nasdaq index down 2.1%. The Dow Jones fell 0.6%. Nvidia, the Californian company at the heart of the revolution in artificial intelligence technology, lost billions of dollars from its market value at the opening bell, with its shares down about 6%. The sell-off, which has spread to semiconductor makers in Asia and Europe, comes after Nvidia said the Trump administration had restricted the sale of its H20 chip in China by means of new licence requirements.