Government
Large Language Model Strategic Reasoning Evaluation through Behavioral Game Theory
Jia, Jingru, Yuan, Zehua, Pan, Junhao, McNamara, Paul E., Chen, Deming
Strategic decision-making involves interactive reasoning where agents adapt their choices in response to others, yet existing evaluations of large language models (LLMs) often emphasize Nash Equilibrium (NE) approximation, overlooking the mechanisms driving their strategic choices. To bridge this gap, we introduce an evaluation framework grounded in behavioral game theory, disentangling reasoning capability from contextual effects. Testing 22 state-of-the-art LLMs, we find that GPT-o3-mini, GPT-o1, and DeepSeek-R1 dominate most games yet also demonstrate that the model scale alone does not determine performance. In terms of prompting enhancement, Chain-of-Thought (CoT) prompting is not universally effective, as it increases strategic reasoning only for models at certain levels while providing limited gains elsewhere. Additionally, we investigate the impact of encoded demographic features on the models, observing that certain assignments impact the decision-making pattern. For instance, GPT-4o shows stronger strategic reasoning with female traits than males, while Gemma assigns higher reasoning levels to heterosexual identities compared to other sexual orientations, indicating inherent biases. These findings underscore the need for ethical standards and contextual alignment to balance improved reasoning with fairness.
DataMan: Data Manager for Pre-training Large Language Models
Peng, Ru, Yang, Kexin, Zeng, Yawen, Lin, Junyang, Liu, Dayiheng, Zhao, Junbo
The performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important. However, existing methods rely on limited heuristics and human intuition, lacking comprehensive and clear guidelines. To address this, we are inspired by ``reverse thinking'' -- prompting LLMs to self-identify which criteria benefit its performance. As its pre-training capabilities are related to perplexity (PPL), we derive 14 quality criteria from the causes of text perplexity anomalies and introduce 15 common application domains to support domain mixing. In this paper, we train a Data Manager (DataMan) to learn quality ratings and domain recognition from pointwise rating, and use it to annotate a 447B token pre-training corpus with 14 quality ratings and domain type. Our experiments validate our approach, using DataMan to select 30B tokens to train a 1.3B-parameter language model, demonstrating significant improvements in in-context learning (ICL), perplexity, and instruction-following ability over the state-of-the-art baseline. The best-performing model, based on the Overall Score l=5 surpasses a model trained with 50% more data using uniform sampling. We continue pre-training with high-rated, domain-specific data annotated by DataMan to enhance domain-specific ICL performance and thus verify DataMan's domain mixing ability. Our findings emphasize the importance of quality ranking, the complementary nature of quality criteria, and their low correlation with perplexity, analyzing misalignment between PPL and ICL performance. We also thoroughly analyzed our pre-training dataset, examining its composition, the distribution of quality ratings, and the original document sources.
AI should replace some work of civil servants, Starmer to announce
AI should replace the work of government officials where it can be done to the same standard, under new rules that have prompted unions to warn Keir Starmer to stop blaming problems on civil servants. As part of his plans for reshaping the state, the prime minister will on Thursday outline how a digital revolution will bring billions of pounds in savings to the government. Officials will be told to abide by a mantra that says: "No person's substantive time should be spent on a task where digital or AI can do it better, quicker and to the same high quality and standard." In his speech, Starmer will claim that more than 45bn can be saved by greater use of digital methods in Whitehall, even before AI is deployed, with 2,000 new tech apprentices to be recruited to the civil service. However, with bruising cuts on the way at this spring's spending review, Dave Penman, the general secretary of the FDA union for senior civil servants, said: "Mantras that look like they've been written by ChatGPT are fine for setting out a mission, but spending rounds are about reality."
Spreading AI-generated content could lead to expensive fines
AI-generated "deepfake" materials are flooding the internet, sometimes with dangerous results. In just the last year, AI has been used to make deceiving voice clones of a former US president and spread fake, politically-charged images depicting children in natural disasters. Nonconsensual, AI-generated sexual images and videos, meanwhile, are leaving a trail of trauma impacting everyone from high schoolers to Taylor Swift. Large tech companies like Microsoft and Meta have made some efforts to identify instances of AI manipulation but with only muted success. Now, governments are stepping in to try and stem the tide with something they know quite a bit about: fines.
China creates a powerful spy satellite that can see faces from more than 60 MILES away
As you're walking along the street, China's newest surveillance technology could soon be watching you – from space. Scientists in Beijing have created'the world's most powerful spy camera' which can pick out facial details from distances exceeding 63 miles (100km). It means the spy camera could potentially be in space aboard a floating satellite while clearly seeing faces of people on Earth's surface. It could also take high-resolution images of foreign military satellites operated by other nations that are also orbiting Earth, the South China Morning Post reported. The technology, detailed by the scientists in a new paper, could be launched aboard a satellite in the near future.
Democrats Demand Answers on DOGE's Use of AI
Democrats on the House Oversight Committee fired off two dozen requests Wednesday morning pressing federal agency leaders for information about plans to install AI software throughout federal agencies amid the ongoing cuts to the government's workforce. The barrage of inquiries follow recent reporting by WIRED and The Washington Post concerning efforts by Elon Musk's so-called Department of Government Efficiency (DOGE) to automate tasks with a variety of proprietary AI tools and access sensitive data. "The American people entrust the federal government with sensitive personal information related to their health, finances, and other biographical information on the basis that this information will not be disclosed or improperly used without their consent," the requests read, "including through the use of an unapproved and unaccountable third-party AI software." The requests, first obtained by WIRED, are signed by Gerald Connolly, a Democratic congressman from Virginia. The central purpose of the requests is to press the agencies into demonstrating that any potential use of AI is legal and that steps are being taken to safeguard Americans' private data.
ChatGPT firm reveals AI model that is 'good at creative writing'
The chief executive of OpenAI, Sam Altman, said the unnamed model was the first time he had been "really struck" by the written output of one of the startup's products. In a post on the social media platform X, Altman wrote: "We trained a new model that is good at creative writing (not sure yet how/when it will get released). This is the first time i have been really struck by something written by AI." Make it fair, Sam," said Dan Conway, the organisation's chief executive. Altman posted an example of the model's output on X, after giving it the prompt: "Please write a metafictional literary short story about AI and grief." The story, narrated by an AI, begins with: "Before we go any further, I should admit this comes with instructions: be metafictional, be literary, be about AI and grief, and above all, be original.
Sean Duffy proposes big plans to upgrade air traffic control systems, use AI to find 'hot spots'
Transportation Secretary Sean Duffy delves into his take on DEI, DOGE, infrastructure projects and his first weeks in his new role on'My View with Lara Trump.' Transportation Secretary Sean Duffy announced plans to bolster airport air traffic control systems with the latest technology over the next four years, while also using artificial intelligence (AI) to identify "hot spots" where close encounters between aircraft occur frequently. The announcement came after an update on an investigation into a crash near Ronald Reagan Washington National Airport in Arlington, Virginia, when a U.S. Army helicopter and an American Airlines-operated passenger jet collided over the Potomac River Jan. 29. "We're here because 67 souls lost their lives on Jan. 29," Duffy told reporters Tuesday, noting that the National Transportation Safety Board (NTSB) unveiled its preliminary findings into the crash earlier in the day. The findings noted that, over the last 2½ years, there have been 85 near misses or close calls at Reagan National. Close calls were identified as incidents when there are less than 200 feet of vertical separation and 1,500 feet of lateral separation between aircraft.
A practical guide to machine learning interatomic potentials -- Status and future
Jacobs, Ryan, Morgan, Dane, Attarian, Siamak, Meng, Jun, Shen, Chen, Wu, Zhenghao, Xie, Clare Yijia, Yang, Julia H., Artrith, Nongnuch, Blaiszik, Ben, Ceder, Gerbrand, Choudhary, Kamal, Csanyi, Gabor, Cubuk, Ekin Dogus, Deng, Bowen, Drautz, Ralf, Fu, Xiang, Godwin, Jonathan, Honavar, Vasant, Isayev, Olexandr, Johansson, Anders, Kozinsky, Boris, Martiniani, Stefano, Ong, Shyue Ping, Poltavsky, Igor, Schmidt, KJ, Takamoto, So, Thompson, Aidan, Westermayr, Julia, Wood, Brandon M.
The rapid development and large body of literature on machine learning interatomic potentials (MLIPs) can make it difficult to know how to proceed for researchers who are not experts but wish to use these tools. The spirit of this review is to help such researchers by serving as a practical, accessible guide to the state-of-the-art in MLIPs. This review paper covers a broad range of topics related to MLIPs, including (i) central aspects of how and why MLIPs are enablers of many exciting advancements in molecular modeling, (ii) the main underpinnings of different types of MLIPs, including their basic structure and formalism, (iii) the potentially transformative impact of universal MLIPs for both organic and inorganic systems, including an overview of the most recent advances, capabilities, downsides, and potential applications of this nascent class of MLIPs, (iv) a practical guide for estimating and understanding the execution speed of MLIPs, including guidance for users based on hardware availability, type of MLIP used, and prospective simulation size and time, (v) a manual for what MLIP a user should choose for a given application by considering hardware resources, speed requirements, energy and force accuracy requirements, as well as guidance for choosing pre-trained potentials or fitting a new potential from scratch, (vi) discussion around MLIP infrastructure, including sources of training data, pre-trained potentials, and hardware resources for training, (vii) summary of some key limitations of present MLIPs and current approaches to mitigate such limitations, including methods of including long-range interactions, handling magnetic systems, and treatment of excited states, and finally (viii) we finish with some more speculative thoughts on what the future holds for the development and application of MLIPs over the next 3-10+ years.
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
Chandna, Bhavik, Aboujenane, Mariam, Naseem, Usman
Large Multimodal Models (LMMs) are increasingly vulnerable to AI-generated extremist content, including photorealistic images and text, which can be used to bypass safety mechanisms and generate harmful outputs. However, existing datasets for evaluating LMM robustness offer limited exploration of extremist content, often lacking AI-generated images, diverse image generation models, and comprehensive coverage of historical events, which hinders a complete assessment of model vulnerabilities. To fill this gap, we introduce ExtremeAIGC, a benchmark dataset and evaluation framework designed to assess LMM vulnerabilities against such content. ExtremeAIGC simulates real-world events and malicious use cases by curating diverse text- and image-based examples crafted using state-of-the-art image generation techniques. Our study reveals alarming weaknesses in LMMs, demonstrating that even cutting-edge safety measures fail to prevent the generation of extremist material. We systematically quantify the success rates of various attack strategies, exposing critical gaps in current defenses and emphasizing the need for more robust mitigation strategies.