threshold
'Limit overshoot, peak, decline': A new global goal for climate change?
The signing of the Paris Agreement by 200 countries in 2015 was a rare moment of global unity, and the euphoria was real. To cheers and tears, the international climate treaty was approved by nearly every country in the world. Fundamental to the agreement was a commitment to keeping temperatures from rising beyond 1.5C (2.7F) above the pre-industrial levels of the mid 19th-century. But 11 years on, it is evident that the Earth is soon going to exceed that threshold because of the continued record burning of fossil fuels. A report from the United Nations Environment Programme (UNEP) in Nairobi, Kenya - called Limiting Overshoot - has found that even under the most optimistic scenario, the Earth will heat up by 1.8C (3.2F) above pre-industrial levels.
OpenAI confirms Astra has reached critical cyber threshold, but will be available soon
Trending Now Say More Look Up Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Switch Off Creator Playbook Mashable Voices Mashable Selects Safety Net Versus Gift Ideas For Everyone On Your List In My Bag All Series OpenAI confirms Astra has reached'critical' cyber threshold, but will be available soon OpenAI previously paused some work on Astra due to safety concerns. Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men's grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he's also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl. OpenAI confirmed on Tuesday that its unreleased Astra model has reached a dangerous new milestone, while simultaneously confirming that it was forging ahead with a public launch.
OpenAI Astra: All about the quantum math-solving model with critical hacking skills
Trending Now Say More Look Up Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Switch Off Creator Playbook Mashable Voices Mashable Selects Safety Net Versus Gift Ideas For Everyone On Your List In My Bag All Series OpenAI provided an update about the release of Astra, and what precautions it's taking. Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men's grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he's also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl. Do you understand quantum parallel repetition?
Bill Gates says we've passed AI's danger thresholds. Now what?
Bill Gates says we've passed AI's danger thresholds. In a new interview, the billionaire philanthropist sounds an alarm on the urgency of getting our AI policies in order. The temperature is in the mid-80s, and the sky is incapable of being any more blue. The view from the Gates Ventures conference room overlooks the Carillon Point Marina, where a flotilla of expensive boats bob in the water, and across the lake to the Olympic Mountains that define the horizon. Because if the scene is placid, the messenger is not. Seated across from me at a conference room table, Bill Gates is rocking back and forth in his chair, totally animated. And the more he has to say--about the threats of terror or economic collapse or just losing control of our AI systems--the more agitated I find myself becoming, too. The philanthropist and former Microsoft CEO says he has been growing increasingly alarmed by the rate of change at which AI technology is advancing, especially since guardrails are not keeping pace. In a new essay published today, Gates argues that we have passed the points where multiple potential dangers should have been checked. "We've crossed the threshold in terms of [AI's] bio-capabilities, cyber-capabilities, psychosocial capabilities, job-market-destruction capabilities, and even the lack of control," he said in an interview with about his new memo "I'm just stunned at the lack of concern and discussion outside of the industry." In an effort to wake the world up to what he sees as a rapidly growing societal disrupter, the 70-year-old tech titan has begun sounding the alarm as a "shrill voice," both publicly with his new essay (the first of multiple he plans on the topic) and in meetings with the press, and privately in conversations with industry, government, and civil society leaders. And while Gates is calling attention to a number of issues, his warnings about the bio-capabilities of the current frontier models are especially chilling. "Any model that can make novel molecules should be monitored," he says.
Bernie Sanders calls on Silicon Valley to 'pause AI development' in interest of humanity
Senator Bernie Sanders speaks to a crowd at the Texas Democratic convention at the Hilliard Center in Corpus Christi, Texas, on 27 June 2026. Senator Bernie Sanders speaks to a crowd at the Texas Democratic convention at the Hilliard Center in Corpus Christi, Texas, on 27 June 2026. Bernie Sanders calls on Silicon Valley to'pause AI development' in interest of humanity Progressive US senator urges Meta, OpenAI and Anthropic to'stop building machines that humans cannot control' Senator Bernie Sanders has called on Meta, OpenAI and Anthropic executives to halt their development of artificial intelligence, warning that the US Senate will implement regulation if the companies continue deploying AI at their current pace. In a new letter addressed to the CEOs of three of the country's leading AI companies, Sanders said the capabilities of these AI models have reached a critical risk threshold and that the companies are losing control over the technology. Altman, Mr. Amodei and Mr. Zuckerberg: In the interest of humanity, stand by your words.
OpenAI pumps the brakes on new Astra model over cybersecurity concerns
When you purchase through links in our articles, we may earn a small commission. The unreleased Astra model may possess "critical" cybersecurity abilities, OpenAI warns. Less than a week after touting the scientific achievements of Astra, its next "major" model, OpenAI says it's "pausing internal activities" related to the model due to its powerful cybersecurity abilities. "Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity," OpenAI stated in a Friday press release . "These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework ."
Conditional Inference Trees and Forests for Feature Selection
Milletich, Robert, Downes, Justin, Goley, Steve, Hirst, Newel
Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive. We study CIT and CIF as top-$k$ feature-ranking methods for downstream prediction using real-data benchmarks, runtime ablations, and synthetic feature-recovery experiments. At a fixed node, if the features and permutation budget do not depend on the node responses, Bonferroni-corrected $+1$ Monte Carlo permutation $p$-values control nodewise rejection under the complete permutation null. CIF ranks 4th among 17 classification methods on 22 datasets and 3rd among 18 regression methods on 8 datasets. With Bonferroni correction held fixed, the CIF runtime ablations indicate that adaptive stopping and the number of thresholds searched have the largest measured effect on runtime: turning off adaptive stopping and using exact threshold search increase fitting time by 4.0--8.4$\times$ and 1.9--10.8$\times$, respectively, while downstream score changes are at most 0.011. Sparse high-$p$ simulations indicate that forest feature sampling can leave informative features out of many split decisions. Overall, the results support CIF as a top-$k$ feature-ranking method in the evaluated downstream prediction benchmarks.
Sequential Structure-Sensitive Residual Diagnostics for PDE Inverse Problems
Computational models in science and engineering are often assessed by checking whether the residual norm is consistent with the assumed noise level. This can be misleading in smoothing inverse problems: structured model errors may be attenuated in observation space, leaving residual magnitudes below practitioner discrepancy thresholds while coherent residual patterns remain. As a result, residual-norm diagnostics can accept fitted models that still give biased parameters, predictions, or quantities of interest. We propose a structure-sensitive sequential diagnostic based on e-processes. The method uses a portfolio of spatial residual-pattern experts, updates their likelihood-ratio wealth as observations are processed, and rejects the fitted model when the aggregate wealth crosses a prescribed threshold, giving anytime-valid type-I error control for a fixed fitted model. We compare the method with Morozov discrepancy checks, fixed-sample residual tests, and batch projection tests. Across three inverse problems (elliptic diffusion, two-dimensional Stokes flow, and a glaciological ice-stream inversion implemented in the community finite-element model icepack) we demonstrate how standard discrepancy checks accept misspecified fits that produce materially wrong quantities of interest. Structure-sensitive batch tests detect these failures using the full dataset, while the e-process detects them earlier from a fraction of the observations. After rejection, the expert wealth attributes the evidence to residual patterns in the chosen dictionary and provides a basis for exploratory model correction.
Online Safety Monitoring for LLMs
Schirmer, Mona, Jazbec, Metod, Timans, Alexander, Naesseth, Christian, Waldron, Maja, Nalisnick, Eric
We deploy a simple into our everyday lives as search engines (Jin et al., 2025; statistical framework based on risk control (Angelopoulos Xiong et al., 2024), coding assistants (Zhao et al., 2023), et al., 2022) that converts any safety signal into a binary and companions (Zhang et al., 2025a). As their applicability grows, so does the potential harm caused by malicious decision rule, and offers statistical guarantees on the false LLM outputs. Despite remarkable performance across a alarm or missed detection rate. The framework is universally applicable to different monitoring purposes and can leverage wide range of tasks, LLMs remain prone to generating halarbitrary proxy signals. Through experiments on mathematlucinated, factually incorrect (Ravichander et al., 2025), or ical problem solving and red teaming conversations, we harmful output (Yu et al., 2025) when deployed.
Decision-Aware Training for Sample-Based Generative Models
Raeth, Kornelius, Ludwig, Nicole
Kornelius Raeth 1 Nicole Ludwig 1 2 Abstractscoring rules distribute the training gradient in proportion to Sample-based generative models are increasingly data density, with no awareness of the decision maker's cost structure. The model's limited capacity is allocated globused for probabilistic forecasting in high-stakes ally, leaving decision-critical regions of the output space decision settings, yet their training objectives are potentially underserved. These models are commonly trained with strictly proper Given a forecast, a decision maker with cost function c(a,y), scoring rules, such as the energy score, which al-of action aand outcome y, selects the action that minimises locate their training signal in proportion to dataexpected cost under the forecast distribution; a point forecast density, with no awareness of where forecast eris insufficient to evaluate this expectation. A good forecast rors are most costly for downstream decisions. Crucially, the energy score objective with a differentiable deci-observed cost of the optimal action is itself a proper scoring sion loss that directly penalises the cost incurredrule (Hartline et al., 2025; Kleinberg et al., 2023), placing by acting on the model's forecast. This combinedit in the same family as the energy score which licenses loss is theoretically grounded, as the decision losstheir combination as a theoretically well-founded training is itself a proper scoring rule. Introduction score acts as that anchor, preventing the model from collapsing outside cost-sensitive regions. Our method is theo-tion based on a temperature forecast, balancing asset loss against the cost of intervention. In the weather domain, retically grounded and leads to better downstream decisions state-of-the-art forecasting systems (Lang et al., 2024; Pricewhile retaining full probabilistic forecasts, as validated on et al., 2023) are trained with strictly proper scoring rulessynthetic and real-world forecasting tasks. A gradient analysis showing which regions benefitscore reduces to the continuous ranked probability score from the decision loss and why, based on the cost (CRPS), widely used in meteorological forecast verificafunction structure. Both model classes introduced above are commonly trained by minimising strictly proper sion calibration.