Goto

Collaborating Authors

 Industry


Classifier Calibration at Scale: An Empirical Study of Model-Agnostic Post-Hoc Methods

arXiv.org Machine Learning

We study model-agnostic post-hoc calibration methods intended to improve probabilistic predictions in supervised binary classification on real i.i.d. tabular data, with particular emphasis on conformal and Venn-based approaches that provide distribution-free validity guarantees under exchangeability. We benchmark 21 widely used classifiers, including linear models, SVMs, tree ensembles (CatBoost, XGBoost, LightGBM), and modern tabular neural and foundation models, on binary tasks from the TabArena-v0.1 suite using randomized, stratified five-fold cross-validation with a held-out test fold. Five calibrators; Isotonic regression, Platt scaling, Beta calibration, Venn-Abers predictors, and Pearsonify are trained on a separate calibration split and applied to test predictions. Calibration is evaluated using proper scoring rules (log-loss and Brier score) and diagnostic measures (Spiegelhalter's Z, ECE, and ECI), alongside discrimination (AUC-ROC) and standard classification metrics. Across tasks and architectures, Venn-Abers predictors achieve the largest average reductions in log-loss, followed closely by Beta calibration, while Platt scaling exhibits weaker and less consistent effects. Beta calibration improves log-loss most frequently across tasks, whereas Venn-Abers displays fewer instances of extreme degradation and slightly more instances of extreme improvement. Importantly, we find that commonly used calibration procedures, most notably Platt scaling and isotonic regression, can systematically degrade proper scoring performance for strong modern tabular models. Overall classification performance is often preserved, but calibration effects vary substantially across datasets and architectures, and no method dominates uniformly. In expectation, all methods except Pearsonify slightly increase accuracy, but the effect is marginal, with the largest expected gain about 0.008%.


Empirical Likelihood-Based Fairness Auditing: Distribution-Free Certification and Flagging

arXiv.org Machine Learning

Machine learning models in high-stakes applications, such as recidivism prediction and automated personnel selection, often exhibit systematic performance disparities across sensitive subpopulations, raising critical concerns regarding algorithmic bias. Fairness auditing addresses these risks through two primary functions: certification, which verifies adherence to fairness constraints; and flagging, which isolates specific demographic groups experiencing disparate treatment. However, existing auditing techniques are frequently limited by restrictive distributional assumptions or prohibitive computational overhead. We propose a novel empirical likelihood-based (EL) framework that constructs robust statistical measures for model performance disparities. Unlike traditional methods, our approach is non-parametric; the proposed disparity statistics follow asymptotically chi-square or mixed chi-square distributions, ensuring valid inference without assuming underlying data distributions. This framework uses a constrained optimization profile that admits stable numerical solutions, facilitating both large-scale certification and efficient subpopulation discovery. Empirically, the EL methods outperform bootstrap-based approaches, yielding coverage rates closer to nominal levels while reducing computational latency by several orders of magnitude. We demonstrate the practical utility of this framework on the COMPAS dataset, where it successfully flags intersectional biases, specifically identifying a significantly higher positive prediction rate for African-American males under 25 and a systemic under-prediction for Caucasian females relative to the population mean.


Incorporating data drift to perform survival analysis on credit risk

arXiv.org Machine Learning

Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk. Unlike most existing methods that implicitly assume a stationary data-generating process, in practise, mortgage portfolios are exposed to various forms of data drift caused by changing borrower behaviour, macroeconomic conditions, policy regimes and so on. This study investigates the impact of data drift on survival-based credit risk models and proposes a dynamic joint modelling framework to improve robustness under non-stationary environments. The proposed model integrates a longitudinal behavioural marker derived from balance dynamics with a discrete-time hazard formulation, combined with landmark one-hot encoding and isotonic calibration. Three types of data drift (sudden, incremental and recurring) are simulated and analysed on mortgage loan datasets from Freddie Mac. Experiments and corresponding evidence show that the proposed landmark-based joint model consistently outperforms classical survival models, tree-based drift-adaptive learners and gradient boosting methods in terms of discrimination and calibration across all drift scenarios, which confirms the superiority of our model design.


Driverless taxis set to launch in UK as soon as September

BBC News

Waymo, the US driverless car firm, said it hopes to be operating a robotaxi service in London as soon as September this year. The UK government has said it plans to change regulations in the second half of 2026 to enable driverless taxis to operate in the city but has not given a specific date. Waymo said a pilot service will launch in April and Local Transport Minister Lilian Greenwood said: We're supporting Waymo and other operators through our passenger pilots, and pro-innovation regulations to make self-driving cars a reality on British roads. The firm, which is owned by Google-parent Alphabet, showed off a fleet of cars it bought to the UK at London's Transport Museum on Wednesday. Waymo's vehicles are currently being operated by a safety driver, mapping the streets.


Rules-based trade with U.S. is 'over': Canada central bank head

The Japan Times

Rules-based trade with U.S. is'over': Canada central bank head Bank of Canada Governor Tiff Macklem speaks with reporters during an interview in Ottawa on Wednesday. Toronto - The era of rules-based trade with the United States is over, Canada's central bank governor said Wednesday, echoing a stark warning from the country's prime minister that President Donald Trump's impact on global trade is permanent. Bank of Canada Governor Tiff Macklem made the comments during an interest rate announcement which held the key rate at 2.25%, citing unpredictable U.S. trade policies. Macklem has repeatedly warned that the bank's efforts to forecast the Canadian economy had grown increasingly difficult given the tariffs imposed and threatened by Trump. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right.


ICE Is Using Palantir's AI Tools to Sort Through Tips

WIRED

ICE Is Using Palantir's AI Tools to Sort Through Tips ICE has been using an AI-powered Palantir system to summarize tips sent to its tip line since last spring, according to a newly released Homeland Security document. United States Immigration and Customs Enforcement is leveraging Palantir's generative artificial intelligence tools to sort and summarize immigration enforcement tips from its public submission form, according to an inventory released Wednesday of all use cases the Department of Homeland Security had for AI in 2025. The AI Enhanced ICE Tip Processing service is intended to help ICE investigators "to more quickly identify and action tips" for urgent cases, as well as translate submissions not made in English, according to the inventory. It also provides a "BLUF," defined as a "high-level summary of the tip," produced using at least one large language model. BLUF, or "bottom line up front," is a military term that's also used internally by some Palantir employees.


Here's the Company That Sold DHS ICE's Notorious Face Recognition App

WIRED

Immigration agents have used Mobile Fortify to scan the faces of countless people in the US--including many citizens. On Wednesday, the Department of Homeland Security published new details about Mobile Fortify, the face recognition app that federal immigration agents use to identify people in the field, undocumented immigrants and US citizens alike. The details, including the company behind the app, were published as part of DHS's 2025 AI Use Case Inventory, which federal agencies are required to release periodically. The inventory includes two entries for Mobile Fortify--one for Customs and Border Protection (CBP), another for Immigration and Customs Enforcement (ICE)--and says the app is in the "deployment" stage for both. CBP says that Mobile Fortify became "operational" at the beginning of May last year, while ICE got access to it on May 20, 2025.


The Doomsday Clock Is Now 85 Seconds to Midnight. Here's What That Means

WIRED

The Doomsday Clock Is Now 85 Seconds to Midnight. Catastrophic risks are increasing, cooperation is declining, and swift action is needed from global leaders to correct course. The Doomsday Clock is closer to midnight than ever. The Doomsday Clock has just been set to 85 seconds to midnight. Nearly 80 years after its creation, this time represents the closest the clock has ever been to midnight.


How drone warfare has changed in Ukraine

Al Jazeera

Could Ukraine hold a presidential election right now? Will Europe use frozen Russian assets to fund war? How can Ukraine rebuild China ties? 'Ukraine is running out of men, money and time' Trump says US ready to attack Iran with'speed and violence'


Anthropic Is at War With Itself

The Atlantic - Technology

The AI company shouting about AI's dangers can't quite bring itself to slow down. T hese are not the words you want to hear when it comes to human extinction, but I was hearing them: "Things are moving uncomfortably fast." I was sitting in a conference room with Sam Bowman, a safety researcher at Anthropic. Worth $183 billion at the latest estimate, the AI firm has every incentive to speed things up, ship more products, and develop more advanced chatbots to stay competitive with the likes of OpenAI, Google, and the industry's other giants. But Anthropic is at odds with itself--thinking deeply, even anxiously, about seemingly every decision. Anthropic has positioned itself as the AI industry's superego: the firm that speaks with the most authority about the big questions surrounding the technology, while rival companies develop advertisements and affiliate shopping links (a difference that Anthropic's CEO, Dario Amodei, was eager to call out during an interview in Davos last week).