Goto

Collaborating Authors

 Country


EXCLUSIVE: DeepL to Release Interpretation Software for Japan

The Japan Times

BERLIN - German technology firm DeepL, known for its artificial intelligence-powered translation software, plans to release a Japanese-language version of its real-time interpretation software by the end of this year, a senior company official has said. The age of machine interpretation has arrived, said Leonardo Doin, head of engineering and research for real-time voice translation service DeepL Voice, in a recent interview. You can just wear an earpiece and ... you can just hear it (foreign-language speech) in your language anytime, Doin said. The interpretation software will integrate DeepL's speech recognition and machine translation technologies, and speech synthesis technology that mimics the tones of the speakers' voices. It will be able to handle multiple languages and speakers, he said, with the software's use in online meetings of multinational companies in mind. DeepL plans to roll out the software on smartphones as well.


Court system on 'brink of collapse', former senior judge warns

BBC News

Court system on'brink of collapse', former senior judge warns The court system is on the brink of collapse as the backlogs for trials reach unprecedented levels, the head of a major review has said. Sir Brian Leveson, a senior retired judge, warned ministers, the police and others that there could not be a pick and mix response to solving the crisis. Last year, in the first stage of the review, Sir Brian called for the right to a jury trial to be scaled back and many intermediate crimes to be dealt with by a judge alone. His second and final report has recommended 130 efficiency changes, from technical measures to allowing prison vans to use bus lanes to hit court appearance deadlines. Sir Brian's two reports were commissioned by ministers as part of an attempt to reverse the backlogs that had reached record levels before Labour came into power, but have continued to worsen since then.


Police told to reinvestigate man's death after suspected blackmail on Grindr

BBC News

Police told to reinvestigate man's death after suspected blackmail on Grindr Police have been told to reopen their investigation into the death of Scott Gough, who allegedly took his own life after being targeted by a gang of men on the gay dating app Grindr. A police Professional Standards Department (PSD) report found failures in the investigation into the 56-year-old's death, which happened the day after a group of men turned up at his home demanding his car keys. His partner, Cameron Tewson accused the police of marking their own homework after his complaint of homophobia was not upheld. Hertfordshire Police, the investigating force, said it remains committed to ensuring members of the LGBTQ+ community feel supported when approaching the force. The report into the police's actions comes after a BBC investigation found multiple cases of suspected blackmail involving victims targeted on Grindr in Gough's local area, with at least four connected to the same gang, which remains at large.


Women in tech and finance at higher risk from AI job losses, report says

The Guardian

The Corporation of London is calling on employers to re-skill female workers not currently in technical roles. The Corporation of London is calling on employers to re-skill female workers not currently in technical roles. 'Mid-career' females also being sidelined by rigid hiring processes, says City of London Corporation Women working in tech and financial services are at greater risk of losing their jobs to increased use of AI and automation than their male peers, according to a report that found experienced females were also being sidelined as a result of "rigid hiring processes". "Mid-career" women - with at least five years' experience - are being overlooked for digital roles in the tech and financial and professional services sectors, where they are traditionally underrepresented, according to the report by the City of London Corporation. The governing body that runs the capital's Square Mile found female applicants were discriminated against by rigid, and sometimes automated, screening of their CVs, which did not take into account career gaps related to caring for children or relatives, or only narrowly considered their professional experience.


Multiparameter Uncertainty Mapping in Quantitative Molecular MRI using a Physics-Structured Variational Autoencoder (PS-VAE)

arXiv.org Machine Learning

Quantitative imaging methods, such as magnetic resonance fingerprinting (MRF), aim to extract interpretable pathology biomarkers by estimating biophysical tissue parameters from signal evolutions. However, the pattern-matching algorithms or neural networks used in such inverse problems often lack principled uncertainty quantification, which limits the trustworthiness and transparency, required for clinical acceptance. Here, we describe a physics-structured variational autoencoder (PS-VAE) designed for rapid extraction of voxelwise multi-parameter posterior distributions. Our approach integrates a differentiable spin physics simulator with self-supervised learning, and provides a full covariance that captures the inter-parameter correlations of the latent biophysical space. The method was validated in a multi-proton pool chemical exchange saturation transfer (CEST) and semisolid magnetization transfer (MT) molecular MRF study, across in-vitro phantoms, tumor-bearing mice, healthy human volunteers, and a subject with glioblastoma. The resulting multi-parametric posteriors are in good agreement with those calculated using a brute-force Bayesian analysis, while providing an orders-of-magnitude acceleration in whole brain quantification. In addition, we demonstrate how monitoring the multi-parameter posterior dynamics across progressively acquired signals provides practical insights for protocol optimization and may facilitate real-time adaptive acquisition.


Preference-based Conditional Treatment Effects and Policy Learning

arXiv.org Machine Learning

We introduce a new preference-based framework for conditional treatment effect estimation and policy learning, built on the Conditional Preference-based Treatment Effect (CPTE). CPTE requires only that outcomes be ranked under a preference rule, unlocking flexible modeling of heterogeneous effects with multivariate, ordinal, or preference-driven outcomes. This unifies applications such as conditional probability of necessity and sufficiency, conditional Win Ratio, and Generalized Pairwise Comparisons. Despite the intrinsic non-identifiability of comparison-based estimands, CPTE provides interpretable targets and delivers new identifiability conditions for previous unidentifiable estimands. We present estimation strategies via matching, quantile, and distributional regression, and further design efficient influence-function estimators to correct plug-in bias and maximize policy value. Synthetic and semi-synthetic experiments demonstrate clear performance gains and practical impact.


Stationarity and Spectral Characterization of Random Signals on Simplicial Complexes

arXiv.org Machine Learning

It is increasingly common for data to possess intricate structure, necessitating new models and analytical tools. Graphs, a prominent type of structure, can encode the relationships between any two entities (nodes). However, graphs neither allow connections that are not dyadic nor permit relationships between sets of nodes. We thus turn to simplicial complexes for connecting more than two nodes as well as modeling relationships between simplices, such as edges and triangles. Our data then consist of signals lying on topological spaces, represented by simplicial complexes. Much recent work explores these topological signals, albeit primarily through deterministic formulations. We propose a probabilistic framework for random signals defined on simplicial complexes. Specifically, we generalize the classical notion of stationarity. By spectral dualities of Hodge and Dirac theory, we define stationary topological signals as the outputs of topological filters given white noise. This definition naturally extends desirable properties of stationarity that hold for both time-series and graph signals. Crucially, we properly define topological power spectral density (PSD) through a clear spectral characterization. We then discuss the advantages of topological stationarity due to spectral properties via the PSD. In addition, we empirically demonstrate the practicality of these benefits through multiple synthetic and real-world simulations.


Rethinking Test-Time Training: Tilting The Latent Distribution For Few-Shot Source-Free Adaptation

arXiv.org Machine Learning

Often, constraints arise in deployment settings where even lightweight parameter updates e.g. parameter-efficient fine-tuning could induce model shift or tuning instability. We study test-time adaptation of foundation models for few-shot classification under a completely frozen-model regime, where additionally, no upstream data are accessible. We propose arguably the first training-free inference method that adapts predictions to the new task by performing a change of measure over the latent embedding distribution induced by the encoder. Using task-similarity scores derived from a small labeled support set, exponential tilting reweights latent distributions in a KL-optimal manner without modifying model parameters. Empirically, the method consistently competes with parameter-update-based methods across multiple benchmarks and shot regimes, while operating under strictly and universally stronger constraints. These results demonstrate the viability of inference-level distributional correction for test-time adaptation even with a fully-frozen model pipeline.


Learning Better Certified Models from Empirically-Robust Teachers

arXiv.org Machine Learning

Adversarial training attains strong empirical robustness to specific adversarial attacks by training on concrete adversarial perturbations, but it produces neural networks that are not amenable to strong robustness certificates through neural network verification. On the other hand, earlier certified training schemes directly train on bounds from network relaxations to obtain models that are certifiably robust, but display sub-par standard performance. Recent work has shown that state-of-the-art trade-offs between certified robustness and standard performance can be obtained through a family of losses combining adversarial outputs and neural network bounds. Nevertheless, differently from empirical robustness, verifiability still comes at a significant cost in standard performance. In this work, we propose to leverage empirically-robust teachers to improve the performance of certifiably-robust models through knowledge distillation. Using a versatile feature-space distillation objective, we show that distillation from adversarially-trained teachers consistently improves on the state-of-the-art in certified training for ReLU networks across a series of robust computer vision benchmarks.


Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

arXiv.org Machine Learning

Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms. On difficult problems, an LLM may fail to produce a correct final answer, yet still provide reliable pairwise comparison signals indicating which of two candidate solutions is better. We leverage this observation to design a statistically efficient evaluation framework that combines standard labeled outcomes with pairwise comparison signals obtained by having models judge auxiliary reasoning chains. Treating these comparison signals as control variates, we develop a semiparametric estimator based on the efficient influence function (EIF) for the setting where auxiliary reasoning chains are observed. This yields a one-step estimator that achieves the semiparametric efficiency bound, guarantees strict variance reduction over naive sample averaging, and admits asymptotic normality for principled uncertainty quantification. Across simulations, our one-step estimator substantially improves ranking accuracy, with gains increasing as model output noise grows. Experiments on GPQA Diamond, AIME 2025, and GSM8K further demonstrate more precise performance estimation and more reliable model rankings, especially in small-sample regimes where conventional evaluation is pretty unstable.