Goto

Collaborating Authors

 Industry


This is Europe's secret weapon against Trump: it could burst his AI bubble Johnny Ryan

The Guardian

Dutch company employees work on a semiconductor lithography tool in Veldhoven, Netherlands, April 2019. Dutch company employees work on a semiconductor lithography tool in Veldhoven, Netherlands, April 2019. This is Europe's secret weapon against Trump: it could burst his AI bubble T he unthinkable has happened. The US is Europe's adversary. The stark, profound betrayal contained in the Trump administration's national security strategy should stop any further denial and dithering in Europe's capitals.


Tesla used deceptive language to market Autopilot, California judge rules

Engadget

The judge recommends suspending Tesla's sales in the state for 30 days. Tesla, Inc. is an American automotive and clean energy company. Tesla's sales in California should be suspended for 30 days because its marketing around Autopilot and Full Self-Driving misled consumers, a California administrative law judge has ruled . Back in 2022, the California DMV accused the automaker of using deceptive language to advertise those products and making it seem like its vehicles are capable of level 5 autonomous driving. Tesla has since added the word "Supervised" to the name of its Full Self-Driving assistance technology.


Russia-Ukraine war: List of key events, day 1,392

Al Jazeera

What is in the 28-point US plan for Ukraine? 'Ukraine is running out of men, money and time' Can the US get all sides to end the war? Why is Europe opposing Trump's peace plan? Kyiv Mayor Vitalii Klitschko said explosions were heard in the Ukrainian capital and warned people to stay in shelters late on Tuesday night as air defences worked to repel a Russian attack. Russian forces launched a "massive" drone attack on Ukraine's Sumy region, targeting energy infrastructure and causing electricity blackouts, Governor Oleh Hryhorov said on Telegram late on Tuesday night.


AI-assisted hiring will drive Indeed's growth, Recruit CEO says

The Japan Times

AI-assisted hiring will drive Indeed's growth, Recruit CEO says Companies embracing artificial intelligence to recruit and hire people won't threaten Indeed.com's Hisayuki "Deko" Idekoba, who leads Indeed and its parent, Tokyo-based Recruit Holdings, said the business is using AI to help companies optimize their talent-acquisition approach based on the pool of candidates, number of applicants per job and other factors, while using the flow of data to set compensation levels or adjust job qualifications. "We're gradually starting to deploy solutions such as AI agents to customers," Idekoba said in an interview in Tokyo. For Recruit, the shift reflects a broader transformation in how employers find and evaluate talent, as AI reshapes recruitment worldwide. Automated tools are speeding up candidate screening, cutting hiring costs and helping businesses respond to labor shortages and changing skill demands.


Essay cheating at universities an 'open secret'

BBC News

A BBC investigation has uncovered claims that essay cheating remains widespread at UK universities despite the introduction of a law designed to stop it. Since April 2022, it has been illegal to provide essays for students in post-16 education in England. But so far there have been no prosecutions. The BBC has spoken to a former lecturer who describes essay cheating as an open secret and to a businessman who claims to have made millions from selling model answer essays to university students. Universities UK, which represents 141 institutions, said there were severe penalties for students caught submitting work that was not their own.


LLmFPCA-detect: LLM-powered Multivariate Functional PCA for Anomaly Detection in Sparse Longitudinal Texts

arXiv.org Machine Learning

Sparse longitudinal (SL) textual data arises when individuals generate text repeatedly over time (e.g., customer reviews, occasional social media posts, electronic medical records across visits), but the frequency and timing of observations vary across individuals. These complex textual data sets have immense potential to inform future policy and targeted recommendations. However, because SL text data lack dedicated methods and are noisy, heterogeneous, and prone to anomalies, detecting and inferring key patterns is challenging. We introduce LLmFPCA-detect, a flexible framework that pairs LLM-based text embeddings with functional data analysis to detect clusters and infer anomalies in large SL text datasets. First, LLmFPCA-detect embeds each piece of text into an application-specific numeric space using LLM prompts. Sparse multivariate functional principal component analysis (mFPCA) conducted in the numeric space forms the workhorse to recover primary population characteristics, and produces subject-level scores which, together with baseline static covariates, facilitate data segmentation, unsupervised anomaly detection and inference, and enable other downstream tasks. In particular, we leverage LLMs to perform dynamic keyword profiling guided by the data segments and anomalies discovered by LLmFPCA-detect, and we show that cluster-specific functional PC scores from LLmFPCA-detect, used as features in existing pipelines, help boost prediction performance. We support the stability of LLmFPCA-detect with experiments and evaluate it on two different applications using public datasets, Amazon customer-review trajectories, and Wikipedia talk-page comment streams, demonstrating utility across domains and outperforming state-of-the-art baselines.


Continual Learning at the Edge: An Agnostic IIoT Architecture

arXiv.org Machine Learning

The exponential growth of Internet-connected devices has presented challenges to traditional centralized computing systems due to latency and bandwidth limitations. Edge computing has evolved to address these difficulties by bringing computations closer to the data source. Additionally, traditional machine learning algorithms are not suitable for edge-computing systems, where data usually arrives in a dynamic and continual way. However, incremental learning offers a good solution for these settings. We introduce a new approach that applies the incremental learning philosophy within an edge-computing scenario for the industrial sector with a specific purpose: real time quality control in a manufacturing system. Applying continual learning we reduce the impact of catastrophic forgetting and provide an efficient and effective solution.


Improving the Accuracy of Amortized Model Comparison with Self-Consistency

arXiv.org Machine Learning

Amortized Bayesian inference (ABI) offers fast, scalable approximations to posterior densities by training neural surrogates on data simulated from the statistical model. However, ABI methods are highly sensitive to model misspecification: when observed data fall outside the training distribution (generative scope of the statistical models), neural surrogates can behave unpredictably. This makes it a challenge in a model comparison setting, where multiple statistical models are considered, of which at least some are misspecified. Recent work on self-consistency (SC) provides a promising remedy to this issue, accessible even for empirical data (without ground-truth labels). In this work, we investigate how SC can improve amortized model comparison conceptualized in four different ways. Across two synthetic and two real-world case studies, we find that approaches for model comparison that estimate marginal likelihoods through approximate parameter posteriors consistently outperform methods that directly approximate model evidence or posterior model probabilities. SC training improves robustness when the likelihood is available, even under severe model misspecification. The benefits of SC for methods without access of analytic likelihoods are more limited and inconsistent. Our results suggest practical guidance for reliable amortized Bayesian model comparison: prefer parameter posterior-based methods and augment them with SC training on empirical datasets to mitigate extrapolation bias under model misspecification.


A variational Bayes latent class approach for EHR-based patient phenotyping in R

arXiv.org Machine Learning

As regulatory agencies increasingly recognise real-world evidence as a complement to traditional clinical trial data, interest has grown in applying Bayesian methods across both interventional and observational research (Boulanger and Carlin (2021). A central objective in many clinical investigations is the delineation of patient subgroups that exhibit comparable disease-related characteristics (He, Belouali, Patricoski, Lehmann, Ball, Anagnostou, Kreimeyer, and Botsis (2023)). Electronic Health Records (EHR) have become an important resource for such phenotypic analyses (Hripcsak and Albers (2013)). Bayesian approaches to patient phenotyping in clinical observational studies have been limited by the computational challenges associated with applying the Markov Chain Monte Carlo (MCMC) approach to real-world data. Hubbard, Huang, Harton, Oganisian, Choi, Utidjian, Eneli, Bailey, and Chen (2019) proposed a Bayes latent class model that could be used in a general context for observational studies that use EHR data. They consider the common clinical context where gold-standard phenotype information, such as genetic and laboratory data, is not fully available. A general model of this form has high potential applicability for use in clinical decision support across disease areas for both primary and secondary clinical databases. Latent Class Analysis (LCA) is widely used when we want to identify patient phenotypes or subgroups given multivariate data (Lanza and Rhoades (2013)). A challenge in clinical LCA is the prevalence of mixed data, where we may have combinations of continuous, nominal, ordinal and count data.


Understanding the Gain from Data Filtering in Multimodal Contrastive Learning

arXiv.org Machine Learning

The success of modern multimodal representation learning relies on internet-scale datasets. Due to the low quality of a large fraction of raw web data, data curation has become a critical step in the training pipeline. Filtering using a trained model (i.e., teacher-based filtering) has emerged as a successful solution, leveraging a pre-trained model to compute quality scores. To explain the empirical success of teacher-based filtering, we characterize the performance of filtered contrastive learning under the standard bimodal data generation model. Denoting $η\in(0,1]$ as the fraction of data with correctly matched modalities among $n$ paired samples, we utilize a linear contrastive learning setup to show a provable benefit of data filtering: $(i)$ the error without filtering is upper and lower bounded by $\frac{1}{η\sqrt{n}}$, and $(ii)$ the error with teacher-based filtering is upper bounded by $\frac{1}{\sqrt{ηn}}$ in the large $η$ regime, and by $\frac{1}{\sqrt{n}}$ in the small $η$ regime.