Goto

Collaborating Authors

 Oceania


Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text

arXiv.org Machine Learning

Modern large language models (LLMs) such as GPT, Claude, and Gemini have transformed the way we learn, work, and communicate. Y et, their ability to produce highly human-like text raises serious concerns about misinformation and academic integrity, making it an urgent need for reliable algorithms to detect LLMgenerated content. In this paper, we start by presenting a geometric approach to demystify rewrite-based detection algorithms, revealing their underlying rationale and demonstrating their generalization ability. Building on this insight, we introduce a novel rewrite-based detection algorithm that adaptively learns the distance between the original and rewritten text. Theoretically, we demonstrate that employing an adaptively learned distance function is more effective for detection than using a fixed distance. Empirically, we conduct extensive experiments with over 100 settings, and find that our approach demonstrates superior performance over baseline algorithms in the majority of scenarios. In particular, it achieves relative improvements from 57.8% to 80.6% over the strongest baseline across different target LLMs (e.g., GPT, Claude, and Gemini). The past few years have witnessed the emergence and rapid development of large language models (LLMs) such as GPT (Hurst et al., 2024), DeepSeek (Liu et al., 2024), Claude (Anthropic, 2024), Gemini (Comanici et al., 2025), Grok (xAI, 2025) and Qwen (Y ang et al., 2025). Their impact is everywhere, from education, academia and software development to healthcare and everyday life (Arora & Arora, 2023; Chan & Hu, 2023; Hou et al., 2024). On one side of the coin, LLMs can support users with conversational question answering, help students learn more effectively, draft emails, write computer code, prepare presentation slides and more. On the other side, their ability to closely mimic human-written text also raises serious concerns, including the generation of biased or harmful content, the spread of misinformation in the news ecosystem, and the challenges related to authorship attribution and intellectual property (Dave et al., 2023; Fang et al., 2024; Messeri & Crockett, 2024; Mahajan et al., 2025; Laurito et al., 2025). Addressing these concerns requires effective algorithms to distinguish between human-written and LLM-generated text, which has become an active and popular research direction in recent literature (see Crothers et al., 2023; Wu et al., 2025, for reviews).


Distributed Causality in the SDG Network: Evidence from Panel VAR and Conditional Independence Analysis

arXiv.org Machine Learning

The achievement of the 2030 Sustainable Development Goals (SDGs) is dependent upon strategic resource distribution. We propose a causal discovery framework using Panel Vector Autoregression, along with both country-specific fixed effects and PCMCI+ conditional independence testing on 168 countries (2000-2025) to develop the first complete causal architecture of SDG dependencies. Utilizing 8 strategically chosen SDGs, we identify a distributed causal network (i.e., no single 'hub' SDG), with 10 statistically significant Granger-causal relationships identified as 11 unique direct effects. Education to Inequality is identified as the most statistically significant direct relationship (r = -0.599; p < 0.05), while effect magnitude significantly varies depending on income levels (e.g., high-income: r = -0.65; lower-middle-income: r = -0.06; non-significant). We also reject the idea that there exists a single 'keystone' SDG. Additionally, we offer a proposed tiered priority framework for the SDGs namely, identifying upstream drivers (Education, Growth), enabling goals (Institutions, Energy), and downstream outcomes (Poverty, Health). Therefore, we conclude that effective SDG acceleration can be accomplished through coordinated multi-dimensional intervention(s), and that single-goal sequential strategies are insufficient.


A Decomposable Forward Process in Diffusion Models for Time-Series Forecasting

arXiv.org Machine Learning

We introduce a model-agnostic forward diffusion process for time-series forecasting that decomposes signals into spectral components, preserving structured temporal patterns such as seasonality more effectively than standard diffusion. Unlike prior work that modifies the network architecture or diffuses directly in the frequency domain, our proposed method alters only the diffusion process itself, making it compatible with existing diffusion backbones (e.g., DiffWave, TimeGrad, CSDI). By staging noise injection according to component energy, it maintains high signal-to-noise ratios for dominant frequencies throughout the diffusion trajectory, thereby improving the recoverability of long-term patterns. This strategy enables the model to maintain the signal structure for a longer period in the forward process, leading to improved forecast quality. Across standard forecasting benchmarks, we show that applying spectral decomposition strategies, such as the Fourier or Wavelet transform, consistently improves upon diffusion models using the baseline forward process, with negligible computational overhead. The code for this paper is available at https://anonymous.4open.science/r/D-FDP-4A29.


SA-PEF: Step-Ahead Partial Error Feedback for Efficient Federated Learning

arXiv.org Machine Learning

Biased gradient compression with error feedback (EF) reduces communication in federated learning (FL), but under non-IID data, the residual error can decay slowly, causing gradient mismatch and stalled progress in the early rounds. We propose step-ahead partial error feedback (SA-PEF), which integrates step-ahead (SA) correction with partial error feedback (PEF). SA-PEF recovers EF when the step-ahead coefficient α = 0 and step-ahead EF (SAEF) when α = 1. For non-convex objectives and δ-contractive compressors, we establish a second-moment bound and a residual recursion that guarantee convergence to stationar-ity under heterogeneous data and partial client participation. To balance SAEF's rapid warm-up with EF's long-term stability, we select α near its theory-predicted optimum. Experiments across diverse architectures and datasets show that SA-PEF consistently reaches target accuracy faster than EF. Modern large-scale machine learning increasingly relies on distributed computation, where both data and compute are spread across many devices. Federated learning (FL) enables model training in this setting without centralizing raw data, enhancing privacy and scalability under heterogeneous client distributions (McMahan et al., 2017; Kairouz et al., 2021). In each synchronous FL round, the server broadcasts the current global model to a subset of clients. These clients perform several steps of stochastic gradient descent (SGD) on their local data and return updates to the server, which aggregates them to form the next global iterate (Huang et al., 2022; Wang & Ji, 2022; Li et al., 2024). Although FL leverages rich distributed data, it faces two key challenges.


Demystifying Prediction Powered Inference

arXiv.org Machine Learning

Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science. However, treating predictions as ground truth introduces bias while ignoring them wastes valuable information. Prediction-Powered Inference (PPI) offers a principled framework that leverages predictions from large unlabeled datasets to improve statistical efficiency while maintaining valid inference through explicit bias correction using a smaller labeled subset. Despite its potential, the growing PPI variants and the subtle distinctions between them have made it challenging for practitioners to determine when and how to apply these methods responsibly. This paper demystifies PPI by synthesizing its theoretical foundations, methodological extensions, connections to existing statistics literature, and diagnostic tools into a unified practical workflow. Using the Mosaiks housing price data, we show that PPI variants produce tighter confidence intervals than complete-case analysis, but that double-dipping, i.e. reusing training data for inference, leads to anti-conservative confidence intervals and coverages. Under missing-not-at-random mechanisms, all methods, including classical inference using only labeled data, yield biased estimates. We provide a decision flowchart linking assumption violations to appropriate PPI variants, a summary table of selective methods, and practical diagnostic strategies for evaluating core assumptions. By framing PPI as a general recipe rather than a single estimator, this work bridges methodological innovation and applied practice, helping researchers responsibly integrate predictions into valid inference.


No Phone, No Social Safety Net: Welcome to the 'Offline Club'

WIRED

No Phone, No Social Safety Net: Welcome to the'Offline Club' Across Europe's largest cities, people are gathering for semi-silent, offline hangouts, in search of an experience that isn't mediated through their smartphones. On cue, the room fell silent. A man seated to my left at a long wooden table began to scratch at a piece of paper with a coloring pencil. To my right, another guy picked up a book. Across the way, someone buried themselves in a puzzle.


Amazon's latest round of layoffs will affect 16,000 workers

Engadget

Apple could unveil Gemini-powered Siri in Feb. Amazon's latest round of layoffs will affect 16,000 workers The news was first leaked in an email mistakenly sent early to workers. Sydney, Australia - 2022-07-22 Amazon prime boxes and envelopes delivered to a front door of residential building. Amazon has confirmed that it's letting go of 16,000 workers and employees across its organization. In an announcement by company SVP Beth Galetti, she explained that Amazon was going through organizational changes to reduce layers and remove bureaucracy. Affected employees in the US will be given 90 days to look for another internal role and will receive severance pay if they do not find any.


Revealed: The outdated British slang terms for sex that have been consigned to history - with 'how's your father' topping the list

Daily Mail - Science & tech

Trump calls Ilhan Omar a'fraud' and suggests she staged shock syringe attack during Minneapolis town hall: 'She probably had herself sprayed' Nicola Peltz is'getting a $1MILLION a month allowance from her father Nelson' as Brooklyn Beckham's billionaire in-laws take him under their wing Iran braces for possible US attack as Trump's'beautiful armada' arrives in Middle East amid claims regime has slaughtered 30,000 protesters America's damning verdict on who's to blame for Minneapolis mayhem between Trump and far-left protesters Melania's shock role in Trump's showdown with Kristi Noem revealed: MARK HALPERIN's fly-on-wall account of Oval Office meeting... and who is ACTUALLY taking the fall for Alex Pretti shooting I've seen possessed children scream like beasts and strung up like puppets... these chilling exorcism cases PROVE hell is real Sickening proof Kanye West's apology was fake: MAUREEN CALLAHAN reveals heinous new Nazi slur everyone missed... and revolting REAL reason he wrote groveling letter Trader Joe's reveals its best products of 2026 in grocery's own Oscars - and there's a surprise winner School principal accused of shoplifting from Walmart using'stacking' method at self-checkout Harper Beckham, 14, puts on a stylish display in a fluffy coat and vintage Chanel bag in Paris with her family - after Nicola Peltz's heartbreaking comments about sister-in-law I was barely eating but kept gaining weight. Then I discovered the'taboo' cancer doctors NEVER talk about. Now sex will never be the same... don't ignore these signs Lost tomb of the mysterious'cloud people' unearthed after 1,400 years in'discovery of the decade' She's All That star Rachael Leigh Cook was a 2000s icon... see her now Revealed: The outdated British slang terms for sex that have been consigned to history - with'how's your father' topping the list Is your sex lingo up-to-date, or are your go-to terms giving away your age? The answer may lie in how many of these slang words and phrases you're still using. A new survey by Perspectus Global has revealed the once-popular terms that have been consigned to history.


I've seen possessed children scream like beasts and strung up like puppets... these chilling exorcism cases PROVE hell is real

Daily Mail - Science & tech

Devastating impact of Minneapolis shooting on Trump is worse than expected: Poll reveals America's crushing verdict... and what he must do next Bodies are STILL in wreckage of private jet that crashed in Maine on Sunday, killing six including powerful lawyer's attorney wife School principal accused of shoplifting from Walmart using'stacking' method at self-checkout Melania's shock role in Trump's showdown with Kristi Noem revealed: MARK HALPERIN's fly-on-wall account of Oval Office meeting... and who is ACTUALLY taking the fall for Alex Pretti shooting I was barely eating but kept gaining weight. Then I discovered the'taboo' cancer doctors NEVER talk about. Now sex will never be the same... don't ignore these signs Harper Beckham, 14, puts on a stylish display in a fluffy coat and vintage Chanel bag in Paris with her family - after Nicola Peltz's heartbreaking comments about sister-in-law Devastating truth about Blind Side actor Quinton Aaron: More to this'than everyone is letting on', friends reveal... as co-star Sandra Bullock'monitors' situation The wild truth about my influencer sons, their psycho dad and how lawsuits nearly left them bankrupt - by Jake and Logan Paul's MOM Trump knifes'little Napoleon' Border Patrol commander over Minnesota mayhem as he declares: 'We'll de-escalate' Lost tomb of the mysterious'cloud people' unearthed after 1,400 years in'discovery of the decade' I've seen possessed children scream like beasts and strung up like puppets... these chilling exorcism cases PROVE hell is real There is a hidden battlefield within our world, where forces of light and darkness collide, believers say, in a conflict that sometimes spills into everyday life. In its most extreme form, the clash is described as possession: a person seemingly seized by demonic beings, their body overtaken, their voice and movements warped into something not quite human. For Anglican reverend Chris Lee, 43, this is not a theological abstraction but a reality he has lived with for nearly two decades.


Intersectional Fairness via Mixed-Integer Optimization

arXiv.org Machine Learning

The deployment of Artificial Intelligence in high-risk domains, such as finance and healthcare, necessitates models that are both fair and transparent. While regulatory frameworks, including the EU's AI Act, mandate bias mitigation, they are deliberately vague about the definition of bias. In line with existing research, we argue that true fairness requires addressing bias at the intersections of protected groups. We propose a unified framework that leverages Mixed-Integer Optimization (MIO) to train intersectionally fair and intrinsically interpretable classifiers. We prove the equivalence of two measures of intersectional fairness (MSD and SPSF) in detecting the most unfair subgroup and empirically demonstrate that our MIO-based algorithm improves performance in finding bias. We train high-performing, interpretable classifiers that bound intersectional bias below an acceptable threshold, offering a robust solution for regulated industries and beyond.