Country
Self-Hinting Language Models Enhance Reinforcement Learning
Liao, Baohao, Dong, Hanze, Xu, Xinxing, Monz, Christof, Bian, Jiang
Group Relative Policy Optimization (GRPO) has recently emerged as a practical recipe for aligning large language models with verifiable objectives. However, under sparse terminal rewards, GRPO often stalls because rollouts within a group frequently receive identical rewards, causing relative advantages to collapse and updates to vanish. We propose self-hint aligned GRPO with privileged supervision (SAGE), an on-policy reinforcement learning framework that injects privileged hints during training to reshape the rollout distribution under the same terminal verifier reward. For each prompt $x$, the model samples a compact hint $h$ (e.g., a plan or decomposition) and then generates a solution $τ$ conditioned on $(x,h)$. Crucially, the task reward $R(x,τ)$ is unchanged; hints only increase within-group outcome diversity under finite sampling, preventing GRPO advantages from collapsing under sparse rewards. At test time, we set $h=\varnothing$ and deploy the no-hint policy without any privileged information. Moreover, sampling diverse self-hints serves as an adaptive curriculum that tracks the learner's bottlenecks more effectively than fixed hints from an initial policy or a stronger external model. Experiments over 6 benchmarks with 3 LLMs show that SAGE consistently outperforms GRPO, on average +2.0 on Llama-3.2-3B-Instruct, +1.2 on Qwen2.5-7B-Instruct and +1.3 on Qwen3-4B-Instruct. The code is available at https://github.com/BaohaoLiao/SAGE.
Fast Sampling for Flows and Diffusions with Lazy and Point Mass Stochastic Interpolants
Damsholt, Gabriel, Frellsen, Jes, Ditlevsen, Susanne
Stochastic interpolants unify flows and diffusions, popular generative modeling frameworks. A primary hyperparameter in these methods is the interpolation schedule that determines how to bridge a standard Gaussian base measure to an arbitrary target measure. We prove how to convert a sample path of a stochastic differential equation (SDE) with arbitrary diffusion coefficient under any schedule into the unique sample path under another arbitrary schedule and diffusion coefficient. We then extend the stochastic interpolant framework to admit a larger class of point mass schedules in which the Gaussian base measure collapses to a point mass measure. Under the assumption of Gaussian data, we identify lazy schedule families that make the drift identically zero and show that with deterministic sampling one gets a variance-preserving schedule commonly used in diffusion models, whereas with statistically optimal SDE sampling one gets our point mass schedule. Finally, to demonstrate the usefulness of our theoretical results on realistic highly non-Gaussian data, we apply our lazy schedule conversion to a state-of-the-art pretrained flow model and show that this allows for generating images in fewer steps without retraining the model.
Online Conformal Prediction via Universal Portfolio Algorithms
Liu, Tuo, Dobriban, Edgar, Orabona, Francesco
Online conformal prediction (OCP) seeks prediction intervals that achieve long-run $1-α$ coverage for arbitrary (possibly adversarial) data streams, while remaining as informative as possible. Existing OCP methods often require manual learning-rate tuning to work well, and may also require algorithm-specific analyses. Here, we develop a general regret-to-coverage theory for interval-valued OCP based on the $(1-α)$-pinball loss. Our first contribution is to identify \emph{linearized regret} as a key notion, showing that controlling it implies coverage bounds for any online algorithm. This relies on a black-box reduction that depends only on the Fenchel conjugate of an upper bound on the linearized regret. Building on this theory, we propose UP-OCP, a parameter-free method for OCP, via a reduction to a two-asset portfolio selection problem, leveraging universal portfolio algorithms. We show strong finite-time bounds on the miscoverage of UP-OCP, even for polynomially growing predictions. Extensive experiments support that UP-OCP delivers consistently better size/coverage trade-offs than prior online conformal baselines.
U.S. jet shoots down Iranian drone near carrier in Arabian Sea
The USS Abraham Lincoln aircraft carrier, is seen at Naval Air Station North Island in San Diego last August. President Donald Trump reiterated that the U.S. and Iran are maintaining diplomatic talks, even after an earlier skirmish in the Arabian Sea spooked oil markets amid heightened tensions between the two countries. We are negotiating with them right now" and they'd like to do something," Trump told reporters at the White House on Tuesday. They had a chance to do something a while ago and it didn't work out, and we did Midnight Hammer," he said, referring to the June U.S. military strike in Iran. Earlier Tuesday, a U.S. F-35C warplane shot down a drone in self-defense as the unmanned aircraft aggressively approached" the USS Abraham Lincoln aircraft carrier with unclear intent," U.S. Central Command said in a statement. The command said no American service members were harmed and no U.S. equipment was damaged.
Paris cybercrime unit searches X office; Musk summoned
Elon Musk attends the 56th annual World Economic Forum meeting in Davos, Switzerland, on Jan. 22. PARIS - French police raided the offices of Elon Musk's social media network X on Tuesday, and prosecutors ordered the tech billionaire to face questions in a widening investigation, amid growing scrutiny of the platform by authorities across Europe. The raid by the Paris prosecutor's cybercrime unit and Musk's summoning -- which could further increase tensions between Europe and the U.S. over Big Tech and free speech -- are linked to a yearlong investigation into suspected abuse of algorithms and fraudulent data extraction by X or its executives. Britain's privacy watchdog, meanwhile, also kicked off a formal investigation into Musk's artificial-intelligence chatbot Grok over the processing of personal data and its potential to produce harmful sexual images and video content. In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right.
An 'Intimacy Crisis' Is Driving the Dating Divide
An'Intimacy Crisis' Is Driving the Dating Divide In his book, sex and relationships researcher Justin Garcia says people have miscalculated their need for human intimacy, which is the real issue at root of the loneliness epidemic. In the US, nearly half of adults are single. A quarter of men suffer from loneliness. Rates of depression are on the rise . And one in four Gen Z adults--the so-called kinkiest generation, according to one study --have never had partnered sex. In an age of endless connection, where hooking up happens with the ease of a swipe and nontraditional relationship structures like polyamory are celebrated, why are people seemingly so disconnected and alone?
Who is in the Epstein files?
Who is in the Epstein files? The list of some of the world's most rich and powerful people with ties to late sex offender Jeffrey Epstein has lengthened with the latest US government release of millions of new files from its investigation into the disgraced financier. The 30 January drop of new material - dubbed the Epstein files - included three million pages, 180,000 images, 2,000 videos, and a number of household names like Richard Branson, Bill Gates and Elon Musk. There is no suggestion that appearing in the documents implies any wrongdoing, and many people who have featured in previous releases have denied any wrongdoing in relation to Epstein. The release came weeks after the deadline set by the Epstein Files Transparency Act which was signed into law by US President Donald Trump in November and required a full release of all Epstein-related documents.