Goto

Collaborating Authors

 rebel


Four days in Oxford: Inside the chaotic return of Lane Kiffin to Ole Miss, that turned into pure cinema

FOX News

Giants' Jaxson Dart tried to reenter game with knee injury but had jersey taken away: reports From the Chargers to the Colts and Falcons, which of the NFL's eight winless teams are already in panic mode Giants' season-opening high comes crashing down as Jaxson Dart's knee injury overshadows Rams' blowout win Chiefs coach suggests Taylor Swift has sparked Travis Kelce's performance in first two weeks of season Michael Penix Jr. officially named Falcons starter for a short-week trial by fire against Packers ESPN deserves credit for pushing back on Dan Le Batard's false accusation of racism Jerry Jones issues'dangerous' warning after Cowboys topple Commanders behind Dak Prescott's historic day Craig Carton roasts Bears' Caleb Williams for emotional reaction to hamstring injury against Vikings Commanders face critical decision on Jayden Daniels' elbow as specialists review MRI results Dr Siegel shares miracles from his book'The Angels Among Us' Sean Duffy reacts to Democratic pushback on FAA's new AI tool to cut flight delays Rubio questions'point' of NATO after US blocked from bases Mamdani touts'productive' meeting with Trump in New York City Latvian official says Ukraine has'frozen the front' in war with Russia Rubio warns that Iran is'run by lunatic clerics' The absolute gem of a weekend in Oxford has come and gone, with a day of college football that could rival anything we've seen in recent memory. Ole Miss versus LSU delivered pure cinema, largely because of Lane Kiffin's dramatic return to Oxford. Now, all sides can move on until the two teams meet again next year in Baton Rouge. While much of the theatrics surrounding Kiffin's return felt like an attempt to sell tickets to a show that had already sold itself, the weekend still delivered an experience that will be remembered for decades. Both Ole Miss and LSU fans will now come to grips with the fact that life goes on, especially during the 2026 season.


Lane Kiffin walked back into Oxford as the main character, but Ole Miss and Pete Golding stole his spotlight

FOX News

Vanderbilt pulled off the most improbable last-second victory you'll ever see in college football ESPN reporter claims she'made it out alive' after vintage Bill Belichick sideline interview Rams coach Sean McVay's '60 Minutes' interview reveals why you shouldn't find your identity in your job Chase Daniel nearly made Tim Tebow walk off the set of'SEC Nation' with'derogatory' comment about Florida Pat McAfee might have made an illegal immigration joke on'College GameDay,' liberals lose their minds Miami QB Darian Mensah continues to set records, but could soft schedule cost him a Heisman? Entire California high school football team suspended after mascot's obscene gesture Caitlin Clark'would sign up' for potential WNBA game overseas after returning from World Cup Piers Morgan shares why he'instantly' left the UK Larry Kudlow: Let's not throw a wrench into the AI movement Mark Levin: AI data centers aren't going to kill you The CCP would use AI to'malevolent' ends: Matthew Continetti Iran was preparing to launch'Armageddon,' Rep Tim Burchett says Iran was preparing to launch'Armageddon,' Rep Tim Burchett says Assistant AG warns'vulnerabilities are severe' amid alleged election fraud crackdown Saudi Arabia's capital of Riyadh on edge after reported explosions Assistant AG warns'vulnerabilities are severe' amid alleged election fraud crackdown Saudi Arabia's capital of Riyadh on edge after reported explosions'Believe in yourself': Miracle Mop creator marks 30 years of invention You can call it karma. You can call it revenge for everything that transpired over the past ten months. But, for those roaming the Ole Miss campus following the cinematic win over LSU, it was more of a sign that this football program is going to be just fine without Lane Kiffin. This page may contain affiliate links to legal sports betting partners.


Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions

Neural Information Processing Systems

Aligning large language models with pointwise absolute rewards has so far required online, on-policy algorithms such as PPO and GRPO. In contrast, simpler methods that can leverage offline or off-policy data, such as DPO and REBEL, are limited to learning from preference pairs or relative signals. To bridge this gap, we introduce Quantile Reward Policy Optimization (QRPO), which learns from pointwise absolute rewards while preserving the simplicity and offline applicability of DPO-like methods. QRPO uses quantile rewards to enable regression to the closed-form solution of the KL-regularized RL objective. This reward yields an analytically tractable partition function, removing the need for relative signals to cancel this term. Moreover, QRPO scales with increased compute to estimate quantile rewards, opening a new dimension for pre-computation scaling. Empirically, QRPO consistently achieves top performance on chat and coding evaluations--reward model scores, AlpacaEval 2, and LeetCode--compared to DPO, REBEL, and SimPO across diverse datasets and 8B-scale models. Finally, we find that training with robust rewards instead of converting them to preferences induces less length bias.


The rebels at the front line of Myanmar's civil war

BBC News

In the five years since Myanmar's military chief led a coup to overthrow the democratically elected government, civil war has torn the country apart. Thousands have been killed and millions displaced by the conflict between the military and an alliance of ethnic and rebel groups. More than two years ago, the rebels made a series of sweeping gains, but things have taken a turn for the worse for them. Forced conscription and increased drone power has put the military on the offensive in most parts of the country. The BBC's Quentin Sommerville travelled to Myanmar without the permission of the authorities - the only way to report from rebel-held territory.


REBEL: Reinforcement Learning via Regressing Relative Rewards

Neural Information Processing Systems

While originally developed for continuous control problems, Proximal Policy Optimization (PPO) has emerged as the work-horse of a variety of reinforcement learning (RL) applications, including the fine-tuning of generative models. Unfortunately, PPO requires multiple heuristics to enable stable convergence (e.g.


REBEL: Reinforcement Learning via Regressing Relative Rewards Zhaolin Gao 1, Jonathan D. Chang

Neural Information Processing Systems

While originally developed for continuous control problems, Proximal Policy Optimization (PPO) has emerged as the work-horse of a variety of reinforcement learning (RL) applications, including the fine-tuning of generative models. Unfortunately, PPO requires multiple heuristics to enable stable convergence (e.g.



Overcoming the Generalization Limits of SLM Finetuning for Shape-Based Extraction of Datatype and Object Properties

arXiv.org Artificial Intelligence

Small language models (SLMs) have shown promises for relation extraction (RE) when extracting RDF triples guided by SHACL shapes focused on common datatype properties. This paper investigates how SLMs handle both datatype and object properties for a complete RDF graph extraction. We show that the key bottleneck is related to long-tail distribution of rare properties. To solve this issue, we evaluate several strategies: stratified sampling, weighted loss, dataset scaling, and template-based synthetic data augmentation. We show that the best strategy to perform equally well over unbalanced target properties is to build a training set where the number of occurrences of each property exceeds a given threshold. To enable reproducibility, we publicly released our datasets, experimental results and code. Our findings offer practical guidance for training shape-aware SLMs and highlight promising directions for future work in semantic RE.



DRO-REBEL: Distributionally Robust Relative-Reward Regression for Fast and Efficient LLM Alignment

arXiv.org Machine Learning

Reinforcement learning with human feedback (RLHF) has become crucial for aligning Large Language Models (LLMs) with human intent. However, existing offline RLHF approaches suffer from overoptimization, where models overfit to reward misspecification and drift from preferred behaviors observed during training. We introduce DRO-REBEL, a unified family of robust REBEL updates with type-$p$ Wasserstein, KL, and $χ^2$ ambiguity sets. Using Fenchel duality, each update reduces to a simple relative-reward regression, preserving scalability and avoiding PPO-style clipping or auxiliary value networks. Under standard linear-reward and log-linear policy classes with a data-coverage condition, we establish $O(n^{-1/4})$ estimation bounds with tighter constants than prior DRO-DPO approaches, and recover the minimax-optimal $O(n^{-1/2})$ rate via a localized Rademacher complexity analysis. The same analysis closes the gap for Wasserstein-DPO and KL-DPO, showing both also attain optimal parametric rates. We derive practical SGD algorithms for all three divergences: gradient regularization (Wasserstein), importance weighting (KL), and a fast 1-D dual solve ($χ^2$). Experiments on Emotion Alignment, the large-scale ArmoRM multi-objective benchmark, and HH-Alignment demonstrate strong worst-case robustness across unseen preference mixtures, model sizes, and data scales, with $χ^2$-REBEL showing consistently strong empirical performance. A controlled radius--coverage study validates a no-free-lunch trade-off: radii shrinking faster than empirical divergence concentration rates achieve minimax-optimal parametric rates but forfeit coverage, while coverage-guaranteeing radii incur $O(n^{-1/4})$ rates.