Industry
Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
Alignment of large language models (LLMs) has predominantly relied on pairwise preference optimization, where annotators select the better of two responses to a prompt. While simple, this approach overlooks the opportunity to learn from richer forms of human feedback, such as multiwise comparisons and top-$k$ rankings. We propose Ranked Choice Preference Optimization (RCPO), a unified framework that bridges preference optimization with (ranked) choice modeling via maximum likelihood estimation. The framework is flexible, supporting both utility-based and rank-based choice models. It subsumes several existing pairwise methods (e.g., DPO, SimPO), while providing principled training objectives for richer feedback formats. We instantiate this framework with two representative ranked choice models (Multinomial Logit and Mallows-RMJ). Empirical studies on Llama-3-8B-Instruct and Gemma-2-9B-it across AlpacaEval 2 and Arena-Hard benchmarks show that RCPO consistently outperforms competitive baselines. RCPO shows how directly leveraging ranked preference data, combined with the right choice models, yields more effective alignment. It offers a versatile and extensible foundation for incorporating (ranked) choice modeling into LLM training.
MARS-M: When Variance Reduction Meets Matrices
Liu, Yifeng, Yuan, Angela, Gu, Quanquan
Matrix-based preconditioned optimizers, such as Muon, have recently been shown to be more efficient than scalar-based optimizers for training large-scale neural networks, including large language models (LLMs). On the other hand, recent benchmarks on optimizers for LLM pre-training have demonstrated that variance-reduction techniques such as MARS can achieve substantial speedups over standard optimizers that do not employ variance reduction. In this paper, to achieve the best of both worlds, we introduce MARS-M, a new optimizer that integrates the variance reduction technique in MARS with Muon. Under standard regularity conditions, we prove that Muon-M converges to a first-order stationary point at a rate of $\tilde{\mathcal{O}}(T^{-1/3})$, which improves upon $\tilde{\mathcal{O}}(T^{-1/4})$ rate attained by Muon. Our empirical results on language modeling and computer vision tasks demonstrate that MARS-M consistently yields lower losses and improved performance across various downstream benchmarks. The implementation of MARS-M is available at https://github.com/AGI-Arena/MARS/tree/main/MARS_M.
Assessing the robustness of heterogeneous treatment effects in survival analysis under informative censoring
Wang, Yuxin, Frauen, Dennis, Schweisthal, Jonas, Schrรถder, Maresa, Feuerriegel, Stefan
Dropout is common in clinical studies, with up to half of patients leaving early due to side effects or other reasons. When dropout is informative (i.e., dependent on survival time), it introduces censoring bias, because of which treatment effect estimates are also biased. In this paper, we propose an assumption-lean framework to assess the robustness of conditional average treatment effect (CATE) estimates in survival analysis when facing censoring bias. Unlike existing works that rely on strong assumptions, such as non-informative censoring, to obtain point estimation, we use partial identification to derive informative bounds on the CATE. Thereby, our framework helps to identify patient subgroups where treatment is effective despite informative censoring. We further develop a novel meta-learner that estimates the bounds using arbitrary machine learning models and with favorable theoretical properties, including double robustness and quasi-oracle efficiency. We demonstrate the practical value of our meta-learner through numerical experiments and in an application to a cancer drug trial. Together, our framework offers a practical tool for assessing the robustness of estimated treatment effects in the presence of censoring and thus promotes the reliable use of survival data for evidence generation in medicine and epidemiology.
The Effects of Flipped Classrooms in Higher Education: A Causal Machine Learning Analysis
Czarnowske, Daniel, Heiss, Florian, Schmitz, Theresa M. A., Stammann, Amrei
This study uses double/debiased machine learning (DML) to evaluate the impact of transitioning from lecture-based blended teaching to a flipped classroom concept. Our findings indicate effects on students' self-conception, procrastination, and enjoyment. We do not find significant positive effects on exam scores, passing rates, or knowledge retention. This can be explained by the insufficient use of the instructional approach that we can identify with uniquely detailed usage data and highlights the need for additional teaching strategies. Methodologically, we propose a powerful DML approach that acknowledges the latent structure inherent in Likert scale variables and, hence, aligns with psychometric principles.
What Elon Musk's Version of Wikipedia Thinks About Hitler, Putin, and Apartheid
What does Elon Musk want the world to know about "white genocide theory"? Because he's been vocal about the issue in the past-- advancing the idea, for example, that Jews are pushing "hatred against whites"--I decided to search for the term on Grokipedia, the competitor to Wikipedia that Musk launched yesterday. First, the site uses just that term,, rather than, as you would see on Wikipedia and elsewhere. Just a few sentences in, Grokipedia provides the "empirical underpinnings" of this supposed campaign to eliminate white people of European descent around the world. And the site argues that conversation about this purported genocide is systematically suppressed by the media and academia, which are "prone to ideological biases favoring multiculturalism" and "relegate the theory to fringe conspiracy status despite the observable data on population trajectories."
In Guillermo del Toro's "Frankenstein," a Vast Vision Gets Netflixed Down to Size
In Guillermo del Toro's "Frankenstein," a Vast Vision Gets Netflixed Down to Size The latest reanimation of Mary Shelley's classic tale, starring Oscar Isaac and Jacob Elordi, is a labyrinthine tour of a filmmaker's career-long obsessions. Earlier this year, Quentin Tarantino, when asked to parse the high points of his filmography in an interview, described the two-part "Kill Bill" (2003-04) as "the movie I was born to make." He added, "I think'Inglourious Basterds' is my masterpiece, but'Once Upon a Time . . . in Hollywood' is my favorite." Might these be distinctions without a difference? I'm generally wary of artistic-birthright narratives, not least because a filmmaker of remarkable talent, consistent vision, and good fortune might well wind up with multiple candidates for the honor.
Dreo space heaters are on sale at Amazon just in time for the cold weather to roll in
If your feet are chilly or your nose feels dry, you're going to want to jump on these limited-time Amazon deals on Dreo heaters and humidifiers. We may earn revenue from the products available on this page and participate in affiliate programs. This is a weird time of year here in Upstate New York and much of the country. I wake up and it's freezing, but then I'm sweating through my hoodie by the time the afternoon rolls around. That's where a space heater comes in handy.
Google warns of rough edges as Gemini for Home arrives
When you purchase through links in our articles, we may earn a small commission. Gemini for Home begins its slow rollout on Google smart devices today, but some of Gemini's smart home features aren't "fully upgraded" yet, Google says. It's finally time to bid adieu to Google Assistant in the smart home world, as Gemini for Home has begun its slow rollout on Google smart speakers and displays. But if you're among the lucky few allowed to take Gemini for Home on a test drive today, you should expect some bumps in the road, the company says. Google took the wraps off Gemini for Home earlier this month, and it's been teasing a "new experience" for Home since last year.
Nvidia will build AI supercomputers for US Department of Energy
Nvidia, the artificial intelligence (AI) chip leader, will build seven new supercomputers for the United States Department of Energy (DOE), CEO Jensen Huang has said. The company has $500bn in bookings for its AI chips, Huang said on Tuesday in a keynote address at the company's GTC event in Washington, DC, the US capital. It is striking deals around the world while also navigating a US-China trade war that could determine which country's technology is most used across the globe. Investors are looking for clarity on what chips the tech company will be able to sell to the vast Chinese market, but Huang in his keynote speech praised policies by US President Donald Trump while announcing new products and deals. These included network technology that will let Nvidia AI chips work with quantum computers.