Goto

Collaborating Authors

 Large Language Model


L0-Reasoning Bench: Evaluating Procedural Correctness in Language Models via Simple Program Execution

arXiv.org Artificial Intelligence

Complex reasoning tasks often rely on the ability to consistently and accurately apply simple rules across incremental steps, a foundational capability which we term "level-0" reasoning. To systematically evaluate this capability, we introduce L0-Bench, a language model benchmark for testing procedural correctness -- the ability to generate correct reasoning processes, complementing existing benchmarks that primarily focus on outcome correctness. Given synthetic Python functions with simple operations, L0-Bench grades models on their ability to generate step-by-step, error-free execution traces. The synthetic nature of L0-Bench enables systematic and scalable generation of test programs along various axes (e.g., number of trace steps). We evaluate a diverse array of recent closed-source and open-weight models on a baseline test set. All models exhibit degradation as the number of target trace steps increases, while larger models and reasoning-enhanced models better maintain correctness over multiple steps. Additionally, we use L0-Bench to explore test-time scaling along three dimensions: input context length, number of solutions for majority voting, and inference steps. Our results suggest substantial room to improve "level-0" reasoning and potential directions to build more reliable reasoning systems.


There's a way to get all your favorite AI tools for life

PCWorld

TL;DR: 1min.AI combines popular AI tools like GPT-4.0 and Midjourney, and lifetime access is only 79.97. AI tools like ChatGPT popped into existence, totally changed the professional world, and then immediately became very expensive. It's hard to get by without them now, but that doesn't mean you have to pay for each one individually. Instead of shelling out for OpenAI, Midjourney, and everything else, now you can get the same AI models all under one umbrella. This platform goes way beyond just text generation.


Small Language Models Are the New Rage, Researchers Say

WIRED

The original version of this story appeared in Quanta Magazine. Large language models work well because they're so large. The latest models from OpenAI, Meta, and DeepSeek use hundreds of billions of "parameters"--the adjustable knobs that determine connections among data and get tweaked during the training process. With more parameters, the models are better able to identify patterns and connections, which in turn makes them more powerful and accurate. But this power comes at a cost.


Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization

arXiv.org Machine Learning

Layer-wise post-training quantization has emerged as a widely used technique for compressing large language models (LLMs) without retraining. However, recent progress in this line of research is saturating, underscoring the need to revisit its core limitation and explore further improvements. This study identifies a critical bottleneck in existing layer-wise PTQ methods: the accumulation of quantization errors across layers significantly degrades performance, particularly in low-bit regimes. To address this, we propose Quantization Error Propagation (QEP), a lightweight and general framework that enhances layer-wise PTQ by explicitly propagating the quantization error which enable compensating for accumulated quantization errors. Additionally, we introduce a tunable propagation mechanism that allows for control over both propagation strength and computational overhead, making the framework adaptable to various architectures and resource constraints. Empirical evaluation on LLaMA2 models (7B, 13B, 70B) demonstrate that incorporating QEP into standard layer-wise PTQ pipelines outperforms standard PTQ methods. Notably, QEP yields substantial performance improvements under extreme low-bit quantization settings.


Ordinary Least Squares as an Attention Mechanism

arXiv.org Machine Learning

I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. This connection offers an alternative perspective on attention beyond the conventional information retrieval framework, making it more accessible to researchers and analysts with a background in traditional statistics. It falls into place when OLS is framed as a similarity-based method in a transformed regressor space, distinct from the standard view based on partial correlations. In fact, the OLS solution can be recast as the outcome of an alternative problem: minimizing squared prediction errors by optimizing the embedding space in which training and test vectors are compared via inner products. Rather than estimating coefficients directly, we equivalently learn optimal encoding and decoding operations for predictors. From this vantage point, OLS maps naturally onto the query-key-value structure of attention mechanisms. Building on this foundation, I discuss key elements of Transformer-style attention and draw connections to classic ideas from time series econometrics.


Netflix is reportedly testing a search function powered by OpenAI

Engadget

Netflix has started testing a new search feature powered by OpenAI that can help customers find movies and shows to watch, according to Bloomberg. The streaming service has reportedly given select users in Australia and New Zealand the option to use the tool. It will allow users to search for terms other than a specific show's title, an actor's name or the genre they want to watch. Bloomberg says it will give them a way to search for content using more specific terms, like their mood. Presumably, that means the service can surface dramatic shows for a search query that says "sad," and seeing as it's powered by generative AI, users will most likely be able to use natural language in their search terms.


A Disaster for American Innovation

The Atlantic - Technology

Nearly three months into President Donald Trump's term, the future of American AI leadership is in jeopardy. Basically any generative-AI product you have used or heard of--ChatGPT, Claude, AlphaFold, Sora--depends on academic work or was built by university-trained researchers in the industry, and frequently both. Today's AI boom is fueled by the use of specialized computer-graphics chips to run AI models--a technique pioneered by researchers at Stanford who received funding from the Department of Defense. They rely on a training method called "reinforcement learning," the foundations of which were developed with National Science Foundation (NSF) grants. "I don't think anybody would seriously claim that these [AI breakthroughs] could have been done if the research universities in the U.S. didn't exist at the same scale," Rayid Ghani, a machine-learning researcher at Carnegie Mellon University, told me.


OpenAI prepares to send GPT-4 out to pasture

Engadget

GPT-4, OpenAI's first big upgrade to ChatGPT months after unleashing it on the world, is on its way out. A changelog the company published on Thursday said the model will be retired from ChatGPT on April 30. GPT-4o, which has been available since last May, will fully replace it. OpenAI says GPT-4o improves on it in writing, coding and STEM. Recent upgrades have boosted the newer model further, enhancing its instruction following, problem-solving and conversational flow.


ChatGPT can now remember all your past conversations

Engadget

The next time you conclude a conversation with ChatGPT, it will save what you said to memory, even if you don't ask it explicitly to do so. "We have greatly improved memory in chatgpt -- it can now reference all your past conversations!" OpenAI CEO Sam Altman wrote on Thursday in an X post spotted by The Verge. "This is a surprisingly great feature imo, and it points at something we are excited about: ai systems that get to know you over your life, and become extremely useful and personalized." OpenAI has been working on improving ChatGPT's memory since 2023 when the company began testing custom instructions, a feature that allows users to set preferences that ChatGPT will consider in future conversations.


Engadget Podcast: Pixel 9a review and bracing for tariffs

Engadget

This week, Engadget's Sam Rutherford dives into his experience with Google's new 499 mid-range smartphone, the Pixel 9a. Is it really the new mid-range king, as we previously predicted? Or is it worth spending more for the Pixel 9? Also, we chat about how the Trump administration's volatile tariff strategy will affect consumer technology (not to mention everything else you buy). Tariff Watch: Switch 2 preorders delayed, Razer pauses laptop sales in the U.S. – 30:27 Samsung's Ballie robot with Google Gemini arrives this Summer (allegedly) – 43:31