Goto

Collaborating Authors

 mathematics



What is OpenAI Astra? Everything we know about the quantum math-solving model.

Mashable

Look Up Say More Versus Creator Hub Switch Off Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Trending Now Safety Net In My Bag VidCon with Mashable Back to School Furtastic All Series Everything we know about the quantum math-solving model. OpenAI confirmed the existence of Astra, a smarter, unreleased model that's already achieving big results in the math world. Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men's grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website.


Something Weird Is Happening in Math

The Atlantic - Technology

Why one of the world's best mathematicians is joining OpenAI One of the winners of this year's Fields Medal is headed to OpenAI. Last Thursday, Jacob Tsimerman was one of four mathematicians awarded the prestigious honor, which is sometimes called the Nobel Prize of mathematics. The same day that he won the Fields, Tsimerman announced that he would be going on leave from the University of Toronto to work on AI safety. As my colleague Rose Horowitch and I wrote last week, top AI companies now employ a range of academics, including physicists, philosophers, economists, and, of course, mathematicians . But given Tsimerman's renown, his decision in particular seemed to catch many people by surprise. "It's like hiring Lionel Messi as project manager," one machine-learning professor posted on X. Tsimerman's expertise is in number theory, and he won the Fields for his work on the André-Oort conjecture, among other contributions.


Mathematicians put AI to work on Fermat's last theorem

New Scientist

Mathematicians put AI to work on Fermat's last theorem At an event in London, mathematicians have made unexpectedly fast progress on formalising Fermat's last theorem using AI In the lobby of a central London hotel, tourists are bracing themselves for a day of sightseeing in a heatwave. Meanwhile, staff are resetting the dining room after breakfast. And in a windowless meeting room, assembled academics are contemplating whether humans have a role to play in the future of mathematics, now that AI can prove theorems by itself. The general mood in the room is one of bewilderment at the recent jump in computer intelligence and excitement about the potential it unlocks - and perhaps a slight unease about what the future holds for them personally. Twenty-five researchers from diverse fields and countries are here to spend a week working on formalising Fermat's last theorem with cutting-edge AI models.


AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?

Neural Information Processing Systems

Despite progress in language model (LM) capabilities, evaluations have thus far focused on models' performance on tasks that humans have previously solved, including in programming (SWE-Bench) and mathematics (FrontierMath). We therefore propose testing models' ability to design and implement algorithms in an open-ended benchmark: We task LMs with writing code that efficiently solves computationally challenging problems in computer science, physics, and mathematics. Our AlgoTune benchmark consists of 120 tasks collected from domain experts and a framework for validating and timing LM-synthesized solution code, which is compared to reference implementations from popular open-source packages.In addition, we develop a baseline LM agent, AlgoTuner, and evaluate its performance across a suite of frontier models.AlgoTuner achieves an average 1.58x speedup against reference solvers, including methods from packages such as SciPy, scikit-learn and CVXPY.However, we find that current models fail to discover algorithmic innovations, instead preferring surface-level optimizations. We hope that AlgoTune catalyzes the development of LM agents exhibiting creative problem solving beyond state-of-the-art human performance.


RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics

Neural Information Processing Systems

Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions---failing to capture the nature of mathematics encountered in actual research environments. We introduce \textsc{RealMath}, a novel benchmark derived directly from research papers and mathematical forums that assesses LLMs' abilities on authentic mathematical tasks. Our approach addresses three critical challenges: sourcing diverse research-level content, enabling reliable automated evaluation through verifiable statements, and designing a continually refreshable dataset to mitigate contamination risks. Experimental results across multiple LLMs reveal surprising capabilities in handling research mathematics compared to competition problems, suggesting current models may already serve as valuable assistants for working mathematicians despite limitations on highly challenging problems.


An AI solution to an 80‑year‑old problem has shocked mathematicians

AIHub

Last week, OpenAI shocked the mathematical community by revealing that one of its internal artificial intelligence (AI) models had found a counterexample to a famous conjecture made by legendary Hungarian mathematician Paul Erdős in 1946. The planar unit distance problem, or Erdős problem 90, has intrigued mathematicians for decades. The new result is no mere curiosity. Canadian mathematician Daniel Litt described it as "the first result produced autonomously by an AI that I find interesting in itself". The breakthrough, produced with a general-purpose AI model rather than one specialised for mathematics, also highlights how AI is changing mathematical research itself.


The maths meme that has been distracting mathematicians for a century

New Scientist

A seemingly simple set of rules kicks off a kind of mathematical magic trick, which has kept great minds busy since the 1930s. Almost a century ago, a mathematician came up with a puzzle that was so seemingly simple and yet so fiendishly difficult that it has been distracting other mathematicians ever since. It has become a meme that jumps from brain to brain, with many people claiming to have solved it, only to have their hopes dashed as the proof unravels. And be warned - once I explain the rules, you will immediately want to start playing around with it yourself, and I take no responsibility for how much of your time you waste. It starts a bit like a magic trick.


A golden age of maths is dawning and mathematicians are freaking out

New Scientist

I am attempting to solve a mathematical conundrum that has stumped many of humanity's greatest thinkers. I have zero mathematical training, apart from a distant undergraduate physics degree, which should put my odds of success at slim to none. But I also have a trick up my sleeve - a kind of mathematical genie that can conjure arcane secrets seemingly out of thin air. I make a short request concerning an esoteric conjecture in number theory, then cross my fingers. Perhaps "genie" is a bit too strong - I'm simply using GPT 5.5 Pro, the latest iteration of OpenAI's flagship model. But for mathematicians, modern AI models appear to have a spark of magic.


How human error became a weapon against large language models

New Scientist

Alan Turing proposed a test for machine intelligence: could a computer convince a human it was human? Recently, a friend told me over coffee about some disheartening feedback she had received. "They said it was good," she said, "but that it read like it was written by AI." Knowing her, I understood immediately what had happened. Her credibility was being questioned not because her work was poor, but because it was too good - too clear, too fluent, too polished. The rapid acceleration of artificial intelligence tools is changing how we think about good writing.