Goto

Collaborating Authors

 AIHub


Congratulations to the #ICML2026 award winners

AIHub

Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility implies a solution space that strictly supersets the fixed autoregressive trajectory, theoretically unlocking superior reasoning potential. Indeed, for specific constraint satisfaction tasks (e.g., sudoku puzzles), this capability has proven to be highly advantageous. However, in this paper, we reveal that for general reasoning tasks (e.g., mathematics and coding), arbitrary order generation may in fact limit the reasoning potential of dLLMs. We find that dLLMs tend to exploit this order flexibility to bypass high-uncertainty tokens that are crucial for exploration, leading to a premature collapse of solution coverage. This observation motivates a rethink of RL approaches for dLLMs, where considerable complexities, such as handling combinatorial trajectories and intractable likelihoods, are often devoted to preserving this flexibility. We demonstrate that effective reasoning can be better elicited by simply forgoing arbitrary order and applying standard Group Relative Policy Optimization (GRPO) instead. Our approach, JustGRPO, is minimalist yet surprisingly effective (e.g., 89.1% accuracy on GSM8K) while fully retaining the parallel decoding ability of dLLMs.


Interactive World Simulator for Robot Policy Training and Evaluation

AIHub

Imagine you want to teach a robot to push an object on a table. The standard recipe in robot learning is to collect hundreds of expert demonstrations on a real robot, train an imitation learning policy on that data, and then evaluate the policy by running it many times on the same real robot. Both stages (data collection and evaluation) are slow, expensive, and hard to reproduce: hardware breaks, lighting changes, objects drift out of place, and every new task means more hours in the lab. A natural question is whether we can replace some of this real-robot work with a simulator. Classical physics-based simulators are powerful, but building one for a new task means manually modeling geometries, contacts, friction, and deformation, and the resulting simulator often still does not match reality closely enough for policies trained inside it to transfer.


#ICML2026 social media round-up

AIHub

The forty-third International Conference on Machine Learning (ICML) took place in Seoul, South Korea from 6-11 July. We take a look at what the participants got up to during the event. The PC Chairs are presenting the welcome remarks. One of the best conference dinners I've ever had. Try TimeChat-Captioner (at #ICML2026) -- a videoLLM that generates dense, time-aware captions for long videos.


AI for science – talk recordings now available to watch

AIHub

On the 31st March, our editorial team headed to the Royal Society for AI for Science . This day-long conference explored how AI is changing the nature of scientific discovery, and was hosted by the Alan Turing Institute. The recordings from the event are now available on YouTube and are well worth a watch. You can read Ella Scallan's blog post about the day here . Lucy Smith is Senior Managing Editor for AIhub.


AAAI presidential panel – factuality and trustworthiness

AIHub

The Future of AI Research report, published in March 2025, aims to clearly identify the trajectory of AI research in a structured way. The report was led by outgoing AAAI President Francesca Rossi and covers 17 different AI topics . Members of the report team, and other selected AI practitioners, are taking part in a series of video panel discussions covering selected chapters from the report. In the sixth discussion in the collection, the three panellists tackle factuality and trustworthiness. Understanding factuality: why preventing false outputs from large language models remains AI's toughest problem Lucy Smith is Senior Managing Editor for AIhub.


The secret to human 'brilliance' that AI just can't match

AIHub

People often make decisions through "satisficing," gathering just enough information to make a satisfactory prediction of a likely outcome. A series of experimental games shows that people also employ satisficing to learn social rules and conventions. This finding offers new insight into social learning and reveals a key difference between how humans and LLMs make predictions. The premise of AI large language models is that any problem can be solved by vacuuming up as much information as possible, running it through probability models, and performing complex calculations to make predictions and come up with the optimal solution. Another premise behind LLMs is that they emulate the way human brains operate.


Pre-training isn't bitter enough

AIHub

Richard Sutton's "Bitter Lesson" is usually read as a warning against building too much human knowledge into AI systems. Over the long run, the methods that win are not the ones that encode our clever intuition most directly, but the ones that scale: search, learning, and other general methods that can absorb more compute and data. We take a general architecture, expose it to massive data, and train it with a simple self-supervised objective. Language models predict the next token. Vision models reconstruct masked patches, align views, or match teacher representations.


Interview with Thi Kieu Khanh Ho: Time-series anomaly detection

AIHub

The latest interview in our series with the AAAI/SIGAI Doctoral Consortium participants features Thi Kieu Khanh Ho who is studying time-series anomaly detection. We found out more about her research, and what inspired her to study AI, and what she plans to work on next. Tell us a bit about your PhD -- where are you studying, and what is the topic of your research? I am doing my PhD at McGill University and Mila - Québec AI Institute, in the Department of Electrical and Computer Engineering, supervised by Professor Narges Armanfard. My research focuses on time-series anomaly detection, the problem of teaching AI systems to recognize when something unusual or abnormal is happening in complex, real-world data streams, without relying on large amounts of labeled examples.


#RoboCup2026 social media round-up

AIHub

This year, RoboCup took place in Incheon, South Korea, from 2-6 July. The event saw teams take part in competitions, training sessions, and a symposium. Take a look at what the participants got up to in our round up from social media. RoboCup 2026 officially begins today! A post shared by RoboCup Federation (@robocup.official)


Congratulations to the 2026 EurAI distinguished service award winners

AIHub

The EurAI Distinguished Service Award started in 2012, and it is presented annually to individuals who have made exceptional contributions to the European AI community. This year, the award goes to two researchers: Jérôme Lang and Luc de Raedt. Find out who won the small, middle and large divisions in Incheon. Find out the latest from day two of the competition. In the first of our round-ups from the humanoid league we introduce the competition, and report some preliminary results.