Instructional Material
Scalable Online Planning via Reinforcement Learning Fine-Tuning
Lookahead search has been a critical component of recent AI successes, such as in the games of chess, go, and poker. However, the search methods used in these games, and in many other settings, are tabular. Tabular search methods do not scale well with the size of the search space, and this problem is exacerbated by stochasticity and partial observability. In this work we replace tabular search with online model-based fine-tuning of a policy neural network via reinforcement learning, and show that this approach outperforms state-of-the-art search algorithms in benchmark settings. In particular, we use our search algorithm to achieve a new state-of-the-art result in self-play Hanabi, and show the generality of our algorithm by also showing that it outperforms tabular search in the Atari game Ms. Pacman.
Windows 11 Notepad gets improved context menus in latest update
Ever since Microsoft killed WordPad in 2024, the much-simpler Notepad app has been receiving several new features--almost as if it's evolving into a better, more modern version of WordPad. Meanwhile, Microsoft is introducing an even simpler text editor called Edit. Some of the recent additions to Notepad include spell check, AI-generated text, and Markdown formatting--and the improvements aren't done yet. The latest news is that Notepad will soon have updated context menus in Windows 11, reports Neowin. In Notepad version 11.2507.26.0, which is currently rolling out to Windows Insiders, the updated context menu now matches the look of Windows 11 24H2's context menus, with quick actions for Copy, Cut, Paste, Select all, and Delete, plus other actions like Write, Rewrite, Summarize, Define with Bing, and more.
Online Meta-Learning via Learning with Layer-Distributed Memory
We demonstrate that efficient meta-learning can be achieved via end-to-end training of deep neural networks with memory distributed across layers. The persistent state of this memory assumes the entire burden of guiding task adaptation. Moreover, its distributed nature is instrumental in orchestrating adaptation.
Supplementary Materials for: Training Feedback Spiking Neural Networks by Implicit Differentiation on the Equilibrium State
Input: Network parameters ฮธ; Input data x; Label y; Time steps T; Other hyperparameters; Output: Trained network parameters ฮธ . Calculate the output o and the loss L based on o and y . Update ฮธ based on the gradient-based optimizer. We first prove Theorem 1. Then Theorem 2 is similarly proved. We omit repetitive details here.
Minimax-Optimal Multi-Agent RL in Markov Games With a Generative Model Gen Li UPenn Y uejie Chi CMU Y uting Wei UPenn Y uxin Chen UPenn
All prior results suffer from at least one of the two obstacles: the curse of multiple agents and the barrier of long horizon, regardless of the sampling protocol in use. We take a step towards settling this problem, assuming access to a flexible sampling mechanism: the generative model. Focusing on non-stationary finite-horizon Markov games, we develop a fast learning algorithm called Q-FTRL and an adaptive sampling scheme that leverage the optimism principle in online adversarial learning (particularly the Follow-the-Regularized-Leader (FTRL) method).