Goto

Collaborating Authors

 Technology





1 Game Dataset 2 Language Dataset Online Game Pro Game General Text Wiki Puzzle Book

Neural Information Processing Systems

When solving decision-making tasks, humans typically depend on information from two key sources: (1) Historical policy data, which provides interaction replay from the environment, and (2) Analytical insights in natural language form, exposing the invaluable thought process or strategic considerations. Despite this, the majority of preceding research focuses on only one source: they either use historical replay exclusively to directly learn policy or value functions, or engaged in language model training utilizing mere language corpus. In this paper, we argue that a powerful autonomous agent should cover both sources. Thus, we propose ChessGPT, a GPT model bridging policy learning and language modeling by integrating data from these two sources in Chess games. Specifically, we build a large-scale game and language dataset related to chess.





where Ns,k(t) = k τs+k τs Ns,k 1(t)

Neural Information Processing Systems

We will prove by the induction. Let's suppose that the formula holds for k up to n. We will prove that this formula also holds for k = n+1. By the definition in Eq. 4 and the chain rule, we can get that: Ns,n+1(t) = t τs A.2 Spline representation In this section, we give error bounds for spline representation. For simplicity, we consider 1D scenario and assume the target function u: [0,1] R is periodic and defined on the unit interval Ω = [0,1].