Goto

Collaborating Authors

 Technology



RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning

Neural Information Processing Systems

Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At its core, RLHF can be viewed as a specialized instance of Preference-based Reinforcement Learning (PbRL), where the preferences specifically originate from human judgments rather than arbitrary evaluators. Despite this connection, most existing approaches in both RLHF and PbRL primarily focus on optimizing a mean reward objective, neglecting scenarios that necessitate risk-awareness, such as AI safety, healthcare, and autonomous driving. These scenarios often operate under a one-episode-reward setting, which makes conventional risk-sensitive objectives inapplicable.


Sequential Memory with Temporal Predictive Coding Supplementary Materials

Neural Information Processing Systems

In Algorithm 1 we present the memorizing and recalling procedures of the single-layer tPC.Algorithm 1 Memorizing and recalling with single-layer tPC Here we present the proof for Property 1 in the main text, that the single-layer tPC can be viewed as a "whitened" version of the AHN. When applied to the data sequence, it whitens the data such that (i.e., Eq.16 in the main text): These observations are consistent with our numerical results shown in Figure 1. MCAHN has a much larger MSE than that of the tPC because of the entirely wrong recalls. In Figure 1 we also present the online recall results of the models in MovingMNIST, CIFAR10 and UCF101. In Fig 4 we show a natural example of aliased sequences where a movie of a human doing push-ups is memorized and recalled by the model.







Supplementary Material Information Geometry of the Retinal Representation ManifoldXuehao Ding

Neural Information Processing Systems

Further experimental details are described in Ref. [4]. Each spatiotemporal stimulus spanned over 400 ms corresponding to the retinal integration timescale. Figure 1: (a) The log-likelihood of empirical data for each PMF averaged over cells. Black line is the identity line. The central 20 20 arrays are shown.