We consider the problem of designing contextual bandit algorithms in the "cross-learning" setting of Balseiro et al., where the learner observes the loss for the action
We consider the problem of designing contextual bandit algorithms in the "cross-learning" setting of Balseiro et al., where the learner observes the loss for the action
Different distribution shifts require different algorithmic and operational interventions. Methodological research must be grounded by the specific shifts they address.
We study the optimal memorization capacity of modern Hopfield models and Kernelized Hopfield Models (KHMs), a transformer-compatible class of Dense Associative Memories.
Therefore, neuro-symbolic RL aims at creating policies that are interpretable in the first place. Unfortunately, interpretability is not explainability. To achieve both, we introduce Neurally gUided Differentiable loGic policiEs (NUDGE).