Conservative Contextual Linear Bandits
Kazerouni, Abbas, Ghavamzadeh, Mohammad, Yadkori, Yasin Abbasi, Roy, Benjamin Van
–Neural Information Processing Systems
Safety is a desirable property that can immensely increase the applicability of learning algorithms in real-world decision-making problems. It is much easier for a company to deploy an algorithm that is safe, i.e., guaranteed to perform at least as well as a baseline. In this paper, we study the issue of safety in contextual linear bandits that have application in many different fields including personalized ad recommendation in online marketing. We formulate a notion of safety for this class of algorithms. We develop a safe contextual linear bandit algorithm, called conservative linear UCB (CLUCB), that simultaneously minimizes its regret and satisfies the safety constraint, i.e., maintains its performance above a fixed percentage of the performance of a baseline strategy, uniformly over time.
Neural Information Processing Systems
Feb-14-2020, 14:11:56 GMT