Multi-Armed Bandits for Adaptive Constraint Propagation
Balafrej, Amine (TASC (INRIA/CNRS), Mines Nantes) | Bessiere, Christian (CNRS, University of Montpellier) | Paparrizou, Anastasia (CNRS, University of Montpellier)
It allows a constraint to play each one. Each machine, after being used, returns a reward solver to exploit various levels of propagation during from a distribution specific to that machine. The goal is search, and in many cases it shows better performance to maximize the sum of rewards obtained through a sequence than static/predefined. The crucial point of plays [Gittins, 1989]. is to make adaptive constraint propagation automatic, We use a MAB model to select the right level of propagation so that no expert knowledge or parameter (also called level of consistency) to enforce at each node specification is required. In this work, we propose during the exploration of the search tree. We specify a simple a simple learning technique, based on multiarmed reward function and the upper confidence bound (UCB) to estimate bandits, that allows to automatically select the best arm, namely the best consistency to apply.
Jul-15-2015
- Technology: