Lifelong Learning in Multi-Armed Bandits
Jedor, Matthieu, Louëdec, Jonathan, Perchet, Vianney
The multi-armed bandit (MAB) is a simple instance of online learning problem with partial feedback (see Bubeck and Cesa-Bianchi [13]; Lattimore and Szepesvári [34]; Slivkins [44], and references therein). In the stochastic bandit framework, a learning agent sequentially pulls arms and obtains noisy rewards. The goal of the agent is to maximize her expected cumulative reward, or equivalently, to minimize her regret, which is the difference between the cumulative reward of an oracle (that knows the mean rewards of arms) and the one of the agent. There is thus a clear trade-off, called "exploration vs exploitation", that arises between gathering information on uncertain arms by pulling them (exploration) and leveraging the information obtained so far by greedily pulling the arm with the highest estimated reward (exploitation). MAB has received a large attention recently due to its wide applications range and the theoretical guarantees associated with learning algorithms (see also Bouneffouf and Rish [9], and references therein for a survey on practical applications of MAB).
Dec-28-2020
- Country:
- North America > United States
- New York > New York County > New York City (0.04)
- Europe > Germany
- Bavaria > Upper Bavaria > Munich (0.04)
- North America > United States
- Genre:
- Research Report (0.64)
- Industry:
- Education > Educational Setting > Continuing Education (0.42)
- Technology: