Position-based Multiple-play Bandit Problem with Unknown Position Bias

Komiyama, Junpei, Honda, Junya, Takeda, Akiko

Dec-31-2017–Neural Information Processing Systems

Motivated by online advertising, we study a multiple-play multi-armed bandit problem with position bias that involves several slots and the latter slots yield fewer rewards. We characterize the hardness of the problem by deriving an asymptotic regret bound. We propose the Permutation Minimum Empirical Divergence (PMED) algorithm and derive its asymptotically optimal regret bound. Because of the uncertainty of the position bias, the optimal algorithm for such a problem requires non-convex optimizations that are different from usual partial monitoring and semi-bandit problems. We propose a cutting-plane method and related bi-convex relaxation for these optimizations by using auxiliary variables.

algorithm, artificial intelligence, big data, (18 more...)

Neural Information Processing Systems

Dec-31-2017

Conferences PDF

Add feedback

Country:
- North America > United States (0.46)

Industry:
- Information Technology > Services (0.34)
- Marketing (0.89)

Technology:
- Information Technology
  - Artificial Intelligence > Machine Learning (1.00)
  - Data Science > Data Mining
    - Big Data (1.00)

Duplicate Docs Excel Report

Title
Position-based Multiple-play Bandit Problem with Unknown Position Bias Junya Honda The University of Tokyo The University of Tokyo / RIKEN
Position-based Multiple-play Bandit Problem with Unknown Position Bias Junya Honda The University of Tokyo The University of Tokyo / RIKEN

Similar Docs Excel Report more

Title	Similarity	Source
None found