Adversarial Attacks on Combinatorial Multi-Armed Bandits
Balasubramanian, Rishab, Li, Jiawei, Tadepalli, Prasad, Wang, Huazheng, Wu, Qingyun, Zhao, Haoyu
Multi-armed bandits (MAB) (Auer, 2002) is a classic framework of sequential decision-making problems that has been extensively studied (Lattimore & Szepesvári, 2020; Slivkins et al., 2019). In each round, the learning agent selects one out of m arms and observes its reward feedback which follows an unknown reward distribution. The goal is to maximize the cumulative reward, which requires the agent to balance exploitation (selecting the arm with the highest average reward) and exploration (exploring arms that have high potential but have not been played enough). Combinatorial multi-armed bandits (CMAB) is a generalized setting of original MAB with many real-world applications such as online advertising, ranking, and influence maximization (Liu & Zhao, 2012; Kveton et al., 2015; Chen et al., 2016; Wang & Chen, 2017). In CMAB, the agent chooses a combinatorial action (called a super arm) over the m base arms in each round, and observes outcomes of base arms triggered by the action as feedback, known as the semi-bandit feedback.
Oct-8-2023
- Country:
- North America > United States
- Oregon (0.04)
- Europe > United Kingdom
- England > Cambridgeshire > Cambridge (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Government > Military (0.51)
- Information Technology
- Security & Privacy (0.66)
- Services (0.48)
- Technology: