Improved Algorithms for Adversarial Bandits with Unbounded Losses

Oct-2-2023–arXiv.org Machine Learning

We consider the Adversarial Multi-Armed Bandits (MAB) problem with unbounded losses, where the algorithms have no prior knowledge on the sizes of the losses. We present UMAB-NN and UMAB-G, two algorithms for non-negative and general unbounded loss respectively. For non-negative unbounded loss, UMAB-NN achieves the first adaptive and scale free regret bound without uniform exploration. Built up on that, we further develop UMAB-G that can learn from arbitrary unbounded loss. Our analysis reveals the asymmetry between positive and negative losses in the MAB problem and provide additional insights. We also accompany our theoretical findings with extensive empirical evaluations, showing that our algorithms consistently out-performs all existing algorithms that handles unbounded losses.

algorithm, exploration, inequality, (16 more...)

arXiv.org Machine Learning

Oct-2-2023

arXiv.org PDF

Add feedback

Country:
- North America > Canada
  - Alberta > Census Division No. 15 > Improvement District No. 9 > Banff (0.04)
- Europe
  - Italy > Sicily (0.04)
  - United Kingdom > England
    - Cambridgeshire > Cambridge (0.04)
  - Spain > Catalonia
    - Barcelona Province > Barcelona (0.04)

Genre:
- Research Report > New Finding (0.46)

Industry:
- Banking & Finance > Trading (0.46)

Technology:
- Information Technology
  - Artificial Intelligence > Machine Learning (1.00)
  - Data Science > Data Mining
    - Big Data (0.88)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found