Online Sign Identification: Minimization of the Number of Errors in Thresholding Bandits

Sep-30-2024, 06:09:20 GMT–Neural Information Processing Systems

In the fixed budget thresholding bandit problem, an algorithm sequentially allocates a budgeted number of samples to different distributions. It then predicts whether the mean of each distribution is larger or lower than a given threshold. We introduce a large family of algorithms (containing most existing relevant ones), inspired by the Frank-Wolfe algorithm, and provide a thorough yet generic analysis of their performance. This allowed us to construct new explicit algorithms, for a broad class of problems, whose losses are within a small constant factor of the non-adaptive oracle ones. Quite interestingly, we observed that adaptive methods empirically greatly out-perform non-adaptive oracles, an uncommon behavior in standard online learning settings, such as regret minimization. We explain this surprising phenomenon on an insightful toy problem.

algorithm, allocation, oracle, (14 more...)

Neural Information Processing Systems

Sep-30-2024, 06:09:20 GMT

Conferences PDF

Add feedback

Country:
- Europe
  - France
    - Auvergne-Rhône-Alpes > Isère
      - Grenoble (0.04)
    - Hauts-de-France > Nord
      - Lille (0.04)
  - United Kingdom > England
    - Cambridgeshire > Cambridge (0.04)
- North America > United States
  - New York > New York County > New York City (0.04)

Genre:
- Research Report (0.67)

Industry:
- Education (0.34)

Technology:
- Information Technology
  - Artificial Intelligence > Machine Learning (1.00)
  - Data Science > Data Mining
    - Big Data (0.68)

Duplicate Docs Excel Report

Title
bb3ea2b28c563b1fd6bc89a32ff4d14b-Paper.pdf
bb3ea2b28c563b1fd6bc89a32ff4d14b-Paper.pdf

Similar Docs Excel Report more

Title	Similarity	Source
None found