Pessimistic Off-Policy Multi-Objective Optimization

Alizadeh, Shima, Bhargava, Aniruddha, Gopalswamy, Karthick, Jain, Lalit, Kveton, Branislav, Liu, Ge

Oct-28-2023–arXiv.org Machine Learning

Multi-objective optimization is a type of decision making problems where multiple conflicting objectives are optimized. We study offline optimization of multi-objective policies from data collected by an existing policy. We propose a pessimistic estimator for the multi-objective policy values that can be easily plugged into existing formulas for hypervolume computation and optimized. The estimator is based on inverse propensity scores (IPS), and improves upon a naive IPS estimator in both theory and experiments. Our analysis is general, and applies beyond our IPS estimators and methods for optimizing them. The pessimistic estimator can be optimized by policy gradients and performs well in all of our experiments.

artificial intelligence, machine learning, optimization, (18 more...)

arXiv.org Machine Learning

Oct-28-2023

arXiv.org PDF

Add feedback

Country:
- Europe
  - United Kingdom > England
    - Oxfordshire > Oxford (0.04)
  - Netherlands > South Holland
    - Leiden (0.04)
  - Germany > North Rhine-Westphalia
    - Münster Region > Münster (0.04)

Genre:
- Research Report (0.64)

Technology:
- Information Technology > Artificial Intelligence
  - Representation & Reasoning > Optimization (1.00)
  - Machine Learning (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found