Country
ContinuousMean-CovarianceBandits
Specifically,inCMCB, there isalearner who sequentially chooses weight vectors on given options and observes random feedback according to the decisions. The agent's objective is to achieve the best trade-off between reward and risk, measured with option covariance. To capture different reward observation scenarios in practice, we considerthreefeedbacksettings,i.e.,full-information,semi-banditandfull-bandit feedback. Wepropose novelalgorithms withoptimal regrets(within logarithmic factors), and provide matching lower bounds to validate their optimalities. The experimental results also demonstrate the superiority of our algorithms.
ContinuousMean-CovarianceBandits
Specifically,inCMCB, there isalearner who sequentially chooses weight vectors on given options and observes random feedback according to the decisions. The agent's objective is to achieve the best trade-off between reward and risk, measured with option covariance. To capture different reward observation scenarios in practice, we considerthreefeedbacksettings,i.e.,full-information,semi-banditandfull-bandit feedback. Wepropose novelalgorithms withoptimal regrets(within logarithmic factors), and provide matching lower bounds to validate their optimalities. The experimental results also demonstrate the superiority of our algorithms.
0503f5dce343a1d06d16ba103dd52db1-Paper-Conference.pdf
Thisproblem of drawing correspondence is easy for humans: we can match object parts not only across different viewpoints, articulations andlighting changes, butevenacross drastically different categories (e.g., betweencatsandhorses)ordifferentmodalities(e.g.,betweenphotosandcartoons).Yet,werarelyif everget explicit correspondence labels fortraining.