Goto

Collaborating Authors

 Country




ContinuousMean-CovarianceBandits

Neural Information Processing Systems

Specifically,inCMCB, there isalearner who sequentially chooses weight vectors on given options and observes random feedback according to the decisions. The agent's objective is to achieve the best trade-off between reward and risk, measured with option covariance. To capture different reward observation scenarios in practice, we considerthreefeedbacksettings,i.e.,full-information,semi-banditandfull-bandit feedback. Wepropose novelalgorithms withoptimal regrets(within logarithmic factors), and provide matching lower bounds to validate their optimalities. The experimental results also demonstrate the superiority of our algorithms.


ContinuousMean-CovarianceBandits

Neural Information Processing Systems

Specifically,inCMCB, there isalearner who sequentially chooses weight vectors on given options and observes random feedback according to the decisions. The agent's objective is to achieve the best trade-off between reward and risk, measured with option covariance. To capture different reward observation scenarios in practice, we considerthreefeedbacksettings,i.e.,full-information,semi-banditandfull-bandit feedback. Wepropose novelalgorithms withoptimal regrets(within logarithmic factors), and provide matching lower bounds to validate their optimalities. The experimental results also demonstrate the superiority of our algorithms.





0503f5dce343a1d06d16ba103dd52db1-Paper-Conference.pdf

Neural Information Processing Systems

Thisproblem of drawing correspondence is easy for humans: we can match object parts not only across different viewpoints, articulations andlighting changes, butevenacross drastically different categories (e.g., betweencatsandhorses)ordifferentmodalities(e.g.,betweenphotosandcartoons).Yet,werarelyif everget explicit correspondence labels fortraining.