Goto

Collaborating Authors

 Agents




Appendix A Lower Bound In this section, we establish a lower bound on the expected regret of any algorithm for our multi-agent

Neural Information Processing Systems

Our goal in this section is twofold. To avoid excessive repetition of notation and proof arguments, we purposefully leave this section not self-contained and only outline these adjustments needed. We refer an interested reader to the work of Auer et al. We leave it to future work to optimize the dependence on N and K . In our MA-MAB problem, an algorithm is allowed to "pull" a distribution over the arms In the proof of their Lemma A.1, in the explanation of their Equation (30), they cite the assumption that given the rewards observed in the first Finally, in the proof of their Theorem A.2, they again consider the probability Thus, we have the following lower bound.





Appendices for No-regret Learning in Price Competitions under Consumer Reference Effects A Expanded Literature Review

Neural Information Processing Systems

There are also very recent works that address the dynamic pricing problem with consumer reference effects under uncertain demand. Nevertheless, these two lines of works are oblivious to consumer reference effects. In contrast to these two papers, our work studies price competitions over an infinite time horizon where reference prices adjust over time, and provides theoretical guarantees for the convergence of pricing strategies under the partial information setting. In their setting, the subgradient for each bidder's objective is a function of all bidders' decisions as well as its budget rate (i.e. total fixed budget divided by a given time horizon), which can be B.1 Proof of Theorem 3.1 (i) By first order conditions, we know that arg max We now follow a similar proof to that of Tarski's fixed point theorem: consider the set Note that convergence is monotonic because U () is nondecreasing. This implies that under Assumption 1, the interior SNE is unique.



We first thank all reviewers for their thoughtful comments, and we wish everyone health during these hard times

Neural Information Processing Systems

We first thank all reviewers for their thoughtful comments, and we wish everyone health during these hard times. We acknowledge the simplicity in our linear demand and reference price update models. These references are also discussed in Section 2 of the paper. The gradient of revenue can be calculated using estimated elasticity, observed sales (i.e. Assumption 1 is invoked in all theorems and lemmas of Section 5, and we will clearly state this in the revised paper. In the proof of Lemma 3.2, we show that This means if firms are willing to consider both prices near zero and those sufficiently large, Assumption 1 holds.