Volkan Cevher, Luca Viano
In this paper, we investigate the existence of online learning algorithms with bandit feedback that simultaneously guarantee O(1) regret compared to a given comparator strategy, and Õ(√ T) regret compared to any fixed strategy, where T is the number of rou ...
2025