2026
CLM-AI-003637<= expected regret = O(sum_i (ln T)/Delta_i)UPPER BOUNDThe cited theorem establishes expected regret = O(sum_i (ln T)/Delta_i) under the recorded model.Source: Finite-time Analysis of the Multiarmed Bandit Problem — UCB1 stochastic-bandit regret bound ↗
PROVEN