Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture introduces the concept of Bandit Problems in Reinforcement Learning, where one has to choose between different actions and immediately receives a reward. The slides cover topics such as one-step horizon games, Q-values, optimal policy, iterative update rules, empirical averaging, and convergence in expectation.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.