Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture covers a two-line proof of the convergence in expectation for the learning rule used in reinforcement learning with a 1-step horizon, demonstrating that the empirical estimate of the Q value converges to the real Q value.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.