Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture introduces the Bellman equation, which connects Q-values of state-action pairs with future rewards. It covers the importance of the discount factor, the concept of total expected discounted reward, and the value consistency of neighboring states. The instructor explains how the Bellman equation is used to determine optimal actions and the implications of different policies on the equation's formulation.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.