Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture introduces variations of the SARSA algorithm, focusing on expected SARSA and Q learning. Expected SARSA updates the policy by averaging over possible next actions, while Q learning updates the policy by considering the maximum possible action. The instructor explains the differences between these variations and how they impact the learning process.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.