Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture introduces the concept of policy gradients, explaining how actions are associated with observations to optimize rewards parametrically using a gradient method, contrasting it with Q-learning.