Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture introduces policy gradient methods using a simple example of a single neuron with binary output, focusing on the disadvantages of Q-learning, SARSA, and TD-learning, and explaining the basic idea of policy gradient methods to optimize rewards directly.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.