Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture delves into actor-critic networks, specifically the advantage actor critic networks, which combine TD learning with policy gradient for optimizing parameters to maximize return. The comparison between actor critic and reinforce with baseline methods is explored, highlighting differences in V value estimation and parameter updates.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.