Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture presents a quiz where the instructor discusses various claims related to reinforcement learning algorithms, such as the use of Q values or V values, the transition from batch to online learning, the optimization of expected total reward, and the intuitive meaning of the derivative of the log policy.
Network: Computation in Neural Systems', Journal of Computational Neuroscience', and `Science'.