Lecture
Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture covers Transformer networks and self-attention layers, explaining how they map sets of inputs and the concept of multi-head attention. It delves into the process of learning weights, the importance of positional encoding, and the interpretability of the heads.