Covers transformer architecture and subquadratic attention mechanisms, focusing on efficient approximations and their applications in machine learning.
Covers subquadratic attention mechanisms and state space models, focusing on their theoretical foundations and practical implementations in machine learning.