Explores kernels for simplifying data representation and making it linearly separable in feature spaces, including popular functions and practical exercises.
Explores learning the kernel function in convex optimization, focusing on predicting outputs using a linear classifier and selecting optimal kernel functions through cross-validation.
Covers transformer architecture and subquadratic attention mechanisms, focusing on efficient approximations and their applications in machine learning.