Explores model selection, evaluation, and generalization in machine learning, emphasizing unbiased performance estimation and the risks of over-learning.
Covers topic models, focusing on Latent Dirichlet Allocation, clustering, GMMs, Dirichlet distribution, LDA learning, and applications in digital humanities.