Topics: Generalization bounds via uniform convergence Theory for deep learning Non-convex optimization Neural tangent kernel Implicit/algorithmic regularization Unsupervised learning and domain adaptation Bandit and online earning https://youtube.com/playlist?list=PLoROMvodv4rP8nAmISxFINlGKSK4rbLKh&si=0BBOR4hJCVQDP51C https://web.stanford.edu/class/stats214/ Francis Bach's book https://www.di.ens.fr/~fbach/ltfp_book.pdf. Maybe not quite entry level, and mostly on classical learning theory a nice one more on NNs is this one https://arxiv.org/abs/2407.18384, it is elementary on the DL but some basic math background is needed, also with some nice pictures :D https://arxiv.org/abs/2106.10165 The Principles of Deep Learning Theory https://www.youtube.com/watch?v=m2bXL5Z5CBM GDL and Categorical DL can be best understood as a principled way of baking in a priori structure/symmetries from your problem/data into your model architecture. e.g. "I want a CNN that is invariant/equivariant under a certain group of transformations, I want an architecture adapted to data structured as binary trees, ...". You trade off some generality for stronger guarantees and possibly fewer parameters https://www.youtube.com/watch?v=bIZB1hIJ4u8 https://geometricdeeplearning.com/ https://www.youtube.com/watch?v=l3O2J3LMxqI https://proceedings.mlr.press/v80/balestriero18b/balestriero18b.pdf https://www.sciencedirect.com/science/article/pii/S0893608024001370 https://arxiv.org/abs/2401.08514 (I've seen other work using homomorphism count/densities as features. It's basically a kernel learning situation) random features: https://people.eecs.berkeley.edu/~brecht/papers/07.rah.rec.nips.pdf