Course Curriculum
17 chapters are available now.
Getting Started
Evaluating Models Properly
Generalization, Validation, and Test Data
Explains why training accuracy is misleading, and introduces the training, validation, and test set paradigm along with overfitting, underfitting, and cross-validation as tools for honestly measuring performance.
Diagnosing Training
Covers learning curves, the bias-variance tradeoff, and metrics beyond accuracy such as precision, recall, and confusion matrices for understanding how and why a model fails.
Digit Classifier: Validation and Diagnostics
A hands-on lab to add data splitting, learning curves, and evaluation metrics to the digit classifier, and use them to diagnose its behavior.
Rethinking the Building Blocks
Activation Functions
Examines the vanishing gradient problem and why sigmoid falls short in deep networks, then explores ReLU, Leaky ReLU, GELU, and Softmax, and how the choice of activation shapes what a network can learn.
Cost Functions
Introduces cross-entropy and log-likelihood as principled loss functions for classification, explaining why mean squared error breaks down and how the right cost function accelerates learning.
Digit Classifier: Upgrading Activations and Cost
A lab to swap in ReLU, softmax, and cross-entropy loss, then measure how each change affects training speed and accuracy.
Stabilizing and Accelerating Training
Normalization & Initialization
Covers techniques for keeping values well-scaled throughout the network, from normalizing inputs and choosing principled weight initializations like Xavier/He to batch normalization, which re-centers and rescales activations to stabilize deep networks.
Improving Gradient Descent
Explores momentum and adaptive learning rate methods such as RMSProp and Adam, which smooth out noisy updates and escape shallow local minima more effectively than vanilla SGD.
Digit Classifier: Better Optimization
A lab to implement better initialization, batch normalization, and Adam, comparing convergence against the original training loop.
Controlling Overfitting
Regularization
Covers L1 and L2 regularization, weight decay, and early stopping, building intuition for how penalizing large weights encourages a network to learn simpler, more generalizable solutions.
Dropout and Data Augmentation
Explores dropout as implicit ensembling and data augmentation as a way to expand the training set, along with other practical techniques for closing the gap between training and validation performance.
Digit Classifier: Regularization
A lab to apply weight decay, dropout, early stopping, and augmentation, using validation curves to measure their impact on generalization.
Putting It All Together
Hyperparameter Tuning
Covers practical tuning strategies including learning rate schedules, batch size, grid and random search, and experiment tracking, along with a checklist for debugging networks that fail to train.
Digit Classifier: The Optimized Network
A capstone lab combining every technique from the course into a tuned, well-validated digit classifier, with a final evaluation on the held-out test set.
Conclusion
Conclusion
Reviews the key ideas of the course and offers guidance on where to go next in your machine learning journey.
Credits
A comprehensive list of credits for the sources that inspired and informed the course material. Useful for further reading.