HomeCourse

Deep Learning and Generative Models: Training, Inference, and Model Evaluation

Learn how deep neural networks are trained, adapted, deployed, and evaluated. Part two in a five-part series, this online course in the Deep Learning and Generative Models program connects core optimization methods with practical approaches to transfer learning, efficient inference, and rigorous model evaluation. You will develop skills for adapting pretrained models to specialized tasks, reducing the computational demands of running them, and testing their performance on unfamiliar data. These capabilities help you assess whether a model is ready for use in a product, research project, or technical workflow.

Course Information

Format: Self-Paced
Estimated: 2 weeks, 6-8 hours per week
Start: AnytimeEnd:

Choose Your Path

Certificate Track

$60
Earn a verified certificate of completion
Access to this course & course materials
Graded assignments & exams
MIT Open Learning certificate of completion

Learn for Free

Free
Audit this course
Access to this course & course materials
Part of a Program

About this Course

Training a deep neural network requires more than choosing an architecture. Follow learning from gradient descent and stochastic gradient descent (SGD) to momentum and gradient clipping, which help stabilize updates. Trace backpropagation through a computation graph to see how PyTorch calculates gradients, and examine how weight initialization, normalization, and residual connections support deeper networks.

A model is not finished when training ends. Adapt pretrained models to new tasks, reduce inference costs with knowledge distillation and model compression, and build reliable evaluations using appropriate metrics. Controlled ablation studies isolate individual design choices, while tests under distribution shift reveal how a model handles data unlike its training set.

A PyTorch lab brings these ideas together. Diagnose unstable training, test targeted fixes, and compare pretrained and randomly initialized ResNet-18 models using the same data and training budget.

This is the second of five courses in the Deep Learning and Generative Models Program:

Show more

What you'll learn

  • Train neural networks using gradient descent and backpropagation
  • Diagnose and correct unstable training
  • Explain how initialization, normalization, and residual connections support optimization
  • Adapt pretrained models and reduce inference costs through distillation and compression
  • Evaluate models using appropriate metrics, controlled ablations, and distribution-shift tests

How you'll learn

  • Real-World Learning

    Learn from MIT faculty and experts who ground their teaching in real-world cases rather than mathematical models, making the material approachable for all.

  • Practical Application

    Apply your new knowledge with hands-on, practical exercises drawn from healthcare, sports, finance, sustainability, and more.

  • Learn On Demand

    Access all course content online with complete flexibility to study at your own pace.

  • AI-Enabled Support

    Deepen your understanding of the course material and get help on assignments from AskTIM, the AI assistant built by MIT researchers.

  • Stackable Credentials

    Earn an MIT Open Learning certificate at each milestone—module, course, and program—demonstrating your AI expertise. Available in paid courses only.

Prerequisites

This program is designed for:

  • Aspiring data scientists and machine learning engineers who want a structured introduction to deep learning
  • Quantitative researchers who want to apply deep learning to technical or scientific problems
  • Working professionals who need to evaluate, adapt, or deploy modern machine learning models
  • Advanced undergraduate and graduate students in computer science, data science, statistics, mathematics, engineering, and related quantitative fields
  • Continuing education learners seeking hands-on experience with deep learning and generative models

You should be comfortable programming in Python and working with linear algebra, multivariable calculus, including partial derivatives and the chain rule, and undergraduate probability and statistics. The Probability and Statistical Data Analysis series or equivalent coursework provides suitable preparation. No prior deep learning experience is required.

The labs run in Google Colab and are designed for its free GPU tier, so no local installation is required.

Meet your instructors

Sara Beery

Assistant Professor, MIT Electrical Engineering and Computer Science

Dr. Sara Beery is the Homer A. Burnell Career Development Professor in the MIT Faculty of Artificial Intelligence and Decision-Making. She received her PhD in Computing and Mathematical Sciences at Caltech, where she was advised by Pietro Perona. Her research focuses on building computer vision methods that enable global-scale environmental and biodiversity monitoring across data modalities, tackling real-world challenges including geospatial and temporal domain shift, learning from imperfect data, fine-grained categories, and long-tailed distributions. Her work has been recognized with a Schmidt Sciences AI2050 Early Career Fellowship, an NSF CAREER Grant, the Amori Doctoral Prize, an Amazon AI for Science Fellowship, a PIMCO Data Science Fellowship, and an NSF GRFP. She partners with industry, nongovernmental organizations, and government agencies to deploy her methods in the wild worldwide. She works to increase access to AI skills through interdisciplinary capacity building and education, and was awarded the MIT EECS Outstanding Educator Award, has founded the AI for Conservation slack community, founded and directs the Workshop on Computer Vision Methods for Ecology, and co-leads the NSF/NSERC Global Center on AI and Biodiversity Change.