HomeCourse

Deep Learning and Generative Models: Neural Network Architectures

Learn how neural network architectures handle different types of data and how to evaluate their performance. Part three in a five-part series, this online course in the Deep Learning and Generative Models program covers convolutional neural networks, transformers, graph neural networks, and the geometry of their parameter spaces. You will learn to match architectures to image, sequential, and relational data, supporting applications such as image recognition, language processing, and analysis of connected systems. Through model comparisons, you will develop the judgment to assess which architectural choices suit a particular problem.

Course Information

Format: Self-Paced
Estimated: 2 weeks, 6-8 hours per week
Start: AnytimeEnd:

Choose Your Path

Certificate Track

$60
Earn a verified certificate of completion
Access to this course & course materials
Graded assignments & exams
MIT Open Learning certificate of completion

Learn for Free

Free
Audit this course
Access to this course & course materials
Part of a Program

About this Course

Neural network architectures are designed around different kinds of structure. Convolutional networks (CNNs) find patterns across images, graph neural networks (GNNs) learn from connections, recurrent networks (RNNs) carry information through sequences, and transformers use attention to connect distant information. Learn what each architecture does well, where it falls short, and how to choose among them.

Then examine parameter space, the possible sets of weights a neural network can learn. Study why different learned settings can produce the same result, how training affects performance on new data, and how targeted changes can steer behavior.

In a hands-on PyTorch lab, compare a CNN and vision transformer on the CIFAR-10 image dataset using the same data and training budget. Test key design choices and measure how they affect accuracy and computational cost.

This is the third of five courses in the Deep Learning and Generative Models Program, which includes:

Show more

How you'll learn

  • Real-World Learning

    Learn from MIT faculty and experts who ground their teaching in real-world cases rather than mathematical models, making the material approachable for all.

  • Practical Application

    Apply your new knowledge with hands-on, practical exercises drawn from healthcare, sports, finance, sustainability, and more.

  • Learn On Demand

    Access all course content online with complete flexibility to study at your own pace.

  • AI-Enabled Support

    Deepen your understanding of the course material and get help on assignments from AskTIM, the AI assistant built by MIT researchers.

  • Stackable Credentials

    Earn an MIT Open Learning certificate at each milestone—module, course, and program—demonstrating your AI expertise. Available in paid courses only.

Prerequisites

This program is designed for:

  • Aspiring data scientists and machine learning engineers who want a structured introduction to deep learning
  • Quantitative researchers who want to apply deep learning to technical or scientific problems
  • Working professionals who need to evaluate, adapt, or deploy modern machine learning models
  • Advanced undergraduate and graduate students in computer science, data science, statistics, mathematics, engineering, and related quantitative fields
  • Continuing education learners seeking hands-on experience with deep learning and generative models

You should be comfortable programming in Python and working with linear algebra, multivariable calculus, including partial derivatives and the chain rule, and undergraduate probability and statistics. The Probability and Statistical Data Analysis series or equivalent coursework provides suitable preparation. No prior deep learning experience is required.

The labs run in Google Colab and are designed for its free GPU tier, so no local installation is required.

Meet your instructors

Phillip Isola

Associate Professor, MIT Electrical Engineering and Computer Science

Dr. Phillip Isola is an associate professor in MIT’s Department of Electrical Engineering and Computer Science. He earned his PhD in Brain and Cognitive Sciences at MIT, advised by Ted Adelson, and his undergraduate degree in computer science at Yale. His research spans computer vision, machine learning, robotics, and artificial intelligence, with contributions to generative AI and self-supervised representation learning. His group investigates how intelligence emerges and whether models trained on different data and modalities develop similar representations of the world. This work aims to uncover principles of human-like intelligence and inform the development of beneficial AI systems. Before joining MIT, he conducted postdoctoral research at UC Berkeley and worked as a visiting research scientist at OpenAI; his research experience also includes Google Research. His honors include Packard and Sloan fellowships, the PAMI Young Researcher Award, a Google Faculty Research Award, and Samsung’s AI Researcher of the Year Award. He received the Ruth and Joel Spira Award for Distinguished Teaching and has taught MIT courses in deep learning, computer vision, and embodied intelligence. Faculty website, official bio