Autumn 2026 ✦ Stanford University
MS&E 319: Efficient Generative Models
Given a modeling goal and a limited computational budget, how should one design the objective, architecture, data pipeline, inference method, and serving system?
- Meets Tuesdays and Thursdays, 3:00–4:20 PM — room TBD
- Units 3 (placeholder)
- Format Seminar. One paper presentation and one project per student.
- Prerequisites CS 324, CS 261, or equivalent
- Instructors Amin Karbasi, Anay Mehrotra, Amin Saberi, and Grigoris Velegkas
- Note May be repeated for credit.
Overview
Modern large language models are powerful partly because of their scale, but equally because of the many careful algorithmic choices behind them. This course studies the foundations of those choices. The central question is: given a modeling goal and a limited computational budget, how should one design the objective, architecture, data pipeline, inference method, and serving system?
We approach this question through recent work on efficient language modeling. After an overview of modern LLMs, the course moves through five themes: efficiency in pre-training; efficiency through architecture; efficiency at inference time; efficiency in post-training; and risks and vulnerabilities. The primary emphasis is on algorithmic and theoretical questions in modern LLMs, though we may also cover other generative paradigms, such as diffusion models and world models, depending on time and interest.
The goal is to understand not only what these methods do, but why they work, what tradeoffs they introduce, and what research questions they open. Students will give one presentation and complete either a survey or a research project.
Who this is for
- Students who want to read recent papers carefully and build a principled view of efficient generative modeling — not just a working knowledge of what the methods do.
- Students comfortable with mathematical and algorithmic reasoning who want to move beyond using LLMs as black boxes, and understand the choices that make them trainable, deployable, and reliable.
Schedule
Course overview
Overview of the course and LLMs
Efficiency in Pre-training
Introduction to the pre-training pipeline
Lecture contents
- Architecture of LLMs (then vs. now)
- Loss function for training
- Causal masking
- Scaling laws
- Optimizers
Advances in pre-training
Lecture contents
- Infilling (fill-in-the-middle)
- Document packing
- Multi-document attention and cross-attention
- Multi-token prediction
- Muon
Efficiency via Architecture
Mixtures of experts
Innovations in Attention I: Alternate Architectures
Innovations in Attention II: Sparser, Approximate, Faster
Efficiency at Inference Time
Efficiency via Speculation and Planning
Lecture contents
- Speculative decoding and learned speculative decoding
- Chunked prefill
- Cached queries?
Efficiency via Quantization
Efficiency in Post-training
Introduction to the post-training pipeline I
Introduction to the post-training pipeline II
Advanced Topics in Post-training: Distillation and Test-time Computation
Risks and Vulnerabilities
Jailbreaking, Watermarking, and Cybersecurity
Looking Ahead
Where do we go from here?
Evaluation
- Paper presentation — one per student 35%
- Course project — survey or research 45%
- Participation and discussion 20%
Presentation
Each student presents one paper from the reading list, situating it within the theme of the surrounding lectures: what tradeoff it makes, what it buys, and what it gives up. Sign-ups open in Week 1.
Project
Choose one of two tracks.
- Survey. A focused survey of one thread running through the course, synthesizing what is settled, what is contested, and what remains open.
- Research. An original project — theoretical, empirical, or both — on a question the course material raises.
Milestones (placeholder): proposal in Week 4, checkpoint in Week 7, final report and presentation during finals week.
Instructors
Amin Karbasi
VP and Chief AI Scientist at Cisco; Adjunct Professor, Stanford University
Office hours: TBD.