🌏 中文版
This series is based on the Fall 2025 edition of Stanford CS149. When I checked on 2026-09-30,
cs149.stanford.edustill redirected to the fall25 site and the fall26 URL returned 404. This is post 0 of the series and its overview; every later post points back here for materials and limits.
Most AI engineers now work on parallel hardware every day. Training runs on GPU clusters, inference means squeezing kernel performance, and even phones ship an NPU. Yet many people have only a vague sense of why GPUs are fast, or why adding cores didn't make their program faster. CS149: Parallel Computing fills that gap.
What the course teaches
The course home page opens with the claim that parallel processing is everywhere, from smartphones, multi-core CPUs, GPUs, and AI accelerators to the largest supercomputers. The course aims to give a deep understanding of the principles and engineering trade-offs behind parallel systems, and to teach the programming techniques needed to use them well. Writing good parallel programs requires understanding a machine's performance characteristics, so the course covers both hardware and software.
Kayvon Fatahalian and Kunle Olukotun co-taught Fall 2025. There were 18 lectures from Sep 23 to Dec 4, an evening midterm on Nov 18, and a final on Dec 11. The lectures fall into five parts:
| Part | Lectures | Topics |
|---|---|---|
| Part 1 | L1–L3 | Why parallelism, multi-core processors, latency vs. bandwidth, ISPC |
| Part 2 | L4–L6 | How to think about parallelizing code, work distribution and scheduling, locality and communication |
| Part 3 | L7–L8 | GPU architecture and CUDA, data-parallel thinking |
| Part 4 | L9–L13 | DNNs on GPUs, hardware specialization, programming systems for specialized hardware, the AI datacenter, DSLs and AI-driven performance optimization |
| Part 5 | L14–L18 | Cache coherence, synchronization and memory consistency, fine-grained locking and lock-free programming, transactional memory |
The first lecture's slides organize the course around three themes: designing parallel programs that scale, how parallel hardware is implemented, and thinking about efficiency. The third theme comes with a line worth remembering from day one: FAST != EFFICIENT. A 2x speedup on 10 processors does make the program faster. Whether it uses the hardware well is a separate question.
Prerequisite self-check
The course info page calls CS111 a "strongly recommended" prerequisite and lists the concepts it expects you to know. Use it as a checklist:
- A compiled program is a sequence of machine instructions. A processor executes them, and each instruction updates state in registers or memory.
- Why a memory hierarchy exists, and how registers, on-chip caches, off-chip memory, and persistent storage implement it.
- You can read, write, and debug C/C++ (classes, STL vectors).
- You have written at least one program that creates threads, for example with
std::threador pthreads.
The page is blunt: the main reason students struggle in CS149 is a lack of debugging experience, because parallel code is hard to debug. Assignments use new C-like languages such as CUDA and ISPC, and the course expects you to pick them up as you go.
If the first two items feel shaky, start with this site's Stanford CS107 guide (machine instructions, assembly, caches, and the memory hierarchy). If threads, locks, and scheduling are new, read the Stanford CS111 guide.
Assignments, exams, and grading
Grading from the course info page:
| Component | Weight |
|---|---|
| 5 programming assignments | 8% + 12% + 12% + 12% + 12% = 56% |
| 4 written assignments | 3% × 4 = 12% |
| Per-lecture participation (in-class quiz) | 4% |
| Midterm | 12% |
| Final | 16% |
Programming assignments can be done in pairs, and one- and two-person teams are graded the same way. Written assignments must be done in groups of three, with random partners assigned by the staff. Each student gets 8 late days for the quarter, for programming assignments only, and not for PA5.
The five programming assignments, with due dates from the home page:
| Assignment | Due | Topic | Environment |
|---|---|---|---|
| PA1 | Oct 6 | Analyzing parallel program performance on a quad-core CPU (threads, SIMD intrinsics, ISPC) | Stanford myth machines (quad-core Intel Core i7) |
| PA2 | Oct 16 | Scheduling task graphs on a multi-core CPU | AWS c7g.4xlarge |
| PA3 | Oct 30 | A circle renderer in CUDA | NVIDIA T4 on AWS |
| PA4 | Nov 13 | Fused conv + maxpool on the Trainium2 accelerator | AWS trn2.3xlarge |
| PA5 | Dec 4 | Make the world's fastest kernels (open-ended) | Course-managed H100 job queue |
The four written assignments are PDFs: Written 1, Written 2, Written 3, and Written 4. The L1 slides say they contain modified versions of past exam questions, so they double as exam practice. Besides the graded questions, each PDF includes several problems marked PRACTICE PROBLEM. For a self-learner, these four PDFs are the closest thing to an exam.
There is no required textbook. The course info page suggests Hennessy & Patterson, Computer Architecture: A Quantitative Approach, 6th edition, as an architecture reference, and notes that plenty of good parallel programming material is free online.
What outside readers can access: A3, with four gaps
Under the grading in this site's global AI/CS course map, CS149 Fall 2025 is A3 (self-study ready). What's public:
- 18 lecture slide decks, each as a PDF and as a slide-by-slide web page
- GitHub repos for all 5 assignments, with starter code, READMEs, and grading rules
- 4 written-assignment PDFs
The gaps need to be stated up front:
- Fall 2025 videos are not public. Course info says lectures are recorded for viewing on Canvas. The home page says outright, "We cannot distribute lecture videos to the public this year," and points to the public 2023 videos instead.
- PA1 and PA2 grading machines are out of reach. The PA1 README asks you to run on the myth machines and report those numbers; PA2 is graded on AWS
c7g.4xlarge. You can do both on your own multi-core CPU, but your speedups won't compare directly with the official reference. - PA4 is effectively A2. PA4's cloud_readme has students boot from a private course AMI and buy a capacity block for
trn2.3xlarge. It lists the upfront price as of 2025-10-31 at $2.25 per hour, about $300 for 7 days. Enrolled students got an extra $400 in AWS credit; outside readers don't. What you can do is read the README and starter code, or build a Neuron environment at your own expense. - PA5's leaderboard requires a SUNet ID. Submitting to the H100 job queue starts with registering in popcorn-cli using a SUNet ID. The README also says you can develop locally on any CUDA-capable NVIDIA GPU and test and profile with
eval.py. That is the only route for outside readers, and there is no leaderboard to compare against.
Also: PA3 needs an NVIDIA GPU (the README uses a T4 as reference); no solutions are published for any assignment or written question; announcements and discussion live on Ed, which outsiders can't see.
Why the 2023 videos are a supplement
The course home page itself sends readers to the 2023 YouTube playlist, which has 19 videos. This series follows two rules:
- Fall 2025 slides are the source of truth. The 2023 videos are a listening supplement. Where the two disagree, 2025 wins, and each post flags the difference.
- The reason is simple. Outsiders can't watch the 2025 recordings, and 2023 is the substitute the course points to. I compared the slide text of L1 and L2 across 2023 and 2025. The technical content is largely the same, and the differences are mostly logistics and a few added slides. Each later post notes the differences for its own lecture.
The two terms don't line up exactly, though. 2023 had lectures on Spark (L9) and Accessing Memory (L19) that Fall 2025 dropped. Fall 2025 added L11 (programming systems for specialized hardware), L12 (the AI datacenter), and the AI-driven optimization half of L13, none of which has a 2023 recording. Those posts rely on slides alone.
| Fall 2025 lecture | Matching 2023 video |
|---|---|
| L1 Why Parallelism? Why Efficiency? | L1 |
| L2 A Modern Multi-Core Processor (Part I) | L2 |
| L3 Multi-Core Architecture (Part II) + ISPC | L3 |
| L4 Parallelizing Code: An Example Thought Process | L4 Parallel Programming Basics |
| L5 Work Distribution and Scheduling | L5 |
| L6 Locality and Communication | L6 |
| L7 GPU Architecture and CUDA Programming | L7 |
| L8 Data-Parallel Thinking | L8 |
| L9 Efficiently Evaluating DNNs on GPUs | L10 |
| L10 Hardware Specialization | L18 |
| L11 Programming Systems for Specialized Hardware | None |
| L12 Mapping AI Applications to the Datacenter Computer | None |
| L13 Domain-Specific Programming Systems + AI-Driven Optimization | DSL half: L15; AI half: none |
| L14 Cache Coherence | L11 |
| L15 Implementing Synchronization + Memory Consistency | L12 |
| L16 Fine-Grained Locking and Lock-Free Programming | L13 |
| L17 Transactional Memory (Part I) | L16 |
| L18 Transactional Memory (Part II) + AMA | L17 |
The assignments changed too. The 2023 L1 slides say "Four programming assignments"; 2025 has five, and the new one is PA5. If a 2023 video mentions assignment details, check the 2025 repos instead.
Series arc
This series keeps the official lecture order, because assignments are tied to specific lectures and reordering would break their prerequisites. Each assignment post sits after the lectures it depends on, and written assignments are folded into the same post.
Part 1: Why parallelism, and what a processor looks like
Part 2: Parallelizing code and making it fast
Part 3: GPUs and data-parallel thinking
Part 4: AI systems
Part 5: Correctness in shared memory
Two reading paths
Full path: read 0 → 22 in order. Part 5 only depends on Parts 1–2, so if you want the shared-memory foundations first, jump to 19–22 after post 8, then come back for GPUs and AI.
AI systems only: 0 → 1–3 → 7 → 9–10 → 12–18. Posts 1–3 build intuition for SIMD, multithreading, and bandwidth limits; post 7 adds arithmetic intensity; then you go straight into GPUs and AI hardware. The cost is skipping the hands-on assignment posts, plus coherence and synchronization.
Things to do tonight
- Open the L1 slides, find the efficiency slide, and decide whether "2x speedup on 10 processors" is a good result.
- Go through the prerequisite checklist above. For any item you're unsure of, read the matching post in the CS107 or CS111 guide.
- Clone the PA1 repo, check how many cores your machine has and whether it supports AVX2, and install ISPC. PA1 is the assignment outsiders can reproduce most completely.
Further reading
These site series overlap with CS149. This series does not cut content because of them; the links are here for reference:
- Reading Stanford CS336: training language models from scratch. For the systems side, see GPUs and TPUs, kernels and Triton, parallelism mechanics, and parallelism strategies
- Reading CMU 11-868 LLM Systems: the systems side of large language models; GPU programming and acceleration and FlashAttention pair with Parts 3–4 here
- CME295 LLM systems: inference and serving systems from an LLM course's angle
- Global AI/CS course map: definitions of the A0–A3 grades
Series navigation: next, L1 Why parallelism, why efficiency
References
- CS149 Fall 2025 home page and schedule
- CS149 Fall 2025 Course Info (prerequisites, grading, late days)
- CS149 Fall 2025 lecture index (18 slide decks)
- L1 slides PDF: Why Parallelism? Why Efficiency? (Fall 2025)
- L1 slides PDF (Fall 2023, for comparison)
- PA1: Analyzing Parallel Program Performance on a Quad-Core CPU
- PA2: Scheduling Task Graphs on a Multi-Core CPU
- PA3: A Circle Renderer in CUDA
- PA4: Fused Conv+MaxPool on Trainium2
- PA4 cloud_readme (private AMI, capacity block pricing)
- PA5: Make the World's Fastest CUDA Kernels
- Written Assignment 1, 2, 3, 4
- CS149 2023 YouTube playlist (Stanford Online)
Loading...