Skip to content

Reading Stanford CS149: A Guide to the Fall 2025 Parallel Computing Course

Sep 30, 20261 min
TL;DRCS149 is Stanford's parallel computing course, taught by Kayvon Fatahalian and Kunle Olukotun. It runs from multi-core CPUs and SIMD through GPUs, AI accelerators, and the datacenter, then returns to cache coherence and lock-free programming. For Fall 2025, all 18 slide decks, the starter code and READMEs for 5 programming assignments, and 4 written-assignment PDFs are public, so this series rates it A3 (self-study ready). There are four gaps: the Fall 2025 lecture videos are Canvas-only; PA1 is graded on Stanford's myth machines; PA4 needs a self-funded AWS Trainium2 instance and a private course AMI; PA5's H100 job queue and leaderboard require a SUNet ID. The public videos are from 2023, and this series treats them as a listening supplement only.

🌏 中文版

This series is based on the Fall 2025 edition of Stanford CS149. When I checked on 2026-09-30, cs149.stanford.edu still redirected to the fall25 site and the fall26 URL returned 404. This is post 0 of the series and its overview; every later post points back here for materials and limits.

Most AI engineers now work on parallel hardware every day. Training runs on GPU clusters, inference means squeezing kernel performance, and even phones ship an NPU. Yet many people have only a vague sense of why GPUs are fast, or why adding cores didn't make their program faster. CS149: Parallel Computing fills that gap.

What the course teaches

The course home page opens with the claim that parallel processing is everywhere, from smartphones, multi-core CPUs, GPUs, and AI accelerators to the largest supercomputers. The course aims to give a deep understanding of the principles and engineering trade-offs behind parallel systems, and to teach the programming techniques needed to use them well. Writing good parallel programs requires understanding a machine's performance characteristics, so the course covers both hardware and software.

Kayvon Fatahalian and Kunle Olukotun co-taught Fall 2025. There were 18 lectures from Sep 23 to Dec 4, an evening midterm on Nov 18, and a final on Dec 11. The lectures fall into five parts:

PartLecturesTopics
Part 1L1–L3Why parallelism, multi-core processors, latency vs. bandwidth, ISPC
Part 2L4–L6How to think about parallelizing code, work distribution and scheduling, locality and communication
Part 3L7–L8GPU architecture and CUDA, data-parallel thinking
Part 4L9–L13DNNs on GPUs, hardware specialization, programming systems for specialized hardware, the AI datacenter, DSLs and AI-driven performance optimization
Part 5L14–L18Cache coherence, synchronization and memory consistency, fine-grained locking and lock-free programming, transactional memory

The first lecture's slides organize the course around three themes: designing parallel programs that scale, how parallel hardware is implemented, and thinking about efficiency. The third theme comes with a line worth remembering from day one: FAST != EFFICIENT. A 2x speedup on 10 processors does make the program faster. Whether it uses the hardware well is a separate question.

Prerequisite self-check

The course info page calls CS111 a "strongly recommended" prerequisite and lists the concepts it expects you to know. Use it as a checklist:

  • A compiled program is a sequence of machine instructions. A processor executes them, and each instruction updates state in registers or memory.
  • Why a memory hierarchy exists, and how registers, on-chip caches, off-chip memory, and persistent storage implement it.
  • You can read, write, and debug C/C++ (classes, STL vectors).
  • You have written at least one program that creates threads, for example with std::thread or pthreads.

The page is blunt: the main reason students struggle in CS149 is a lack of debugging experience, because parallel code is hard to debug. Assignments use new C-like languages such as CUDA and ISPC, and the course expects you to pick them up as you go.

If the first two items feel shaky, start with this site's Stanford CS107 guide (machine instructions, assembly, caches, and the memory hierarchy). If threads, locks, and scheduling are new, read the Stanford CS111 guide.

Assignments, exams, and grading

Grading from the course info page:

ComponentWeight
5 programming assignments8% + 12% + 12% + 12% + 12% = 56%
4 written assignments3% × 4 = 12%
Per-lecture participation (in-class quiz)4%
Midterm12%
Final16%

Programming assignments can be done in pairs, and one- and two-person teams are graded the same way. Written assignments must be done in groups of three, with random partners assigned by the staff. Each student gets 8 late days for the quarter, for programming assignments only, and not for PA5.

The five programming assignments, with due dates from the home page:

AssignmentDueTopicEnvironment
PA1Oct 6Analyzing parallel program performance on a quad-core CPU (threads, SIMD intrinsics, ISPC)Stanford myth machines (quad-core Intel Core i7)
PA2Oct 16Scheduling task graphs on a multi-core CPUAWS c7g.4xlarge
PA3Oct 30A circle renderer in CUDANVIDIA T4 on AWS
PA4Nov 13Fused conv + maxpool on the Trainium2 acceleratorAWS trn2.3xlarge
PA5Dec 4Make the world's fastest kernels (open-ended)Course-managed H100 job queue

The four written assignments are PDFs: Written 1, Written 2, Written 3, and Written 4. The L1 slides say they contain modified versions of past exam questions, so they double as exam practice. Besides the graded questions, each PDF includes several problems marked PRACTICE PROBLEM. For a self-learner, these four PDFs are the closest thing to an exam.

There is no required textbook. The course info page suggests Hennessy & Patterson, Computer Architecture: A Quantitative Approach, 6th edition, as an architecture reference, and notes that plenty of good parallel programming material is free online.

What outside readers can access: A3, with four gaps

Under the grading in this site's global AI/CS course map, CS149 Fall 2025 is A3 (self-study ready). What's public:

  • 18 lecture slide decks, each as a PDF and as a slide-by-slide web page
  • GitHub repos for all 5 assignments, with starter code, READMEs, and grading rules
  • 4 written-assignment PDFs

The gaps need to be stated up front:

  1. Fall 2025 videos are not public. Course info says lectures are recorded for viewing on Canvas. The home page says outright, "We cannot distribute lecture videos to the public this year," and points to the public 2023 videos instead.
  2. PA1 and PA2 grading machines are out of reach. The PA1 README asks you to run on the myth machines and report those numbers; PA2 is graded on AWS c7g.4xlarge. You can do both on your own multi-core CPU, but your speedups won't compare directly with the official reference.
  3. PA4 is effectively A2. PA4's cloud_readme has students boot from a private course AMI and buy a capacity block for trn2.3xlarge. It lists the upfront price as of 2025-10-31 at $2.25 per hour, about $300 for 7 days. Enrolled students got an extra $400 in AWS credit; outside readers don't. What you can do is read the README and starter code, or build a Neuron environment at your own expense.
  4. PA5's leaderboard requires a SUNet ID. Submitting to the H100 job queue starts with registering in popcorn-cli using a SUNet ID. The README also says you can develop locally on any CUDA-capable NVIDIA GPU and test and profile with eval.py. That is the only route for outside readers, and there is no leaderboard to compare against.

Also: PA3 needs an NVIDIA GPU (the README uses a T4 as reference); no solutions are published for any assignment or written question; announcements and discussion live on Ed, which outsiders can't see.

Why the 2023 videos are a supplement

The course home page itself sends readers to the 2023 YouTube playlist, which has 19 videos. This series follows two rules:

  • Fall 2025 slides are the source of truth. The 2023 videos are a listening supplement. Where the two disagree, 2025 wins, and each post flags the difference.
  • The reason is simple. Outsiders can't watch the 2025 recordings, and 2023 is the substitute the course points to. I compared the slide text of L1 and L2 across 2023 and 2025. The technical content is largely the same, and the differences are mostly logistics and a few added slides. Each later post notes the differences for its own lecture.

The two terms don't line up exactly, though. 2023 had lectures on Spark (L9) and Accessing Memory (L19) that Fall 2025 dropped. Fall 2025 added L11 (programming systems for specialized hardware), L12 (the AI datacenter), and the AI-driven optimization half of L13, none of which has a 2023 recording. Those posts rely on slides alone.

Fall 2025 lectureMatching 2023 video
L1 Why Parallelism? Why Efficiency?L1
L2 A Modern Multi-Core Processor (Part I)L2
L3 Multi-Core Architecture (Part II) + ISPCL3
L4 Parallelizing Code: An Example Thought ProcessL4 Parallel Programming Basics
L5 Work Distribution and SchedulingL5
L6 Locality and CommunicationL6
L7 GPU Architecture and CUDA ProgrammingL7
L8 Data-Parallel ThinkingL8
L9 Efficiently Evaluating DNNs on GPUsL10
L10 Hardware SpecializationL18
L11 Programming Systems for Specialized HardwareNone
L12 Mapping AI Applications to the Datacenter ComputerNone
L13 Domain-Specific Programming Systems + AI-Driven OptimizationDSL half: L15; AI half: none
L14 Cache CoherenceL11
L15 Implementing Synchronization + Memory ConsistencyL12
L16 Fine-Grained Locking and Lock-Free ProgrammingL13
L17 Transactional Memory (Part I)L16
L18 Transactional Memory (Part II) + AMAL17

The assignments changed too. The 2023 L1 slides say "Four programming assignments"; 2025 has five, and the new one is PA5. If a 2023 video mentions assignment details, check the 2025 repos instead.

Series arc

This series keeps the official lecture order, because assignments are tied to specific lectures and reordering would break their prerequisites. Each assignment post sits after the lectures it depends on, and written assignments are folded into the same post.

Part 1: Why parallelism, and what a processor looks like

Part 2: Parallelizing code and making it fast

Part 3: GPUs and data-parallel thinking

Part 4: AI systems

Part 5: Correctness in shared memory

Two reading paths

Full path: read 0 → 22 in order. Part 5 only depends on Parts 1–2, so if you want the shared-memory foundations first, jump to 19–22 after post 8, then come back for GPUs and AI.

AI systems only: 0 → 1–3 → 7 → 9–10 → 12–18. Posts 1–3 build intuition for SIMD, multithreading, and bandwidth limits; post 7 adds arithmetic intensity; then you go straight into GPUs and AI hardware. The cost is skipping the hands-on assignment posts, plus coherence and synchronization.

Things to do tonight

  1. Open the L1 slides, find the efficiency slide, and decide whether "2x speedup on 10 processors" is a good result.
  2. Go through the prerequisite checklist above. For any item you're unsure of, read the matching post in the CS107 or CS111 guide.
  3. Clone the PA1 repo, check how many cores your machine has and whether it supports AVX2, and install ISPC. PA1 is the assignment outsiders can reproduce most completely.

Further reading

These site series overlap with CS149. This series does not cut content because of them; the links are here for reference:

Series navigation: next, L1 Why parallelism, why efficiency

References