Skip to content

CMU 11-868 HW5: Writing Data Parallelism and Pipeline Parallelism Yourself on Two GPUs

Sep 30, 20261 min
TL;DRCMU 11-868's fifth assignment switches to PyTorch and Hugging Face GPT-2. Using only torch.distributed and torch.multiprocessing, you write data parallelism (partition the data, set up a process group, average gradients; 50 points), then a GPipe-style pipeline (split the model, generate a clock schedule, run micro-batches on worker threads; 50 points). Both parts need benchmarks and plots on at least two GPUs: data parallelism must reach at least 1.5x speedup on 2 GPUs, and the pipeline must beat plain model parallelism. The Spring 2026 deadline was 3/25.

🌏 中文版

This guide follows the Spring 2026 edition of CMU 11-868 LLM Systems. It is part 14 of Reading CMU 11-868 LLM Systems. The handout and repo are described as seen on 2026-09-30. The homework site is shared across semesters and may be changed for Fall 2026.

This assignment puts parts 11 through 13 into practice: data parallelism (L14–L15) and pipeline parallelism (first half of L16). ZeRO from the previous post isn't part of it; you run ZeRO through DeepSpeed in HW6.

At a glance

ItemDetails
HandoutAssignment 5: Distributed Training and Parallelism
Starter codeThe page links llmsys_f25_hw5, which 301-redirects to llmsystem/llmsys_hw5
Due3/25 in Spring 2026 (Week 11 on the Syllabus); the Syllabus doesn't list a release date
PointsProblem 1 Data Parallel 50, Problem 2 Pipeline Parallel 50
HardwareAt least 2 GPUs; the handout strongly recommends PSC
EnvironmentA conda env with Python 3.9; requirements.txt pins torch 2.2.0, transformers 4.37.2, datasets 3.6.0
LecturesL14–L15 (data parallelism), L16 (pipelines); Recitation 6 on 3/20 covers Distributed Training

The biggest change from the first four assignments: this one doesn't use MiniTorch. The starter project/run_data_parallel.py loads Hugging Face's pretrained GPT2LMHeadModel on bbaaaa/iwslt14-de-en-preprocess, takes only the first 5,000 training examples, and runs 10 epochs by default. The focus shifts from building a framework to building the distributed layer yourself.

Rules: two packages for communication

Problem 1 allows only torch.distributed and torch.multiprocessing.Process for GPU communication, and forbids adding new imports. Problem 2 allows only packages the starter code already imports. Every region you implement is marked with BEGIN_HW5_* and END_HW5_* comments, and all edits must stay inside them. The handout says staff will inspect the code by hand to check the package restrictions.

Problem 1: data parallelism (50 points)

1.1 Partition the data. Implement three things in data_parallel/dataset.py:

  • Partition: a dataset class that returns items by a list of indices
  • DataPartitioner: shuffles the indices and splits them by fractions (for four GPUs, [0.25, 0.25, 0.25, 0.25]); use(rank) returns one partition
  • partition_dataset: computes the per-GPU batch size (a total batch of 128 on four GPUs gives 32 each), takes this rank's partition, and wraps it in a DataLoader

Test a5_1_1 only checks that partitions don't overlap.

1.2 Set up the process group and average gradients. In project/run_data_parallel.py, implement setup: set MASTER_ADDR to localhost and MASTER_PORT to 11868, then call init_process_group. The main block must create, start, and cleanly shut down world_size processes. Then implement average_gradients, which walks the model's parameters and aggregates gradients with torch.distributed, and call it after backward in train in project/utils.py.

To test, run one batch with world_size 2 so each rank saves its gradients as model{rank}_gradients.pth, then run a5_1_2 to check that both GPUs hold the same gradients.

1.3 Benchmark. Compare training time and tokens per second on one GPU (batch 64) and two GPUs (total batch 128). The handout specifies the math: average training time across GPUs, and sum tokens per second across GPUs for throughput. It suggests dropping the first epoch or at least doing one warmup run. Plot each metric separately and save the figures in submit_figures. Full marks require at least 1.5x speedup on 2 GPUs in both training time and throughput.

Problem 2: pipeline parallelism (50 points)

This part implements the GPipe-style schedule from L16: split the model by layer across GPUs, split the input into micro-batches, and let stages work on different micro-batches at the same time.

2.1 Split the model and build the schedule. In pipeline/partition.py, implement _split_module, which cuts an nn.Sequential into layer-wise partitions, one per GPU. In pipeline/pipe.py, implement _clock_cycles(num_batches, num_partitions), which produces, for each time step, which stage processes which micro-batch. Test a5_2_1.

2.2 The Pipe module. Read worker.py first, then implement Pipe.forward and Pipe.compute. Pipe is a generic wrapper that turns any nn.Sequential into a pipelined module. create_workers starts one worker thread per GPU; each takes tasks from in_queue and puts results in out_queue. Your job is to wrap computations as Task objects, put them on the right device's queue, and collect the results. The handout stresses that forward must put its output on the last device; leaving it on the device of the input x breaks pipeline training. Test a5_2_2.

2.3 Wire it into GPT-2. In pipeline/model_parallel.py, implement _prepare_pipeline_parallel. GPT2ModelCustom.parallelize has already placed GPT-2's blocks on different GPUs; you extract the transformer blocks in self.h and package them as an nn.Sequential for Pipe. The handout warns that GPT2Block returns a tuple, not a tensor, and you only need the hidden states. If you add a helper module with no parameters, wrap it in WithDevice, or it will be treated as living on the CPU.

Finally, train on two GPUs with --model_parallel_mode='model_parallel' and 'pipeline_parallel' and plot the comparison. Full marks require faster training time and higher throughput with the pipeline than with plain model parallelism.

Submission and grading

The handout provides scripts/create_submission_zip.sh for packaging and scripts/run_benchmarks.py for performance logs, submit_figures/performance_summary.json, and plots. Staff compile and run the code, check the figures, and manually check the package restrictions.

Where the handout and starter code disagree

Comparing the handout with the llmsys_hw5 main branch on 2026-09-30 turned up three mismatches:

  1. partition_dataset signature: the handout shows partition_dataset(dataset, batch_size=128, collate_fn=None); the starter code has partition_dataset(rank, world_size, dataset, batch_size=128, collate_fn=None). Go with the starter code.
  2. Submission name: the Submission section says to "submit the whole llmsys_s25_hw5 as a zip on canvas," still using the Spring 2025 repo name.
  3. Speedup threshold: Problem 1.3 asks for 1.5x on both training time and throughput, but the grader-mode example uses --dp-time-threshold 1.2 --dp-throughput-threshold 1.5 (and 1.0 for both pipeline metrics). The official pages don't say which values grading actually uses.

Also, the last commit on main is dated 2026-04-30, so it still reflects spring. The repo also has a merged-hw5-hw6 branch. If Fall 2026 changes the assignment, the changes may land there or on main. To match the spring version, checking out a commit from before 2026-03-25 is the safest bet.

Where outside readers get stuck

  • Two GPUs: this is the first assignment in the series that strictly needs more than one GPU. The Logistics page describes PSC as a resource provided to enrolled students, so outside readers need to rent a two-GPU cloud machine. Logistics also warns that PSC uses job-based scheduling with no guarantee of when a job will start
  • 1.5x doesn't come for free: syncing gradients across two GPUs has a cost (the subject of L14–L15). How much speedup you get from batch 64 on one GPU versus 128 on two depends on your interconnect and your implementation
  • No Canvas, no manual grading: the tests only check correctness (no overlapping data, matching gradients, a correct schedule). You have to check performance and the package rules yourself

Try it yourself: three exercises without enrolling

  1. Get data parallelism working on CPU first. With the gloo backend and two processes, init_process_group runs on a laptop, so you can verify DataPartitioner and average_gradients before measuring speedup on GPUs.
  2. Write a clock schedule by hand. On paper, draw the GPipe forward schedule for 3 stages and 4 micro-batches, listing the (micro-batch, stage) pairs at each time step, then compare with the _clock_cycles tests.
  3. Measure the bubble. Once the pipeline works, vary the split with run_pipeline.py's --n_chunk (the number of micro-batches, default 4), record how throughput changes, and compare with L16's bubble formula O((K−1)/(M+K−1)).

Further reading

References