🌏 中文版
Version note: Dates follow the Spring 2026 offering of CMU 11-868 LLM Systems. Assignment content follows the Assignment 3 page and the llmsys_hw3 repo as seen on 2026-09-30. The assignment site is shared across semesters, and llmsys_hw3 already contains Fall 2026 changes (see "Version note" at the end). Access level A3: the problems, starter code, and local tests are public; the private tests, Canvas submission, and recordings are not. This post contains no solutions or reference code.
Series: previous L08–L09: Tokenization, decoding, and speculative decoding | next L10: Accelerating Transformers on GPU (LightSeq) | Series overview
The first two assignments built the foundation: HW1 wrote CUDA kernels for map, zip, reduce, and matmul, and HW2 wrote autodiff and a sentiment classifier. HW3 is the first time they come together as a real language model.
The assignment page opens with two sentences. The first: implement a decoder-only GPT-2 architecture in MiniTorch, train it on IWSLT14 German-English translation, and benchmark it. The second is a warning: training for Problem 4 takes at least 10 hours.
Timeline and dependencies
From the Spring 2026 Syllabus:
| Date | Event |
|---|---|
| Feb 2 | L06 Transformer |
| Feb 4 | L07 Pre-trained LLMs; HW2 due, HW3 released |
| Feb 6 | Recitation 3: The Annotated Transformer |
| Feb 9, Feb 11 | L08 Tokenization, L09 Decoding |
| Feb 16 | L10 Accelerating Transformer on GPU Part 1 |
| Feb 18 | L10 Part 2; HW3 due |
Conceptually you need L06–L07; Problem 4's generate relates to greedy decoding from L09.
In code, HW3 consumes your previous two assignments directly. The setup section asks you to:
- Copy
llmsys_hw2/minitorch/autodiff.pyover, and renamerun_sentiment.pytoproject/run_sentiment_linear.py - From
llmsys_hw1/src/combine.cu, extract only the implementations ofMatrixMultiplyKernel,mapKernel,zipKernel, andreduceKernelinto the newsrc/combine.cu - Run
bash compile_cuda.shto compile the kernels
The page explains why you copy only the functions: HW3's combine.cu and cuda_kernel_ops.py have changed. GPU memory allocation, deallocation, and host/device copies moved into combine.cu, and the tensor storage type changed from numpy.float64 to numpy.float32.
In other words, a bug in HW1 or HW2 will surface in HW3 in strange ways. FAQ Q3 on the assignment page is one example: forward tests pass but gradient assertions fail, and the official pointer is the backpropagate function in autodiff.py.
The four problems
| Problem | Content | File | Points |
|---|---|---|---|
| 1 | Tensor functions: logsumexp, softmax_loss | minitorch/nn.py | 20 |
| 2 | Basic modules: Linear, Dropout, LayerNorm1d, Embedding | minitorch/modules_basic.py | 20 |
| 3 | Decoder-only Transformer LM: MultiHeadAttention, TransformerLayer, DecoderLM | minitorch/transformer.py | 40 |
| 4 | Machine translation pipeline: generate | project/run_machine_translation.py | 20 |
Each problem's code region is marked with BEGIN ASSIGN3_x / END ASSIGN3_x, and comes with pytest commands (for example python -m pytest -l -v -k "test_softmax_loss_student").
Problem 1: softmax loss
The page gives the formula ℓ(z, y) = log Σ exp(z_i) − z_y and asks you to build it from logsumexp, one_hot, and other existing ops. Inputs are (minibatch, C) logits and (minibatch,) labels; the output has shape (minibatch,), with no reduction.
Problem 2: four basic modules
Linear: reuse HW2, adapted for the newbackendargumentDropout: whenself.trainingis false, leave the input untouched; to match the autograder's random seed, the page requiresnp.random.binomialfor the maskLayerNorm1d: layer normalization over a 2D tensorEmbedding: maps one-hot word vectors to embeddings
Problem 3: assembling GPT-2
This problem carries the most points and gets the most detail. The architecture follows the GPT-2 paper; of the four modules, FeedForward is already written for you.
The key to MultiHeadAttention is shapes. The page spells it out: the input X is B×S×D (batch, sequence length, hidden dimension); project to Q, K, and V, split into h heads, permute to B×h×S×D_h, and transpose K's last two dimensions. After computing softmax(QKᵀ/√D_h + M)V (M is the causal mask), permute and reshape back to B×S×D and apply the output projection. Batched matrix multiplication is provided.
This is the full version of the two shapes on L06 page 12 (len × dim and len × len), with batch and head dimensions added.
TransformerLayer must use pre-LN. The page shows post-LN and pre-LN side by side and cites On Layer Normalization in the Transformer Architecture.
DecoderLM runs: look up token and positional embeddings and add them, apply dropout, pass through every Transformer layer, apply a final LayerNorm, and project to vocabulary size with a linear layer.
Problem 4: the translation pipeline
You implement generate: for each source sentence, produce the target with argmax decoding, one example at a time, without batching. The page suggests reading collate_batch and loss_fn in the same file first to understand data processing and loss computation.
After python project/run_machine_translation.py, outputs and BLEU scores land in ./workdir_vocab10000_lr0.02_embd256. Reference numbers from the page:
- BLEU around 7 after the first epoch, around 20 after 10 epochs
- About one hour per epoch on a PSC V100
- The default hyperparameters are not guaranteed to be stable (you may see nan); you can tune learning rate, vocabulary size, embedding dimension, number of layers, number of heads, and dropout
Grading and submission
Submit the whole llmsys_hw3 as a zip on Canvas, containing the full codebase, one workdir with your best result, and a screenshot of training progress (or the slurm log if you used sbatch).
Grading has two parts: private MiniTorch test cases and the IWSLT evaluation results. According to the page, full marks require passing all tests and a BLEU of about 20 ± 2.
Where self-learners get stuck
- You need an NVIDIA GPU. The whole assignment sits on CUDA kernels you compile yourself. The page's instructions assume PSC (
module load cuda/12.4.0, Python 3.12+, auvvirtual environment); outside CMU you need your own CUDA machine or a cloud GPU - Training time. Going by the V100 reference, 10 epochs is roughly ten hours, longer on a slower card. Confirm the loss drops with small hyperparameters before starting a long run
- CUDA and driver versions. FAQ Q1 covers nvcc being newer than the CUDA version the driver supports, which makes PyTorch abort with
Aborted (core dumped). The official fix is to align CUDA and PyTorch to the same version - Get the earlier assignments right first. Without HW1 and HW2, HW3 is missing its kernels and autodiff and cannot be started on its own
- No private tests. The local pytest suite is only the public half. A BLEU of 20 ± 2 is something you can measure yourself, and it is the most reliable self-check available outside CMU
Version note
The assignment site and repos are shared across semesters. Checking the llmsys_hw3 commit history on 2026-09-30:
- Several fixes landed during Spring 2026 (Feb 4–6: Embedding docstring,
datasetsversion, Adam second-moment coefficient, transformer module filename, view backward) - A commit on 2026-08-28, "add unit test for machine translation; cleanup code comment inconsistency", was merged into main on Sep 9 as part of Fall 2026
To match the spring version, the last commit before the Feb 18 deadline is 376bf97 (2026-02-06). Problem statements and points reflect what the page showed on the check date; Fall 2026 is in progress and may change them again.
Further reading
- The Annotated Transformer: the Recitation 3 material, useful for comparing shapes and masking
- CS336 series overview: CS336's first assignment also has you write a Transformer LM from scratch, but in PyTorch; in 11-868 the framework and kernels underneath are yours too
- CMU 11-785 Lecture 18: attention and Transformers
References
- 11-868 Assignment 3: Transformer Architecture
- llmsys_hw3 starter code repo
- 11-868 assignment site overview
- 11-868 Spring 2026 Syllabus
- L06 Transformer slides (PDF)
- Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2, 2019)
- Xiong et al., On Layer Normalization in the Transformer Architecture (2020)
- The Annotated Transformer (Harvard NLP)
Loading...