Skip to content

CMU 10-423 HW3: Fine-Tuning GPT-2 with LoRA — Written Questions, Files to Edit, and Compute

Sep 30, 20261 min
TL;DRHW3 in CMU 10-423 Spring 2026 is worth 66 points and was due 2026-03-12 (Slot A). The written part covers in-context learning (14 points), parameter-efficient fine-tuning (10), and the DPO derivation (15). The programming part (25) has you write LoRALinear from scratch, wire it into GPT-2's attention, and instruction-tune the model for sentiment classification on Rotten Tomatoes reviews. Every experiment uses gpt2-medium; the handout estimates 25–30 minutes per training run on a Colab T4, and you need a WandB account.

🌏 中文版

This post is based on the Spring 2026 offering of CMU 10-423/623/723 Generative AI. It's post 12 in the Reading CMU 10-423 series and closes the "adapting foundation models" unit. The previous two posts, L10–L11: PEFT and in-context learning and Instruction tuning, RLHF, and DPO, cover the ideas. This one looks at how the homework tests them.

Official materials used: hw3.zip from the Coursework page (the 23-page hw3.pdf, starter code, unit tests, and a LaTeX template), the read-only Overleaf template, the schedule, and the homework rules in the syllabus. I checked every fact against these materials on 2026-09-30. This post covers the structure and setup only. It contains no solutions.

At a glance

ItemDetails
NameHomework 3: Applying and Adapting LLMs
Lectures coveredL9–L12 (per the schedule)
Released2026-02-21
Due2026-03-12, 11:59 pm (Slot A)
SubmissionGradescope: one written PDF, plus 5 .py files
Total66 points

Point breakdown (hw3.pdf, page 1):

SectionPoints
LaTeX Template Alignment0
In-Context Learning14
Parameter Efficient Fine-Tuning10
Direct Preference Optimization15
Programming: LoRA for GPT-225
Code Upload0
Collaboration Questions2

The syllabus gives every homework two deadlines. Slot A takes human work only, with no AI allowed. After grading, you learn which questions you got wrong. Slot B is due three days after that feedback. AI help and full collaboration are allowed, but only the questions you missed in Slot A are regraded, and each question keeps the higher of the two scores. The schedule lists HW3 feedback on March 17 and Slot B on March 20, both marked tentative.

One detail in the zip echoes this rule. The handout folder ships a .cursorignore, an .aiderignore, and a .vscode/settings.json that turns off GitHub Copilot and inline suggestions.

What the written questions test

In-Context Learning (14 points)

All six parts revolve around one word problem about an electricity bill. The price per kilowatt-hour rises by 5 cents. A family used 150 kWh last month and 100 kWh this month, for $45 in total. What was the old rate? You write, in order:

  1. The difference between ICL and few-shot chain-of-thought prompting
  2. A prompt that would encourage in-context learning
  3. A one-shot CoT version
  4. A zero-shot CoT version
  5. One advantage and one disadvantage of zero-shot CoT compared with CoT
  6. One similarity and one difference between ICL and meta learning

These questions test whether you can tell prompt formats apart. You don't run a model.

Parameter Efficient Fine-Tuning (10 points)

Question 3.1 is about counting parameters. You get a 10-layer fully connected network with sigmoid activations, where layer l has width 2 to the power (L−l). First count the hidden units and total parameters. Then compute the fraction of parameters fine-tuned under two PEFT setups:

  • Bias-only tuning, in the style of BitFit (Ben-Zaken et al., 2021)
  • A rank-4 bottleneck adapter after every layer, following Houlsby et al. (2019), with GELU as the nonlinearity

Question 3.2 is about prefix tuning. The question writes out attention with K̃ = [P_K; K] and Ṽ = [P_V; V]. You express a single attention weight with exp, then prove that adding the prefix leaves the relative weight between any two content tokens unchanged.

Direct Preference Optimization (15 points)

This section walks you through the derivation in the DPO paper. The question suggests reading the paper all the way to the end:

  • 4.1.a: Show that the optimum of the KL-constrained objective has the form π_ref · exp(r/β) / Z(x), and explain why a closed-form solution means you don't need a policy-gradient loop
  • 4.1.b: Rewrite the reward as a policy ratio plus a partition term, and explain why Z(x) doesn't depend on y
  • 4.1.c: Plug into the Bradley–Terry model to get the DPO loss, and explain how it turns PPO's RL objective into supervised learning
  • 4.1.d: If the closed form exists, why doesn't PPO-based RLHF just use it? Give two reasons
  • 4.1.e: Derive the gradient of the DPO loss
  • 4.1.f: Explain which direction each term of the gradient pushes the parameters, and when the weighting term gets large

These questions line up with the DPO slides from the second half of L11 and the first half of L12.

Programming: LoRA from scratch, attached to GPT-2

Data and task

The data is the Rotten Tomatoes movie-review sentiment dataset on Hugging Face, with balanced positive and negative labels. It downloads automatically when you run train.py. The code uses the old short name rotten_tomatoes, which now redirects to cornell-movie-review-data/rotten_tomatoes.

  • train.py shuffles the 8,530 training examples and keeps the first 5,000. Validation uses the full validation split.
  • generate.py computes accuracy on the 1,066-example test split.

GPT-2 only knows how to continue text, so the assignment uses instruction fine-tuning to make it emit labels. get_sentiment_prompt() in dataloader.py wraps each review as Instruction: …\nText: …\nLabel: …, with labels converted to positive/negative. During collation, labels on the instruction part are set to −100, so the loss ignores the instruction.

What you write for LoRA

hw3.pdf gives the forward pass as h = W₀x + (α/r)·BAx. W₀ stays frozen; only the low-rank matrices A and B train. B starts at zero, so ΔW = BA is also zero when training begins. The original paper adds LoRA only to the query and value matrices. This assignment asks for query, key, and value.

FileWhat you doUpload
lora.pyFinish LoRALinear: __init__, reset_parameters (A with kaiming_uniform_, B with zeros), forward (apply dropout to the input when weights aren't merged), train (un-merge the LoRA weights), eval (merge BA into the main weight), and mark_only_lora_as_trainableYes
model.pyThe plain Transformer from HW1 (no GQA, no RoPE). Replace attention's c_attn and c_proj with LoRALinearYes
train.pyOne TODO: after loading the pretrained model, make only the LoRA parameters trainableYes
generate.pyFinish predict_labels: truncate the generated text and check for positive/negativeYes
dataloader.pyAlready complete. You come back to edit the prompt for question 5.11Yes

generate.py has one rule that's easy to miss. If the model's output isn't one of the two labels, record it as −1 and keep it in the denominator. Don't drop it. The handout also notes that small models rarely emit EOS, so checking the first few characters of the generation is enough.

Defaults and flags

Defaults in train.py (used whenever an experiment question doesn't say otherwise):

ParameterDefault
LoRA rank r128
LoRA alpha512
LoRA dropout0.05
dropout0.0
learning rate2.5e-4
max_iters80
batch size × gradient accumulation4 × 32

generate.py defaults to max_new_tokens=5, temperature=0.6, and top_k=1. Every experiment question uses --init_from="gpt2-medium". The hints say the smallest GPT-2 won't show much benefit from LoRA fine-tuning.

The 12 experiment questions

QContentPoints
5.1Does LoRA add inference latency?2
5.2What fraction of parameters is fine-tuned at r=128, α=512?2
5.3Accuracy of gpt2-medium with no fine-tuning1
5.4Accuracy for r ∈ {16, 128, 196} (α=4r) and for full fine-tuning4
5.5WandB validation-loss curves for those four runs4
5.6LoRA vs. full fine-tuning, and how r affects performance3
5.7Is anything unexpected about the validation loss at r=196?2
5.8(16, 64) vs. (16, 256) vs. full fine-tuning2
5.9Effect of raising α with r fixed1
5.10Why tuning both learning rate and α may be redundant1
5.11Rewrite the prompt and explain your motivation1
5.12Accuracy with the old and new prompt at default settings2

Always report accuracy from the "Best Val Checkpoint", not the last-iteration checkpoint. generate.py evaluates both and labels each in its output.

Environment and compute: can you still run it?

hw3.pdf recommends a free Colab T4. When the quota runs out, you can wait, switch Google accounts, buy credits, or move to Kaggle, GCP, or AWS. The handout gives three time estimates:

  • 5.3, the zero-shot evaluation: about 2 minutes on a T4
  • 5.4, 5.8, and 5.12: about 25–30 minutes per training run
  • The four runs in 5.4, the new run in 5.8, and the new prompt in 5.12 add up to 6 training runs, roughly an afternoon of T4 time

run_in_colab.ipynb already has these commands. One thing to watch: the notebook's full fine-tuning cell passes --lora_rank=0 --dropout=.05, while hw3.pdf's default dropout is 0.0. Follow what the question asks.

WandB is a hard requirement. train.py reads a key from wandb_api.json and calls wandb.login right at the start, and 5.5 and 5.8 ask for WandB curves, so sign up first. requirements.txt lists torch, transformers, datasets, tiktoken, wandb, and tqdm, none of them pinned. transformers downloads the GPT-2 weights from Hugging Face; no extra access request is needed.

I downloaded hw3.zip on 2026-09-30 and tested it (macOS, PyTorch 2.13, CPU). Running python test_lora.py on the unmodified starter code, all 7 tests error out. That's expected: placeholder lines in lora.py such as self.lora_A = # TODO are syntax errors on their own, so the module won't import until you fill them all in.

The zip also has two things hw3.pdf doesn't spell out:

  • hw3.pdf's file list leaves out test_lora.py, data.pt, and chargpt.py, but all three are in the zip. test_lora.py has 7 tests, each weighted 1, covering reset_parameters, __init__, mark_only_lora_as_trainable, forward, train, and eval. The test data lives in data.pt.
  • train.py does more than train. When training finishes, it calls ModelSampler from generate.py to compute test-split accuracy and logs it to WandB.

What you can't get from outside CMU

  • HW3 recitation slides: The February 20 row on the schedule links to Google Slides. From outside CMU on 2026-09-30, it returned 401, and so did the PDF export. The HW1 and HW2 recitation slides are public; this one isn't.
  • Gradescope grading: The autograder, manual grading rubric, and Slot A feedback are all closed.
  • Recordings: Lecture videos are on Panopto and need a CMU login.
  • Official solutions: Not public.

The schedule lists two more HW3-related assessments: Quiz 3 on February 25 (L9–L12) and the HW3/HW4 programming test on March 27. Neither is available.

How to self-study this assignment

  1. Do written questions 3.1 and 4.1 first. 3.1 gives you an order-of-magnitude feel for how many parameters LoRA saves. After 4.1, DPO is more than a one-line loss.
  2. While writing lora.py, run test_lora.py locally on CPU until all 7 tests pass, then move to Colab. Don't spend T4 quota on debugging.
  3. Run the 5.3 zero-shot baseline first. Seeing how untuned gpt2-medium does with this prompt gives every later number a reference point.

One thing to do tonight: download hw3.zip, open lora.py, and draw how forward, train, and eval relate. When is BA merged into W₀? When is it split back out? In which state does dropout apply? Once that diagram is clear, the code is about half done.

Further reading

Series navigation: previous L11–L12: Instruction tuning, RLHF, and DPO | next L12–L13: Text-to-image, latent diffusion, and vision-language models | Series overview

References