Skip to content

NTU ADL 2025 Lecture 7.5: PEFT — Adapter, LoRA, Prompt Tuning, and HW2

Sep 30, 20261 min
TL;DRWhen an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.

🌏 中文版

This is post 10 of Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall. The course is ADL Fall 2025 (NTU semester 114-1, 2025/09/01–12/15). The course page puts LLM Adaptation on the same day as Post-Training (9/22). That week's TA recitation is LLM LoRA Training, and the homework column shows HW 2.

Sources: the LLM Adaptation slides (14 pages), video 7.5 Parameter-Efficient Fine-Tuning (Adaptor, LoRA) (19:21), and the HW2 video ADL 2025 Fall Homework 2 (13:05, uploaded 2025-10-06). All checked on 2026-09-30. The slides run only 14 pages and are mostly diagrams, so this post is shorter than the previous one.

Series position: previous Post-Training: Instruction Tuning, RLHF, and InstructGPT | next RAG and HW3 | Series overview

The problem: full fine-tuning is too expensive

The slides open with the same map as the previous lecture (pp.2–4). To do well on known tasks, you can do prompt tuning/engineering or tune the LM itself. Page 4 adds a note next to the second option: fine-tuning LLMs may be expensive and impractical. Page 5 therefore names the topic Parameter-Efficient LM Tuning, a more practical way to adapt LLMs.

Page 6 gives three reasons efficient adaptation matters (the slide credits Benji Xie and Regina Wang):

  1. The current AI paradigm emphasizes accuracy over efficiency.
  2. Training and fine-tuning LLMs carries hidden environmental costs.
  3. As training costs rise, AI development concentrates in well-funded organizations, especially in industry.

Page 7 states the shared idea in one line: slightly modify the hidden representations instead of the whole model.

Three approaches

Adapter

Page 8 (the slide cites He et al., 2022) inserts small trainable submodules into each Transformer block, after multi-head attention and after the feed-forward layer. The key line is at the bottom: all tasks share the same original pre-trained model, and the adapters are task-specific modules. The result is better robustness and less storage.

LoRA

Pages 9–11 cover LoRA (Hu et al., 2021):

  • The idea is low-rank adaptation. The diagram places LoRA modules beside attention and feed-forward, added in parallel to the original weights.
  • The justification is on p.10: weight updates for downstream fine-tuning have a low intrinsic rank.
  • Page 11 closes with results on GPT-3 175B and concludes that LoRA shows better scalability and task performance.
The low-rank update as a formula (from the LoRA paper; the slides show it as a diagram)

The original weight W₀ is a d×k matrix. LoRA freezes W₀ and learns two small matrices:

W = W₀ + ΔW = W₀ + B·A
B ∈ R^(d×r),  A ∈ R^(r×k),  r ≪ min(d, k)

Trainable parameters drop from d·k to r·(d+k). Before inference you can merge B·A back into W₀, so there is no extra latency. These details come from the LoRA paper; the slides state only the core observation about low intrinsic rank.

Prompt tuning

Page 12 has one line of text: prefix tuning and soft prompt tuning are also parameter-efficient adaptation. Both already appeared in post 8 on pre-training and prompt learning. Here they are reclassified under PEFT: leave the model weights alone and learn only a sequence of continuous vectors placed before the input.

Which one is best?

Page 13 cites a comparison by Mao et al. 2022. The conclusion is one line: no one can fit all tasks.

ApproachWhat it changesThe slides' reason
AdapterInserts new modules into Transformer layersShared base model, per-task adapters: robust and storage-efficient
LoRAAdds a low-rank update beside existing weightsWeight updates are low-rank anyway; better scalability and performance on GPT-3 175B
Prompt tuningLearns only a prefix or soft prompt before the inputAlso parameter-efficient

Hands-on work is in the TA recitation

That week's recitation is LLM LoRA Training. The course page's slide link, f114-adl/doc/w5-LoRA.pdf, returned 404 on 2026-09-30. A file with the same name opens under the Fall 2024 path. The linked video is ADL TA Recitation: LLM LoRA Training (18:15, in Mandarin), uploaded 2023-11-16, so it is a recording reused from an earlier year. Implementation details are left for post 18 of this series, on the TA recitations.

HW2: only the title is public

The HW 2 button in the course page's 9/22 row links straight to ADL 2025 Fall Homework 2 on YouTube. What can be confirmed:

  • Title: the video description reads "LLM Tuning and Prompt Tuning for Classical Chinese Translation", that is, translating Classical Chinese using LLM tuning and prompt tuning.
  • Length 13:05, uploaded 2025-10-06.
  • Page 11 of Course Logistics lists the second assignment's topic as LLM Tuning.

What is missing: the video has no subtitles, and there are no public spec slides or written instructions. This post cannot confirm the dataset, base model, baseline, metric, submission format, or deadline, so it states none of them. Submission goes through NTU COOL, which needs an NTU account.

How to use it from outside NTU: take HW2 as a direction. Build a small parallel set of Classical and modern Chinese yourself and compare three approaches: prompting alone, prompt tuning, and LoRA. You will have to define your own metric, and you cannot compare results with the official grading.

After this lecture you should be able to

  • Explain in one sentence why PEFT helps: only a few parameters change, so the base model can be shared.
  • Say where Adapter, LoRA, and prompt tuning each change the model.
  • Explain why LoRA works: weight updates during fine-tuning have a low intrinsic rank.

One thing to try tonight: open the config of any Transformer model you have, take the d×k of one attention projection matrix, and compute the r·(d+k) parameters LoRA would train at r = 8. The ratio is what the three reasons on p.6 look like in practice.

Further reading

Next: RAG and HW3

References