🌏 中文版
Version note: This post is based on W9_PEFT.pdf (61 pages) from the Fall 2025 (114-1) edition of Prof. Hung-Yu Kao's Natural Language Processing course at National Tsing Hua University. The W9 row of the 2025 schedule attaches both this deck and the GPT-2 / T5 TA-session deck, with recordings Week 9 Tue. and Week 9 Thu. (in Mandarin). I did not watch them to confirm which recording covers PEFT and which is the TA session, so skim both when you study. The W9 Topics column says "ELMo, BERT, GPT, and T5"; it's a syllabus template that doesn't match the attached slides. Facts checked on 2026-09-30. Access grade A3: slides and recordings are public.
Series: Previous GPT-3, InstructGPT, and RLHF | Next RAG (Part 1): Hallucination and Retrievers | Series overview
Every step of InstructGPT and Llama-2 in the last post touches every parameter in the model. A university lab, a small company, or a student in this course (the syllabus says "No GPU provided") who wants to adapt a 7B model to their own task hits GPU memory first.
This post answers one question: how do you fine-tune a large model without an A100 cluster?
Opening: what's left for NLP in the LLM era
The deck opens with Eduard Hovy (CMU) from his ROCLING 2024 talk. The first slide is self-mockery repeated three times: "Look what an LLM can do! Why can it do that? I have no idea / that's future work / I've never thought about it." The second gives three directions: make LLMs usable (NLP engineering), make them useful (NLP applications), and make them understandable (NLP research). The first item under the first direction is tuning LLMs to domains and building smaller, cheaper models.
Next is an October 2024 TechCrunch story in which OpenAI's CEO says a lack of compute is delaying the company's products. If OpenAI is short on compute, the case for PEFT makes itself.
First, the budget: how much memory full fine-tuning takes
The slides list PaLM 540B, MT-NLG 530B, and GPT-3 175B, then work through a real budget for Llama 2-7B (16-bit float, sequence length 4096, batch size 1):
| Item | Formula | Full fine-tuning | Training 0.2M params |
|---|---|---|---|
| CUDA | Fixed overhead | ~1GB | ~1GB |
| Model weights | size(float) × N_parameter | 13.03GB | 13.03GB (same) |
| Gradients | size(float) × N_trainable | 13.03GB | 0.4MB |
| Hidden states | Grow with layers, sequence length, heads | 3.16GB | 3.16GB (same) |
| Optimizer states | 2 × size(float) × N_trainable | 26.06GB | 0.8MB |
| Total | 56.28GB | 17.19GB |
The key point: gradients and optimizer states both scale with the number of trainable parameters, and optimizer states (counted as twice the trainable parameters on the slide) are the biggest chunk. Shrink the trainable parameters and both nearly disappear. Weights and hidden states stay, so you never get to zero.
The slides also give a formula for hidden states and cite the hardware table from LLaMA-Factory, noting that inference at batch 1 with a 4k context needs only 24GB.
Hidden-state estimate (from the slides)
During training, roughly: 3·h·seq·bs + 18·L·h·seq·bs + 3·L·heads·seq² + vocab·seq·bs
L is the number of layers (32 for Llama 2-7B), heads the number of attention heads (32), h the hidden size, bs the batch size. The seq² term comes from attention scores, probabilities, and dropout.
Not a new idea
The slides point out that computer vision has long updated only the last layer, that NLP experimented with static and non-static word embeddings, and that ELMo didn't fine-tune its contextualized embeddings at all. PEFT carries that old intuition over to LLMs.
Five benefits of PEFT
- Lower compute and storage costs.
- Portability: each task stores a small set of parameters, and the general pretrained parameters are shared.
- Less catastrophic forgetting: most parameters don't move, so language knowledge from pretraining is less likely to be overwritten.
- Less overfitting when data is scarce.
- Performance close to full fine-tuning. The slides' example: adding a small adapter lands within 1% of fully fine-tuned BERT on several NLU benchmarks.
There's also a comparison on RTE (DeBERTa-v3-base). Full fine-tuning scores 83.75% training 184M parameters. LoRA scores 86.60% training 0.8M (0.43%). AdaLoRA scores 88.09% with 1.27M (0.69%). In this example, the methods with fewer parameters score higher.
Why a small slice is enough: intrinsic dimensionality
The definition from Li et al. (ICLR 2018): in a D-dimensional parameter space, optimize only within a random d-dimensional subspace and ask how large d must be to reach a given fraction of the full result. d90 is the dimension needed for 90%.
The numbers on the slide are striking. A fully connected network on MNIST has 199,210 parameters, yet d90 is only 750 (0.38%). A ConvNet on Atari Pong has 1,005,974 parameters and d90 is 6,000 (0.60%). For many problems, the effective dimension is two to three orders of magnitude smaller than the parameter count.
Aghajanyan et al. (ACL 2021) carried the idea to language model fine-tuning. The slides draw three findings:
- Many problems have small intrinsic dimensions.
- For RoBERTa-base on six datasets (MRPC, QQP, Yelp Polarity, SST-2, MNLI, ANLI), the intrinsic dimension of fine-tuning drops the longer pretraining runs.
- At a fixed number of pretraining updates, larger models need a lower intrinsic dimension to fine-tune on MRPC.
Put together: pretraining already moves the model to a spot where a small adjustment adapts it to a new task, and more so for larger models. That's the theory behind PEFT.
Method map: three families plus hybrids
The slides follow the taxonomy of the Lialin et al. 2023 survey:
| Type | Approach | Examples |
|---|---|---|
| Additive | Add new trainable parameters; freeze the original model | Adapters, Prompt Tuning, Prefix Tuning |
| Selective | Train only a chosen subset of the original parameters | BitFit |
| Reparametrization | Represent the weight update with low-rank matrices | LoRA |
| Hybrid | Combine the above | MAM Adapters, S4 |
Adapters
Houlsby et al. 2019 insert a small bottleneck network after attention and after the FFN. It projects d-dimensional features down to a smaller m, applies a nonlinearity, and projects back to d. With m much smaller than d, each task adds few parameters. The slides also list later variants: Bottleneck Adapter (2019), Parallel Adapter (2020), and Compact Adapter (2021).
Prompt Tuning
Lester et al. 2021 prepend a sequence of trainable vectors (a soft prompt) to the input embeddings. The model stays frozen and only those vectors update. That makes multi-task serving easy: each task is a prompt, not a model, and one batch of inputs can carry prompts for different tasks.
# Pseudocode from the slides
soft_prompt = torch.nn.Parameter(torch.rand(num_tokens, embedding_dim))
def input_soft_prompt(x, soft_prompt):
return concatenate([soft_prompt, x], dim=seq_len)
train(model(input_soft_prompt(x, soft_prompt))) # only soft_prompt is updated
BitFit
Ben Zaken et al. 2022 fine-tune only the bias terms (in LayerNorm, FFN, and attention) and freeze everything else. In code, you hand the optimizer every parameter whose name contains "bias".
LoRA
In Hu et al. 2021, the original weight W (d×d) is frozen and a side path B·A is added: A compresses the input to r dimensions and B expands it back to d, with r much smaller than d. The slide shows the initialization. A is sampled from N(0, σ²) and B is set to 0, so the side path outputs zero at the start and the model behaves exactly like the original. α is a scaling factor.
In the comparison table, LoRA's edge is no inference overhead: after training you can add B·A back into W, and the architecture doesn't change.
Hybrids: MAM Adapters and S4
He et al. (ICLR 2022) put adapters, prefix tuning, and LoRA in one framework and assembled the MAM Adapter: a scaled parallel adapter on the FFN layer plus a soft prompt.
Chen et al. (ICLR 2023) treat PEFT design as a search problem, decided in four steps: how to group layers, how to allocate trainable parameters, which groups to tune, and which method each group uses (Adapter, Prefix, BitFit, LoRA). Under a 0.1% extra-parameter budget, the search found spindle grouping, uniform allocation, tuning every group, and then picked method combinations for the four groups in sequence.
Comparison table and selection criteria
The slides' comparison, following Lialin et al. (excerpt):
| Method | Type | Inference overhead | Trainable parameters |
|---|---|---|---|
| Adapters | A | Extra FFN | 0.1%–6% |
| Prompt Tuning | A | Longer input | 0.1% |
| BitFit | S | — | 0.5% |
| LoRA | R | None | 0.01%–0.5% |
| MAM Adapters | A | Extra FFN and input | 0.5% |
| S4 | A+S+R | Extra FFN and input | 0.5% |
To choose, the slides ask four questions. How many parameters? How efficient is training (do you backpropagate through the original network, can you keep the GPU busy)? How efficient is inference (are parameters added, and at what cost)? And how accurate is the result?
Second half: prompt-based learning
The last section takes a different angle: instead of changing the model, change what the input looks like. There are hard prompts (discrete, real text) and soft prompts (continuous vectors).
Fine-tuning with task descriptions
The idea already appears in the GPT-2 paper: put a task description like "translated to English:" in the input and the model knows what to do.
Schick & Schütze (EACL 2021) apply it to MLM classification. To rate a Yelp review from 1 to 5 stars, rewrite the input as "review [SEP] In summary, the restaurant is [MASK]." and use a verbalizer to map labels to words: 1→terrible, 2→bad, 3→okay, 4→good, 5→great. Gao et al. (ACL 2021) add SST-2 (positive→great, negative→terrible) and MNLI (entailment→Yes, neutral→Maybe, contradiction→No).
Two benefits. Classification becomes generation, so you reuse the MLM's output layer instead of adding a classifier. And you need fewer training examples to match standard fine-tuning.
The slides note this helps GPT-3 too. Adding task descriptions without fine-tuning is in-context learning; fine-tuning with them is prompt-based fine-tuning.
Problems with discrete prompts
- There are too many possible descriptions to find the best one.
- Searching for the best description per task is expensive.
- Discrete text can't be optimized directly with gradients during training.
Those three points lead to soft prompts.
Prefix Tuning and Soft Prompt Tuning
Prefix Tuning (Li & Liang, ACL 2021) was designed for generation. It prepends p virtual hidden states at every layer, mimicking virtual outputs of self-attention. The pretrained model (GPT-2 / BART) is frozen and only these states train. In practice they're first reparametrized through an MLP, which stabilizes training. Initialization matters: with little data, random initialization gives low, high-variance scores, and initializing with task words like "summarization" or "table-to-text" works better. With full data it makes no difference. P-Tuning v2 does the same thing, tested on NLU tasks.
Soft Prompt Tuning (Lester et al. 2021) prepends p trainable vectors only at the input layer, with a frozen T5. For initialization, words from the class labels work best. Small models show large gaps between initializations; at XXL size the gaps disappear.
The slides' comparison:
| Prefix Tuning | Soft Prompt Tuning | |
|---|---|---|
| Trainable parameters | More (prefix length × hidden size × layers) | Fewer (prompt length × hidden size) |
| Inference speed | Slower | Faster |
| Performance | Better | Worse |
| Use case | Beats full fine-tuning in few-shot settings; no difference with abundant data | Same |
How to self-study this lecture
- Redo the memory budget table for a 13B model. If you can, you understand why PEFT saves on gradients and optimizer states rather than weights.
- Use Hugging Face's PEFT library to run LoRA and prompt tuning on a small model. Print
print_trainable_parameters()and compare with the parameter ratios in the slides' table. - Read Figure 1 of the LoRA paper and check that you see why the A and B initialization makes training start from the original model.
- The series post on HW3 multi-output learning specifies bert-base-uncased. Once you finish it, try wrapping it in LoRA and compare memory and scores.
One thing to do tonight: open the training script you're using, compute the ratio of trainable to total parameters, and use the slide's formula to estimate how much memory the optimizer states take.
Further reading
- Efficient adaptation from prompting to LoRA: CS224N Lecture 8: Efficient Adaptation
- Memory and parameter budgets in training: Reading CME295: LLM Training
- Fine-tuning without losing old abilities: Reading NTU Hung-yi Lee ML 2026: HW5 Fine-tuning Without Forgetting
References
- IKMLab/NTHU_Natural_Language_Processing (GitHub) — course repo
- 2025 schedule README — slides and recordings attached to the W9 row
- W9_PEFT.pdf — source of every table, formula, and category in this post
- [Fall 2025] Week 9 Tue. recording (in Mandarin)
- [Fall 2025] Week 9 Thu. recording (in Mandarin)
- hiyouga/LLaMA-Factory hardware requirements
- Li et al., Measuring the Intrinsic Dimension of Objective Landscapes (ICLR 2018)
- Aghajanyan et al., Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning (ACL 2021)
- Lialin et al., Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning (2023)
- Houlsby et al., Parameter-Efficient Transfer Learning for NLP (ICML 2019)
- Lester et al., The Power of Scale for Parameter-Efficient Prompt Tuning (EMNLP 2021)
- Ben Zaken et al., BitFit (ACL 2022)
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (ICLR 2022)
- He et al., Towards a Unified View of Parameter-Efficient Transfer Learning (ICLR 2022)
- Chen et al., Parameter-Efficient Fine-Tuning Design Spaces (ICLR 2023)
- Schick & Schütze, Exploiting Cloze Questions for Few Shot Text Classification and NLI (EACL 2021)
- Gao et al., Making Pre-trained Language Models Better Few-shot Learners (ACL 2021)
- Li & Liang, Prefix-Tuning (ACL 2021)
- Liu et al., P-Tuning v2 (ACL 2022)
- huggingface/peft (GitHub)
Loading...