Skip to content

Reading NTHU Kao's NLP: Parameter-Efficient Fine-Tuning — Fine-Tuning Large Models Without an A100 Cluster

Sep 30, 20261 min
TL;DRHung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.

🌏 中文版

Version note: This post is based on W9_PEFT.pdf (61 pages) from the Fall 2025 (114-1) edition of Prof. Hung-Yu Kao's Natural Language Processing course at National Tsing Hua University. The W9 row of the 2025 schedule attaches both this deck and the GPT-2 / T5 TA-session deck, with recordings Week 9 Tue. and Week 9 Thu. (in Mandarin). I did not watch them to confirm which recording covers PEFT and which is the TA session, so skim both when you study. The W9 Topics column says "ELMo, BERT, GPT, and T5"; it's a syllabus template that doesn't match the attached slides. Facts checked on 2026-09-30. Access grade A3: slides and recordings are public.

Series: Previous GPT-3, InstructGPT, and RLHF | Next RAG (Part 1): Hallucination and Retrievers | Series overview

Every step of InstructGPT and Llama-2 in the last post touches every parameter in the model. A university lab, a small company, or a student in this course (the syllabus says "No GPU provided") who wants to adapt a 7B model to their own task hits GPU memory first.

This post answers one question: how do you fine-tune a large model without an A100 cluster?

Opening: what's left for NLP in the LLM era

The deck opens with Eduard Hovy (CMU) from his ROCLING 2024 talk. The first slide is self-mockery repeated three times: "Look what an LLM can do! Why can it do that? I have no idea / that's future work / I've never thought about it." The second gives three directions: make LLMs usable (NLP engineering), make them useful (NLP applications), and make them understandable (NLP research). The first item under the first direction is tuning LLMs to domains and building smaller, cheaper models.

Next is an October 2024 TechCrunch story in which OpenAI's CEO says a lack of compute is delaying the company's products. If OpenAI is short on compute, the case for PEFT makes itself.

First, the budget: how much memory full fine-tuning takes

The slides list PaLM 540B, MT-NLG 530B, and GPT-3 175B, then work through a real budget for Llama 2-7B (16-bit float, sequence length 4096, batch size 1):

ItemFormulaFull fine-tuningTraining 0.2M params
CUDAFixed overhead~1GB~1GB
Model weightssize(float) × N_parameter13.03GB13.03GB (same)
Gradientssize(float) × N_trainable13.03GB0.4MB
Hidden statesGrow with layers, sequence length, heads3.16GB3.16GB (same)
Optimizer states2 × size(float) × N_trainable26.06GB0.8MB
Total56.28GB17.19GB

The key point: gradients and optimizer states both scale with the number of trainable parameters, and optimizer states (counted as twice the trainable parameters on the slide) are the biggest chunk. Shrink the trainable parameters and both nearly disappear. Weights and hidden states stay, so you never get to zero.

The slides also give a formula for hidden states and cite the hardware table from LLaMA-Factory, noting that inference at batch 1 with a 4k context needs only 24GB.

Hidden-state estimate (from the slides)

During training, roughly: 3·h·seq·bs + 18·L·h·seq·bs + 3·L·heads·seq² + vocab·seq·bs

L is the number of layers (32 for Llama 2-7B), heads the number of attention heads (32), h the hidden size, bs the batch size. The seq² term comes from attention scores, probabilities, and dropout.

Not a new idea

The slides point out that computer vision has long updated only the last layer, that NLP experimented with static and non-static word embeddings, and that ELMo didn't fine-tune its contextualized embeddings at all. PEFT carries that old intuition over to LLMs.

Five benefits of PEFT

  1. Lower compute and storage costs.
  2. Portability: each task stores a small set of parameters, and the general pretrained parameters are shared.
  3. Less catastrophic forgetting: most parameters don't move, so language knowledge from pretraining is less likely to be overwritten.
  4. Less overfitting when data is scarce.
  5. Performance close to full fine-tuning. The slides' example: adding a small adapter lands within 1% of fully fine-tuned BERT on several NLU benchmarks.

There's also a comparison on RTE (DeBERTa-v3-base). Full fine-tuning scores 83.75% training 184M parameters. LoRA scores 86.60% training 0.8M (0.43%). AdaLoRA scores 88.09% with 1.27M (0.69%). In this example, the methods with fewer parameters score higher.

Why a small slice is enough: intrinsic dimensionality

The definition from Li et al. (ICLR 2018): in a D-dimensional parameter space, optimize only within a random d-dimensional subspace and ask how large d must be to reach a given fraction of the full result. d90 is the dimension needed for 90%.

The numbers on the slide are striking. A fully connected network on MNIST has 199,210 parameters, yet d90 is only 750 (0.38%). A ConvNet on Atari Pong has 1,005,974 parameters and d90 is 6,000 (0.60%). For many problems, the effective dimension is two to three orders of magnitude smaller than the parameter count.

Aghajanyan et al. (ACL 2021) carried the idea to language model fine-tuning. The slides draw three findings:

  • Many problems have small intrinsic dimensions.
  • For RoBERTa-base on six datasets (MRPC, QQP, Yelp Polarity, SST-2, MNLI, ANLI), the intrinsic dimension of fine-tuning drops the longer pretraining runs.
  • At a fixed number of pretraining updates, larger models need a lower intrinsic dimension to fine-tune on MRPC.

Put together: pretraining already moves the model to a spot where a small adjustment adapts it to a new task, and more so for larger models. That's the theory behind PEFT.

Method map: three families plus hybrids

The slides follow the taxonomy of the Lialin et al. 2023 survey:

TypeApproachExamples
AdditiveAdd new trainable parameters; freeze the original modelAdapters, Prompt Tuning, Prefix Tuning
SelectiveTrain only a chosen subset of the original parametersBitFit
ReparametrizationRepresent the weight update with low-rank matricesLoRA
HybridCombine the aboveMAM Adapters, S4

Adapters

Houlsby et al. 2019 insert a small bottleneck network after attention and after the FFN. It projects d-dimensional features down to a smaller m, applies a nonlinearity, and projects back to d. With m much smaller than d, each task adds few parameters. The slides also list later variants: Bottleneck Adapter (2019), Parallel Adapter (2020), and Compact Adapter (2021).

Prompt Tuning

Lester et al. 2021 prepend a sequence of trainable vectors (a soft prompt) to the input embeddings. The model stays frozen and only those vectors update. That makes multi-task serving easy: each task is a prompt, not a model, and one batch of inputs can carry prompts for different tasks.

# Pseudocode from the slides
soft_prompt = torch.nn.Parameter(torch.rand(num_tokens, embedding_dim))

def input_soft_prompt(x, soft_prompt):
    return concatenate([soft_prompt, x], dim=seq_len)

train(model(input_soft_prompt(x, soft_prompt)))  # only soft_prompt is updated

BitFit

Ben Zaken et al. 2022 fine-tune only the bias terms (in LayerNorm, FFN, and attention) and freeze everything else. In code, you hand the optimizer every parameter whose name contains "bias".

LoRA

In Hu et al. 2021, the original weight W (d×d) is frozen and a side path B·A is added: A compresses the input to r dimensions and B expands it back to d, with r much smaller than d. The slide shows the initialization. A is sampled from N(0, σ²) and B is set to 0, so the side path outputs zero at the start and the model behaves exactly like the original. α is a scaling factor.

In the comparison table, LoRA's edge is no inference overhead: after training you can add B·A back into W, and the architecture doesn't change.

Hybrids: MAM Adapters and S4

He et al. (ICLR 2022) put adapters, prefix tuning, and LoRA in one framework and assembled the MAM Adapter: a scaled parallel adapter on the FFN layer plus a soft prompt.

Chen et al. (ICLR 2023) treat PEFT design as a search problem, decided in four steps: how to group layers, how to allocate trainable parameters, which groups to tune, and which method each group uses (Adapter, Prefix, BitFit, LoRA). Under a 0.1% extra-parameter budget, the search found spindle grouping, uniform allocation, tuning every group, and then picked method combinations for the four groups in sequence.

Comparison table and selection criteria

The slides' comparison, following Lialin et al. (excerpt):

MethodTypeInference overheadTrainable parameters
AdaptersAExtra FFN0.1%–6%
Prompt TuningALonger input0.1%
BitFitS—0.5%
LoRARNone0.01%–0.5%
MAM AdaptersAExtra FFN and input0.5%
S4A+S+RExtra FFN and input0.5%

To choose, the slides ask four questions. How many parameters? How efficient is training (do you backpropagate through the original network, can you keep the GPU busy)? How efficient is inference (are parameters added, and at what cost)? And how accurate is the result?

Second half: prompt-based learning

The last section takes a different angle: instead of changing the model, change what the input looks like. There are hard prompts (discrete, real text) and soft prompts (continuous vectors).

Fine-tuning with task descriptions

The idea already appears in the GPT-2 paper: put a task description like "translated to English:" in the input and the model knows what to do.

Schick & Schütze (EACL 2021) apply it to MLM classification. To rate a Yelp review from 1 to 5 stars, rewrite the input as "review [SEP] In summary, the restaurant is [MASK]." and use a verbalizer to map labels to words: 1→terrible, 2→bad, 3→okay, 4→good, 5→great. Gao et al. (ACL 2021) add SST-2 (positive→great, negative→terrible) and MNLI (entailment→Yes, neutral→Maybe, contradiction→No).

Two benefits. Classification becomes generation, so you reuse the MLM's output layer instead of adding a classifier. And you need fewer training examples to match standard fine-tuning.

The slides note this helps GPT-3 too. Adding task descriptions without fine-tuning is in-context learning; fine-tuning with them is prompt-based fine-tuning.

Problems with discrete prompts

  • There are too many possible descriptions to find the best one.
  • Searching for the best description per task is expensive.
  • Discrete text can't be optimized directly with gradients during training.

Those three points lead to soft prompts.

Prefix Tuning and Soft Prompt Tuning

Prefix Tuning (Li & Liang, ACL 2021) was designed for generation. It prepends p virtual hidden states at every layer, mimicking virtual outputs of self-attention. The pretrained model (GPT-2 / BART) is frozen and only these states train. In practice they're first reparametrized through an MLP, which stabilizes training. Initialization matters: with little data, random initialization gives low, high-variance scores, and initializing with task words like "summarization" or "table-to-text" works better. With full data it makes no difference. P-Tuning v2 does the same thing, tested on NLU tasks.

Soft Prompt Tuning (Lester et al. 2021) prepends p trainable vectors only at the input layer, with a frozen T5. For initialization, words from the class labels work best. Small models show large gaps between initializations; at XXL size the gaps disappear.

The slides' comparison:

Prefix TuningSoft Prompt Tuning
Trainable parametersMore (prefix length × hidden size × layers)Fewer (prompt length × hidden size)
Inference speedSlowerFaster
PerformanceBetterWorse
Use caseBeats full fine-tuning in few-shot settings; no difference with abundant dataSame

How to self-study this lecture

  1. Redo the memory budget table for a 13B model. If you can, you understand why PEFT saves on gradients and optimizer states rather than weights.
  2. Use Hugging Face's PEFT library to run LoRA and prompt tuning on a small model. Print print_trainable_parameters() and compare with the parameter ratios in the slides' table.
  3. Read Figure 1 of the LoRA paper and check that you see why the A and B initialization makes training start from the original model.
  4. The series post on HW3 multi-output learning specifies bert-base-uncased. Once you finish it, try wrapping it in LoRA and compare memory and scores.

One thing to do tonight: open the training script you're using, compute the ratio of trainable to total parameters, and use the slide's formula to estimate how much memory the optimizer states take.

Further reading

References