Skip to content

Reading NTU ADL 2025 Fall: Three Pre-training Families and Prompt Learning — From BERT, GPT, and T5 to Prompts Only Machines Understand

Sep 30, 20261 min
TL;DRLecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.

🌏 中文版

This guide covers the 9/15 week of NTU Yun-Nung Chen's Applied Deep Learning (ADL), Fall 2025 (114-1, 2025/09/01–12/15). It is part 8 of the Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall series. Part 6 covered BERT and its family, and the previous post on HW1 applied BERT to Chinese extractive QA. This one zooms out: how do encoder-only, decoder-only, and encoder-decoder models differ, and why can a large enough model do a task from a prompt alone?

Official materials used:

Access level follows the series rating of A2. The slides and videos are public, and there is no new assignment this week (HW2's topic is in part 10, with only an explainer video).

What pre-training is

Page 2's analogy: learn general knowledge from textbooks before being tested on a specific subject. Pre-training trains a model on a large, diverse dataset before fine-tuning it for a task. The three key steps are large-scale diverse data, self-supervised learning, and general representations; the payoff is scalability, generalizability, and transferability.

Where does the data come from? Page 3 uses BookCorpus (free books from smashwords.com). Page 4 raises the "fair use vs. copyright" question for web-scale data: regulations are still evolving and differ across countries.

Three pre-training architectures

The table on page 5 is the lecture's map, and it keeps coming back:

TypePropertyExamples
EncoderBidirectional contextBERT and its variants
DecoderLanguage modeling, better for generationGPT, GPT-2, GPT-3
Encoder-DecoderSequence-to-sequenceTransformer, BART, T5

Encoders: the BERT family

Pages 7–8 recap quickly. BERT's masked LM hides 15% of tokens; RoBERTa mainly trains BERT on more data for longer; SpanBERT masks contiguous spans, which makes a harder and more useful pre-training task. Details are in part 6.

Why decoders are still needed

Page 9 spells out the encoder's limit: BERT and other pre-trained encoders don't naturally generate one word at a time. With "Vivian goes to [MASK] tasty tea," an encoder can only fill the blank, while a decoder can continue from "Vivian goes to" and write "make tasty tea."

Decoders: the GPT series

  • GPT (Radford et al. 2018): a Transformer decoder pre-trained on BooksCorpus (about 7,000 books, 5GB), with 12 layers, 768-dim hidden states, 3072-dim feed-forward layers, and BPE with 40,000 merges. Downstream it uses supervised fine-tuning and keeps next-word prediction during fine-tuning.
  • GPT-2 (Radford et al. 2019): more data (WebText from Reddit, 40GB), good for natural language generation.
  • GPT-3 (Brown et al. 2020): more data again, from Common Crawl, WebText2, Books1 and Books2, and English Wikipedia.

Page 15 lines up the generations: GPT at 117M parameters, GPT-2 at 1.5B, GPT-3 at 175B. The GPT-4 (2023) and GPT-5 (2025) rows show "?" for both parameters and data, an honest note that those figures are not public.

Encoder-decoders: BART and T5

Page 17 explains the split: the encoder gets bidirectional context, and the decoder trains the whole model through language modeling. The pre-training objective is span corruption (denoising), done during preprocessing.

Page 18 shows the difference on one sentence (Thank you for inviting me to your party last week):

  • BART: takes the sentence with gaps and outputs the whole original sentence.
  • T5: replaces the gaps with markers like <X> and <Y> and outputs only the missing parts: <X> for inviting <Y> last <Z>.

Fine-tuning for classification also differs (page 19): BART repeats the input in the decoder and predicts the label from its output, while T5 treats classification as seq2seq and generates the label text. Page 22 notes that T5's multi-task pre-training "learns multiple tasks via seq2seq," and the slide adds that this resembles today's post-training, which sets up the next post on post-training.

Page 23 compares the two: BART uses roughly twice T5's training data; BART uses learnable absolute positions and T5 relative positions; and in the slide's understanding and summarization tables, BART scores higher on most columns.

Two obstacles of the pre-trained-model era

Video 6.3 is subtitled "the obstacles of the PLM era." Page 24 defines the standard recipe first: initialize from a pre-trained LM, then tune its parameters for the downstream task. Two problems follow.

Obstacle 1: scarce labeled data

Page 25 lists GLUE dataset sizes, from 391K for MNLI down to 2.5K for RTE, two orders of magnitude apart. The slide's conclusion: more practical cases are few-shot, one-shot, or even zero-shot.

This is where GPT-3's in-context learning comes in (pages 26–28). Compared with fine-tuning:

  • Fine-tuning: after pre-training, update parameters on labeled task data.
  • In-context learning: after pre-training, no more learning; the input carries instructions and a few examples.

The slide uses a local example: a Taiwanese English-proficiency test's instructions for a vocabulary section plus one worked question ("the answer is D") is exactly a few-shot prompt. Zero-, one-, and few-shot differ only in how many examples you give.

Obstacle 2: models are huge

Page 33 tabulates model sizes, from ELMo at 93M, BERT-Base at 110M, and BERT-Large at 340M, through eight GPT-3 sizes up to 175B with 96 layers. Larger models do better (pages 34–35), but they cost:

  • Training compute: pages 37–38 cite Sevilla et al. 2022 on compute trends.
  • Scaling laws: page 39's Kaplan et al. 2020 describes how model quality changes with model size N, data D, and compute C, which helps predict performance and allocate resources. Page 40's Hoffmann et al. 2022 says model size and training tokens should scale equally. Chinchilla is smaller than other large models yet better, which the slide notes is good for fine-tuning and inference.
  • Storage: page 41, each task needs its own copy. With an 11B-parameter model, three tasks mean three 11B copies.

Page 42 folds both obstacles into one line: the solution is prompt learning.

Hard prompts: steering with natural language

Pages 45–46 contrast two setups on NLI. The usual one puts a classifier after [CLS] premise [SEP] hypothesis [SEP] to predict neutral, contradiction, or entailment. The prompt version rewrites the input as "Vivian likes dancing. Is it true that Vivian loves singing?" and lets the model answer maybe, no, or yes.

Pages 47–50 break prompt-tuning into three parts:

  1. Prompt template: a manually designed natural-language input format, e.g. Premise? [MASK], Hypothesis.
  2. PLM: performs language modeling (masked or autoregressive).
  3. Verbalizer: maps vocabulary back to labels, e.g. yes → entailment, maybe → neutral, no → contradiction.

Pages 51–52: a few labeled examples are enough to fine-tune, and in zero-shot settings no parameters change at all. The slides cite Le Scao and Rush 2021: under data scarcity, prompt-tuning works better because it uses and preserves pre-trained knowledge.

LM-BFF (Gao et al. 2021, pages 53–54) adds demonstrations to the prompt and generates templates automatically; the slides show results with RoBERTa-Large.

Soft prompts: only the machine needs to understand

Page 55 names the problems with hard prompts:

  • Prompts that look reasonable to humans are not necessarily effective for LMs (the slide cites Liu et al. 2021).
  • Pre-trained LMs are sensitive to the choice of prompt (Zhao et al. 2021).

So skip the words and optimize vectors:

MethodWhat it doesSlide takeaway
P-Tuning (Liu et al. 2021)Optimizes prompt embeddings directly instead of prompt tokensExample: prompt search for "The capital of Britain is [MASK]"
Prefix-Tuning (Li and Liang 2021)Optimizes only the prefix embeddings, at every layerBetter training time and space efficiency
Prompt Tuning (Lester et al. 2021)Stores only a small task-specific prompt (one layer) per task and allows mixed-task inference on the same frozen PLMCompetitive performance, better space efficiency

Page 59 compares fine-tuning, prefix-tuning, and hard and soft prompt-tuning on performance and space. This thread leads directly to part 10 on PEFT: adapters and LoRA answer the same question of tuning a small part instead of the whole model.

The prompting paradigm: a map of the field

Pages 60–66 follow Liu et al.'s 2021 survey and sort prompting research along five dimensions. Video 6.6 calls it a grab bag of prompt-based research:

  • Pre-trained models: left-to-right LMs (the GPT series), masked LMs (BERT, RoBERTa), prefix LMs (UniLM), encoder-decoders (T5, BART).
  • Prompt engineering: shape (cloze vs. prefix), and human-designed vs. automated (discrete like LM-BFF, continuous like Prefix-Tuning).
  • Answer engineering: answer shape (token, span, sentence) and how it is designed.
  • Multi-prompt learning: ensembling, augmentation, composition, decomposition, sharing.
  • Training strategies: which parameters get tuned. Promptless fine-tuning (BERT), tuning-free prompting (GPT-3), fixed-LM prompt tuning (Prefix-Tuning), fixed-prompt LM tuning (T5), and prompt plus LM tuning (P-Tuning).

The last dimension is the most useful in practice. When you meet a new method, ask first: does it tune the model's parameters, the prompt's parameters, or neither?

What you can do after reading

  • Given a model name, say whether it is an encoder, decoder, or encoder-decoder, and what its pre-training objective is.
  • Explain how fine-tuning differs from in-context learning, and what resource question scaling laws answer.
  • Tell hard prompts (template plus verbalizer, LM-BFF) from soft prompts (P-Tuning, Prefix-Tuning, Prompt Tuning) by what gets tuned and what gets stored.

Try this: take paragraph selection from HW1 and rewrite it as a prompt. Design a template (e.g. "Question: … Paragraph: … Can this paragraph answer the question? [MASK]") and a verbalizer (yes/no), then use pages 61–66 to place your approach in each of the five dimensions.

What this guide can and cannot confirm

Confirmed: every slide title, bullet, and table text in the 67-page deck; video titles and lengths (checked against the playlist); every arXiv paper title above (checked on arXiv).

Not confirmed: the videos were not transcribed, so the instructor's spoken examples and comments are not included. Figures in the slides (GPT-3 task results, training cost, scaling-law plots, LM-BFF and prompt-tuning results) are described by title and conclusion only; take numbers from the original papers. The slides give GPT-3's data as 45TB; this guide did not check how the original paper defines that figure before and after filtering.

Further reading on this site: CS224N Lecture 7: Pretraining covers the same three architectures and in-context learning; CS336 scaling-law foundations goes deeper on Kaplan and Chinchilla; and CS224U on in-context learning extends the story through prompt design and DSPy.

Series navigation: Series overview | Previous: HW1 Chinese Extractive QA | Next: Post-training: Instruction Tuning, RLHF, and InstructGPT

References