Skip to content

NTHU NLP Guide 8: ELMo, BERT, T5, BART, GPT and the Three Roads to Pretraining

Sep 30, 20261 min
TL;DRA guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.

🌏 中文版

This guide is based on the public Fall 2025 (114-1) materials of Prof. Hung-Yu Kao's Natural Language Processing course at NTHU. It is part 8 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. The previous part is Sub-word Tokenization.

The official material for this lecture is W4_bert_and_its_family.pdf (58 slides), titled "ELMo, BERT, GPT, and T5 (BERT and its Family)", with recordings Week 6 Tue. and Week 6 Thu. (in Mandarin). The 2025 schedule places it in W6. The W9 row's Topics column does say "ELMo, BERT, GPT, and T5", but the file attached there is the PEFT deck. The Topics column is a syllabus template; this guide goes by the attached files.

The outline: recap of word embeddings and RNNs, then from word embeddings to pretrained language models (ELMo), then pretraining with Transformers (encoder BERT, encoder-decoder T5, decoder GPT), then GPT-3's in-context learning and large language models.

Starting point: one "record", two meanings

Slide 6 uses "I record the record": the first record is a verb, the second a noun. Word2Vec and GloVe give both the same vector, because static embeddings ignore context. That is the problem this lecture solves, and it picks up the contextualized embeddings that already appeared in the part 2 slides.

Slide 8 uses fill-in-the-blank examples to show how much next-token prediction can teach:

  • National Tsing Hua University is in ___ (Hsinchu): world knowledge
  • I put ___ bag on the table (a): grammar
  • The movie was ___ (bad): you must understand the complaint before it
  • 1, 1, 2, 3, 5, 8, 13, 21, ___ (34): reasoning

ELMo: read the whole sentence, then embed each word

ELMo (Embeddings from Language Models, Peters et al., 2018) does not give each word a fixed representation. It reads the entire sentence first, then produces a vector for each word in it.

Slides 9–11 draw the backbone: characters go through a CNN to get token representations, topped by a two-layer bidirectional language model (biLM), each layer with a left-to-right and a right-to-left RNN.

Slides 12–14 build the ELMo vector from the biLM:

  1. Concatenate: at each layer, join the forward and backward hidden states (the token layer is joined with itself)
  2. Weight: multiply each layer by a softmax-normalized weight stask
  3. Sum: add the weighted vectors and scale by a scalar γtask, letting the downstream model rescale the whole ELMo vector. The slides say this matters for optimization

The weights are learned with the downstream task, so the slides define ELMo as a task-specific combination of the biLM's intermediate layer representations.

Formula: ELMo layer weighting (slides 12–14)

$$\mathrm{ELMo}k^{task} = \gamma^{task} \sum{j=0}^{L} s_j^{task}, h_{k,j}^{LM},\qquad h_{k,j}^{LM} = \left[\overrightarrow{h}{k,j}^{LM};\ \overleftarrow{h}{k,j}^{LM}\right]$$

j = 0 is the token layer and j = 1…L the biLM layers; stask is softmax-normalized and γtask is a scalar.

Enter the Transformer: three ways to pretrain

Slide 16's timeline: Word2Vec in 2013, GloVe in 2014, ELMo, BERT and GPT in 2018, T5 in 2019. The Transformer's machine translation results led researchers to see it as a better alternative to RNNs.

Slide 17 is the backbone of the lecture, splitting Transformer pretraining into three roads:

ArchitectureHow the slides describe itExamples
EncoderBidirectional, can look into the future; suits downstream tasksBERT
Encoder-decoderWants the strengths of both; the slides leave a question: what does having both cost during pretraining?T5, BART
DecoderThe kind of LM seen so far; suits generation; not bidirectionalGPT

Road one: encoders (BERT)

Two pretraining tasks

Slides 19–21 introduce the two tasks of BERT (Devlin et al., 2018):

MLM (Masked Language Model): pick 15% of tokens for prediction, of which:

  • 80% become [MASK]
  • 10% become a random token
  • 10% stay unchanged

Why not mask them all? The slides explain that [MASK] never appears during fine-tuning. If only masked positions required careful representation, the model would get lazy about the unmasked words and fail to build robust representations.

NSP (Next Sentence Prediction): the input is "[CLS] sentence A [SEP] sentence B", and the [CLS] output classifies whether B follows A (IsNext / NotNext). The goal is to learn sentence relationships, and any monolingual corpus can generate the training data automatically.

Pretrain, then fine-tune

Slides 22–24 describe today's standard recipe: pretrain a general model of language understanding on a lot of unlabeled text, then fine-tune on a specific task. Fine-tuning adds one output layer on top of BERT, for example:

  • Sentence-pair classification: a classifier on the [CLS] output
  • Sequence tagging (e.g. NER): a classifier on each token's output, producing labels such as O and B-PER

Slide 25's details:

BERT-baseBERT-large
Parameters110M340M
Layers1224
Hidden size7681024
Attention heads1216

Training data was BooksCorpus (800M words) and English Wikipedia (2,500M words). Pretraining took 64 TPU chips for 4 days, which is impractical on a single GPU; fine-tuning on one GPU is common. That is what makes HW3, multi-output learning with bert-base-uncased, feasible.

Extensions of BERT

Slide 27's table:

ModelMLM changeNSP changeRelease
BERTStatic maskingSentence relationship2018/10
RoBERTaDynamic maskingNSP removed2019/07
SpanBERTSpan maskingNSP removed2019/07
ALBERTN-gram maskingSentence order2019/09
DistilBERTCompression by knowledge distillation2019/10
TinyBERTDistillation for NLU2019/09
DeBERTaDecoding-enhanced BERT with disentangled attention2020/06

Slide 28 adds SpanBERT's Span Boundary Objective: each word in a masked span is predicted by ordinary MLM and also from only the two boundary words plus its position. The example is "Generative AI is [MASK] [MASK] [MASK] at NTHU", where the model must recover "a NLP course".

Slides 29–30 list RoBERTa's four changes: train longer with bigger batches on more data (over 160GB); remove NSP; train on longer sequences; use dynamic masking that changes every iteration. The slides' conclusion: with the right training strategy, BERT's MLM objective is highly competitive.

Slide 31 lists domain-specific BERTs from around 2019: SciBERT, FinBERT, BioBERT, PubMedBERT, BlueBERT, ClinicalBERT and LegalBERT.

The limit of encoders

Slide 33: encoders do well across understanding tasks but poorly on generation. If your task produces output one word at a time, use a pretrained decoder instead.

Road two: encoder-decoders (T5, BART)

T5: every task is text in, text out

Slide 34 first lists design options for encoder-decoder pretraining: language modeling, BERT-style, deshuffling, and corruption strategies (replace with a mask, replace spans, drop tokens).

T5 (Text-to-Text Transfer Transformer, Raffel et al., Google, 2019) casts classification, similarity, sequence tagging and generation into one format. Slides 36–38 compare two objectives on one sentence:

Original: Thank you for inviting me to your party last week.

InputTarget
BERT-style maskingThank you <M> <M> me to your party apple weekThe full original sentence
Replace spansThank you <X> me to your party <Y> week<X> for inviting <Y> last <Z>

Replace spans turns each run of corrupted tokens into one unique placeholder, so both input and target get shorter. Slide 39's setting corrupts 15% of tokens with an average span length of 3, and notes that below a 50% corruption rate, results are not very sensitive to these two parameters.

Slide 37 has an easy-to-miss line: the BERT-style objective, originally designed for encoder-only models, was found in T5's comparison to outperform the alternatives.

Slide 40's model sizes:

VariantParametersLayersHidden sizeHeads
T5-small60M65128
T5-base220M1276812
T5-large770M24102416
T5-3B3B24102432
T5-11B11B241024128

Training data was C4 (the Colossal Clean Crawled Corpus, extracted from Common Crawl; the slide says 34B words). Slide 41 shows fine-tuning for closed-book QA: asked when Franklin D. Roosevelt was born, with no supporting text, T5 answers 1882 from what it learned in pretraining.

BART: break it, then fix it

Slides 42–44 cover BART (Lewis et al., 2019; the slide labels it Meta). It is a denoising autoencoder: corrupt text with an arbitrary noise function, then learn to reconstruct the original. Five noise types:

NoiseHowWhat the model learns
Token MaskingRandom tokens become [MASK]Predict masked tokens
Token DeletionRandom tokens are deletedPredict deleted tokens and their positions
Text InfillingLike SpanBERT, but a 0-length span inserts a [MASK]Predict how many tokens a span is missing
Sentence PermutationSplit on full stops and shuffle sentencesUnderstand how sentences relate
Document RotationPick a token at random and rotate the document to start thereFind the start of the document

For fine-tuning, classification feeds the same input to encoder and decoder and uses the final output representation. For generation tasks such as translation, a newly trained encoder is placed in front of BART and can use a different vocabulary from the original.

Road three: decoders (GPT)

Slide 46 gives GPT-1 (Radford et al., 2018): a 12-layer Transformer decoder with 117M parameters, 768-dimensional hidden states, 3072-dimensional feed-forward layers, and BPE with 40,000 merges, trained on BooksCorpus (over 7,000 books), whose long runs of contiguous text help learn long-distance dependencies. A bit of trivia from the slide: the acronym "GPT" never appears in the original paper.

Slide 48's pretraining task is next-token prediction, factoring a sentence's probability into a chain of conditional probabilities. Slide 49 shows fine-tuning: convert any structured input into a token sequence followed by a linear classifier. Entailment becomes "Start premise Delim hypothesis Extract"; multiple choice runs each answer through GPT separately and compares them.

Slide 50's GPT family table:

ModelSlide descriptionParametersDataRelease
GPT-1Transformer decoder followed by linear-softmax117MBooksCorpus 4.5 GB2018/06
GPT-2Same as GPT-1 with differently placed normalization; the text generation era starts1.5BWebText 40 GB2019/02
GPT-3A larger GPT-2175BCommonCrawl 45 TB2020/05
GPT-4Trained with RLHF> 1.5TUndisclosed2023/03

Treat the GPT-4 parameter count with caution: the GPT-4 Technical Report states that it withholds architecture details including model size, so the slide's "> 1.5T" is not a figure OpenAI published.

GPT-3, in-context learning and scaling laws

Slides 51–53 cite the GPT-3 paper (Brown et al., 2020). Before it, there were two ways to use a pretrained model: sample from the distribution it defines, or fine-tune it on task data. Large enough models show an emergent ability: they learn a task from a few examples in the context, without any gradient steps. This is in-context learning, and 175-billion-parameter GPT-3 is the example.

Slide 55 covers scaling laws (Kaplan et al., 2020): increase model size, data and training compute together and performance improves smoothly, following simple predictive rules.

Slides 56–57 list large models and community models: GPT-3 (175B), BLOOM (176B), Flan-PaLM (540B); Llama 2 (7B/13B/70B), Mistral (7B, 8×7B), Phi 2 (2.7B) and Gemma (2B/7B), the last four all decoder-only.

Slide 58's takeaways close the three roads in three lines: the BERT family pretrains with MLM and NSP and is not good at generation; T5 and BART boost MLM with span corruption and target a text-to-text format; the GPT family is pure language modeling and the most popular backbone today.

Going deeper

  • Load bert-base-uncased in a transformers fill-mask pipeline, feed it "I [MASK] NLP", and look at the top five candidates.
  • Run "I record the record" through BERT, take the last-layer vectors of both records, compute their cosine similarity, and compare with GloVe's single vector.
  • Write a span corruption function: given a sentence, a 15% rate and average length 3, output a T5-format input and target, and check it against slide 36.
  • Take any decoder-only model and run the same classification task with zero-shot and 3-shot prompts to see slide 52's in-context learning.

Further reading: CS224N guide: pretraining, CS224U contextual representations II: GPT, BERT, RoBERTa, ELECTRA, seq2seq and distillation, CS224U guide: in-context learning, CME295 guide: LLM training.

Gaps in the materials

  • Slides 53–54 (GPT-3 ICL, GPT-3 family) and 26 (why bidirectional) are paper figures; the extracted text has only their titles, so this guide does not restate numbers from them.
  • This lecture has no dedicated assignment. BERT hands-on work is in the next part, HF BERT tutorial and HW3; GPT-2/T5 hands-on work is in GPT-2/T5 Chinese summarization.
  • What came after GPT-3 (InstructGPT, RLHF) is in GPT-3, InstructGPT and RLHF.
  • Solutions, quizzes and class discussion live on NTU COOL and are not available to outside readers. The Fall 2026 version of this lecture is not yet public.

Series navigation: previous Sub-word Tokenization | next HF BERT Tutorial and HW3: Multi-output Learning | series overview

References