🌏 中文版
This post is based on the Spring 2023 offering of CS224U. It is part 4 of the Stanford CS224U guide series. The Transformer block and positional encoding are covered in the previous post; this one starts directly with the model families.
The last seven sections of the same contextual representations slide deck (GPT, BERT, RoBERTa, ELECTRA, seq2seq, Distillation, Wrap-up) match videos 07 through 13 in the XCS224U playlist. Each runs about 6 to 14 minutes.
The best way to read these seven sections is to keep asking one question: which part of pretraining does this family change?
The big picture
| Family | What it changes | Evidence the course cites |
|---|---|---|
| GPT | Objective: autoregressive, left context only | Scaling table from GPT to GPT-3 |
| BERT | Objective: bidirectional MLM plus NSP | The two downsides the original paper names itself |
| RoBERTa | Data size, batch size, masking, dropping NSP | Ablations on static vs. dynamic masking, batch size, and data size |
| ELECTRA | Learning signal: from "guess the masked word" to "was each word replaced?" | Ablations on generator size, compute efficiency, and prediction coverage |
| T5, BART | Architecture: encoder–decoder; BART also changes how inputs are corrupted | T5's three-architecture figure; BART's corruption combinations |
| Distillation | Model size: a large teacher trains a small student | GLUE results for DistilBERT and others |
I built this table from the videos; it isn't a slide from the deck.
GPT: looking left only
Objective. Video 07 starts with the autoregressive loss. At position t, take the embedding of the token to predict and dot it with the hidden representation the model has built up through t−1. The rest is softmax normalization over the whole vocabulary; you take the log and look for parameters that maximize it.
Formula: autoregressive loss (slide 29)
$$\max_{\theta} \sum_{t=1}^{T} \log \frac{\exp\left(e(x_t)^\top h_\theta(x_{1:t-1})\right)}{\sum_{x' \in V} \exp\left(e(x')^\top h_\theta(x_{1:t-1})\right)}$$
Masking. Inside a Transformer, attention has to hide the future: position a sees only itself, b can see a, and c can see a and b.
Teacher forcing. During training, the correct token goes in at the next step no matter what the model predicted. Potts stresses a point that's easy to miss: the model never predicts tokens. It predicts scores over the whole vocabulary. Choosing a token is a separate decision rule. Taking the highest score is one rule, beam search is another, and neither is part of the model itself.
Fine-tuning. The standard approach puts task parameters on the final output state. Potts believes the first GPT paper's fine-tuning was based entirely on that state. You can also mean- or max-pool over all the output states.
Scale. The OpenAI models on the slide:
| Model | Layers | d_k | Parameters |
|---|---|---|---|
| GPT (2018) | 12 | 768 | 117M |
| GPT-2 (2019) | 48 | 1,600 | ~1.5B |
| GPT-3 (2020) | 96 | 12,288 | 175B |
Another slide lists open models: GPT-Neo, GPT-J, GPT-NeoX, OPT-66B, and BLOOM (176B parameters). The slide itself warns that "this table will be out of date by the time anyone reads it."
BERT: both directions, but learning from 15%
Input. Every sequence starts with [CLS] and uses [SEP] as a separator. Besides word and position embeddings, there's a segment embedding (SentA, SentB). This is the hierarchical position from the previous post, such as premise and hypothesis in NLI.
MLM. Mask some tokens and have the model recover them from bidirectional context. The slides show three treatments: no masking, replacing with [MASK], and replacing with a random word ("rules" becomes "every"). Only a small share is masked so the other positions supply enough context. The loss includes an indicator m_t that is 1 for masked positions and 0 otherwise.
Formula: MLM loss (slide 41)
$$\max_{\theta} \sum_{t=1}^{T} m_t \log \frac{\exp\left(e(x_t)^\top h_\theta(\hat{x})t\right)}{\sum{x' \in V} \exp\left(e(x')^\top h_\theta(\hat{x})_t\right)}$$
$\hat{x}$ is the masked sequence, and $m_t = 1$ means position $t$ was masked. The difference from GPT is that $h_\theta$ can use the whole sequence (minus position $t$), not just what comes before $t$.
NSP. The second objective is binary next sentence prediction: two sentences that really follow each other are labeled IsNext, and random pairs are labeled NotNext. Potts says the motivation was to help the model learn some discourse-level information.
Fine-tuning. The lightest approach adds a few dense layers on top of the [CLS] output. Because [CLS] is always in the first position, it becomes a constant element that carries a lot of information about the sequence. Pooling over all outputs is the alternative.
Releases. The original release had only base and large, each cased and uncased (Potts recommends always using cased). Google and others later released smaller ones. BERT-tiny has 2 layers and 4M parameters; large has 24 layers and 340M. All of them use absolute positional encoding, so the maximum length is 512 tokens.
Four known limitations. This slide is the hinge of the whole unit, because the next three families respond to it:
- The original paper's ablation and optimization studies are "admirably detailed but still partial"
- "We are creating a mismatch between pre-training and fine-tuning, since the [MASK] token is never seen during fine-tuning" (the original paper's own words)
- "Only 15% of tokens are predicted in each batch" (the original paper's own words)
- "BERT assumes the predicted tokens are independent of each other given the unmasked tokens" (from the XLNet paper). Potts's example: mask both "New" and "York," and the model guesses each one independently
RoBERTa: answering limitation 1
RoBERTa stands for Robustly Optimized BERT Approach. The slide compares them item by item:
| BERT | RoBERTa |
|---|---|
| Static masking | Dynamic masking |
| Input: two concatenated document segments | Input: sentence sequences that may cross document boundaries |
| NSP | No NSP |
| Batches of 256 | Batches of 2,000 |
| WordPiece | Character-level byte-pair encoding |
| BooksCorpus + English Wikipedia | Plus CC-News, OpenWebText, Stories |
| 1M steps | Up to 500K steps (with much bigger batches, so more examples overall) |
| Short sequences first | Full-length sequences throughout |
The most interesting evidence is the trade-off on input format. Using sentences from a single document (DOC-SENTENCES) scored slightly better on their benchmarks, but the team chose FULL-SENTENCES, which can cross documents, because it makes efficient batching easier. Potts likes the decision: in this era, accuracy isn't the only thing that matters; resources do too.
The data table is just as direct. Going from 16GB to 160GB and from 100K to 500K steps, every step helps.
Potts also points out a change in method. RoBERTa is far more thorough than BERT, but nowhere near the exhaustive hyperparameter searches of the pre-deep-learning era. The reason is simple: it's too expensive, so even RoBERTa is a heuristic, partial exploration. For more on how to set up BERT-style models, he recommends A Primer in BERTology.
ELECTRA: answering limitations 2 and 3
ELECTRA (the slides cite Clark et al.'s ICLR version) is shown with "the chef cooked the meal":
- Mask about 15% of tokens, as in BERT: "the chef [MASK] the meal"
- A small, BERT-like generator fills in the masked positions by sampling from its own distribution. Sometimes it restores the original word; sometimes it picks another, such as "ate" for "cooked"
- The discriminator, which is ELECTRA itself, decides for every token whether it's the original or a replacement
The two are trained jointly. Afterward the generator is dropped and the discriminator is kept. [MASK] only appears in the generator's input, so the discriminator never sees it, which removes limitation 2. The discriminator makes a decision at every position, which removes limitation 3.
Keep the generator small. When the generator and discriminator are the same size, they can share parameters, and more sharing helps. But the best results come from a generator much smaller than the discriminator: with a 768-dimensional discriminator, a 256-dimensional generator is best, and the curve is an inverted U. Potts's intuition is that a weaker generator leaves the discriminator more interesting work to do.
Ablation on prediction coverage (GLUE scores):
| Variant | GLUE |
|---|---|
| ELECTRA | 85.0 |
| All-tokens MLM | 84.3 |
| Replace MLM | 82.4 |
| ELECTRA 15% (the discriminator judges only replaced positions) | 82.4 |
| BERT | 82.2 |
The lesson: more predictions are better. Even All-tokens MLM, which stays within the BERT architecture and simply predicts every token, clearly beats the original BERT.
There were three releases: Small, Base, and Large. Small was designed to be "quickly trained on a single GPU," which Potts reads as another sign of the growing focus on efficiency.
seq2seq: T5 and BART
Tasks. The slides first list tasks with natural seq2seq structure: machine translation, summarization, free-form question answering, dialogue, semantic parsing, and code generation. The broader class is encoder–decoder, which doesn't have to involve sequences.
From RNNs to the Transformer. Potts adds some history. RNN seq2seq models first gained lots of attention mechanisms so the decoder could look back at the encoder (the slides cite Luong et al. 2015). The Transformer then embraced attention fully and dropped recurrence.
Three architectures. The slides use Figure 4 from the T5 paper: encoder–decoder; a standard language model (causal mask throughout); and a prefix LM (full attention over the input, causal over the output). Potts notes that the last two became more common as GPT models got bigger.
T5. An encoder–decoder with extensive multi-task supervised and unsupervised training. Its most forward-looking idea is the task prefix: putting a natural-language instruction like "translate English to German:" before the input. Potts says this gave an early glimpse of in-context learning. Releases range from 60M to 11B parameters; FLAN-T5 is the instruction-tuned version.
BART. BART has a BERT-like encoder and a GPT-like decoder. The interesting part is pretraining: corrupt the input, then learn to restore it. Corruptions include text infilling, sentence shuffling, token masking, token deletion, and document rotation. The video says the best combination was text infilling plus sentence shuffling. Fine-tuning uses no corruption. For classification, uncorrupted input goes to both the encoder and decoder, and the final decoder state is used for the prediction. For seq2seq tasks, you feed the input and output as usual.
Distillation: a large teacher trains a small student
As models keep growing, distillation is one way to shrink them: train a student whose input-output behavior matches the teacher's but is more efficient to run.
Levels of objectives. The slides list them from lightest to heaviest; in practice people often use weighted combinations:
- Gold labels for the task, if available
- The teacher's output labels. The lightest option: you don't even need access to the teacher during distillation, just one earlier pass
- The teacher's output score vectors (Hinton et al. 2015)
- The teacher's final output states, pulled together with a cosine loss (DistilBERT)
- Other hidden states and embeddings
- Training the student to mimic the teacher's counterfactual behavior under internal interventions (Wu et al. 2022, work Potts was involved in)
Training modes. The standard setup freezes the teacher and updates only the student. There's also multi-teacher distillation, co-distillation (training both at once, also called online distillation), and self-distillation (aligning some parts of a model with other parts of the same model).
Results. Using GLUE as the yardstick, the slides list three consistent results. DistilBERT distilled 12-layer BERT-base into 6 layers while keeping 97% of GLUE performance. Sun et al. 2019 distilled to 3 and 6 layers. Jiao et al. 2020 distilled to 4 layers.
Wrap-up: what got left out, and the 2023 outlook
Making amends. Video 13 covers three architectures there wasn't time for:
- Transformer-XL: caches states from earlier parts of a long sequence and links them into the current computation with recurrent connections
- XLNet: uses an autoregressive loss but samples many permutations of the input order, so it still gets bidirectional context. This is the answer to BERT's limitation 4
- DeBERTa: separates word and position representations, each with its own attention
Pretraining data. Potts says he feels "guilty" that the series didn't cover pretraining data, so he lists OpenBookCorpus, The Pile, BigScience data, Wikipedia processing tools, and Pushshift Reddit. The point isn't to train your own large model. It's to audit these datasets and understand where the models you have are likely to succeed and where they may be seriously problematic.
Four trends as of 2023 (as stated on the slide; this post doesn't update them):
- Autoregressive architectures seem to have taken over, possibly just because the field is focused on generation
- Bidirectional models may still have the edge for representation
- seq2seq is still a dominant choice for tasks with that structure
- People are still obsessed with scaling up, but there's a counter movement toward "smaller" models (still around 10B parameters)
Hands-on: static word vectors from BERT
The notebook for this unit is vsm_03_contextualreps.ipynb. Two caveats first: its version string says "Spring 2022," and the README marks all the vsm_* material as background.
It asks an interesting question: can a model that only supplies contextual representations give us good static word vectors? It builds on Bommasani et al. 2020.
What the notebook does:
- Loads
bert-base-casedand shows how the tokenizer splits "Bert knows Snuffleupagus" - Uses
output_hidden_states=Trueto get 13 sets of hidden states: the embedding layer plus 12 layers - Flags an easy mistake: don't use
pooler_outputfor [CLS], becausetransformersadds randomly initialized parameters on top of it for fine-tuning. Uselast_hidden_state[:, 0]instead - The decontextualized approach: feed a single word through the model and pool its pieces if it's split. Bommasani et al. found mean pooling best overall
- The aggregated approach: find every occurrence of the word in a corpus and average its representations
Running the whole notebook requires the course's data.tgz (still downloadable when checked on 2026-09-29). The first half, on tokenizing and hidden states, needs no data.
What you can do tonight: run up to the len(reps.hidden_states) cell and confirm you get 13. Then change bert_weights_name to roberta-base and compare how the same sentence gets tokenized.
Further reading: CS224N guide: pretraining covers encoder, decoder, and encoder–decoder pretraining from another course.
Gaps in the materials
- Slides 50–51 paste table screenshots from the RoBERTa paper, and the extracted PDF text has misaligned columns. This post only uses numbers the video confirms and that can be read from the slide tables.
- ELECTRA's efficiency and generator-size curves are figures from the paper; this post only relays how the video describes them.
- This unit has no homework of its own. The first homework is where you actually fine-tune these models.
Series navigation: Previous: Contextual representations I: guiding ideas, the Transformer, and positional encoding | Next: HW1: multi-domain sentiment analysis and the bake-off
References
- CS224U course website (Spring 2023)
- Contextual word representations slides (handout PDF)
- XCS224U Spring 2023 YouTube playlist (videos 07–13)
- vsm_03_contextualreps.ipynb (cgpotts/cs224u)
- Radford et al. (2018). Improving Language Understanding by Generative Pre-Training
- Devlin et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Liu et al. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Clark et al. (2020). ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
- Raffel et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)
- Lewis et al. (2020). BART
- Sanh et al. (2019). DistilBERT
- Hinton, Vinyals & Dean (2015). Distilling the Knowledge in a Neural Network
- Wu et al. (2022). Causal Distillation for Language Models
- Sun et al. (2019). Patient Knowledge Distillation for BERT Model Compression
- Jiao et al. (2020). TinyBERT
- Rogers, Kovaleva & Rumshisky (2020). A Primer in BERTology
- Luong, Pham & Manning (2015). Effective Approaches to Attention-based Neural Machine Translation
- Chung et al. (2022). Scaling Instruction-Finetuned Language Models (FLAN-T5)
- Dai et al. (2019). Transformer-XL
- He et al. (2021). DeBERTa
- Gao et al. (2020). The Pile
- Bommasani, Davis & Cardie (2020). Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings
Loading...