Skip to content

CS224N Lecture 7: Pretraining, Subwords, and In-Context Learning

Aug 22, 2026 1 min
TL;DR Lecture 7 decomposes pretraining into scalable data, subword tokenization, three model objectives, and in-context learning. A general self-supervised objective yields reusable representations; downstream signals specify their use.
Table of Contents
  1. Why pretraining scales
  2. Subwords repair the fixed-vocabulary break
  3. Three pretraining forms
  4. Knowledge, capability, or pattern continuation?
  5. From static embeddings to contextual models
  6. BPE walkthrough and vocabulary trade-offs
  7. Decoder pretraining
  8. Encoder pretraining
  9. Encoder–decoder pretraining
  10. Scale data, model, and compute together
  11. What pretraining teaches and what probes show
  12. Design an in-context learning experiment
  13. Material gap and numbering note
  14. References

🌏 中文版

The official CS224N Winter 2026 schedule places lecture 7 on January 27, but does not name a lecturer; this article therefore attributes it only to the course staff. The official deck, Pretraining (Scaling, Systems, Data), has six agenda parts: motivation, subwords, the move from word embeddings to model pretraining, three architectures, what pretraining teaches, and large models with in-context learning.

Why pretraining scales

Supervised tasks depend on human labels, limiting both volume and task coverage. Pretraining creates prediction targets from text itself, allowing a model to use large, diverse, unlabelled corpora. Smaller labelled sets, instructions, or prompts can then specify downstream use.

This is not free knowledge. What the model learns depends on the corpus, tokenizer, objective, parameter count, and compute budget. More data means more training signal; it does not guarantee balanced sources, correct content, or reliable downstream behavior.

Subwords repair the fixed-vocabulary break

A whole-word vocabulary maps unseen words to UNK. The byte-pair encoding subword method begins with characters and repeatedly merges the most frequent adjacent units until reaching a target vocabulary size. Common words may remain whole while rare and novel words decompose into known subwords.

The method trades vocabulary size against sequence length. A small vocabulary produces longer sequences; a large one leaves rare tokens undertrained. The segmentation is not linguistic analysis, and words can split into unintuitive pieces.

Three pretraining forms

A decoder uses a left-to-right language-model objective and naturally supports generation. An encoder, as exemplified by BERT, recovers masked tokens from bidirectional context and is suited to representation learning. An encoder–decoder conditionally generates a target from a source, using span corruption or text-to-text formulations.

These forms differ not merely in size but in visible information and training interface. Choose according to whether the downstream task needs open generation, bidirectional representation, or an explicit input-to-output transformation.

Knowledge, capability, or pattern continuation?

The deck gives “what is pretraining teaching?” its own interlude. Behavioral evidence may show factual recall, syntax, or new-task performance, but one output cannot identify an internal representation. In-context learning adds another layer: parameters remain fixed while instructions or examples in the prompt change behavior for the current context.

Evaluation should separate what pretraining already supplied, what fine-tuning added, and what a prompt temporarily elicited. One benchmark score cannot locate the source of capability.

From static embeddings to contextual models

Pretraining transfers an entire composition function rather than one word table. Token states vary by context; layer, pooling, subword alignment, and fine-tuning must be specified when comparing representations.

BPE walkthrough and vocabulary trade-offs

BPE repeatedly merges the most frequent adjacent symbols and freezes the learned merge order. Large vocabularies shorten sequences but undertrain rare rows; small vocabularies lengthen attention. Tokenizer and checkpoint must stay paired.

Decoder pretraining

Causal next-token prediction gives dense loss and a generation-native interface. It sees only left context, and likelihood alone does not supply factuality or instruction following.

Encoder pretraining

Masked language modeling restores tokens from bidirectional context. Masking policy affects learning and creates pretrain–finetune mismatch. Encoders suit representation tasks but are not naturally autoregressive generators.

Encoder–decoder pretraining

Corrupted sources are encoded bidirectionally and reconstructed by a decoder. Span corruption and text-to-text formulations unify conditional tasks while retaining source/target roles.

Scale data, model, and compute together

The Llama 3 technical report provides a concrete large-scale pretraining case. Compute-aware scaling balances parameters and tokens. Preserve source mixture, deduplication, and filters; total token count alone cannot describe data.

What pretraining teaches and what probes show

A probe shows information is decodable, not necessarily used. Interventions provide stronger evidence but can have broad side effects. Memorization and generalization require deduplication, temporal splits, and near-neighbor checks.

Design an in-context learning experiment

Hold the checkpoint fixed and vary instructions, demonstrations, order, label mapping, and format. Report variation across orders, not one prompt. More examples can add cost and interference rather than monotonically improving performance.

Add a length-matched no-demonstration control and preserve every prompt and raw output so label-mapping and formatting errors remain inspectable.

Material gap and numbering note

Winter 2026 recordings are not public. The deck cover retains a stale “Lecture 6: Pretraining” label, while the official schedule, date, filename, and sequence establish it as regular lecture 7. This article follows the schedule and does not speculate about the stale label.

References