Skip to content

CME295 Lecture 1: From Tokens to Transformer, or How One Sentence Gets Translated into Another Language

Sep 29, 20261 min
TL;DRCME295 Lecture 1 threads a single sentence, "A cute teddy bear is reading.", through the whole class: split it into tokens, turn them into vectors, see why an RNN can't hold on to long sentences, then translate it into French with self-attention and an encoder-decoder. The 2026 edition drops the entire section on NLP tasks and evaluation metrics and opens instead with a timeline running from 2017 to the agent era.

🌏 中文版

This post covers Lecture 1, "Transformer," of Stanford's CME295. The main sources are the 2025 recording and its 135-page slide deck, cross-checked against the new edition released on September 25, 2026 (recording, slides).

The whole lecture runs on one example sentence: "A cute teddy bear is reading." It gets split into tokens, turned into vectors, handed to an RNN, and then we see why the RNN can't cope, until finally a Transformer translates it into French: "Un ours en peluche mignon lit." Follow that sentence from start to finish and you've covered the entire lecture.

Step 1: Split the sentence into tokens

Models don't see text; they see tokens. There are three granularities for splitting, and the slides list the trade-offs of each:

GranularityExampleUpsideCost
word-levelcute teddy bearSimple, easy to interpretHuge vocabulary; can't tell that reading and read share a root; helpless with unseen words
character-levelc u t e …Small vocabulary; handles random capitalization and typosMuch longer sequences, slower compute; a single letter's vector carries no meaning
subword-levelted ##dy read ##ingShares roots; learned from dataNeeds an extra training step; quality depends on the training corpus

A single typo exposes the problem with word-level tokenization: in "tedi bear," tedi isn't in the vocabulary, so it can only become [UNK] (unknown). Subword tokenization breaks it into smaller pieces it does recognize, which is why nearly every LLM today uses subwords (for example BPE).

Besides [UNK], the slides introduce three special tokens: [BOS] marks the beginning of a sequence, [EOS] marks the end, and [PAD] pads sentences of different lengths to the same size. The slides also point out that real models have many more special tokens, such as the markers chat models use to separate user and assistant turns, and every vendor writes them differently. If you want to see how BPE merges bytes step by step, read the CS336 Lecture 1 guide.

Step 2: Turn tokens into vectors

The most direct representation is one-hot: with a vocabulary of V words, each token is a length-V vector with a single 1. The problem is that any two one-hot vectors are orthogonal, so "cute" is exactly as far from "soft" as it is from "laughing."

word2vec sets up a proxy task that forces a neural network to learn meaningful vectors on its own:

  • CBOW: predict the middle word from its surrounding context
  • Skip-gram: predict the surrounding context from the middle word

The network has only three layers: a V-dimensional one-hot input, a d-dimensional hidden layer, and an output that maps back to a V-dimensional probability distribution. After training, we throw away its predictions and keep only the middle layer, which becomes each word's embedding. The slides walk through predicting the next word in "a cute teddy bear is reading" step by step, with toy numbers small enough to compute by hand.

word2vec has two fundamental limits: it ignores word order, and a word gets the same vector no matter which sentence it appears in, so it can't adapt to context. Those two gaps are exactly what the next step sets out to fix. For the full derivation of word2vec, see the CS224N Lecture 2 guide.

Step 3: RNNs remember order but lose track of long sentences

An RNN reads one token at a time and passes its understanding so far along in a hidden state. That solves the word-order problem, and the same architecture can handle classification, sequence labeling, text generation, and translation. In 1997, the LSTM added more structured gates to the hidden state and went on to set the state of the art at the time.

The slides list two drawbacks. First, vanishing gradients: as a sentence gets longer, information from the beginning fades by the time it reaches the end. Second, slow computation: tokens have to be processed one after another, with no way to parallelize.

Translation magnifies the first problem. A seq2seq model has to compress the entire English sentence into one fixed-length vector and then produce French from that vector, so with longer sentences it "forgets" what came earlier. In 2014, Bahdanau et al. proposed a fix: when generating each French word, look back over the original English sentence and decide which words to align with right now. That was the beginning of attention.

Step 4: Attention lets each token find the words that matter to it

The 2017 paper Attention Is All You Need went a step further: if attention works this well, drop the RNN entirely and build the whole model out of attention.

The slides explain self-attention through three roles: Query, Key, and Value. The intuition is a lookup:

  • Query: what I'm looking for right now (e.g. "bear" wants to know what's describing it)
  • Key: the label each token puts out so others can judge whether it's relevant to them
  • Value: the actual content that gets taken

Each token compares its Query against every token's Key to get relevance scores, then takes a weighted average of all the Values using those scores. The result is that the new vector for "bear" now carries information from "cute" and "teddy," which is exactly the context-dependent representation word2vec couldn't produce. And since every token can be computed at the same time, the RNN's speed bottleneck disappears too.

Formula: scaled dot-product attention
Attention(Q, K, V) = softmax( Q Kᵀ / √d_k ) V
  • Q Kᵀ: relevance scores for every pair of tokens, computed in one matrix multiplication
  • √d_k: scaling. Question I.5 of the 2025 midterm tests exactly this: when the vector dimension is large, dot products get large, softmax saturates, and gradients go nearly to zero
  • softmax: turns the scores into weights that sum to 1
  • Multiplying by V: takes the weighted average of the Values

Step 5: Assemble a Transformer

The original Transformer was designed for translation and has two halves:

  • Encoder: in the slides' words, "compute meaningful embeddings." It reads the whole English sentence and computes a context-aware vector for every token
  • Decoder: "generate next token." It produces one French token at a time
flowchart LR
  A["A cute teddy bear is reading."] --> T["Tokenize<br/>add [BOS] [EOS]"]
  T --> E["Embedding<br/>+ positional encoding"]
  E --> ENC["Encoder × N<br/>self-attention → FFN"]
  ENC --> DEC["Decoder × N<br/>masked self-attention<br/>→ encoder-decoder attention<br/>→ FFN"]
  D0["[BOS] Un ours ..."] --> DEC
  DEC --> S["Linear layer + softmax<br/>next-token probabilities"]
  S --> O["Un ours en peluche mignon lit."]

Because attention on its own can't see order, the input first gets a positional encoding, either a learned vector or a fixed sinusoidal function. The decoder has one extra layer compared with the encoder, encoder-decoder attention: while generating French, it uses Queries from the French side to look up Keys and Values from the English side, which does the same job as Bahdanau's attention. The final output layer is really a classification problem, where the classes are the words in the vocabulary.

The slides spend a good deal of space on the tricks that make this machine trainable at all:

TechniqueWhat it doesWhy it's needed
Residual connectionAdds the sublayer's input directly to its outputGives gradients a shortcut to flow backward
Layer normalizationNormalizes each token's hidden vectorKeeps numerical scale stable across layers, faster convergence
Masking (causal)Hides future tokens not yet generated during trainingStops the model from peeking at the answer, and lets the whole sentence be computed in one vectorized pass
Multi-head attentionRuns several attention computations in parallelEach head captures a different relationship, much like multiple filters in a CNN
DropoutRandomly switches off some connectionsBetter generalization
Label smoothingLowers the correct answer's probability slightly below 1Prevents overconfidence and improves BLEU

Connecting back to the models you use

Today's mainstream chat LLMs (the GPT series, Llama, Qwen, and other models whose architectures have been made public) aren't this full encoder-decoder machine; they keep only the decoder. There's no "English sentence" for them to read, only "the conversation so far," and the task is always to predict the next token. Masking, multi-head attention, residual connections, and layer norm all carry over unchanged.

So although this lecture uses translation as its example, what it's really covering are the parts inside every LLM today. Lecture 2 picks up from here to show how the same Transformer split into encoder-only BERT and decoder-only GPT, and how attention was reworked to use less memory.

What changed in 2026

Comparing the two slide decks (135 pages in 2025, 118 in 2026), the core is nearly identical, including the closing "Stitching all the pieces together" example that walks the example sentence through the entire Transformer one cell at a time, which appears in both. The differences are at the beginning and end:

  • The whole "NLP overview" section is gone: the 2025 edition opened with three task types (sentiment analysis, named entity recognition, translation), along with evaluation metrics such as BLEU, ROUGE, F1, and perplexity, plus datasets. The 2026 edition covers these tasks in a single timeline slide.
  • The timeline gains an "Agentic era": it features Claude Code, Cursor, Codex, and Antigravity, and the final slide says the course will cover both the "conversational" and "agentic" eras.
  • A new batch of acronyms: MLA, SWA, QKNorm, GSPO, RLVR, SWE-bench, HLE, and others are added, while GloVe, CoT, ToT, RAG, and others are dropped, which also hints at where later lectures will focus.

Self-check

These questions are adapted from Part I of the 2025 midterm; answers are in the solutions PDF:

  1. Compared with word-level tokenization, what is the main advantage of subword tokenization? (Question 1)
  2. Which word2vec proxy task predicts the middle word from its surrounding context? (Question 2)
  3. Which sublayer appears only in the decoder and not in the encoder? (Question 4)
  4. Why does scaled dot-product attention divide by √d_k? (Question 5)
  5. Write out the self-attention formula and explain the role of Q, K, and V. (Question 9)
  6. What does label smoothing optimize for, and why does it help generalization? (Question 10)

Going deeper

References