🌏 中文版
This guide is based on the public Fall 2025 (114-1) materials of Prof. Hung-Yu Kao's Natural Language Processing course at NTHU. It is part 7 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. The previous part is Transformer and Self-Attention.
The official material for this lecture is W3_subword.pdf (43 slides), with recordings Week 5 Tue. and Week 5 Thu. (in Mandarin). As with the previous lecture, the W3 in the file name is an old week number; the 2025 schedule places it in W5, the same row where HW2 is released. That row's Topics column is a syllabus template and is not cited here.
The slides have three parts: recap, word segmentation, and sub-word tokenization.
Recap: the model outputs a probability over the whole vocabulary
Slides 4–5 revisit a language model's output layer. After an RNN reads "I love an", the hidden state goes through a classification layer that outputs a distribution as long as the vocabulary: apple 0.6, elephant 0.3, eraser 0.05, and so on. So how you build the vocabulary decides which words the model can say at all.
Slide 6 lays out the basic NLP pipeline in three steps, build the vocabulary, learn the representations (training), perform predictions (testing), and lists three vocabulary sizes:
| Model | Vocabulary size (slide 6) |
|---|---|
| BERT | 30522 |
| GPT-2 / GPT-3 | 50257 |
| T5 | 32,128 |
These match the vocab_size in the Hugging Face configs for bert-base-uncased, gpt2 and t5-base. Note that section 3.1.3 of the T5 paper says "32,000 wordpieces"; the config's 32,128 is 128 larger, and the official materials do not explain the difference.
Where white-space splitting breaks
Slide 7 shows the obvious approach: split "I love apples. I like apples and pineapples." on white space, collect and, apples, I, like, love, pineapples and ".", and add <UNK> for unseen words. Stemming can follow.
Slides 8–9 list three problems:
- It only works for Western languages: Chinese and Japanese have no spaces
- It cannot handle unseen words: a misspelled word still carries morphological information but becomes
<UNK> - Translation does not line up: source and target words are not always one-to-one. The slide's example maps English "sewage water treatment plant" to the single German word Abwasserbehandlungsanlage
The conclusion is that sub-word units are preferable. Slide 11 adds a distinction: all tokenization is segmentation, but not the reverse. Segmentation also covers splitting a document into sentences, while tokenization can mean splitting into words or splitting words further into sub-words. Slide 12 points students to the OpenAI Tokenizer to try it themselves.
Three common algorithms
Slide 10 lists three:
- BPE (Sennrich et al., 2016): the GPT series
- WordPiece (Schuster & Nakajima, 2012): BERT
- Unigram Language Model (Kudo, 2018)
Slide 40 has a fuller table: WordPiece for BERT, ALBERT and MT-DNN; BPE for RoBERTa, XLM and GPT-1/2/3; Unigram for XLNet, T5 and mT5.
The slides disagree with themselves in one place: slide 10 puts T5 under WordPiece, slide 40 under Unigram. The T5 paper says it uses SentencePiece "to encode text as WordPiece tokens", mentioning both terms, so both labels have a source. It is enough to know the discrepancy exists.
BPE: one merge at a time
Slides 13–27 run the whole algorithm on a toy corpus: low ×5, lower ×2, newest ×6, widest ×3. Each word is split into characters with </w> appended, an end-of-word symbol that lets you restore the original split later:
| Word | Count |
|---|---|
l o w </w> | 5 |
l o w e r </w> | 2 |
n e w e s t </w> | 6 |
w i d e s t </w> | 3 |
The initial vocabulary is the set of characters seen: </w>, d, e, i, l, n, o, r, s, t, w. Then three actions repeat: find the most frequent adjacent pair, add it to the vocabulary, and merge that pair everywhere in the corpus.
| Merge | Pair found | Frequency | Added to vocabulary |
|---|---|---|---|
| 1 | e s | 6 + 3 = 9 | es |
| 2 | es t | 6 + 3 = 9 | est |
| 3 | est </w> | 6 + 3 = 9 | est</w> |
| 4 | l o | 5 + 2 = 7 | lo |
After four merges (num_merges = 4) it stops. num_merges is a hyperparameter you set.
To tokenize a new word (slide 28), split it into characters plus </w> the same way, then merge according to the learned vocabulary: low becomes lo w </w>, and widest becomes w i d est</w>.
Slide 29 draws out two properties:
- Final vocabulary size = initial size + num_merges. Here that is 11 + 4 = 15
- It is statistical: the more frequent a sub-word is in the corpus, the more likely it enters the vocabulary
BPE's problem: greedy is not always best
Slide 30 notes that BPE by default splits with the largest sub-words available, a greedy, deterministic, left-to-right procedure. "Hello world" may become Hell / o / world, but H / ello / world, He / llo / world and others are also valid, and BPE's choice may be sub-optimal.
Slide 31 gives each sub-word an occurrence probability and compares the product for each split: BPE's Hell / o / world comes to about 4.2×10−6, while H / ello / world reaches about 7.9×10−6. That motivates the next method.
Unigram Language Model: choosing splits by probability
Slide 32 positions ULM: like BPE, it splits sentences into sub-words, but it does so by the joint probability of the whole sentence, and every token in the vocabulary has its own probability learned from the corpus.
Slides 33–37 give three steps:
- Set a vocabulary size and build the initial vocabulary: take all sub-words (characters included) from the corpus, keep the most frequent, and compute each one's frequency. The original paper speeds this up with an Enhanced Suffix Array, which the course skips
- Train a unigram language model: not a neural network but a probabilistic model, fit with the EM algorithm to maximize corpus likelihood; the best split per sentence is found with the Viterbi algorithm, also skipped in class
- Prune the vocabulary: repeatedly compute how much the loss would rise if each sub-word were removed, and keep the top η% (for example η = 80)
The slides link two resources for details: SentencePiece and the Hugging Face NLP Course chapter on Unigram.
Formula: sub-word frequency and subword sampling (slides 35, 38)
$$\mathrm{frequency}(x_i) = \frac{\mathrm{Count}(x_i)}{\sum \text{all counts}}$$
Sample one of the l best segmentations from:
$$P(\mathbf{x}_i \mid X) \approx \frac{P(\mathbf{x}i)^{\alpha}}{\sum{i=1}^{l} P(\mathbf{x}_i)^{\alpha}},\qquad P(\mathbf{x}i) = \prod{j=1}^{n_i} P(t_j)$$
X is the sentence, xi the i-th segmentation, ni its token count, and α controls how smooth the distribution is.
Subword sampling: same word, different splits
Slides 38–39 explain that the original ULM paper does not always choose the most probable split at inference; it samples. The slides run the T5-base tokenizer on "internationalization" five times and get five splits, for example:
- ▁ / inter / national / ization
- ▁ / international / ization
- ▁in / tern / at / i / o / n / ali / z / ation
The leading "▁" is the boundary marker the T5 tokenizer adds before a word.
BPE or ULM?
Slide 41's comparison:
| Aspect | ULM | BPE |
|---|---|---|
| Algorithm | EM algorithm | Greedy, no probabilistic model |
| Training | Slower | Faster |
| Consistency | Can split the same word differently | Same split every time |
| Inference speed | Slower | Faster |
| Downstream tasks | Maybe better for machine translation | Suits most tasks |
Slides 42–43 wrap up. Sub-word tokenization handles unknown, misspelled and compound words, eases compound-word mismatches in translation, and models like GPT-3 and BERT apply it before pretraining. The downsides are few: num_merges must be tuned, and once the vocabulary is built it is fixed, so new data means rerunning the algorithm. The last line asks: what about Chinese? The slide's answer is character-level encoding.
Going deeper
- Code the four merges from slides 13–27 in Python: a
Counterfor adjacent pairs and a merge function. Compare your vocabulary with the slides. - Set
num_mergesto 10 and see whether newest and widest end up as single tokens. - Load the
t5-basetokenizer intransformers, enable its sampling option, and run "internationalization" a few times against slide 39. - Paste a Chinese paragraph and an English paragraph with the same meaning into the OpenAI Tokenizer and compare token counts.
Further reading: CS224N guide: tokenization and multilinguality covers tokenization from a multilingual-fairness angle; CS336 guide: overview and tokenization has you implement byte-level BPE from scratch.
Gaps in the materials
- This lecture has no dedicated assignment. HW2, released the same week, generates arithmetic at the character level and does not use BPE.
- WordPiece appears only in the lists; the slides do not explain its algorithm.
- Slides 10 and 40 classify T5 differently; see above.
- Solutions, quizzes and class discussion live on NTU COOL and are not available to outside readers. The Fall 2026 version of this lecture is not yet public.
Series navigation: previous Transformer and Self-Attention | next ELMo, BERT, T5, BART, GPT: Three Roads to Pretraining | series overview
References
- IKMLab/NTHU_Natural_Language_Processing (course GitHub repo)
- 2025 schedule README
- W3_subword.pdf (Sub-word Tokenization slides)
- Recording: [Fall 2025] Natural Language Processing - Prof. Hung-Yu Kao - Week 5 Tue. (in Chinese)
- Recording: [Fall 2025] Natural Language Processing - Prof. Hung-Yu Kao - Week 5 Thu. (in Chinese)
- Sennrich, Haddow & Birch (2016). Neural Machine Translation of Rare Words with Subword Units
- Kudo (2018). Subword Regularization
- Raffel et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)
- google/sentencepiece
- Hugging Face NLP Course: Unigram tokenization
- OpenAI Tokenizer
- bert-base-uncased config.json
- gpt2 config.json
- t5-base config.json
Loading...