Skip to content

NTHU NLP Guide 7: Sub-word Tokenization, or Why a Model's Vocabulary Is Made of Word Pieces

Sep 30, 20261 min
TL;DRA guide to the Sub-word Tokenization unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). Splitting on white space only works for Western languages and cannot handle unseen words or German-style compounds. BPE starts from characters and repeatedly merges the most frequent adjacent pair into the vocabulary, so each merge adds one entry; its weakness is that the greedy split is not always the best one. The Unigram LM picks splits by probability and can even sample different ones. The slides give vocabulary sizes of 30522 for BERT, 50257 for GPT-2/GPT-3 and 32,128 for T5.

🌏 中文版

This guide is based on the public Fall 2025 (114-1) materials of Prof. Hung-Yu Kao's Natural Language Processing course at NTHU. It is part 7 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. The previous part is Transformer and Self-Attention.

The official material for this lecture is W3_subword.pdf (43 slides), with recordings Week 5 Tue. and Week 5 Thu. (in Mandarin). As with the previous lecture, the W3 in the file name is an old week number; the 2025 schedule places it in W5, the same row where HW2 is released. That row's Topics column is a syllabus template and is not cited here.

The slides have three parts: recap, word segmentation, and sub-word tokenization.

Recap: the model outputs a probability over the whole vocabulary

Slides 4–5 revisit a language model's output layer. After an RNN reads "I love an", the hidden state goes through a classification layer that outputs a distribution as long as the vocabulary: apple 0.6, elephant 0.3, eraser 0.05, and so on. So how you build the vocabulary decides which words the model can say at all.

Slide 6 lays out the basic NLP pipeline in three steps, build the vocabulary, learn the representations (training), perform predictions (testing), and lists three vocabulary sizes:

ModelVocabulary size (slide 6)
BERT30522
GPT-2 / GPT-350257
T532,128

These match the vocab_size in the Hugging Face configs for bert-base-uncased, gpt2 and t5-base. Note that section 3.1.3 of the T5 paper says "32,000 wordpieces"; the config's 32,128 is 128 larger, and the official materials do not explain the difference.

Where white-space splitting breaks

Slide 7 shows the obvious approach: split "I love apples. I like apples and pineapples." on white space, collect and, apples, I, like, love, pineapples and ".", and add <UNK> for unseen words. Stemming can follow.

Slides 8–9 list three problems:

  1. It only works for Western languages: Chinese and Japanese have no spaces
  2. It cannot handle unseen words: a misspelled word still carries morphological information but becomes <UNK>
  3. Translation does not line up: source and target words are not always one-to-one. The slide's example maps English "sewage water treatment plant" to the single German word Abwasserbehandlungsanlage

The conclusion is that sub-word units are preferable. Slide 11 adds a distinction: all tokenization is segmentation, but not the reverse. Segmentation also covers splitting a document into sentences, while tokenization can mean splitting into words or splitting words further into sub-words. Slide 12 points students to the OpenAI Tokenizer to try it themselves.

Three common algorithms

Slide 10 lists three:

Slide 40 has a fuller table: WordPiece for BERT, ALBERT and MT-DNN; BPE for RoBERTa, XLM and GPT-1/2/3; Unigram for XLNet, T5 and mT5.

The slides disagree with themselves in one place: slide 10 puts T5 under WordPiece, slide 40 under Unigram. The T5 paper says it uses SentencePiece "to encode text as WordPiece tokens", mentioning both terms, so both labels have a source. It is enough to know the discrepancy exists.

BPE: one merge at a time

Slides 13–27 run the whole algorithm on a toy corpus: low ×5, lower ×2, newest ×6, widest ×3. Each word is split into characters with </w> appended, an end-of-word symbol that lets you restore the original split later:

WordCount
l o w </w>5
l o w e r </w>2
n e w e s t </w>6
w i d e s t </w>3

The initial vocabulary is the set of characters seen: </w>, d, e, i, l, n, o, r, s, t, w. Then three actions repeat: find the most frequent adjacent pair, add it to the vocabulary, and merge that pair everywhere in the corpus.

MergePair foundFrequencyAdded to vocabulary
1e s6 + 3 = 9es
2es t6 + 3 = 9est
3est </w>6 + 3 = 9est</w>
4l o5 + 2 = 7lo

After four merges (num_merges = 4) it stops. num_merges is a hyperparameter you set.

To tokenize a new word (slide 28), split it into characters plus </w> the same way, then merge according to the learned vocabulary: low becomes lo w </w>, and widest becomes w i d est</w>.

Slide 29 draws out two properties:

  • Final vocabulary size = initial size + num_merges. Here that is 11 + 4 = 15
  • It is statistical: the more frequent a sub-word is in the corpus, the more likely it enters the vocabulary

BPE's problem: greedy is not always best

Slide 30 notes that BPE by default splits with the largest sub-words available, a greedy, deterministic, left-to-right procedure. "Hello world" may become Hell / o / world, but H / ello / world, He / llo / world and others are also valid, and BPE's choice may be sub-optimal.

Slide 31 gives each sub-word an occurrence probability and compares the product for each split: BPE's Hell / o / world comes to about 4.2×10−6, while H / ello / world reaches about 7.9×10−6. That motivates the next method.

Unigram Language Model: choosing splits by probability

Slide 32 positions ULM: like BPE, it splits sentences into sub-words, but it does so by the joint probability of the whole sentence, and every token in the vocabulary has its own probability learned from the corpus.

Slides 33–37 give three steps:

  1. Set a vocabulary size and build the initial vocabulary: take all sub-words (characters included) from the corpus, keep the most frequent, and compute each one's frequency. The original paper speeds this up with an Enhanced Suffix Array, which the course skips
  2. Train a unigram language model: not a neural network but a probabilistic model, fit with the EM algorithm to maximize corpus likelihood; the best split per sentence is found with the Viterbi algorithm, also skipped in class
  3. Prune the vocabulary: repeatedly compute how much the loss would rise if each sub-word were removed, and keep the top η% (for example η = 80)

The slides link two resources for details: SentencePiece and the Hugging Face NLP Course chapter on Unigram.

Formula: sub-word frequency and subword sampling (slides 35, 38)

$$\mathrm{frequency}(x_i) = \frac{\mathrm{Count}(x_i)}{\sum \text{all counts}}$$

Sample one of the l best segmentations from:

$$P(\mathbf{x}_i \mid X) \approx \frac{P(\mathbf{x}i)^{\alpha}}{\sum{i=1}^{l} P(\mathbf{x}_i)^{\alpha}},\qquad P(\mathbf{x}i) = \prod{j=1}^{n_i} P(t_j)$$

X is the sentence, xi the i-th segmentation, ni its token count, and α controls how smooth the distribution is.

Subword sampling: same word, different splits

Slides 38–39 explain that the original ULM paper does not always choose the most probable split at inference; it samples. The slides run the T5-base tokenizer on "internationalization" five times and get five splits, for example:

  • ▁ / inter / national / ization
  • ▁ / international / ization
  • ▁in / tern / at / i / o / n / ali / z / ation

The leading "▁" is the boundary marker the T5 tokenizer adds before a word.

BPE or ULM?

Slide 41's comparison:

AspectULMBPE
AlgorithmEM algorithmGreedy, no probabilistic model
TrainingSlowerFaster
ConsistencyCan split the same word differentlySame split every time
Inference speedSlowerFaster
Downstream tasksMaybe better for machine translationSuits most tasks

Slides 42–43 wrap up. Sub-word tokenization handles unknown, misspelled and compound words, eases compound-word mismatches in translation, and models like GPT-3 and BERT apply it before pretraining. The downsides are few: num_merges must be tuned, and once the vocabulary is built it is fixed, so new data means rerunning the algorithm. The last line asks: what about Chinese? The slide's answer is character-level encoding.

Going deeper

  • Code the four merges from slides 13–27 in Python: a Counter for adjacent pairs and a merge function. Compare your vocabulary with the slides.
  • Set num_merges to 10 and see whether newest and widest end up as single tokens.
  • Load the t5-base tokenizer in transformers, enable its sampling option, and run "internationalization" a few times against slide 39.
  • Paste a Chinese paragraph and an English paragraph with the same meaning into the OpenAI Tokenizer and compare token counts.

Further reading: CS224N guide: tokenization and multilinguality covers tokenization from a multilingual-fairness angle; CS336 guide: overview and tokenization has you implement byte-level BPE from scratch.

Gaps in the materials

  • This lecture has no dedicated assignment. HW2, released the same week, generates arithmetic at the character level and does not use BPE.
  • WordPiece appears only in the lists; the slides do not explain its algorithm.
  • Slides 10 and 40 classify T5 differently; see above.
  • Solutions, quizzes and class discussion live on NTU COOL and are not available to outside readers. The Fall 2026 version of this lecture is not yet public.

Series navigation: previous Transformer and Self-Attention | next ELMo, BERT, T5, BART, GPT: Three Roads to Pretraining | series overview

References