Skip to content

NTHU Hung-Yu Kao NLP, Week 2: Word Embeddings and Language Models, from Counting N-grams to RNNs

Sep 30, 20261 min
TL;DRWeek 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.

🌏 中文版

Version note: this post is based on the Fall 2025 run of NTHU Hung-Yu Kao's Natural Language Processing, specifically W2_Word embeddings and Language Modeling (RNN).pdf (62 pages). The matching recordings are Week 2 Tue. and Week 2 Thu. (in Mandarin). Facts were checked against the slides on 2026-09-30. The post follows the slides only; I did not transcribe the recordings. Access rating A3 (see the series overview for why).

Series position: previous: Intro to NLP and classical text processing | next: HW1 Word Analogy | series overview

"Please turn your homework …" What comes next? Over? In? You can probably guess, and guess well. Week 2's slides start from that question: how did predicting the next word go from counting to neural networks?

The previous post turned text into vectors from the retrieval side. This one switches to language models. The first slide is titled "GAI Motivation" and lists two problems with supervised learning (text classification, QA systems): not enough training data, and limited domain knowledge. The next slide says the most common way to generate a sentence is to write the words down one after another.

Statistical language models: counting n-grams

The slides recall Markov (1913, on how the chance of a letter depends on the letter before it) and Shannon (1951, "Prediction and Entropy of Printed English"). Then they use a collage of Chinese song lyrics: in that text, the probability that one particular character follows "you" is 1/4, and the probability that it follows a longer three-character history is 0. A language model learns exactly these conditional probabilities.

An n-gram is a sequence of n words: the unigram "please," the bigram "please turn," the trigram "please turn your." The slides slip in an application: the J.K. Rowling pen-name case reported by Scientific American. The stylometry software JGAAP used four-character sequences (four-grams), the frequency of the most common words, the distribution of word lengths, and frequent word pairs to link The Cuckoo's Calling to the author of Harry Potter.

From counts to probabilities

A bigram model is counting. The slides first use three sentences (C(I want) = 2, C(want to) = 3), then give a bigram count table for eight words: "I" followed by "want" appears 827 times, and "want" followed by "to" appears 608 times.

The table has many zeros. Unseen does not mean impossible, so the slides apply add-k smoothing (k=1): add 1 to every cell, then divide by the totals to get relative frequencies. Afterward P(want | I) is about 0.21 and P(to | want) about 0.26.

A sentence's probability is split with the chain rule, then simplified with the Markov assumption: a word's probability depends only on the previous word (bigram) or the previous n−1 words. That step makes the computation feasible, and it plants every limitation that follows.

Perplexity: evaluating a language model

The slides say lower perplexity means a better language model, and give four readings of it:

  1. A measure of uncertainty: how unsure the model is when predicting.
  2. An average branching factor: how many choices the model picks among at each step, on average.
  3. A quantification of performance: how well the model has captured the rules and structure of the language.
  4. A compression indicator: lower perplexity means the model assigns higher probability to the test data, which is better compression.

The slide ends with a question it does not answer: can perplexity tell whether a text was written by AI? Keep it in mind. Perplexity comes back in the decoding and evaluation post.

Four limits of n-grams

  • Limited context: cannot capture dependencies much longer than N.
  • Data sparsity: parameters grow exponentially as N grows.
  • Ignoring word order and context: words are assumed independent.
  • Low flexibility: weak at synonyms and at adapting to different settings such as dialogue.

Sparse vectors: TF-IDF and PPMI

Before neural networks, the slides cover the sparse ways to represent a word by its context. Each dimension is a word, and most values are 0.

TF-IDF is a week 1 recap. The new one is PPMI. It starts from the distributional hypothesis: words in similar contexts have similar meanings ("I enjoy coding" and "I like coding").

PMI compares two things: how often word w and context word c actually co-occur, and how often they would if they were independent, then takes log2. PPMI sets every negative value to 0.

The example counts four words across five contexts: cherry appears with pie 442 times, digital with computer 1,670 times. After converting to probabilities and computing PPMI, cherry's vector is (0, 0, 0, 4.38, 3.30), nonzero only on pie and sugar, while digital and information concentrate on computer, data, and result. For the cherry–sugar cell, the slide's arithmetic is log2(0.0021 / (0.0415 × 0.0052)) ≈ 3.30.

Dense vectors: Word2Vec and contextualized embeddings

Dense vectors place words in a continuous space where similar words sit close together. The slides show two properties:

  • Analogy: Washington − U.S. + U.K. = London. This is exactly what HW1 in the next post asks you to test.
  • Semantic change: embeddings can track how meaning shifts over time. In the slide's figure, broadcast in the 1850s sits near sow, seed, and scatter, and by the 1900s it has moved next to newspapers and television. Network sits near telegraph and wires in the 1920s, near Internet and Email in the 1990s, and near cloud computing, blockchain, and IoT in the 2020s.

Word2Vec with negative sampling

Week 1 covered the Skip-gram network. This week explains training in four steps:

  1. Treat the target word and a context word from its window as a positive example.
  2. Randomly sample other words from the vocabulary as negative examples.
  3. Train a logistic-regression classifier to tell the two apart.
  4. Use the learned weights as the embeddings.

The slides also flag efficiency: each update touches only the vectors in the window. With window size m, a window holds just 2m+1 words, so the gradient is sparse.

Why one word needs more than one vector

Word2Vec gives each word one fixed vector. The slides use the Chinese word for "apple" to show why that falls short. Apple the company and apple pie are different things. "An apple changed his life" means one thing for Newton and another for Steve Jobs.

Contextualized embeddings give the same word different vectors depending on context; the slides cite BERT and GPT. They learn from long contexts rather than a small window, and use every layer of a deep neural language model. The slides close with a practical question: in a downstream task, should the embedding layer be frozen or trained along with the rest?

Neural language models: from FFN to RNN

Processing these vectors needs a model. The slides fill in a few pages of deep-learning basics:

  • Three components: the model (the structure from inputs to outputs), the optimizer (the algorithm that adjusts parameters to reduce error), and the loss function (how far predictions are from targets).
  • Six training steps: prepare data, build the model in a framework (TensorFlow, PyTorch), pick a loss (cross-entropy and others), pick an optimizer (Adam, SGD, and others), train, evaluate.
  • Activation functions: softmax outputs a distribution summing to 1, for multi-class classification; sigmoid outputs 0 to 1, for binary classification; tanh outputs −1 to 1, zero-centered, common in hidden layers; ReLU adds nonlinearity and avoids vanishing gradients.

Three weaknesses of FFNs

A feedforward network is a multilayer network with no cycles between units. The slides list three problems for language:

  • No sequence modeling: word order and dependencies matter.
  • Fixed input size: sentences vary in length.
  • Limited context: many tasks need long-range dependencies.

RNNs

RNNs are built for sequences. The slides describe them as an "advanced moving average." Each step's hidden state is computed from the previous hidden state and the current input through learnable weights and a nonlinearity, and every time step shares the same weights.

A timeline runs from the Hopfield network, Elman RNN, BPTT, LSTM, bidirectional RNNs, GRU, and Seq2seq to attention (2015) and the Transformer (2017). The slides list four properties of RNNs:

  • Sequential processing, which models dependencies over time.
  • Recurrent connections, which keep an internal memory.
  • Parameter sharing across time steps, which makes learning efficient.
  • Vanishing gradients: plain RNNs struggle to learn long-range dependencies.

This week does not expand on the last point. It is where post 4, Seq2seq and attention, begins.

What RNNs can do

  • Named entity recognition (NER): find countries, organizations, and people in a sequence. It is token classification: each step's output goes through an FFN to a one-hot label.
  • Sentence classification: classify the whole sequence, not each token. Take the last token's hidden state and feed it to an FFN and softmax.
  • Stacked RNNs: several RNN layers, each taking the previous layer's output. They usually beat a single layer because layers learn representations at different levels of abstraction, but training cost rises quickly with depth.
  • Bidirectional RNNs: many applications need the whole input. Two independent RNNs read start-to-end and end-to-start; for sentence classification, the final hidden states from both directions are combined and passed to the classifier.

The week's map

The final summary slide is this post's skeleton:

CategoryMethodsCore idea
Statistical LMn-gramCount, Markov assumption
Sparse vectorsTF-IDF, PPMIEncode a word by its contexts
Dense vectorsWord2Vec, contextualized embeddingsSelf-supervised training
Neural LMFFN, RNNFully connected network; recurrent structure that keeps a hidden state

The Fall 2026 counterpart

The 2026 main README links the v2 of this deck in W3, with the Fall 2026 Week 3 recording, and HW1 is released the same week. v2 is also 62 pages, and its extracted text is nearly identical to the 2025 deck apart from layout.

What to do after reading

  • Tonight: take a text you know well, count every bigram with Python's collections.Counter, pick a word, and see what most often follows it. That is an n-gram model learning only what it has seen.
  • One step further: hand-compute one cell of the slide's PPMI table and confirm where 3.30 comes from.
  • Next post: HW1 Word Analogy has you run analogy questions on pretrained vectors and on a Word2Vec you train yourself, to check whether relations like Washington − U.S. + U.K. = London were really learned.

Further reading

References