Skip to content
Series
20 posts

Reading NTHU Hung-Yu Kao Natural Language Processing

Reading NTHU Prof. Hung-Yu Kao's Mandarin TAICA NLP course through its complete Fall 2025 materials on the official IKMLab GitHub: lecture slides, 36 public W1–W16 recordings, HW1–HW4 specs and notebooks, and TA tutorials on PyTorch, Hugging Face, LLM APIs and RAG. The series moves from classic text processing, word embeddings, seq2seq, Transformers and the BERT family through decoding and evaluation to RLHF, PEFT, RAG and reasoning, and ends with what changes in Fall 2026.

Reading NTHU Hung-Yu Kao's Natural Language Processing: What Outsiders Can Get from a 1,200-Seat TAICA Course

Hung-Yu Kao's Natural Language Processing at National Tsing Hua University is a graduate-level flagship course in the TAICA alliance. The syllabus caps it at 1,200 students, it is taught in Mandarin, and it runs from TF-IDF and word vectors to RLHF, PEFT, and RAG. For Fall 2025, the slides, 32 class recordings, and 4 assignments with starter notebooks are all on GitHub, which rates A3. Solutions, grading, and the term-project spec are not public. Fall 2026 is in progress and only goes up to W3, so it rates A2. Grading changed to 75% assignments plus a 25% in-person midterm, and a Reasoning/Agent unit was added.

NTHU Hung-Yu Kao NLP, Week 1: Why Language Is Hard, and How Text Became Numbers Before LLMs

Week 1 of Hung-Yu Kao's NLP course at NTHU is a 91-page deck, W1_NLP_brief. It opens with 'Watch for kids', five readings of the telescope sentence, and a Chinese tongue-twister about eleven uncles to show why language is hard. Then, from an information-retrieval angle, it builds one pipeline: inverted index, tokenization, stemming, TF-IDF, BM25. The second half hits that pipeline's dead ends (synonyms, polysemy, vocabulary mismatch), moves to SVD-based LSA, and closes with a preview of dense vectors through Skip-gram, GloVe, and FastText.

NTHU Hung-Yu Kao NLP, Week 2: Word Embeddings and Language Models, from Counting N-grams to RNNs

Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.

NTHU NLP HW1: Testing Word Vectors on Google Analogy — Pretrained GloVe vs. Word2Vec Trained on 20% of Wikipedia

HW1 tests word vectors on the 19,544 Google Analogy questions (8,869 semantic, 10,675 syntactic). You first answer them with pretrained glove-wiki-gigaword-100 loaded through Gensim, then train your own Word2Vec on a 20% sample of a pre-cleaned Wikipedia dump, and plot t-SNE for the family subcategory both times. Seven TODOs are worth 55%, the report 45%. Fall 2026 keeps the same TODOs but asks for an .ipynb with outputs.

NTHU NLP W3: When Input and Output Lengths Differ — Seq2seq, LSTM, and Attention

"Look over there" is 3 tokens, the Chinese version 4, the Japanese version 12. When output length does not track input length, the classifier trick of adding an FFN at the end fails. This 33-slide deck starts from encoder-decoder models, works through vanishing gradients in RNNs and the three LSTM gates, and ends with attention fixing both long-range memory and parallelism, including why scores are divided by √d.

NTHU NLP HW2: Arithmetic as a Language — PyTorch TA Session, a Two-Layer LSTM, and Teacher Forcing

HW2 treats expressions like "14*(43+20)=882" as character sequences and asks a two-layer LSTM to generate the answer one character at a time after it sees "=". The training set has 2,369,250 rows and the eval set 263,250, with every number in 0–49. Six TODOs run from building a vocabulary and batching with loss only after "=" to a generator, teacher-forced training, and exact-match evaluation. The W4 PyTorch TA session is the toolbox for it.

NTHU NLP Guide 6: Without the RNN, How Does a Transformer Know How Words Relate and Where They Sit?

A guide to the Transformer unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). The slides start from two RNN problems: words interact only across O(N) steps, and time steps cannot run in parallel. Self-attention lets every word look at every other word and computes it all in one matrix multiplication, fixing both at once. The cost is that the model can no longer tell word order, so sinusoidal positional encoding is added. The lecture then assembles multi-head attention, Add & Norm, feed forward, cross-attention, masked attention and teacher forcing, and closes with GPT-2, ViT and four variants to show how far the architecture went.

NTHU NLP Guide 7: Sub-word Tokenization, or Why a Model's Vocabulary Is Made of Word Pieces

A guide to the Sub-word Tokenization unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). Splitting on white space only works for Western languages and cannot handle unseen words or German-style compounds. BPE starts from characters and repeatedly merges the most frequent adjacent pair into the vocabulary, so each merge adds one entry; its weakness is that the greedy split is not always the best one. The Unigram LM picks splits by probability and can even sample different ones. The slides give vocabulary sizes of 30522 for BERT, 50257 for GPT-2/GPT-3 and 32,128 for T5.

NTHU NLP Guide 8: ELMo, BERT, T5, BART, GPT and the Three Roads to Pretraining

A guide to the BERT and its Family unit of Prof. Hung-Yu Kao's NLP course at NTHU (Fall 2025). It starts from "I record the record": Word2Vec and GloVe give both records the same vector. ELMo fixes this with a bidirectional LSTM language model. With Transformers, pretraining splits into three roads: encoders (BERT: MLM plus NSP, strong at understanding, weak at generation), encoder-decoders (T5's span corruption, BART's five noise types), and decoders (GPT: pure next-token prediction). The last part covers GPT-3's in-context learning and scaling laws, and why decoders became today's dominant backbone.

NTHU NLP Guide 9: One BERT That Scores and Classifies at Once — The Hugging Face Tutorial and HW3 Multi-Output Learning

A guide to the Hugging Face BERT tutorial and HW3 in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The tutorial walks through binary IMDb sentiment classification: AutoTokenizer, the input_ids / token_type_ids / attention_mask fields, AutoModelForSequenceClassification, and Trainer. HW3 applies the same tools to SemEval 2014 Task 1: one bert-base-uncased with two heads, one regressing a 1–5 relatedness score and one classifying entailment into three classes. You add the two losses and write the training loop yourself, because Trainer is not allowed.

NTHU NLP Guide 10: Why the Same Model Sounds Stiff or Rambles — Decoding Strategies and NLG Evaluation

A guide to the decoding and evaluation unit in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). The first half covers how to pick a word once the model outputs a probability distribution: greedy decoding can't take back a mistake, beam search keeps several candidates but favors short outputs, and top-k / top-p trade determinism for diversity (missing from the slides; the professor covers it verbally in class). The second half covers scoring generated text: BLEU's modified precision and brevity penalty, ROUGE-N and ROUGE-L, perplexity, and what GLUE, SQuAD 2.0, MTEB, and MMLU each measure.

NTHU NLP Guide 11: GPT-2 vs. T5 for Chinese Summarization — Left Padding, −100, and ROUGE

A guide to the GPT-2 / T5 tutorial in NTHU Prof. Hung-Yu Kao's NLP course (Fall 2025). One task, LCSTS Chinese summarization, is solved twice. Decoder-only GPT-2 is written in native PyTorch: you join article and summary into one sequence, switch to left padding, and set padding labels to −100. Encoder-decoder mT5 uses Seq2SeqTrainer: no left padding needed, and DataCollatorForSeq2Seq handles the −100 for you. Both segment with jieba and score word-level ROUGE.

Reading NTHU Kao's NLP: GPT-3, InstructGPT, and RLHF — How a Model That Continues Text Becomes an Assistant That Follows Instructions

Hung-Yu Kao's Fall 2025 W8 slides walk from GPT-1 to GPT-3, explain how the Sparse Transformer behind GPT-3 cuts attention cost, and then use InstructGPT to show the gap between continuing text and following instructions. The maximum likelihood objective can't tell a fabricated fact from a slightly wrong synonym, so three extra stages are added: SFT learns how humans write, a reward model learns how humans grade, and PPO optimizes against that grade while a KL penalty keeps the model from drifting too far. The lecture closes with Llama-2: separate safety and helpfulness reward models, context distillation, and GQA for faster inference.

Reading NTHU Kao's NLP: Parameter-Efficient Fine-Tuning — Fine-Tuning Large Models Without an A100 Cluster

Hung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.

Reading NTHU Kao's NLP: RAG (Part 1) — Hallucination and Retrievers, from BM25 to DPR and GTR

An LLM will confidently answer that Oppenheimer was born in 1967 (the year he died). RAG retrieves first and generates second. The first 60 pages of Hung-Yu Kao's Fall 2025 RAG deck are all about finding the right material. Sparse vectors (bag-of-words, TF-IDF, BM25) are cheap and dependable; dense vectors catch paraphrases. BERT's [CLS] isn't a good sentence vector as-is, hence Sentence-BERT pooling and bi-encoders. A cross-encoder is accurate but needs nearly 50 million passes for 10,000 sentences, while a bi-encoder needs 20,000. SimCSE uses dropout as data augmentation, DPR beats BM25 with only 1,000 training examples, and GTR shows that scaling up a dual encoder improves out-of-domain retrieval.

NTHU NLP RAG, Part 2: Connecting the Retriever to a Reader, from ORQA and REALM to Self-RAG

The second half of W11_RAG.pdf starts at the "From Retrievers to QA" slide and turns a retriever plus a reader into a full QA system. It begins with ORQA and REALM (2019–2020), where BERT is the reader, then covers the first paper named RAG and REPLUG, which keeps the LLM frozen. A single table then sorts seven recent fixes into three groups: rewrite the query (Query Rewriting, HyDE), make the generator robust to noise (RetRobust, RAFT, RAAT), and decide when to retrieve (FLARE, Self-RAG). It ends with noise types, four abilities an LLM needs inside RAG, and generative retrieval (GR) with reliable response generation (RRG). On the recording side, the W11 Thursday lecture stops at the RAG paper; I could not find a recording that covers the later slides.

NTHU NLP LLM API Lab: NLI Classification with Gemini, OpenAI, and Claude, Using prompts.yaml, JSON Output, Few-Shot, and Token Counts

This 34-slide TA session answers a practical question. Pasting data into the ChatGPT web page one row at a time is slow and hits hourly limits, so research and homework should use the API. The notebook runs one SemEval 2014 entailment example through Gemini, Claude, and OpenAI in turn: prompts live in prompts.yaml, output is forced into JSON, then few-shot and token counting. The material is from 2024. The slide cover says 2024/11/21, and the notebook uses gemini-1.5-pro, gpt-4o, and claude-3-5-sonnet-20241022. That Claude model was retired on 2025-10-28, and Google's old Gemini SDK reached end of support on 2025-11-30.

NTHU NLP RAG Labs + HW4: Building a Cat-Facts RAG Two Ways, with LangChain and by Hand

Each of the two RAG TA sessions builds one version. The first installs Ollama on Colab to run llama3.2:1b and wires up a minimal RAG with LangChain's Chroma, MMR, and retrieval chain. The second uses LangChain only for data prep and writes the rest by hand: chunking, text and vector stores, hybrid BM25 + cosine retrieval merged with RRF, then generation with Llama-3.2-1B-Instruct. HW4 applies the first session's skeleton to 150 cat facts and 150 GPT-5-generated QA pairs. The generator must be Llama3.2-1b and the embedding model jina-embeddings-v2-base-en, and you report recall@1, recall@5, and exact match. Code is 45% of the grade and the report 55%; the report analyzes how prompts, data format, document order, and counterfactual information change the results.

NTHU NLP Course Summary and LLM Reasoning Notes: A Three-Column Map of the Semester, Then Denny Zhou's Question of Whether Pretrained Models Can Reason

In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.

NTHU NLP Term Project and the Fall 2026 Redesign: A 30% Group Project Becomes an In-Person W14 Midterm, Plus Reasoning/Agentic AI, an AI-TA and TAICA Compute

Fall 2025 was graded 70% assignments + 30% term project. Projects were done in groups of 3–4 and split into Proposal 6%, Progress 6%, Poster 6% and Report 12%, with no GPUs provided. The repo has no project spec, only the syllabus structure, an end-of-term reminder and the W15–W16 recordings. Fall 2026 switches to 75% assignments (4 of them) + a 25% in-person midterm in W14. The schedule drops the presentation weeks, adds a Reasoning/Agent unit, and brings in an AI-TA for grading support and TAICA compute credits. The official materials don't say why.