Hung-Yu Kao's Fall 2025 W8 slides walk from GPT-1 to GPT-3, explain how the Sparse Transformer behind GPT-3 cuts attention cost, and then use InstructGPT to show the gap between continuing text and following instructions. The maximum likelihood objective can't tell a fabricated fact from a slightly wrong synonym, so three extra stages are added: SFT learns how humans write, a reward model learns how humans grade, and PPO optimizes against that grade while a KL penalty keeps the model from drifting too far. The lecture closes with Llama-2: separate safety and helpfulness reward models, context distillation, and GQA for faster inference.
Week 1 of Hung-Yu Kao's NLP course at NTHU is a 91-page deck, W1_NLP_brief. It opens with 'Watch for kids', five readings of the telescope sentence, and a Chinese tongue-twister about eleven uncles to show why language is hard. Then, from an information-retrieval angle, it builds one pipeline: inverted index, tokenization, stemming, TF-IDF, BM25. The second half hits that pipeline's dead ends (synonyms, polysemy, vocabulary mismatch), moves to SVD-based LSA, and closes with a preview of dense vectors through Skip-gram, GloVe, and FastText.
Hung-Yu Kao's Natural Language Processing at National Tsing Hua University is a graduate-level flagship course in the TAICA alliance. The syllabus caps it at 1,200 students, it is taught in Mandarin, and it runs from TF-IDF and word vectors to RLHF, PEFT, and RAG. For Fall 2025, the slides, 32 class recordings, and 4 assignments with starter notebooks are all on GitHub, which rates A3. Solutions, grading, and the term-project spec are not public. Fall 2026 is in progress and only goes up to W3, so it rates A2. Grading changed to 75% assignments plus a 25% in-person midterm, and a Reasoning/Agent unit was added.
Hung-Yu Kao's Fall 2025 PEFT slides open with a budget: full fine-tuning of Llama 2-7B in 16-bit needs about 56GB of GPU memory, while training only 0.2M parameters brings it down to about 17GB, because gradients and optimizer states nearly vanish. Intrinsic dimensionality then explains why tuning a small slice is enough: the longer a model is pretrained and the larger it is, the fewer effective dimensions fine-tuning needs. Methods fall into additive (Adapters, Prompt Tuning), selective (BitFit), reparametrization (LoRA), and hybrid (MAM Adapters, S4). The second half runs from GPT-2's task descriptions and verbalizers to the trade-offs between prefix tuning and soft prompt tuning.
An LLM will confidently answer that Oppenheimer was born in 1967 (the year he died). RAG retrieves first and generates second. The first 60 pages of Hung-Yu Kao's Fall 2025 RAG deck are all about finding the right material. Sparse vectors (bag-of-words, TF-IDF, BM25) are cheap and dependable; dense vectors catch paraphrases. BERT's [CLS] isn't a good sentence vector as-is, hence Sentence-BERT pooling and bi-encoders. A cross-encoder is accurate but needs nearly 50 million passes for 10,000 sentences, while a bi-encoder needs 20,000. SimCSE uses dropout as data augmentation, DPR beats BM25 with only 1,000 training examples, and GTR shows that scaling up a dual encoder improves out-of-domain retrieval.
In week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.
Fall 2025 was graded 70% assignments + 30% term project. Projects were done in groups of 3–4 and split into Proposal 6%, Progress 6%, Poster 6% and Report 12%, with no GPUs provided. The repo has no project spec, only the syllabus structure, an end-of-term reminder and the W15–W16 recordings. Fall 2026 switches to 75% assignments (4 of them) + a 25% in-person midterm in W14. The schedule drops the presentation weeks, adds a Reasoning/Agent unit, and brings in an AI-TA for grading support and TAICA compute credits. The official materials don't say why.
Week 2 of Hung-Yu Kao's NLP course at NTHU is a 62-page deck about predicting the next word. The first part covers statistical language models with a bigram count table, add-one smoothing, and perplexity, then sparse vectors through a PPMI example with cherry and digital. The middle returns to Word2Vec's negative sampling and uses the Chinese word for 'apple' to show why contextualized embeddings are needed. The last part derives RNNs from three weaknesses of feedforward networks and shows how they handle NER, sentence classification, and stacked and bidirectional variants.
Most public AI courses in Taiwan outside NTU come from TAICA, an alliance set up by the Ministry of Education. Each semester's course list says where every flagship course streams, and courses that stream on YouTube are usually watchable by anyone. Two courses are complete enough to self-study: Hung-Yu Kao's Natural Language Processing at NTHU (Fall 2025) and Yen-Lung Tsai's Generative AI at NCCU (Spring 2025), both A3. Wei-Ta Chu's Introduction to AI at NCKU, Ping-Hsuan Han's Human-AI Interaction at NTUT, and Min-Chun Hu's Robotic Navigation and Exploration at NTHU have full recordings but keep assignments on NTU COOL, so they rate A2. NYCU's TAICA courses are taught in English; the Deep Learning recordings are not publicly listed and Physical AI has just started, so neither made the main table.