Skip to content
Series
19 posts

Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall

Reading NTU Yun-Nung (Vivian) Chen's Applied Deep Learning (ADL) Fall 2025 through its 17 lecture decks, 77-video playlist, TA recitations and the HW1 spec: neural nets, RNNs, Transformers, BERT, pretraining, RLHF, LoRA, RAG, decoding, safety and alignment, language agents, and reasoning.

Reading NTU Yun-Nung Chen's Applied Deep Learning 2025 Fall: Course Map, A2 Rating, and How to Read It

Applied Deep Learning (ADL) Fall 2025, taught by Yun-Nung (Vivian) Chen in NTU's CSIE department, is a deep learning course built around NLP. It runs from neural network basics through Transformers, BERT, pretraining and prompting, post-training, LoRA, RAG, generation and evaluation, alignment issues, and language agents. Lectures L0–L11 come with slide PDFs and segmented videos, and the playlist adds videos for L12–L14. On the assignment side only the HW1 spec is public; HW2, HW3, and the final project have explainer videos only. That makes it A2.

NTU ADL 2025 Lecture 1: What Machine Learning and Deep Learning Are

The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.

NTU ADL 2025 Lecture 2: Neural Networks and Backpropagation

ADL Fall 2025's NN Basics and Backpropagation decks break model training into three questions. What is the model? Layers of neurons, each computing z = Wa + b and then a nonlinearity. What makes a function good? A smaller loss. How do we pick the best one? Gradient descent, in practice mini-batch SGD. Backpropagation computes gradients for millions of parameters efficiently: the forward pass stores each layer's output, the backward pass sends an error signal δ back from the output layer, and multiplying the two gives each weight's gradient.

NTU ADL Lecture 3: Word Representations, Language Models, and RNNs

The first real lecture of ADL Fall 2025 has one through-line: a language model predicts the next word. It starts from one-hot vectors and co-occurrence matrices, moves through the zero-probability problem of n-grams and the smoothing that a neural LM gets for free, and ends at the RNN LM, which folds all previous words into a hidden state. BPTT and vanishing/exploding gradients are the training cost, and LSTM and GRU patch it with gating. The lecture closes by splitting applications into sequence input versus sequence output, which separates tagging from encoder-decoder models.

NTU ADL Lecture 4: Attention and the Transformer

An RNN translator has to squeeze the whole source sentence into one vector, and long sentences don't fit. Attention lets the decoder look back at every input position each time it produces a word, score each one, and take a weighted sum. Call the scorer the query, the thing being scored the key, and the thing being averaged the value, and you have dot-product attention. The Transformer goes one step further: the input attends to itself, recurrence disappears, and in return you get parallelism and a constant path length. The price is that position information has to be added back by hand.

NTU ADL Lecture 5: Tokenization and BPE, or Where the Vocabulary Comes From

With whole words as units, any unseen word becomes UNK. With single characters, meaning is hard to reassemble. This 22-page ADL deck explains the mainstream compromise, subwords, and the most common way to build them, BPE. The core is a tiny corpus of 4 words and 16 occurrences: start from characters, merge the most frequent adjacent pair each round, and after 9 merges you have units like newest</w> and low</w>, which then segment the unseen words lowest and powest. It ends with a GPT-3 tokenizer screenshot where the Chinese version of a sentence takes more than twice as many tokens as the English.

Reading NTU ADL 2025 Fall: BERT and Its Family — From the Polysemy Problem to XLNet, RoBERTa, and mBERT

A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.

Reading NTU ADL 2025 Fall: HW1 Chinese Extractive QA — Finding the Answer Span Among Four Paragraphs

HW1 in ADL Fall 2025 gives a question and four Chinese paragraphs. The model first picks the relevant paragraph (paragraph selection, framed as four-way multiple choice), then marks the answer's start and end inside it (span selection), scored by Exact Match. The spec slides point you straight at Hugging Face's run_swag_no_trainer.py and run_qa_no_trainer.py. The simple baseline uses bert-base-chinese, length 512, effective batch size 2, and learning rate 3e-5, and both stages together take under three hours on an 8GB RTX 3070. The Kaggle leaderboard closed 9/29, and code plus report were due 10/1 on NTU COOL. Outside readers cannot get the Kaggle data or grading, but the task design, baseline settings, and five report questions are all usable for practice.

Reading NTU ADL 2025 Fall: Three Pre-training Families and Prompt Learning — From BERT, GPT, and T5 to Prompts Only Machines Understand

Lecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.

NTU ADL 2025 Lecture 7: Post-Training — Instruction Tuning, RLHF, and InstructGPT

A pre-trained model can continue text, but that does not mean it follows instructions. The Post-Training slides of NTU ADL Fall 2025 fix this in two steps. Instruction tuning (FLAN, T0) teaches the model to read task descriptions. RLHF then pulls its outputs toward human preference. Three limits of instruction tuning connect the two steps, a reward model and pairwise comparisons solve two practical RL problems, and InstructGPT's SFT → reward model → PPO pipeline ties it all together. ChatGPT runs the same pipeline on multi-turn dialogue.

NTU ADL 2025 Lecture 7.5: PEFT — Adapter, LoRA, Prompt Tuning, and HW2

When an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.

NTU ADL 2025 Lecture 8: RAG — From Retrieval and Reranking to Search-R1, Plus HW3

LLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".

NTU ADL 2025 Lecture 9: Decoding, Generation Control, and Evaluation for NLG

Lecture 9 of ADL Fall 2025 answers two questions. The model gives you a probability distribution at every step, so how do you pick a word from it? And once you have a sentence, how do you judge it? The slides start with teacher forcing and exposure bias to show the gap between training and generation, then compare greedy, beam search, sampling, top-k, and nucleus sampling, and file temperature and the penalties under 'control' rather than decoding algorithms. The evaluation half covers BLEU, ROUGE, perplexity, and LLM-Eval, then explains why you would use RL to optimize whole-sentence quality directly.

NTU ADL 2025 Lecture 10: Bias, Safety, Hallucination, and Alignment, Plus the Jailbreaking Olympics Final Project

Lecture 10 of ADL Fall 2025 sorts the problems of pretrained models into four groups, each paired with a goal: bias with fairness, toxicity with safety, hallucination with factuality, and finally alignment. The slides argue that bias can enter at any stage of the ML pipeline, that safeguards belong at four layers (data, input, training, output), that hallucination can be checked atomic fact by atomic fact, and that over-optimizing a reward model produces familiar symptoms: verbosity, excessive apologies, over-refusal. The final project announced that week is called Jailbreaking Olympics, but all that is public is the titles and one-line descriptions of two videos.

NTU ADL 2025 Lecture 11: Reasoning, Memory, Planning, and Multi-Agent Systems in Language Agents

Lecture 11 of ADL Fall 2025 builds on the EMNLP 2024 Language Agents tutorial. It defines an agent as an entity that perceives and acts, then names what is new about language agents: reasoning itself counts as an internal action. The lecture is organized around three concepts. Reasoning covers CoT and ReAct; memory covers Generative Agents and its recency / importance / relevance retrieval; planning goes from greedy reactive planning to tree search and world models. It closes with multi-agent systems in three steps: initialization, orchestration, and team optimization.

Reading NTU ADL 2025 Fall: Reasoning — A Video-Only Lecture, Five Steps from CoT to RL

The Reasoning lecture of NTU ADL Fall 2025 has no public slides. It exists only as five videos in the course playlist: 12.1 What is Reasoning?, 12.2 Short CoT, 12.3 Test-Time Scaling, 12.4 Learning to Reason (imitating others), and 12.5 RL for Reasoning (evolving reasoning through exploration). This post lays out that route from the video titles alone, then pairs it with the CoT, ReAct, and 'reasoning enlarges the action space' pages of the previous Language Agents deck. Technical detail is left to the site's CS224N and CME295 reasoning posts.

Reading NTU ADL 2025 Fall: Conversational AI and Tool Use — From LU/DST/Policy/NLG to LaMDA, WebGPT, and Toolformer

Dialogue systems split into chit-chat and task-oriented. Task-oriented systems were traditionally built from four modules: language understanding (LU) turns a sentence into domain, intent, and slots; dialogue state tracking (DST) accumulates the user's goal; the dialogue policy picks the next system action; and NLG turns that action back into a sentence. An LLM can act out all four steps by itself, but it cannot actually make the booking, so it needs external tools. LaMDA learns to call a search engine, calculator, and translator. BlenderBot 2.0 adds internet search and long-term memory. WebGPT learns to drive a browser from human demonstrations, a reward model, and PPO. Toolformer has the model generate and filter its own tool-use training data. The lecture ends with evaluation: automatic metrics, four kinds of human evaluation, and LLM-Eval. ADL Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

Reading NTU ADL 2025 Fall: Beyond Supervised Learning and Multimodality — Auto-Encoders, VAE, Dual Learning, Contrastive Learning, and CLIP

Big data is not big annotated data. The last ADL lecture asks how to learn good representations without labels, and answers: find the latent factors that control the data. An auto-encoder squeezes the input into a short code and reconstructs it. The denoising version adds noise or masks 15% of tokens first, which is exactly the idea behind BERT's masked LM. A VAE forces the code to follow a distribution, so you can sample from it to generate. Dual learning lets paired tasks, such as translation and back-translation or understanding and generation, act as feedback for each other. Self-supervised learning has two camps: self-prediction (hide part, guess it back) and contrastive learning (pull similar pairs together, push dissimilar ones apart). CLIP runs contrastive learning on 400 million image-text pairs, making zero-shot image classification possible, and DALL·E 2 uses CLIP's representations to generate images. Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

NTU ADL 2025 TA Recitations: From PyTorch and Hugging Face to LoRA, Quantization, and vLLM Deployment

The ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.