Skip to content

NTU ADL 2025 TA Recitations: From PyTorch and Hugging Face to LoRA, Quantization, and vLLM Deployment

Sep 30, 20261 min
TL;DRThe ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.

🌏 中文版

This guide is based on ADL Fall 2025 (114-1, 2025/09/01–12/15). It is post 18, the last one, in the Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall series. The lecture posts only link to the recitations; the recitation content lives here.

The ADL Fall 2025 course page splits each week into a Lecture column and a Recitation column. Lectures cover the ideas. Recitations cover how to actually get things running. Page 6 of the Course Logistics slides lists seven recitation topics: dev infra and tooling (Colab, GPU, PyTorch), the DL workflow, Hugging Face basics, LLM architecture, LLM evaluation, LLM training, and LLM inference.

This post answers one question: the lectures teach the principles, so what hands-on skills do the recitations add? You will come away knowing the order of the seven sessions, which tools each one teaches, which files still open, and how they line up in time with the three homework assignments.

Sources: the Recitation column of the course page, the ten recitation videos (YouTube titles and descriptions), five slide PDFs (same-name files on the Fall 2024 path), two TA Colab notebooks, and YouTube's auto-generated captions for the Deployment video. I opened and checked all of them on 2026-09-30. The lectures are taught in Mandarin; the slides mix Mandarin and English.

The big picture

WeekRecitationVideos (length)SlidesLectures that week
9/01Dev Infra & ToolingPyTorch (17:45), Debugging (11:46)None; a Colab insteadSequence Modeling
9/08NLP LifecycleStep by Step (36:10)w2-ProjLife.pdf (46 pages) + ColabAttention, Transformer, Tokenization, BERT; HW1 released
9/15Underlying Logics of ProjectsStep by Step (24:00)w3-UnderlyLogic.pdf (43 pages)Pretraining & Prompt Learning
9/22LLM LoRA TrainingLoRA (18:15)w5-LoRA.pdf (30 pages)Post-Training, LLM Adaptation; HW2 released
10/13LLM Basics & MoEBasics (20:16), Architecture (11:15), MoE (23:17)w4-LLMBasicsMOE.pdf (33 pages)RAG; HW3 released
10/27LLM Inference & EvaluationInfer & Eval (17:00)w6-LLMInferenceEval.pdf (37 pages)NLG Decoding, NLG Evaluation
11/03LLM DeploymentDeployment (22:34)NoneIssues in Pre-Trained Models; final project announced

Two notes on ordering:

  • The week numbers in the file names (w2–w6) come from an older term. Fall 2025 puts LoRA (w5) on 9/22, ahead of LLM Basics & MoE (w4) on 10/13. This post follows the Fall 2025 course page.
  • 9/29 and 10/06 were holidays (Teacher's Day and Mid-Autumn Festival), and 10/20 was the midterm break, so those weeks had no recitation. The schedule lists no recitations after 11/10.

Before you start: the videos and slides come from older terms

This is the most important context for the whole track. Every one of the ten video descriptions says "Lectured by Yen-Ting Lin 林彥廷 @ NTU CSIE", and the recording dates fall into two batches:

  • Fall 2023: PyTorch and Debugging (2023/09/07), NLP Lifecycle (2023/09/21), Underlying Logic (2023/10/05), LoRA (2023/11/16), Inference & Evaluation (2023/11/30)
  • Fall 2024: LLM Basics, Architecture, MoE (2024/10/09), and Deployment (2024/12/04)

The slides match. The w2 cover says "ADL 2023 Fall - Recitation 2", the w4 cover says "Sep 30, 2024", and the w5 and w6 covers are dated November 16 and 30, 2023. Fall 2025 reused these materials without re-recording.

That has two practical consequences. First, any homework mentioned in the slides is that year's homework. Page 4 of w5 says "Homework 3 Instruction tuning", which is the 2023 HW3, not the Fall 2025 HW3 (RAG). Second, the library versions in the demos date from 2023–2024, so expect to fix some API calls if you run them. The sections below point out where.

1. Dev Infra & Tooling: Colab, tensors, and ipdb (9/01)

This session has no slides. The material is a 30-cell Colab notebook titled "Recitation on Development Infrastructure and PyTorch Tutorial", in four parts:

  1. Google Colab basics: mount Google Drive, list your Drive files, check the PyTorch version.
  2. The GPU environment in Colab: confirm you have a GPU with torch.cuda.is_available().
  3. Intro to PyTorch: 1D tensors (arange, rand, randn, ones, zeros, from a list, clone), requires_grad and torch.no_grad(), shapes of multi-dimensional tensors, and conversion to and from NumPy.
  4. Python debugging tools: mostly ipdb.

The debugging section sorts errors bluntly: send syntax errors to ChatGPT, use ipdb for runtime and tensor errors, and leave GPU errors for the next recitation. The notebook explains six ipdb commands in Mandarin (n, c, q, p, l, s) plus a few usage patterns: python -m ipdb -c continue, a conditional ipdb.set_trace(), and ipdb.launch_ipdb_on_exception(). It ends with a SimpleMLP regression example to practice on.

Treat that example as a debugging exercise. The model output has shape (batch, 1) while the label y has shape (batch,). Set a breakpoint before loss = criterion(outputs, batch_y) and run p outputs.shape to compare the two. That is exactly the kind of tensor error the session wants you to catch.

The two videos are PyTorch Tutorial and PyTorch Debugging. If you have never written PyTorch, do this session before or alongside post 2: neural networks and backpropagation.

2. The life of an NLP project: data → model → demo (9/08)

w2-ProjLife.pdf opens with two goals: the workflow from data through model development to a demo, and the common tools and libraries. The whole deck keeps returning to a three-box diagram: data → model → demo.

Data has three steps: collection, cleaning and validation, and labeling. The listed collection sources are web crawling, customer data, self-generated data, and GPT-4.

Model starts from "what can NLP do?" and narrows tasks down to a few shapes:

  • Classify a whole sentence: sentiment analysis, spam detection, intent detection
  • Classify each word: part-of-speech tags, named entities
  • Generate text: auto-replies, fill in the blank
  • Extract an answer from text: given a question and a context, find the answer span

Splitting the data (pages 31–38) is the page worth remembering. You may look at the training set. You may not look at the validation set; it checks the model during training. You must never look at the test set; it is only for evaluation after training. The training set is the largest, and validation and test sets are each about 10–30%.

The hands-on part uses intent-classification data from an old assignment (the slides link to data/intent from the 2021 ADL HW1 on GitHub) and only does whole-sentence classification. For tools, the slides say to use Hugging Face for everything, with Gradio for the demo; the example site is twllm.com.

The companion NLP Lifecycle Colab (68 cells) adapts Hugging Face's "Fine-tuning a model on a text classification task" example:

  1. Load data with load_dataset("yentinglin/ntu_adl_recitation"), rename the intent column to label, and encode it as classes
  2. Preprocess with the distilbert-base-uncased AutoTokenizer, applied to every split via dataset.map(..., batched=True)
  3. Train AutoModelForSequenceClassification with Trainer: learning rate 2e-5, 5 epochs, evaluation every epoch, best model picked by accuracy
  4. Upload with trainer.push_to_hub(); the last cell is a Gradio demo template

The notebook targets older library versions, so a few spots may need changes today: load_metric in datasets (newer setups use the separate evaluate library) and evaluation_strategy in TrainingArguments (renamed eval_strategy in newer releases). The final Gradio cell is only a template. The model name literally reads "你的模型" ("your model"), fn=... is left for you to fill in, and the pipeline task needs to match your own task.

3. The underlying logic of NLP projects: data prep for four task types (9/15)

The cover of w3-UnderlyLogic.pdf calls it the third hands-on session of Applied Deep Learning. Its goals are to prepare data, train, and predict for four task types, using the Hugging Face ecosystem. The previous session only did sentence classification; this one fills in the rest:

TaskSlide pagesFocus
Sentence classificationRecapLast session's approach
Token classification10–25Find data on Hugging Face Datasets, read the per-token label fields, train, build a Gradio demo
Text generation (summarization)33–36Data format and training
Extractive QA38–42Data format and training

Most of the "how do we train?" pages are code screenshots with no extractable text, so the details are in the video.

The extractive QA part is especially useful for Fall 2025 readers. The span selection step in HW1: Chinese extractive QA has exactly this task shape. Watch this session before you read the HW1 spec.

4. LLM LoRA Training: tuning a model that is too big (9/22)

w5-LoRA.pdf draws LLM development in three stages (pretraining → instruction tuning → learning from feedback) and places LoRA/QLoRA in the instruction-tuning stage. The deck follows a memory storyline:

  1. Can you even run it? Pages 5–6 point to the Hugging Face Space "Can it run LLM" to estimate how much memory a model needs.
  2. Why low rank is enough: language models have low intrinsic dimensionality, which is why they fine-tune well on little data. Two pages then review matrix rank.
  3. LoRA itself: the original weight W is d×d; add two matrices A (d×r) and B (r×d), with r usually between 1 and 32. The listed benefits: less memory, no gradients for the pretrained weights, and plug-and-play LoRA weights.
  4. Choices: which weight matrices to adapt, what rank to use, and how LoRA compares with other PEFT methods on GPT-3, all taken from the LoRA paper.
  5. LoRA's limit: the pretrained weights still take a lot of memory.
  6. QLoRA: floating-point formats (including FP8) first, then how QLoRA quantizes the frozen weights and trains LoRA on top.

After the last page, "How to use (Q)Lora?", the hands-on demo is not in the PDF. Watch the video.

In timing, this session shares a week with the LLM Adaptation lecture (7.5) and the HW2 release. The Fall 2025 HW2 topic is "LLM Tuning and Prompt Tuning for Classical Chinese Translation", but the spec is not public, so we cannot confirm that this session is the HW2 recipe. For the lecture side of LoRA, see post 10: PEFT + HW2.

5. LLM basics, Transformer architecture, and MoE (10/13)

w4-LLMBasicsMOE.pdf pairs with three videos and has three parts.

LLM Basics (video, pages 2–14) is about scale: scaling laws, emergent abilities ("an ability is emergent if it is not present in smaller models but is present in larger models"), and the compute estimate "training FLOPs ≈ 6 × model size × number of tokens", with Taiwan-LLM and GPT-4 as examples. Next comes the Chinchilla question: with a fixed compute budget, how should you split it between model size and training tokens? Figures from the Llama 3 paper then raise whether pretraining loss predicts downstream performance.

Transformer Architecture (video, pages 15–16) is a single checklist: RMSNorm, rotary positional encoding, KV-cache, grouped-query attention, SwiGLU. Next to "Vanilla Transformers vs LLaMA" the slide says to watch the previous year's recitation video, so the details live in the videos.

MoE (video, pages 17–33) takes up most of the deck:

  • A table of open-weight MoE models (Mixtral, Grok-1, DBRX, Arctic) and a dense-vs-MoE comparison
  • Shared experts (from DeepSeekMoE) and MoE in the attention layer (from JetMoE)
  • Token-choice routing: when too many tokens pick one expert, some get dropped (token dropping), which is why a load-balancing loss is needed (from Switch Transformers; page 27 has a worked example)
  • Problems with token-choice routing: load imbalance, an unstable balancing loss, and every token getting the same compute
  • Two responses: expert-choice routing, and Mixture-of-Depths, which lets tokens skip layers

6. LLM Inference & Evaluation: quantization and three kinds of evaluation (10/27)

The first half of w6-LLMInferenceEval.pdf covers inference speedups; the second half covers evaluation.

The speedup outline lists five items: quantization, AWQ, GPTQ, PagedAttention, and FlashAttention. Only quantization gets expanded; the last two appear only on the outline page.

  • Post-training quantization has two paths: AWQ and GPTQ on GPU, GGUF/GGML on CPU.
  • Pages 7–11 work through quantizing to 2 bits by hand to contrast round-to-nearest with scaling before quantizing.
  • AWQ: pick the scaling factor that minimizes activation error; it needs calibration data.
  • GPTQ: compress layer by layer, minimizing each layer's reconstruction loss; it also needs data.

Evaluation comes in three kinds:

KindExamples in the slides
Traditional benchmarksMMLU, TruthfulQA
Model as judgeMT-Bench, AlpacaEval
Human evaluationChatbot Arena

The same week's lectures cover NLG metrics (BLEU, ROUGE, perplexity, LLM-Eval); see post 12: NLG decoding and evaluation. Read together, one side asks whether a generated sentence is good, and the other asks how capable the whole model is.

7. LLM Deployment: serving a model with vLLM (11/03)

This session has no slides, only the video LLM Deployment. The summary below comes from YouTube's auto-generated Mandarin captions, which contain recognition errors, so I only keep the main thread.

The video uses vLLM to deploy Llama 3 at 8B and 70B on two 80GB H100s. One card is plenty for 8B. At BF16, 70B needs 2 bytes per parameter and exceeds one card, so it serves as the parallelism demo. In order:

  1. Install and parameters: on a recent NVIDIA GPU, installation is straightforward. The speaker only tunes three things: dtype (older cards without BF16 need FP16, the first common pitfall), enforce_eager (turning off CUDA Graph saves some memory at some speed cost), and quantization. His advice: when memory runs short, quantize to FP8 or AWQ instead of fiddling with many parameters.
  2. Offline vs online: offline means batch inference where all prompts are known in advance, common in research and data processing, with higher overall throughput. Online means something like twllm.com, where you don't know when a user will arrive or what they will type.
  3. Sampling pitfalls: the default max_tokens is small, so always raise it. Temperature, top-p, top-k, and min-p are also available.
  4. Check the output first: run one or two prompts offline and confirm the output makes sense. Otherwise a wrong tokenizer, model, or float type can quietly hurt results.
  5. Multiple GPUs: inference usually uses tensor parallelism (splitting matrices across cards and merging results via communication). Pipeline parallelism is usually reserved for multi-node setups when one machine is not enough. As a user, you just set the tensor parallel size to your GPU count.
  6. Online serving: start a server with vllm serve. A web server in front accepts HTTP requests; the LLM engine sits behind it. Clients keep using the OpenAI library and only change the API base to your own address, which is what an OpenAI-compatible API means. The speaker also notes that online serving trades off latency against throughput.

This session does not map to a homework. It is more like the last mile you need for the final project and for real work.

How recitations line up with homework

The course page does not say which recitation supports which assignment. The TA table only lists duties: two TAs for HW1/HW2, two for HW3, and two for the final project. The table below lines them up by week. It is a timing alignment, not an official mapping:

AssignmentRelease weekRecitations in or before that weekPublic material
HW1 Chinese extractive QA9/08Dev Infra, NLP Lifecycle; the following week's Underlying Logic covers extractive QA dataFull spec
HW2 LLM tuning + prompt tuning (Classical Chinese translation)9/22LLM LoRA TrainingIntro video only
HW3 Retriever & Reranker Training for RAG10/13LLM Basics & MoEIntro video only
Final project (Jailbreaking Olympics)11/03LLM DeploymentIntro video only

Access and gaps

The series as a whole is graded A2 (per the definition in the global AI/CS course map: materials partially open). On the recitation track alone, the videos and Colabs are complete, but the slides need a detour:

  1. All five slide links on the course page return 404. The course page points to paths such as f114-adl/doc/w2-ProjLife.pdf, and all returned 404 on 2026-09-30. This post cites the same-name files under the Fall 2024 path (f113-adl/doc/); all five open.
  2. Deployment has no slides, only the video.
  3. Dev Infra has no slides, only the Colab.
  4. The videos and slides are 2023–2024 versions, so the homework numbers and contents they mention belong to those years.
  5. Most code in the slides is screenshots; watch the videos alongside.

How to use this post

  • Never written PyTorch: run the Dev Infra Colab first, then read post 1 and post 2.
  • Getting ready for HW1: after post 6 on BERT, watch NLP Lifecycle and Underlying Logic in that order, then move to HW1.
  • Only want LLM engineering: jump to the last four sessions. You can reorder them as Basics & MoE → LoRA → Inference & Eval → Deployment, which matches the older w4→w5→w6 file order.

One thing to do tonight: open the NLP Lifecycle Colab, change evaluation_strategy to eval_strategy, replace load_metric with evaluate.load, and run through trainer.train(). If it finishes, your environment is ready for HW1.

Further reading


Previous: Post 17: beyond supervised learning and multimodality Series overview: Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall

References