🌏 中文版
This post is based on the L12 videos in the playlist of NTU Applied Deep Learning (ADL), Fall 2025 (114-1, 2025/09/01–12/15), taught by Yun-Nung (Vivian) Chen. It is post 15 of the Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall series. The previous post, Language Agents, treated reasoning as one of three key concepts for agents. This one pulls it out on its own: how does a model learn to think before it answers?
First, the limit of this post: L12 has no public slides. The 12/01 row on the course page just says "Reasoning," with no slides and no video links. The five videos appear only in the 2025 Fall playlist. I did not transcribe the videos, so all I can report are their titles, lengths, and descriptions, plus the reasoning pages in the previous lecture's deck.
The five videos
The lectures are in Mandarin; each title carries a Chinese subtitle, translated here.
| # | Video | Chinese subtitle (translated) | Length |
|---|---|---|---|
| 12.1 | What is Reasoning? | Can machines reason too? | 10:21 |
| 12.2 | Short CoT | Reason briefly, then answer | 12:06 |
| 12.3 | Test-Time Scaling | Thinking more during the exam helps | 29:17 |
| 12.4 | Learning to Reason | Imitating how others reason | 18:25 |
| 12.5 | RL for Reasoning | Evolving reasoning behavior through exploration | 12:14 |
All five descriptions read "2025/11/17 Applied Deep Learning" and note "Slides credited from Hung-Yi Lee," meaning the slides were borrowed from Prof. Hung-yi Lee. One thing does not line up. The course page labels 11/17 "Knowledge, Multimodality" and puts "Reasoning" on 12/01, but the video descriptions give 11/17 as the lecture date. Public information cannot settle which is right, so this post records both.
The route the titles lay out
The five titles already form a path from basic to advanced. This section only restates what the titles and subtitles say. It adds nothing I could not verify from the videos.
1. What is reasoning (12.1). The subtitle is a question: "Can machines reason too?" This video defines the problem; the next four cover methods.
2. Short CoT (12.2). "Reason briefly, then answer" makes two points: reasoning comes before the answer, and it is short. That sets up a contrast with the "think more" direction from 12.3 on.
3. Test-Time Scaling (12.3). "Thinking more during the exam helps." The trained model stays fixed, and more compute is spent at inference. At 29 minutes it is the longest of the five and carries the most weight.
4. Learning to Reason (12.4). "Imitating how others reason": the model learns from existing reasoning traces. This is the turn from "how to use it at inference" to "how to teach it in training."
5. RL for Reasoning (12.5). "Evolving reasoning behavior through exploration": instead of only imitating, the model uses reinforcement learning to discover its own ways of reasoning.
So the whole arc is: define → let it think at inference (short, then long) → teach it to think in training (imitate first, then explore).
Try this: before watching, copy down the five titles. After each video, write one sentence next to it saying which question it answered. Together the five sentences summarize the lecture, and they show whether you caught the inference-to-training turn between 12.3 and 12.4.
Reasoning in the previous lecture's deck
L12 has no slides, but pages 8–20 of the Language Agents deck (251110_LangAgent.pdf) are entirely about reasoning. They are the closest official Fall 2025 material. Page 1 says the deck draws on the EMNLP 2024 tutorial on language agents.
Reasoning as an internal action (pages 5, 9). The deck draws a language agent as: perceive the environment → reason in an inner monologue → act on the environment. Page 5 says that generating tokens to reason can be viewed as an internal action; self-reflection is a "meta" action that reasons over the reasoning process; and reasoning exists to act better. Page 9 splits the action space three ways: reasoning updates short-term memory (the context window), retrieval/learning reads and writes long-term memory, and planning chooses an external action at inference time.
CoT (page 10). Cites Wei et al., 2022, in one line: intermediate generation imitates human mental processes.
Reasoning helps acting, and acting helps reasoning (pages 11–12). Page 12 uses a Chinese example: asked "Do you know NTU's Yun-Nung Chen?", the model searches first and then answers from the results. The search gives the reasoning something to stand on.
ReAct (pages 13–17). Cites Yao et al., 2022. The deck's takeaways are "Reasoning + Act are both essential" and "reasoning provides explanations for controlling actions."
Reasoning enlarges the action space (pages 18–19). A larger action space means more capacity but harder decisions, since the space of reasoning and language is infinite. The last point on page 19 echoes the 12.4 subtitle directly: LLMs learn reasoning priors by imitating many human reasoning traces.
Page 20 is titled "Action Planning for Improving Reasoning (Yao et al, 2023)" and has no further text. I do not guess which paper it refers to.
What this post can and cannot confirm
Confirmed: the five videos' titles, Chinese subtitles, lengths, upload date (2025-11-20), and description text (checked with YouTube oEmbed and yt-dlp); the titles and bullets of pages 5–20 of the Language Agents deck; the arXiv titles of the CoT and ReAct papers.
Not confirmed: the content of the videos themselves. Without slides or a transcript, I do not say which test-time scaling methods 12.3 covers, what data 12.4 imitates, or which RL algorithm or model 12.5 uses. The descriptions credit the slides to Prof. Hung-yi Lee, but which deck and which pages cannot be confirmed from public information.
Further reading
For detail beyond the video titles, two site series cover the same ground from courses with slides:
- CS224N Lecture 12: Decoding, DeepSeek-R1, and Reasoning Training: R1-Zero/R1, PPO, GRPO, DAPO, matching the RL route of 12.5.
- CS224N Lecture 13: Speculative Decoding and Test-Time Scaling: matching the test-time scaling of 12.3.
- CME295: LLM Reasoning, plus CME295: RL with LLMs on RL training.
- From the same university, the Hung-yi Lee Machine Learning 2026 Spring guide: all five L12 descriptions credit the slides to Prof. Lee, so his own course is a natural companion.
Series navigation: Series overview | Previous: Language Agents | Next: Conversational AI and Tool Use
References
- NTU Applied Deep Learning Fall 2025 course page — the 12/01 row carries only the title "Reasoning"
- 2025 Fall NTU CSIE ADL playlist (lectures in Mandarin)
- Videos: 12.1 What is Reasoning?, 12.2 Short CoT, 12.3 Test-Time Scaling, 12.4 Learning to Reason, 12.5 RL for Reasoning
- 251110_LangAgent.pdf (Language Agents, Fall 2025) — pages 5–20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (arXiv 2201.11903)
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv 2210.03629)
Loading...