Skip to content

NTHU NLP Course Summary and LLM Reasoning Notes: A Three-Column Map of the Semester, Then Denny Zhou's Question of Whether Pretrained Models Can Reason

Sep 30, 20261 min
TL;DRIn week 14 of Fall 2025, Hung-Yu Kao closed the course with two slide decks. Course_summary sorts the semester into three columns (NLP Fundamentals, NLP Models, NLP Advances), adds two columns for the TA labs, and lists five directions for further study. The second deck is his notes on Denny Zhou's (Google DeepMind) April 2025 Stanford talk. Its claim: pretrained models can already reason, and decoding is what brings it out. CoT decoding, self-consistency and retrieval + reasoning add up to four inequalities. The Fall 2026 W16 'Reasoning / Agent' unit has not been released yet.

🌏 中文版

This post is based on the official materials for NTHU Prof. Hung-Yu Kao's Natural Language Processing course, Fall 2025 (114-1). It is part 18 of the Reading NTHU Hung-Yu Kao Natural Language Processing series and follows the RAG labs and HW4.

Three official materials are used here, all listed in the W14 row of the 2025 schedule: Course_summary.pdf (7 pages), Note_from_Google_DeepMind's_Reasoning_Talk.pdf (13 pages), and the recording Week 14 Tue. (about 77 minutes). The Fall 2025 term is rated A3, enough for self-study. The rating scale is defined in the global AI/CS course map.

Two limits up front. First, the recording has no captions, and I did not watch it segment by segment. Everything below comes from the slides. Second, the reasoning deck is a secondhand summary: the professor turned Denny Zhou's talk into notes, and many of its figures are marked "Generated by NotebookLM." For the original claims, go to the talk itself (YouTube title: "Stanford CS25: V5 I Large Language Model Reasoning") and the papers.

One timeline: where the course started and where it ended

Page 2 of Course_summary is an NLP milestone timeline, the same figure used in the first-week Syllabus. On the left, around 2000, sit bag of words, vector space, TF-IDF and parsing trees. Next come Bengio's neural probabilistic language model and Google's word2vec in 2013. From 2018 the figure is labeled "Transformer Era": GPT, BERT, GPT-2, GPT-3, and finally InstructGPT, ChatGPT and LLMs in 2022–2024.

In week 1 this figure is a list of names. By the end of the term, each point maps to something you have read.

The three-column map, mapped back to this series

The lectures are grouped into three columns. The table follows the slide's original grouping; the right column points to the matching posts in this series:

Slide columnTopics listed on the slideThis series
NLP FundamentalsText representation: one-hot, TF-IDF, Word2VecPart 1, intro and classic text processing, Part 2, word embeddings and language models
NLP FundamentalsLanguage model: n-gramPart 2
NLP ModelsBasic models: RNN, LSTMPart 2, Part 4, seq2seq and attention
NLP ModelsCore models: seq2seq, attention, Transformer (self-attention, multi-head attention)Part 4, Part 6, Transformers
NLP ModelsPretrained language models: encoder/decoder, BERT, GPT, T5Part 8, the BERT family
NLP ModelsSubwordPart 7, sub-word tokenization
NLP AdvancesInstruct fine-tuningPart 12, GPT-3, InstructGPT and RLHF
NLP AdvancesPEFT: prompt tuning, LoRAPart 13, PEFT
NLP AdvancesRAG: sparse/dense vectors, dual encoder, RAAT, Self-RAGPart 14, RAG retrievers, Part 15, RAG QA and advanced methods

The TA labs get two more columns. NLP Basics covers Python (NumPy, Pandas, PyTorch) and Hugging Face (BERT, GPT-2, T5), which match Part 5, PyTorch + HW2, Part 9, HF BERT + HW3 and Part 11, GPT-2/T5 summarization. NLP Advances covers the LLM API and RAG with and without LangChain, which match Part 16, LLM API and Part 17, RAG lab + HW4.

One topic is missing from the slide: decoding strategies and NLG evaluation. That deck was taught mid-semester, and this series covers it in Part 10. It is also the closest relative of the reasoning notes below, because CoT decoding is about not settling for the greedy path.

What to do: Use this table as an end-of-term self-check. For each row, say in one sentence what problem it solves for the row above it. If you can't, reread that post.

Five directions for further study

The "Future study" page of Course_summary lists five areas. The slide uses bullet points; the gist is:

  • Advanced Applications of LLMs: the architecture and applications of MoE, GPT-4 and ChatGPT; fine-tuning pretrained models for specific domains.
  • Semantic Understanding and Logical Reasoning: strengthening reasoning with knowledge graphs and symbolic reasoning. After "new methods for improving language understanding and inference," the slide adds "test-time scaling" in parentheses.
  • Multimodal NLP: models that combine text and images, such as CLIP and BLIP, plus understanding and generation for video and audio (VLM, VLA).
  • Ethics and Societal Impact: model bias and mitigation, explainable AI (XAI), energy efficiency.
  • Cutting-edge Technologies: RL in NLP such as RLHF, zero-shot and few-shot learning, coherence and control in long-form generation, with "lost in the middle" in parentheses.

A few paths on this site pick up from here: for multimodal work, see the Stanford CS231N guide; for test-time scaling, see HW8 of Hung-yi Lee's ML 2026.

The reasoning notes: can pretrained models already reason?

Page 3 of the notes poses the question. The common belief is that a pretrained LLM cannot reason without prompt engineering or fine-tuning. Page 4 gives the talk's answer: pretrained LLMs are ready to reason; all we need is decoding. The cited source is Wang and Zhou's Chain-of-Thought Reasoning Without Prompting (NeurIPS 2024).

Same question, greedy vs. the other candidates

Page 5 uses a grade-school word problem: "I have 3 apples. My dad has 2 more apples than me. How many apples do we have in total?"

Greedy decoding picks the most likely first token, "5," and ends with the wrong answer, 5 apples. If the first token comes from another candidate, such as "I" or "You," the model writes out the reasoning ("he has 5 apples, 3+5=8") on its own and reaches the correct 8. The path that starts with "The" goes back to the wrong 5.

The point: the reasoning path is already in the model's output distribution. It just isn't the most likely one.

CoT decoding in two steps

Page 6 reduces the method to two steps:

  1. Look past the greedy path and check more generation candidates.
  2. Pick the candidate with the highest confidence in its final answer.

The figure draws the output space as terrain. The highest peak is "5 apples"; a lower path writes out "3+5=8," and its final answer carries 98% confidence. The figure is a NotebookLM illustration, not an experimental plot from the paper.

This extends the question from Part 10 on beam search and top-k/top-p. There the question was whether the generated text reads naturally. Here it is whether the answer is right.

Prompting shifts the output distribution

Page 7 puts two common CoT prompting styles side by side. Few-shot chain-of-thought prompting gives a worked example first, and the catch is that it needs task-specific examples. Adding "Let's think step by step" after the question is the generic option; the slide's verdict is generic, but worse than few-shot.

Read through the CoT decoding lens, both push the output distribution so that reasoning paths rise to the top.

SFT imitates one path; iterative fine-tuning explores many

Page 9 contrasts two kinds of fine-tuning. SFT trains on human-written solutions, and the figure describes it as imitating a single path. The other side is labeled "IO fine-tuning" on the slide, and its figure is a four-step loop:

  1. Generate: have the model produce several different solution paths for a batch of problems.
  2. Verify: use a simple verifier, such as checking the final answer, to keep the solutions that got it right.
  3. Fine-tune: train on these machine-generated "problem + correct solution" pairs.
  4. Repeat: go back to step 1 with the stronger model.

A red box on the same page adds a reminder: LLMs are probabilistic models trained to predict the next token. They are not humans. Page 8, titled "LLM reasoning vs. human reasoning," shows Gemini 2.0 thinking mode (December 2024) working through "use each of the numbers 1 to 10 once, with only + and *, to make 2025." The notes don't state a separate conclusion for it.

Self-consistency: ignore the steps, vote on the answer

Page 10 covers self-consistency, and the steps are blunt:

  1. Use random sampling to generate many reasoning paths (for example, 40).
  2. Ignore the intermediate reasoning entirely.
  3. Look only at the final answers and pick the most frequent one.

The figure shows five paths answering 8, 8, 5, 8 and 9; the vote returns 8. A NotebookLM bar chart in the corner shows PaLM on GSM8K going from 58% with single-path CoT to 75% with self-consistency, roughly a 30% relative gain. Treat the original paper, Self-Consistency Improves Chain of Thought Reasoning in Language Models, as the authority on these numbers; its abstract reports a 17.9% gain on GSM8K.

Retrieval or reasoning? Both

Page 11 answers "Retrieval + Reasoning." The example is finding the area of a quadrilateral. A retrieval-style prompt says "before solving, recall a related problem." The model recalls the distance formula, uses it to get each side, and solves the problem.

Retrieval here means having the model pull a related problem out of its own knowledge. It is not the external document search of Part 14. What the two share: get the right material in front of the model before it reasons.

Four inequalities

Page 12 condenses the talk into four lines:

  • Reasoning > no reasoning
  • RL fine-tuning > SFT
  • Aggregating multiple answers > one answer
  • Retrieval + reasoning > reasoning only

The accompanying figure draws reasoning ability in five stages. It is a latent ability in the pretrained model, elicited by decoding or prompting, hard-coded by SFT (though brittle), generalized by iterative fine-tuning, and amplified by aggregation and retrieval. The last page reads "Next-token prediction → emerging intelligence," with a Feynman line: the truth always turns out to be simpler than you thought.

What to do: Tonight, take any model API that lets you set temperature and run the apple problem from page 5 ten times. Record each final answer and take a majority vote. Then set temperature to 0, run it once, and compare. The notebook in Part 16, LLM API has calling code you can adapt.

The Fall 2026 reasoning unit isn't out yet

The Fall 2026 Syllabus-115 adds a new block, "NLP issues in LLM era: Reasoning, Agentic AI," and schedules W16 as "Reasoning / Agent." As of 2026-09-30, the schedule on the repo's front page is filled in only through W3, and the W16 slide and video cells are empty. The official materials don't say whether this 2025 deck is a precursor to the 2026 unit, and I won't guess. The rest of the redesign is covered in the next post.

Further reading

These posts on this site each go further; only links here:

What this post can and cannot confirm

Confirmed: the text and figures of both decks, the files and video ID in the W14 row, the video's title and length, and the titles and arXiv IDs of both papers. Not confirmed: what the professor actually said in the W14 recording, or whether he covered anything beyond the slides (no captions, not watched segment by segment); whether the numbers in the NotebookLM figures match the original papers exactly; the content of Fall 2026 W16.

Series navigation: previous, RAG labs + HW4 | next, term project and the Fall 2026 redesign | series overview

References