Skip to content

Reading NTU ADL 2025 Fall: Conversational AI and Tool Use — From LU/DST/Policy/NLG to LaMDA, WebGPT, and Toolformer

Sep 30, 20261 min
TL;DRDialogue systems split into chit-chat and task-oriented. Task-oriented systems were traditionally built from four modules: language understanding (LU) turns a sentence into domain, intent, and slots; dialogue state tracking (DST) accumulates the user's goal; the dialogue policy picks the next system action; and NLG turns that action back into a sentence. An LLM can act out all four steps by itself, but it cannot actually make the booking, so it needs external tools. LaMDA learns to call a search engine, calculator, and translator. BlenderBot 2.0 adds internet search and long-term memory. WebGPT learns to drive a browser from human demonstrations, a reward model, and PPO. Toolformer has the model generate and filter its own tool-use training data. The lecture ends with evaluation: automatic metrics, four kinds of human evaluation, and LLM-Eval. ADL Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

🌏 中文版

This post is based on the L13 videos in the playlist of NTU Applied Deep Learning (ADL), Fall 2025 (114-1, 2025/09/01–12/15), with the Fall 2024 slides filling in. It is post 16 of the Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall series. The previous two posts, Language Agents and Reasoning, were about how a model thinks and plans. This one goes back to an older problem, talking with people: what happened between modular task-oriented dialogue systems and LLMs that use tools on their own?

Official materials used:

#Video (2025)Chinese subtitle (translated)LengthMatching 2024 pages
13.1Learning to Converse and InteractHow machines learn to converse and interact15:022–37
13.2Tool Use in LLMs – LaMDAGoogle's dialogue model that made employees think it was conscious!?13:1538–49
13.3Tool Use in LLMs – BlenderBotA dialogue model that remembers past interactions and external knowledge9:4950–58
13.4Tool Use in LLMs – WebGPTGPT with search engine abilities7:1459–68
13.5ToolformerGenerating training data that teaches GPT to use tools7:4369–72
13.6Plan-and-ExecutePlan a strategy, then execute15:11none found
13.7User InteractionInteracting with users beats working alone15:11none found
13.8Theory-of-MindUnderstanding what users are thinking12:47none found
13.9Conversation EvaluationJudging how good a dialogue system is13:0799–106

The "matching pages" column is my own topic match. Public information cannot confirm that the videos actually show these pages.

Two branches of dialogue systems

Page 3 sorts why people want dialogue systems into four sentences. "I want to chat" is social chit-chat, the Turing-test kind of human-likeness. "I have a question" is information lookup. "I need to get this done" is task completion, such as booking a train from Kaohsiung to Taipei or a table at Din Tai Fung for five at 7 PM tonight. "What should I do?" is decision support. Page 4 folds these into two branches: chit-chat and task-oriented.

Task-oriented dialogue: four modules

The architecture on page 5 cites Young (2000). After speech recognition, the input passes through LU → DST → Dialogue Policy → NLG, then speech synthesis. The whole deck runs on one example: the user says "Can you help me book a 5-star hotel on Sunday?" and the system replies "For how many people?"

Language understanding, LU (pages 10–17) has three steps: identify the domain (hotel), detect the intent (Hotel_Book), then fill slots, using B-/O tags to mark fields such as star=5 and day=sunday. All three need a predefined ontology or schema. Pages 15–16 cover joint models that handle intent and slots together, including the Slot-Gated model of Goo et al. (2018). Evaluation (page 17) has two levels: domain/intent accuracy and slot F1, plus frame accuracy, which checks whether the whole frame is right.

Dialogue state tracking, DST (pages 18–22) accumulates the conversation into a state. When the user adds "For two people, thanks!" in the second turn, the state grows from Hotel_Book(star=5, day=sunday) to include people_num=2. Pages 20–21 deal with slot values that are not in a fixed list (Xu & Hu 2018, TripPy). Evaluation uses slot accuracy and joint accuracy.

Dialogue policy (pages 23–29) decides the next system action from the state, such as request(people_num) or inform(hotel_name=B&B). Page 25 contrasts two ways to learn it. Supervised learning is "learning from a teacher": the teacher says to answer Hello with Hi. Reinforcement learning is "learning from critics": you only learn at the end whether the whole dialogue went well. Pages 27–28 illustrate with the Deep Q-Network dialogue manager of Li et al. (2017) and the end-to-end TC-Bot. Evaluation works at the turn level (system action accuracy) and the dialogue level (task success rate, reward).

NLG (pages 30–33) turns inform(name=B&B) back into a sentence like "I have book a hotel B&B for you." Page 32 fine-tunes a pre-trained GPT-2 for conditional generation, because pre-trained models write more fluent sentences. Evaluation uses automatic metrics and human judgment.

Try this: pick a booking or customer-service chatbot you use. For one turn, fill in page 5's four boxes with a line each: which intent and slots it understood, what state it has accumulated, what system action comes next, and what sentence it finally wrote. The box you cannot fill is usually where it answers the wrong question.

The LLM plays both roles, but books nothing

Page 36 is titled with a Chinese phrase meaning roughly "directing and starring in its own show." A user asks an LLM to book a restaurant atop Taipei 101. The LLM naturally asks for the date and party size, says "let me check availability," and after a while reports there are no seats. It looks like task-oriented dialogue, but it is not connected to any booking system, so the availability result is made up. The page's conclusion: access to external tools is necessary. Page 37 redraws the four modules to show that one LLM can act out LU, DST, policy, and NLG. What it lacks is the step that reaches the outside world.

The next four models are four ways to add that step.

LaMDA: learning to look things up and correct itself (pages 38–49)

LaMDA (Thoppilan et al., 2022).

  • Pre-training: public dialogue data, 1.56T words according to the deck. The input is the conversation history; the output is the current utterance. Page 39's example asks for its opinion of a Jolin Tsai concert.
  • Quality and safety fine-tuning (page 40): one model both generates and discriminates. Training data is written as "context + RESPONSE + response + attribute name + rating," with attributes SENSIBLE, INTERESTING, and UNSAFE. The same model can then generate and score its own output.
  • Groundedness (pages 42–49): teach LaMDA to use a search engine to validate or fix its claims. The system has three roles. LaMDA-Base is the original pre-trained model. LaMDA-Research decides whether to use an external tool and how to phrase the query. The Tool Set holds the tools: a calculator ("135+7721" → "7856"), a translator, and an information retrieval system. In page 49's example, Base drafts "He is 31 years old right now," Research looks up Nadal's age, finds 35, and the response is rewritten to 35. The deck also notes that 40K dialog turns were labeled correct or incorrect to train the ranking.

Page 49's closing line: LaMDA already combined retrieval-augmented generation, tool use, and factual alignment in one system.

BlenderBot: search, memory, and safety (pages 50–58)

  • BlenderBot 1 (Roller et al., 2020, page 51): pre-trained on 1.5B conversations, in three sizes: 90M, 2.7B, and 9.4B. The fine-tuning data, Blended Skill Talk, mixes three skills: personality (PersonaChat), knowledge (Wizard of Wikipedia), and empathy (Empathetic Dialogues). Generation uses a retrieve-and-refine strategy.
  • BlenderBot 2.0 (pages 52–55): adds internet search and long-term memory, backed by the Wizard of the Internet and Multi-Session Chat datasets. For safety, it learns on the BAD dataset to emit a _POTENTIALLY_UNSAFE_ token after an unsafe response.
  • BlenderBot 3.0 (Shuster et al., 2022, pages 56–58): two training techniques. SeeKeR generates a search query, then a knowledge sequence, then the final response. Director learns to avoid undesirable sequences; the deck lists contradiction, repetition, and toxicity. It also keeps improving by collecting feedback from real interactions.

WebGPT: learning to use a browser from people (pages 59–68)

The three steps of WebGPT (Nakano et al., 2021) resemble InstructGPT's:

  1. Supervised fine-tuning (page 60): questions come from ELI5, such as "Which has more words, the Harry Potter series or The Lord of the Rings?" Human demonstrators write answers with references, and GPT-3 is fine-tuned on them.
  2. Reward model (page 64): people rank two model outputs for the same question.
  3. Reinforcement learning with PPO (page 65): the reward model scores outputs, and the generation policy is updated.

Pages 62–63 use a Chinese example to show that generating a search query is itself just token continuation. Page 66 evaluates truthfulness on TruthfulQA. Page 68 lists WebGPT's action set and leaves a question: how can a model learn to use these actions without human demonstrations? That is the problem Toolformer takes on.

Toolformer: generating its own tool-use data (pages 69–72)

Toolformer (Schick et al., 2023) is subtitled "Language Models Can Teach Themselves to Use Tools." The deck walks through two steps with a Chinese example:

  1. Prompt the model to generate candidates (page 70): given the sentence "The district with the highest housing prices in Taipei is Da'an District," the model inserts a tool call: "The district with the highest housing prices in Taipei is [QA("Which Taipei district has the highest average price per unit?")]."
  2. Keep only verified data for fine-tuning (page 71): actually call the QA tool. If it returns "Da'an District," matching the rest of the sentence, the example is kept for fine-tuning. The deck uses this example only as an illustration; see the paper for the exact filtering rule.

Page 72 lists the tool set, QA, WikiSearch, Calculator, Calendar, and MT, on a task of completing a short statement with a missing fact such as a date or place. The conclusion on page 108 puts the two side by side: WebGPT learns from human steps; Toolformer learns from data it generates itself.

Pages 73–98 go on to the GPT Store, ChatGPT Plugins, recommendation, and SalesBot (steering chit-chat naturally into a task-oriented dialogue). None of these match a 2025 video title, so I do not expand on them.

Three new 2025 videos: names only

The three videos below have no matching pages in the 2024 deck, and Fall 2025 published no slides. This post gives only their titles and subtitles:

  • 13.6 Plan-and-Execute: plan a strategy, then execute
  • 13.7 User Interaction: interacting with users beats working alone
  • 13.8 Theory-of-Mind: understanding what users are thinking

The planning thread connects back to the planning section of the deck in post 14, Language Agents. For a fuller treatment, see CMU 11-768 under Further reading.

How to evaluate dialogue (pages 99–106)

Automatic evaluation (page 100) compares the model's response with a gold response. The deck lists four measures: perplexity (how likely the model is to generate the gold response), n-gram overlap (BLEU and similar), slot error rate (whether the required slots are mentioned), and distinct n-grams (response diversity).

Human evaluation comes in four forms (pages 101–104), varying along two axes: score one model or compare two, and judge a single response or a whole conversation.

Single responseWhole conversation
Rate 0–5LikertDynamic Likert
Pick one of twoA/BA/B Dynamic

All four judge humanness, fluency, and coherence. The two dynamic forms cite ACUTE-EVAL (Li et al., 2019), and A/B Dynamic amounts to dialogue-level evaluation.

LLM-Eval (pages 105–106) comes from Prof. Chen's own lab (Lin & Chen, 2023). The deck says LLMs are reasonably capable of judging dialogue responses, that LLM-Eval works well on both single-turn and multi-turn evaluation, and that it correlates better with human scores than all existing metrics. The conclusion: LLM-Eval scores can serve as a proxy for human evaluation.

Try this: next time you compare two chatbots, don't judge from one reply. Follow A/B Dynamic: take the same task, hold a full conversation with each, then pick a side on humanness, fluency, and coherence. Of the four human evaluations, it is the closest to real use.

What this post can and cannot confirm

Confirmed: the nine videos' titles, Chinese subtitles, lengths, and descriptions (checked with YouTube oEmbed and yt-dlp); the titles and bullets of all 108 pages of the Fall 2024 ConvAI deck; the arXiv titles of the papers cited above.

Not confirmed: whether the Fall 2025 videos use this 2024 deck, or how much it changed. The content of videos 13.6–13.8. I did not transcribe the videos, so the lecturer's spoken asides and examples are not included. Figures such as 1.56T words and 1.5B conversations are restated from the deck; check the original papers for exact numbers.

Further reading

Series navigation: Series overview | Previous: Reasoning | Next: Beyond Supervised Learning and Multimodality

References