Skip to content

NTU ADL 2025 Lecture 11: Reasoning, Memory, Planning, and Multi-Agent Systems in Language Agents

Sep 30, 20261 min
TL;DRLecture 11 of ADL Fall 2025 builds on the EMNLP 2024 Language Agents tutorial. It defines an agent as an entity that perceives and acts, then names what is new about language agents: reasoning itself counts as an internal action. The lecture is organized around three concepts. Reasoning covers CoT and ReAct; memory covers Generative Agents and its recency / importance / relevance retrieval; planning goes from greedy reactive planning to tree search and world models. It closes with multi-agent systems in three steps: initialization, orchestration, and team optimization.

🌏 中文版

This is post 14 of Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall. ADL Fall 2025 (114-1, 2025/09/01–12/15) taught this lecture on 11/10, and the course page marks the week as Virtual. It is the last row on the course page with slides attached; the next three rows (Knowledge / Multimodality, Personalization, Reasoning) have titles only.

Sources: the slide deck Language Agents (251110_LangAgent.pdf) (65 pages) and five videos: 11.1 Language Agents Introduction (21:11), 11.2 Reasoning (21:40), 11.3 Memory (19:28), 11.4 Planning (23:19), and 11.5 Multi-Agent Systems (19:38). The videos are taught in Mandarin. I checked the slides on 2026-09-30, and all page numbers below refer to the PDF.

Version note: The slide cover says November 10th, 2025 and names the EMNLP 2024 Language Agents tutorial as its reference. The five videos were uploaded to the Fall 2025 playlist on 2025/11/10, but their descriptions carry the date 2024/12/04. I did not compare the video frames against the 2025 slides page by page; where they differ, this post follows the slides.

Series: Previous: Bias, Safety, Hallucination, and Alignment + Final Project | Next: Reasoning (videos only) | Series overview

There is no homework for this lecture. Slides and videos are public, with no extra gaps beyond the series-wide A2 grade. The question it answers: everyone talks about agents, but what exactly is one, and what do reasoning, memory, planning, and multiple agents each solve?

What an agent is, and what language agents add

Slide 2 lays out both camps. Bill Gates, Andrew Ng, and Sam Altman are bullish on agents; the other side says current agents are thin wrappers around LLMs and that autoregressive LLMs can never reason or plan. The slides don't pick a side. They go back to definitions.

Slides 3–4 quote Russell and Norvig's AI: A Modern Approach: an agent is anything that perceives its environment through sensors and acts on it through actuators. In one line, an agent is an entity that perceives and acts, and a rational agent picks actions that maximize its (expected) utility.

Slide 5 explains what is new in a language agent:

  • Generating reasoning tokens can be viewed as an internal action, taking place in an internal environment in the manner of an inner monologue.
  • Self-reflection is a "meta" reasoning action: reasoning over the reasoning process.
  • Reasoning serves better acting: inferring environment states, replanning, and so on.
  • Percepts and external actions are represented in language.

The table on slide 6 compares three generations of agents:

Logical agentNeural agentLanguage agent
ExpressivenessLow: bounded by the logical languageMedium: whatever a (small) NN can encodeHigh: almost anything verbalizable
ReasoningLogical inference: sound, explicit, rigidParametric inference: stochastic, implicit, rigidLanguage-based: fuzzy, semi-explicit, flexible
AdaptivityLow: bounded by knowledge curationMedium: data-driven but sample-inefficientHigh: strong LLM priors plus language use

Slide 7 sets the three key concepts for the rest of the lecture: reasoning, memory, and planning.

Reasoning: CoT and ReAct

Slide 9 lays out a language agent's action space in three kinds: reasoning (update short-term memory, i.e. the context window), retrieval and learning (read and write long-term memory), and planning (choose an external action at inference time).

  • Chain-of-Thought (Wei et al. 2022, slide 10): have the model generate intermediate steps that imitate human thinking.
  • Reasoning helps acting, and acting helps reasoning (slides 11–12). The example on slide 12 asks, in Chinese, "Do you know NTU's Yun-Nung Chen?"; the model searches first, then answers from the results.
  • ReAct (Yao et al. 2022, slides 13–17): interleave reasoning and acting. The slides draw two conclusions: both are essential, and reasoning provides explanations for controlling actions.

Slides 18–19 add a more abstract but important point: reasoning enlarges the action space. The space of language and reasoning is infinite. A bigger action space means more capacity but harder decisions, and LLMs pick up reasoning priors by imitating many human reasoning traces. Slide 20 follows with Yao et al. 2023 on using action planning to improve reasoning.

The next post in this series, Reasoning, has videos but no slides, so it points back to the CoT and ReAct material on slides 8–20 here.

Memory: short-term, long-term, and Generative Agents

The comparison on slide 22 is easy to remember:

Short-term memoryLong-term memory
FormInstruction, Thought, Action, Obs appended in orderRead and write
ContentsContext for the current taskExperience, knowledge, skills
LimitsAppend-only; limited context; limited attention—
PersistenceDoes not persist across new tasksPersists over new experience

Generative Agents (Park et al. 2023, slides 23–25) faces a two-part problem: the context window can't hold the whole event stream, and it's hard to attend to the relevant events. The approach has two steps: simulate a series of events to build episodic memory, then retrieve from it. The point of slide 25 is that retrieval should weigh recency, importance, and relevance together, not relevance alone.

Slides 26–30 cover social simulation agents (Zhang et al. 2024). Given the same line, "I passed the bar exam!", an agent that was just promoted suggests a party, while one that took the exam and failed replies half-heartedly. The slides use this work to show that an agent's own emotion changes its response. In simulated group discussions, negative emotion leans toward objections and positive emotion toward agreement, and groups in a positive mood tend to reach more peaceful decisions.

Planning: from reactive to tree search to world models

Slide 32's definition: given a goal G, decide on a sequence of actions (a0, a1, …, an) that leads to a state passing the goal test g(·).

The slides build up with a few examples:

  • Commonsense-inferred planning (Kuo & Chen 2023, slide 33): the user only says "I want to plan a trip to SF", the agent infers the implicit flight and hotel intents, and calls an airline bot and a hotel bot in turn.
  • Web planning agents (Deng et al. 2024, slide 34): decompose a task into several web actions.

Slides 35–37 compare three planning paradigms:

ParadigmStrengthsWeaknesses
Reactive (decide each step directly)Fast, easy to implementGreedy, short-sighted
Tree search with real interactions (Koh et al. 2024)Systematic explorationIrreversible actions, unsafe, slow
Model-based planning with a world modelFaster, safer, systematic explorationHow do you get a world model?

World models

Slide 38: a world model is an environment simulator. It answers "if I take action a_t in state s_t, what happens next?" Slides 39–42 illustrate it with two papers on dialogue policy learning (slide 41 lists the D3Q authors, Chen among them):

  • Deep Dyna-Q (Peng et al. 2018): while interacting with real users, also learn a world model that generates simulated experience for planning. The catch is that low-quality fake experience drags policy learning down.
  • D3Q (Su et al. 2018): add a discriminator to filter out bad simulated experience. Policy learning becomes more robust, and human evaluation improves.

Slides 43–45 carry the thread to LLMs: an LLM can serve directly as a user simulator. Slide 44 shows role-play prompts for an extroverted and an introverted persona, and the slides conclude that this makes diverse user simulators easy to build for training assistants. Slide 45 adds that LLMs can predict state transitions in some cases. Slides 46–47 cite Gu et al. 2024 on web agents: on VisualWebArena, model-based planning is more accurate than reactive planning and more efficient than tree search.

Multi-agent systems

Slide 49 lists the motivations: a single agent isn't strong enough, multiple agents scale easily in parallel, different agents represent different expertise, and control can be decentralized and privacy-preserving. Slide 50 splits construction into three steps, each with two examples:

  1. Agent initialization: by persona description (slide 52, the long profile of pharmacy shopkeeper John Lin from Generative Agents), or by roles and actions (Chen et al. 2024, slides 53–54).
  2. Orchestration: multi-agent debate (Du et al. 2023, slides 56–58) improves factuality and reasoning; AutoGen (Wu et al. 2023, slides 59–60) lets agents interact through conversation, which the slides call conversational programming.
  3. Team optimization: optimize the team by selecting agents (Liu et al. 2024, slides 62–63). After optimization, a team of the same size gets better results with fewer API calls.

The summary on slide 64 boils each concept down to a line or two. Reasoning is an internal action for agents: reasoning guides acting, and acting updates reasoning. Language agents interact with both external environments and internal memories. Reasoning in language enables new planning abilities.

How to self-study this lecture

  1. Start with 11.1 and get slide 5 ("reasoning is an internal action") and slide 9's three kinds of actions clear. The other four videos build on them.
  2. Pair 11.2 with the ReAct paper, and 11.3 with the memory retrieval section of the Generative Agents paper.
  3. The planning paradigms table in 11.4 (slide 37) is the most useful single slide in the lecture. Afterward, try sorting the agent frameworks you know into its three rows.

One thing you can do tonight: take an agent you use or are building, and following slide 22's table, write down what goes into its short-term and long-term memory. Then ask whether its long-term retrieval considers recency and importance, or relevance only.

Further reading

Previous: Bias, Safety, Hallucination, and Alignment + Final Project Next: Reasoning (videos only)

References