🌏 中文版
Lecture 8 of CMU 11-768 AI Agents (Sep 17, 2026) opens the training module. The speaker is Yueqi Song from CMU's LTI, who built the Agent Data Protocol with Graham Neubig. The first seven lectures tuned the harness: context, tools, memory, planning. This one pulls a different lever: training recorded agent trajectories into the model's weights.
The slide subtitle is "Turning recorded trajectories into weights." The lecture has six parts: trajectories as tokens, choosing trajectories, one format for many datasets, running the training, knowing it worked, and handing off to RL. It is aimed at engineers who use LLM APIs but have not trained an agent model themselves — though more students raised their hands for "done agent SFT" than the speaker expected.
- Course page: cmu-agents.com schedule (slides and recording for Lecture 8)
- The schedule lists no required readings for this lecture, only references. Numbers from technical reports below start from the slides and were then checked against the full reports; mismatches are noted where they occur.
Starting point: what a trajectory looks like
The slides open with a trajectory from SWE-Gym: the agent gets an issue and two tools (bash and a file editor) in a sandbox, writes thoughts and commands, reads back the outputs, and the repository's own tests decide whether the run succeeded. There are four roles — system, user, assistant, observation — and only a handful of assistant messages were actually written by the model.
Why update weights
Song lists three places to update an agent; the first two were earlier lectures' topics:
| Location | Pros | Cons |
|---|---|---|
| Context window (Lecture 3) | High fidelity to what happened | Costly, noisy, doesn't decide what matters |
| External artifacts (Lecture 4's skills and memory) | Inspectable, editable, retrievable, portable | Must be induced, selected, and maintained |
| Model weights (this lecture) | Faster inference, broad behavior change | Slower updates, opaque, model-specific |
Change the weights once and every later call has the new behavior with nothing extra in the prompt. The price: every change costs a training run, you cannot open the weights to see what changed, and you cannot move the change to another model.
What trajectories teach
- How to write a tool call this harness can actually run.
- How to keep going when the conversation gets long, because the training trajectories were long.
- What a tool result looks like, so the model waits for one instead of writing one itself.
- And whatever else was in the data: extra commentary, repeated commands, the one solution that happened to be recorded when several would work.
That last point foreshadows the whole lecture: SFT does not separate good from bad. It learns whatever is there.
Where SFT sits in the pipeline
The slide draws the pipeline to token scale; each step down is fewer tokens, more heavily curated:
| Stage | Scale (slide examples) |
|---|---|
| Pre-training | 20T tokens (Nemotron 3 Ultra) |
| Mid-training | 1.43T (slide, K2 Horizon) |
| SFT (this lecture) | 332B (K2 Horizon) |
| RL (from Sep 22) | No fixed corpus — rollouts plus rewards |
Song notes that definitions of mid-training vary, and some people count SFT as part of it. The first three stages all have fixed corpora; only RL generates its own data.
One number does not check out. On the K2 Horizon model card, the three SFT phases do sum to 332B (81B + 201B + 50B), but the four mid-training stages sum to about 1.91T (1.09T + 503B + 117B + 201B), not the slide's 1.43T. K2 Horizon also runs mid-training → RL → SFT, with SFT after RL, which differs from the generic ordering in the table above.
SFT as the cold start for RL
"Cold start" is the checkpoint RL begins from. Kimi K3 runs post-training in three stages: SFT establishes baseline agent capabilities → RL develops domain experts at varying reasoning effort → on-policy distillation consolidates them into one model. Tool calling and long tasks are learned in SFT, before any RL.
So why not RL alone? DeepSeek-R1 is the answer: R1-Zero ran RL directly on the base model, and its answers mixed languages and were hard to read. The next model, R1, put SFT back, starting from a few thousand curated examples. The Kimi K3 report puts it this way: the SFT stage establishes a high-quality cold-start policy for the subsequent RL stage.
SFT and RL side by side
| SFT | RL | |
|---|---|---|
| Learns from | Recorded trajectories, whoever produced them | Its own rollouts, scored by a reward |
| Signal | A target for every token | One number per rollout |
| Needs | Data and GPUs | Live environments, sandboxes, a checkable reward |
| Good at | Tool calls, formats, long tasks | Sharpening what the model can already attempt |
| Stuck when | Demonstrations are wrong, narrow, or used up | The start is too weak, or the reward can't be checked |
RL gets the next three lectures (Sep 22, Sep 29, Oct 1), starting with L9 RL Basics.
Decision one: which tokens carry the loss
A trajectory becomes one token string
Run the trajectory through the chat template and you get one string, for example:
<|im_start|>tool
3 failed, 41 passed
<|im_end|>
<|im_start|>assistant
The test compares floats. Fix it.
<tool_call>
str_replace(...)
</tool_call>
<|im_end|>
Each <|…|> marker is a single special token; everything else is what the agent wrote and saw.
Cross-entropy on assistant tokens only
$$ \mathcal{L}(\theta) = -\sum_{t:,m_t=1} \log p_\theta(x_t \mid x_{<t}) $$
$m_t$ is the mask, one bit per token. Minimizing this raises the probability of each recorded action given everything before it. Observations (tool outputs) are conditioning context, not prediction targets.
A student asked: why not train on tool outputs too? Song says assistant-only is the general rule, since observations come from the environment or the user, not from the model. But some work does train models to predict execution results — Meta's CWM, for instance — which helps on tasks that require understanding code execution. The cost is model capacity; a small model might regress elsewhere. It remains an open question.
How the mask is built
- Render: the template writes the conversation as one string and records where each assistant span starts and ends.
- Tokenize: character positions become token positions.
- Mask: tokens inside an assistant span get 1, everything else 0.
By the time training starts, the roles are gone; the trainer sees only the string and the bits. Three common frameworks:
| Framework | Setting | How the mask is produced |
|---|---|---|
| TRL | assistant_only_loss=True | The template marks assistant spans as it renders (needs {% generation %} markers) |
| Axolotl | roles_to_train, train_on_eos | Finds each turn's role header in the rendered string |
| LLaMA-Factory | train_on_prompt, mask_history | Tags whole turns as train-or-ignore while encoding |
Same decision in all three; only TRL puts it in the template. Song often uses LLaMA-Factory for her own agent SFT.
The one people miss: the stop token
In <|im_start|>assistant ... <|im_end|>, the role header should not carry loss (the harness supplies it at inference), but the closing <|im_end|> must, because stopping is something the model has to do itself. Leave it out of the mask and the model may not learn to stop.
One level deeper: "end of turn" is really three events — end this message, hand off to a tool, or finish the task. The slide says Kimi K3's template gives each its own token. Section 4.1.1 of the Kimi K3 report describes something slightly different: its XTML template uses a single [end_of_msg] token as the generation stop marker, and splits the assistant message into think, response, and tools channels, with tool calls in the tools channel. The three events are told apart by channel structure plus the stop token, not by three separate stop tokens.
Reasoning blocks: two separate switches
- Does the reasoning stay in the context? If it does, it explains the action that follows.
- Is it trained on? If it is, it becomes the style the model writes in.
Nemotron 3 Ultra keeps budget-truncated reasoning in the context but masks the artificial cut-off out of the loss.
Weighting within the assistant turn, and samples versus tokens
- Long reasoning dominates the loss, yet the short tool call is what actually has to be right. BalanceSFT (Hao et al., 2025) uses a learned weight (the SSB loss) to rebalance reasoning tokens against tool-call tokens and resamples hard examples (HDR). The schedule's link is correct; the arXiv listing still carries the old title, FunReason, and only the v3 PDF is titled BalanceSFT. The slide says rebalancing helps multi-turn more than single-turn "on both base models," but the paper's Table 5 supports only half of that: adding SSB alone gives Qwen2.5-Coder-7B +0.5 single-turn and +3.4 multi-turn, while Llama-3.2-3B goes the other way, +3.2 single-turn and +1.4 multi-turn.
- MAI-Thinking-1 set its mixture by counting samples, but loss is computed per token. STEM and coding traces run far longer and ended up taking nearly all the tokens. Lesson: after balancing by trajectory count, check what that turned into in tokens.
An experiment nobody has run
Song is blunt: she could not find a published experiment that changes the mask for agent SFT and measures the effect. The careful masking experiments are single-turn instruction tuning (Instruction Modelling, Weighted Instruction Tuning). The Instruction Modelling paper found unmasking the prompt helps when prompts are long and answers short, or when training data is scarce — agent trajectories are the opposite on both counts. At 8B scale the experiment is cheap, and it is sitting there unrun.
Decision two: which trajectories to keep
Where trajectories come from
| Keep everything | Keep the runs that passed | |
|---|---|---|
| A stronger model ran it | Plain distillation | SWE-Gym, Nemotron, OpenThoughts-Agent (most agent SFT lives here) |
| The model ran it itself | MAI, on its own reasoning runs | Expert iteration (Sep 22) |
Most data sits in the top right: a stronger model writes better trajectories than the student could, and the task's own tests make keeping only the good ones cheap. The bottom row is the model making its own data; loop it and you get expert iteration.
The "MAI" cell needs a note, because the recording's auto-generated captions say "Anthropic uses their own reasoning runs" at this point. The slide says MAI, and the MAI-Thinking-1 report (§3.1.4) matches: Microsoft AI collects rollouts during RL and runs SFT on a mid-trained checkpoint with them (they call it self-distillation), which becomes the starting point for the next RL run. This guide follows the slide and the report; the caption is most likely a transcription error.
Distillation and saturation
In SWE-Gym, GPT-4o and Claude ran the tasks; 491 runs passed the repository tests, and fine-tuning a 32B Qwen model on those 491 alone raised both SWE-Bench splits by more than ten points. With the task set held fixed and only the number of sampled trajectories growing, accuracy was still rising when the sampling budget ran out — they had tasks to spare and ran out of compute.
Successful runs can still teach bad habits
The slide shows several runs that all passed their tests: one spent most of its turns in an edit–test loop, one left print( in the final patch, one read files and never edited. The loss copies every action, good or not, so the model learns the mess along with the fix. The Nemotron report says such runs "would teach undesirable behaviors if used directly as SFT data."
Filters, teachers, and task sources
OpenThoughts-Agent (Raoof et al., 2026) ran over a hundred controlled ablations. The lecture pulls four findings:
- The best filter is crude: count turns and drop every run shorter than five. Turn count measures how much work a run took, not whether it was good — a clean four-turn fix is dropped and a forty-turn mess is kept. On average it still helped.
- The strongest model is not always the best teacher: GPT-5.3-Codex was the strongest of five candidates on the benchmarks themselves, yet its runs produced the weakest student. Song's explanation: its way of thinking or solving may be too different from the student's. Leaderboard rank misleads here; trying two teachers and keeping the better student is cheap by comparison.
- Task source matters most: different sources help different benchmarks (SWE-smith, for example, helps SWE-Bench Verified a lot and Terminal-Bench little). Spend effort here before tuning anything else, and mix the best four to eight sources — going wider did not help.
- New tasks beat more runs: starting from the same 10K dataset, one curve adds more runs of the same tasks and flattens; the other adds new tasks and keeps climbing.
Class discussion: which of three runs do you keep?
Three runs of the same task came back:
- A: a short run that solved it in four turns.
- B: a long run that made a wrong edit, saw the test fail, and recovered.
- C: a run that solved it after eleven attempts that changed nothing.
Students leaned toward B because it teaches recovery. Song added B's risk: the model might learn the wrong edit rather than the recovery. Her answer was to keep both A and B, so the data covers the capabilities your target benchmark needs most. And note that a minimum-five-turns filter would throw A away.
Decision three: one format for many datasets
This section is Song's own work, the Agent Data Protocol (ADP).
- The problem: each agent dataset was built by a different group in its own format. For web pages, some save HTML and others the accessibility tree. To train OpenHands on both Go-Browse and Mind2Web, you write two custom converters; N datasets times M harnesses explodes.
- The approach: convert each dataset once into typed actions and observations, keeping both HTML and the accessibility tree instead of picking one. Each harness then needs only one converter out of ADP. The paper unifies 13 existing datasets.
- Rendering out: OpenHands, SWE-Agent, and AgentLab take different actions, so the same trajectory looks different in each. This is also where you set the system prompt and decide what happens when context gets long. Then check the result: do tool calls parse, does each call come with a thought, does the conversation end the way it should.
- Diversity versus task-specific data: same harness, same model, same evaluation; only the training mixture changes. The mixture beats the matching single-domain set — even on that domain's own benchmark.
Decision four: running the training
Real configurations
| Model | Sequence | Batch | Learning rate |
|---|---|---|---|
| Nemotron 3 Ultra | Packed, 294,912 then 515,000 | 64 | 1.5e-5 → 1e-6 (stage 1; stage 2 is 1e-5 → 2e-6) |
| MAI self-distillation | Packed 128k | 2,048 | 1.7e-5 → 5.2e-6, 2% warmup |
| K2 Horizon | 512K, three phases | — | Decayed on a high-quality subset |
A student asked whether this is full-parameter training. Yes — essentially the pre-training loss on chat-formatted data, computed on a subset of tokens. Well-resourced labs prefer full fine-tuning; Song has used LoRA herself for efficiency.
MoE expert routing
The frontier models in the lecture are mostly mixtures of experts, while the models public papers fine-tune are mostly dense. Fine-tune an MoE on narrow agent data and a few experts take most of the tokens. MAI's fix: weight routing balance a thousand times more heavily during SFT than during RL, and raise dropout to 0.15.
Packing whole conversations
Training sequences have fixed length; packing fills each with several conversations instead of padding. Nemotron assigns each conversation to the pack whose remaining capacity it fits most tightly; no conversation is split or truncated, a deliberate choice against hallucination; identical prompts stay out of the same pack; and packs are shuffled afterwards.
Template drift
The template you train with has to be the template you serve with, down to the whitespace. The slide's trap: render the same trajectory twice, once with a tool message after it and once without, and you can get two different strings — some templates conditionally rewrite the previous assistant turn's thinking block, so the prefix no longer matches. Inside a tool loop, that is fatal. TRL's Chat Templates doc uses exactly this case: the original Qwen3 template decides whether to emit the thinking block based on loop.last, so appending a tool message changes how the previous assistant turn renders and breaks prefix preservation; TRL's qwen3_training.jinja always emits the thinking block. Hugging Face's Writing a chat template doc puts it simply: the chat template should always match the format the model was trained with.
Read the released training artifacts
The K2 Horizon model card has a stage-by-stage training table (steps, tokens, sequence length, purpose of each phase) — the source of the configuration numbers above. It also links the Weights & Biases run with real loss curves and ships intermediate checkpoints for every stage. Nemotron released post-trained checkpoints together with the training data and recipe. Song recommends reading at least one before doing SFT yourself.
Decision five: knowing it worked
The train/deploy mismatch
Training always predicts the next step while continuing the recorded run. Deployed, there is no recording: every action changes what the model sees next. One edit that differs from the demonstration and it is in a state no training run visited, where low loss promises nothing. This is the problem DAgger addresses, and why recovery ability mattered in the discussion above. Loss curves monitor training; they do not replace evaluation.
An evaluation checklist
- First, check you can overfit 20 examples. If the loss won't drop, the bug is upstream (mask, template, data).
- Evaluate in the target harness, on tasks with checkable outcomes.
- Read published numbers carefully: some are the best of several harnesses.
- Match the split to the claim: keep all runs of one task on the same side; split by repository if you want to claim generalization to new repos.
- Robustness across harnesses: trajectories collected in one system teach its conventions along with the task. Data collected only through OpenHands may not transfer to another harness; labs collect across several.
- Regressions outside the target domain: training one capability moves others whether you measure them or not. Your mixture is the control variable, and balancing by examples is not balancing by tokens. Evaluate the things you were not trying to improve.
Finally: handing off to RL
OpenThoughts-Agent's Table 11 gives three conclusions:
- RL with no SFT at all barely moves the base model (the R1-Zero lesson).
- SFT alone falls well short of what SFT plus RL reaches.
- The winning recipe starts RL from a checkpoint that was deliberately not trained to its own best score.
As the slide quotes: RL provides the most gains when the SFT model is selected with RL in mind. Song's overall view is that SFT gives an agent its most general capabilities, while RL trains for specific domains.
Q&A
- What if there is no stronger teacher? Distill from yourself, or rely on RL. Song believes frontier labs such as OpenAI and Anthropic get agent capabilities mainly through RL, since SFT is fundamentally capped by the teacher.
- Should incorrect intermediate steps be supervised? Yes, because the model needs to learn to recover from mistakes — but the final result must be correct.
- How do you pick mixture ratios? Mostly empirically: try on a small model and small data first. MAI's 56% coding and STEM was found that way. Also, RL reweights implicitly: higher-reward trajectories get larger gradient weight.
- Any imitation learning without teacher forcing? SFT uses teacher forcing — fixed context, no sampling from the model during training. Methods where the model produces its own training data come in the RL lectures.
Something to try tonight
Take one trajectory from your own agent, render it with your target model's tokenizer, and print the tokens where the mask is 1:
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("your-model")
out = tok.apply_chat_template(
messages, tools=tools, tokenize=True,
return_dict=True, return_assistant_tokens_mask=True,
)
ids, mask = out["input_ids"], out["assistant_masks"]
print(tok.decode([i for i, m in zip(ids, mask) if m]))
Check three things: did tool output leak in, is the closing stop token included, is the role header wrongly counted. If the output is empty, the template has no {% generation %} markers: the Transformers tokenizer docs say return_assistant_tokens_mask only works with templates that support it via the {% generation %} keyword. TRL's Chat Templates doc lists the model families it ships training templates for (Qwen3, Llama 3, DeepSeek-V3, GPT-OSS, and others) and swaps them in automatically when assistant_only_loss=True; for other models you have to add the markers yourself, or assistant_only_loss won't work correctly. Then confirm you can overfit 20 examples before anything else.
Further reading
- Same series: L7 Computer Use Agents (SFT data for CUAs), L9 RL Basics
- SFT and RLHF fundamentals: CS336 Lecture 15: SFT and RLHF
- What comes next, RL: CS336 Lecture 16: RLVR and GRPO
- The training picture from another course: CME295 Lecture 4: the bill for pretraining, SFT, and LoRA
- How other teams do post-training RL: Three RL Post-Training Playbooks: Ornith, Nous Research, MiniMax
The schedule's references also list the Kimi K2 technical report. The full text was opened, but this guide does not cite it; where it discusses templates and the post-training pipeline, it uses the later Kimi K3 report.
References
- Course: CMU 11-768 AI Agents site and schedule, Lecture 8 recording, speaker Yueqi Song's homepage
- Technical reports: Kimi K3, DeepSeek-R1, Nemotron 3 Ultra, MAI-Thinking-1, K2 Horizon model card
- Data and methods: OpenThoughts-Agent, SWE-Gym, Agent Data Protocol, BalanceSFT (arXiv:2505.20192), Instruction Modelling, Weighted Instruction Tuning, DAgger, CWM
- Tools: Transformers chat templating, Transformers: Writing a chat template, Transformers
apply_chat_templatesource, TRL SFTTrainer, TRL Chat Templates ({% generation %}markers, prefix preservation, bundled training templates and supported families), Transformers tokenizer docs (thereturn_assistant_tokens_maskparameter ofapply_chat_template), Axolotl conversation datasets, LLaMA-Factory
Loading...