🌏 中文版
This is post 11 of Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall. The course is ADL Fall 2025 (NTU semester 114-1, 2025/09/01–12/15). Class was off on 9/29 (Teachers' Day) and 10/06 (Mid-Autumn Festival), so RAG came after the break, on 10/13. That week's TA recitation is LLM Basics & MoE, and the homework column shows HW 3.
Sources: the RAG slides (60 pages) and six videos: 8.1 What is RAG? (27:32), 8.2 RAG Framework (39:52), 8.3 Advanced RAG (10:26), 8.4 RAG Roadmap (28:17), 8.5 RAG Pros & Cons (17:23), and 8.6 RAG Q&A (24:19, an in-class Q&A with no slides). The HW3 video is ADL 2025 Fall Homework 3 (20:34, uploaded 2025-10-12). All checked on 2026-09-30. Videos are in Mandarin; page numbers refer to the slide PDF.
Series position: previous PEFT: Adapter, LoRA, Prompt Tuning, and HW2 | next NLG: Decoding, Control, and Evaluation | Series overview
Why look things up before answering
Four examples
The slides open by asking an LLM questions in Chinese (pp.2–5):
- "Do you know Tsai Ing-wen?" The answer is thorough. She is famous, so the training data has plenty on her.
- "Do you know NTU's Yun-Nung Chen?" The model says she is an electrical engineering professor who works on semiconductors, which is almost entirely wrong. Page 3's takeaway: LLMs cannot memorize all details about long-tail information.
- Ask again with search results attached, and the answer becomes correct: computer science department, language understanding, dialogue systems. Page 4's takeaway: knowledge grounding helps reduce hallucinations.
- "Who is Taiwan's current president?" The model still says Tsai Ing-wen. Page 5: LLM knowledge goes stale easily and is hard to update. Knowledge editing is another possible fix.
Page 9 adds one more layer: generation follows the distribution of the pre-training data. After "my good friend is a nurse", the model continues with "she is very patient with patients".
Four limits of parametric memory
Page 10 turns this into a list. LLMs store a lot in their parameters, but:
- They cannot memorize everything.
- The world changes over time.
- Private documents are not reachable through the web.
- With a black-box model, it is hard to verify whether an answer is correct.
Page 11 sorts the knowledge RAG targets into three kinds: long-tail knowledge, dynamically changing knowledge, and knowledge absent from pre-training.
A precursor: WebGPT
Pages 6–8 first present WebGPT (Nakano et al., 2021): fine-tune GPT-3 on human demonstrations so it writes answers with references. The model treats search as token continuation. It generates [SEARCH] ... [END] as a query, then [CLICK] 1 [END] to open a document, learning human browsing actions.
The RAG framework: index, retrieve, generate
The framework diagram on p.12 has four parts. Documents are indexed. When a query arrives, the system retrieves, puts the results into the prompt with the query, and passes it to the LLM for output. The diagram marks the LLM as either open- or closed-source.
Indexing and retrieval
Page 13: documents are split into chunks or passages and encoded; relevance is computed when a query arrives. Retrieval comes in two kinds:
| Type | How | Methods on the slides |
|---|---|---|
| Sparse retrieval (p.14) | Term matching | n-gram (TF-IDF), BM25 |
| Dense retrieval (p.15) | A neural encoder produces vectors; score by cosine similarity | Off-the-shelf embeddings (pre-trained BERT, GPT); embeddings learned with contrastive learning: DPR, Contriever |
Training a dense retriever
Page 16 uses "Which U.S. state has the largest area?" as the example. A query encoder and a passage encoder each produce a vector, the inner product s(q, d) is the score, and training maximizes the softmax probability of the correct passage d⁺ among all candidates.
The contrastive loss on p.16
L = − log( exp(s(q, d⁺)) / Σᵢ exp(s(q, dᵢ)) )
The dᵢ include one correct passage d⁺ and several negatives d⁻. The slide adds a note at the bottom: paired training data is difficult to collect.
Since paired data is hard to get, pp.17–19 flip the direction. Use a language model to estimate P(q | d), the probability of generating the query after reading the document, and treat it as relevance (query likelihood; the slide cites Sachan et al., 2023). Pages 18–19 present the course group's own extension (Huang & Chen, 2024): an unsupervised multilingual dense retriever trained with query likelihood. On XOR-TyDi QA, the slide concludes that training with query likelihood beats training with paired data.
Generation
Page 21 is a RAG prompt example (an image). Page 23 lists the issues in RAG:
- Retrieval: misaligned or irrelevant chunks; redundant information from multiple sources.
- Generation: hallucinations not supported by the retrieved context; irrelevant, toxic, or biased output; over-reliance on retrieved text, simply echoing it.
Advanced RAG: one stage before retrieval, one after
Pages 24–25 add a pre-retrieval and a post-retrieval stage to the framework:
- Pre-retrieval: decide whether to search, rewrite the query, enrich context by expanding more chunks for the LLM, and run sparse and dense retrieval in parallel (hybrid search).
- Post-retrieval: rerank to move the most relevant content up, and compress the context to keep only the essentials.
Reranking
Page 26 compares two prompts for reranking with an LLM:
- Pointwise: ask the model to rate the relevance of the query and one passage from 1 to 5.
- Pairwise: give two passages and ask which is more relevant (A or B). The slide concludes that pairwise is more accurate.
Pages 27–29 are again the course group's research (Huang & Chen, 2024). InstUPR does zero-shot instruction-based reranking, first pointwise and then pairwise. Page 28 concludes it performs comparably with supervised rerankers. PairDistill iteratively trains the retriever using a loss from the previous iteration's reranking. Page 30's ITER-RETGEN (Shao et al., 2023) concatenates the previous output with the query for the next retrieval round.
The RAG roadmap: what, how, and when to retrieve
Pages 31–48 are the most structured part of the lecture. The slides use three questions as three columns and fill methods in step by step. The table comes from the ACL 2023 retrieval-based LM tutorial slides, extended with 2025 methods. The table below fills only the cells where the slides place a method; the rest are "—":
| Method | What to retrieve | How to use it | When to retrieve | Slide note |
|---|---|---|---|---|
| RAG (Lewis et al., 2020) | Text chunks | Input layer | Once, at the start of generation | Trains retriever and generator end to end (p.33) |
| RETRO (Borgeaud et al., 2022) | Text chunks | Intermediate layers | — | More efficient than the input layer, but requires training (p.35) |
| Retrieval-in-context (Ram et al.; Shi et al., 2023) | — | — | Every n tokens | Retrieving more often helps, but inference gets slower (p.38) |
| FLARE (Jiang et al., 2023) | — | — | Adaptively | Generates first, retrieves when model certainty is low (p.40) |
| Search-R1 (Jin et al., 2025) | — | — | Adaptively | RL learns when to search; reward is answer correctness (p.42) |
| AdaSearch (Lin et al., 2025) | — | — | Adaptively | RL for "knowing what you know"; reward is answer correctness plus correct decisions (p.43) |
| kNN-LM (Khandelwal et al., 2020) | Tokens | Output layer | Every token | Retrieves similar examples and uses their next token; finer-grained and compute-efficient, but space-expensive (p.45) |
| Efficient NN-LM (He et al., 2021) | Tokens | Output layer | Adaptively | Fixes kNN-LM's per-token retrieval cost (p.47) |
How to read the table: every step right or down trades effectiveness against cost at a new point. RETRO gains efficiency but needs retraining. Retrieving every n tokens helps quality but slows inference. Search-R1 and AdaSearch let the model learn for itself whether to look something up.
Retrieval sources and choosing an approach
Pages 49–50 cover retrieval sources. Unstructured data is text. Semi-structured data is text plus tables, such as PDFs; the problems are data corruption and extracting table content, and converting tables to text for text-only RAG is suboptimal. Structured data means knowledge graphs. Page 50 cites the GraphRAG survey (Peng et al., 2024): knowledge graphs help more on reasoning-intensive tasks.
Pages 51–54 cite the RAG survey of Gao et al., 2023 and use an analogy. A CS student is asked which legal rules apply to the contract between NVIDIA's headquarters and Shin Kong Life:
| Route | Analogy | Pros | Cons |
|---|---|---|---|
| Prompting | Answer from what's already in your head | Simple, efficient | Relies entirely on internal knowledge |
| RAG | Look it up in law textbooks, then answer | Suits dynamic environments, highly interpretable | High latency, depends on retrieval quality |
| Fine-tuning | Take a related course first, then answer | Deep customization of behavior and style | High cost of retraining |
Page 54 concludes that choosing between RAG and fine-tuning depends on data dynamics, customization needs, and available compute.
Pages 55–58 add four ways to train retrieval-augmented LLMs (training-free, independent, sequential, and joint, citing the RA-LLM survey), the range of applications, and six advantages of RAG: accuracy, up-to-date knowledge, interpretability, customization, safety and privacy, and cost (no model update).
HW3: only the title is public
The HW 3 button in the course page's 10/13 row links straight to ADL 2025 Fall Homework 3 on YouTube. What can be confirmed:
- Title: the video description reads "Retriever & Reranker Training for RAG".
- Length 20:34, uploaded 2025-10-12.
- Page 11 of Course Logistics lists the third assignment's topic as RAG.
What is missing: the video has no subtitles, and there are no public spec slides or written instructions. This post cannot confirm the dataset, corpus, base model, baseline, metric, submission format, or deadline, so it states none of them. Submission goes through NTU COOL, which needs an NTU account.
How to use it from outside NTU: the title lines up with pp.15–29 of the slides. Pick a public QA dataset, start with a BM25 baseline, fine-tune a dense retriever with the p.16 contrastive loss, then add a pointwise or pairwise reranker and compare retrieval metrics at each stage. You will define your own grading and cannot compare with the official one.
After this lecture you should be able to
- Name the four limits of parametric memory and the three kinds of knowledge RAG targets.
- Draw RAG's four components and explain the difference between sparse and dense retrieval.
- Explain pointwise vs. pairwise reranking, and why the slides consider pairwise more accurate.
- Place RAG, RETRO, FLARE, and Search-R1 in the right cells of the what/how/when table.
One thing to try tonight: pick an obscure person or an internal document you know well and ask any LLM about it twice, once directly and once with the relevant passage pasted into the prompt. You will see the difference between pp.3 and 4 yourself. Then swap the pasted passage for an irrelevant one, and you will see what p.23 means by over-relying on retrieved text.
Further reading
- CS224N Lecture 10: Six Components of RAG and Language Agents
- The Complete Guide to RAG System Patterns
- RAG vs Fine-tuning: It's Not Either/Or
Next: NLG: Decoding, Control, and Evaluation
References
- ADL Fall 2025 (114-1) course page — the 10/13 session, break weeks, recitation, and HW 3 link
- RAG slides (251013_RAG.pdf) — source of all page numbers
- ADL 8.1: Retrieval-Augmented Generation (RAG) (YouTube, in Mandarin)
- ADL 8.2: RAG Framework (YouTube, in Mandarin)
- ADL 8.3: Advanced RAG (YouTube, in Mandarin)
- ADL 8.4: RAG Roadmap (YouTube, in Mandarin)
- ADL 8.5: RAG Pros & Cons (YouTube, in Mandarin)
- ADL 8.6: RAG Q&A (YouTube, in Mandarin)
- ADL 2025 Fall Homework 3 (YouTube) — the description is a single line with the title
- Course Logistics slides (250901_Course.pdf) — p.11 lists the three assignment topics
- 2025 Fall NTU CSIE ADL playlist (in Mandarin)
- ACL 2023 Tutorial: Retrieval-based Language Models and Applications — source of the roadmap table
- Nakano et al., WebGPT
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Karpukhin et al., Dense Passage Retrieval for Open-Domain Question Answering
- Izacard et al., Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)
- Borgeaud et al., Improving language models by retrieving from trillions of tokens (RETRO)
- Jiang et al., Active Retrieval Augmented Generation (FLARE)
- Jin et al., Search-R1
- Khandelwal et al., Generalization through Memorization: Nearest Neighbor Language Models
- Shao et al., Enhancing Retrieval-Augmented LLMs with Iterative Retrieval-Generation Synergy (ITER-RETGEN)
- Gao et al., Retrieval-Augmented Generation for Large Language Models: A Survey
- Peng et al., Graph Retrieval-Augmented Generation: A Survey
- Fan et al., A Survey on RAG Meeting LLMs
Loading...