Skip to content

NTU ADL 2025 Lecture 8: RAG — From Retrieval and Reranking to Search-R1, Plus HW3

Sep 30, 20261 min
TL;DRLLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".

🌏 中文版

This is post 11 of Reading NTU Yun-Nung Chen Applied Deep Learning 2025 Fall. The course is ADL Fall 2025 (NTU semester 114-1, 2025/09/01–12/15). Class was off on 9/29 (Teachers' Day) and 10/06 (Mid-Autumn Festival), so RAG came after the break, on 10/13. That week's TA recitation is LLM Basics & MoE, and the homework column shows HW 3.

Sources: the RAG slides (60 pages) and six videos: 8.1 What is RAG? (27:32), 8.2 RAG Framework (39:52), 8.3 Advanced RAG (10:26), 8.4 RAG Roadmap (28:17), 8.5 RAG Pros & Cons (17:23), and 8.6 RAG Q&A (24:19, an in-class Q&A with no slides). The HW3 video is ADL 2025 Fall Homework 3 (20:34, uploaded 2025-10-12). All checked on 2026-09-30. Videos are in Mandarin; page numbers refer to the slide PDF.

Series position: previous PEFT: Adapter, LoRA, Prompt Tuning, and HW2 | next NLG: Decoding, Control, and Evaluation | Series overview

Why look things up before answering

Four examples

The slides open by asking an LLM questions in Chinese (pp.2–5):

  • "Do you know Tsai Ing-wen?" The answer is thorough. She is famous, so the training data has plenty on her.
  • "Do you know NTU's Yun-Nung Chen?" The model says she is an electrical engineering professor who works on semiconductors, which is almost entirely wrong. Page 3's takeaway: LLMs cannot memorize all details about long-tail information.
  • Ask again with search results attached, and the answer becomes correct: computer science department, language understanding, dialogue systems. Page 4's takeaway: knowledge grounding helps reduce hallucinations.
  • "Who is Taiwan's current president?" The model still says Tsai Ing-wen. Page 5: LLM knowledge goes stale easily and is hard to update. Knowledge editing is another possible fix.

Page 9 adds one more layer: generation follows the distribution of the pre-training data. After "my good friend is a nurse", the model continues with "she is very patient with patients".

Four limits of parametric memory

Page 10 turns this into a list. LLMs store a lot in their parameters, but:

  1. They cannot memorize everything.
  2. The world changes over time.
  3. Private documents are not reachable through the web.
  4. With a black-box model, it is hard to verify whether an answer is correct.

Page 11 sorts the knowledge RAG targets into three kinds: long-tail knowledge, dynamically changing knowledge, and knowledge absent from pre-training.

A precursor: WebGPT

Pages 6–8 first present WebGPT (Nakano et al., 2021): fine-tune GPT-3 on human demonstrations so it writes answers with references. The model treats search as token continuation. It generates [SEARCH] ... [END] as a query, then [CLICK] 1 [END] to open a document, learning human browsing actions.

The RAG framework: index, retrieve, generate

The framework diagram on p.12 has four parts. Documents are indexed. When a query arrives, the system retrieves, puts the results into the prompt with the query, and passes it to the LLM for output. The diagram marks the LLM as either open- or closed-source.

Indexing and retrieval

Page 13: documents are split into chunks or passages and encoded; relevance is computed when a query arrives. Retrieval comes in two kinds:

TypeHowMethods on the slides
Sparse retrieval (p.14)Term matchingn-gram (TF-IDF), BM25
Dense retrieval (p.15)A neural encoder produces vectors; score by cosine similarityOff-the-shelf embeddings (pre-trained BERT, GPT); embeddings learned with contrastive learning: DPR, Contriever

Training a dense retriever

Page 16 uses "Which U.S. state has the largest area?" as the example. A query encoder and a passage encoder each produce a vector, the inner product s(q, d) is the score, and training maximizes the softmax probability of the correct passage d⁺ among all candidates.

The contrastive loss on p.16
L = − log(  exp(s(q, d⁺))  /  Σᵢ exp(s(q, dᵢ))  )

The dᵢ include one correct passage d⁺ and several negatives d⁻. The slide adds a note at the bottom: paired training data is difficult to collect.

Since paired data is hard to get, pp.17–19 flip the direction. Use a language model to estimate P(q | d), the probability of generating the query after reading the document, and treat it as relevance (query likelihood; the slide cites Sachan et al., 2023). Pages 18–19 present the course group's own extension (Huang & Chen, 2024): an unsupervised multilingual dense retriever trained with query likelihood. On XOR-TyDi QA, the slide concludes that training with query likelihood beats training with paired data.

Generation

Page 21 is a RAG prompt example (an image). Page 23 lists the issues in RAG:

  • Retrieval: misaligned or irrelevant chunks; redundant information from multiple sources.
  • Generation: hallucinations not supported by the retrieved context; irrelevant, toxic, or biased output; over-reliance on retrieved text, simply echoing it.

Advanced RAG: one stage before retrieval, one after

Pages 24–25 add a pre-retrieval and a post-retrieval stage to the framework:

  • Pre-retrieval: decide whether to search, rewrite the query, enrich context by expanding more chunks for the LLM, and run sparse and dense retrieval in parallel (hybrid search).
  • Post-retrieval: rerank to move the most relevant content up, and compress the context to keep only the essentials.

Reranking

Page 26 compares two prompts for reranking with an LLM:

  • Pointwise: ask the model to rate the relevance of the query and one passage from 1 to 5.
  • Pairwise: give two passages and ask which is more relevant (A or B). The slide concludes that pairwise is more accurate.

Pages 27–29 are again the course group's research (Huang & Chen, 2024). InstUPR does zero-shot instruction-based reranking, first pointwise and then pairwise. Page 28 concludes it performs comparably with supervised rerankers. PairDistill iteratively trains the retriever using a loss from the previous iteration's reranking. Page 30's ITER-RETGEN (Shao et al., 2023) concatenates the previous output with the query for the next retrieval round.

The RAG roadmap: what, how, and when to retrieve

Pages 31–48 are the most structured part of the lecture. The slides use three questions as three columns and fill methods in step by step. The table comes from the ACL 2023 retrieval-based LM tutorial slides, extended with 2025 methods. The table below fills only the cells where the slides place a method; the rest are "—":

MethodWhat to retrieveHow to use itWhen to retrieveSlide note
RAG (Lewis et al., 2020)Text chunksInput layerOnce, at the start of generationTrains retriever and generator end to end (p.33)
RETRO (Borgeaud et al., 2022)Text chunksIntermediate layers—More efficient than the input layer, but requires training (p.35)
Retrieval-in-context (Ram et al.; Shi et al., 2023)——Every n tokensRetrieving more often helps, but inference gets slower (p.38)
FLARE (Jiang et al., 2023)——AdaptivelyGenerates first, retrieves when model certainty is low (p.40)
Search-R1 (Jin et al., 2025)——AdaptivelyRL learns when to search; reward is answer correctness (p.42)
AdaSearch (Lin et al., 2025)——AdaptivelyRL for "knowing what you know"; reward is answer correctness plus correct decisions (p.43)
kNN-LM (Khandelwal et al., 2020)TokensOutput layerEvery tokenRetrieves similar examples and uses their next token; finer-grained and compute-efficient, but space-expensive (p.45)
Efficient NN-LM (He et al., 2021)TokensOutput layerAdaptivelyFixes kNN-LM's per-token retrieval cost (p.47)

How to read the table: every step right or down trades effectiveness against cost at a new point. RETRO gains efficiency but needs retraining. Retrieving every n tokens helps quality but slows inference. Search-R1 and AdaSearch let the model learn for itself whether to look something up.

Retrieval sources and choosing an approach

Pages 49–50 cover retrieval sources. Unstructured data is text. Semi-structured data is text plus tables, such as PDFs; the problems are data corruption and extracting table content, and converting tables to text for text-only RAG is suboptimal. Structured data means knowledge graphs. Page 50 cites the GraphRAG survey (Peng et al., 2024): knowledge graphs help more on reasoning-intensive tasks.

Pages 51–54 cite the RAG survey of Gao et al., 2023 and use an analogy. A CS student is asked which legal rules apply to the contract between NVIDIA's headquarters and Shin Kong Life:

RouteAnalogyProsCons
PromptingAnswer from what's already in your headSimple, efficientRelies entirely on internal knowledge
RAGLook it up in law textbooks, then answerSuits dynamic environments, highly interpretableHigh latency, depends on retrieval quality
Fine-tuningTake a related course first, then answerDeep customization of behavior and styleHigh cost of retraining

Page 54 concludes that choosing between RAG and fine-tuning depends on data dynamics, customization needs, and available compute.

Pages 55–58 add four ways to train retrieval-augmented LLMs (training-free, independent, sequential, and joint, citing the RA-LLM survey), the range of applications, and six advantages of RAG: accuracy, up-to-date knowledge, interpretability, customization, safety and privacy, and cost (no model update).

HW3: only the title is public

The HW 3 button in the course page's 10/13 row links straight to ADL 2025 Fall Homework 3 on YouTube. What can be confirmed:

  • Title: the video description reads "Retriever & Reranker Training for RAG".
  • Length 20:34, uploaded 2025-10-12.
  • Page 11 of Course Logistics lists the third assignment's topic as RAG.

What is missing: the video has no subtitles, and there are no public spec slides or written instructions. This post cannot confirm the dataset, corpus, base model, baseline, metric, submission format, or deadline, so it states none of them. Submission goes through NTU COOL, which needs an NTU account.

How to use it from outside NTU: the title lines up with pp.15–29 of the slides. Pick a public QA dataset, start with a BM25 baseline, fine-tune a dense retriever with the p.16 contrastive loss, then add a pointwise or pairwise reranker and compare retrieval metrics at each stage. You will define your own grading and cannot compare with the official one.

After this lecture you should be able to

  • Name the four limits of parametric memory and the three kinds of knowledge RAG targets.
  • Draw RAG's four components and explain the difference between sparse and dense retrieval.
  • Explain pointwise vs. pairwise reranking, and why the slides consider pairwise more accurate.
  • Place RAG, RETRO, FLARE, and Search-R1 in the right cells of the what/how/when table.

One thing to try tonight: pick an obscure person or an internal document you know well and ask any LLM about it twice, once directly and once with the relevant passage pasted into the prompt. You will see the difference between pp.3 and 4 yourself. Then swap the pasted passage for an irrelevant one, and you will see what p.23 means by over-relying on retrieved text.

Further reading

Next: NLG: Decoding, Control, and Evaluation

References