🌏 中文版
Version note: This post is based on slides 60–125 of W11_RAG.pdf from Hung-Yu Kao's Natural Language Processing course at National Tsing Hua University (NTHU), Fall 2025, plus the Chinese caption tracks of the W11 Tuesday and W11 Thursday recordings. The lectures are in Mandarin. I checked every fact against the official materials on 2026-09-30. Access level A3: slides and recordings are public. This unit has no assignment of its own; the hands-on part is in the RAG labs and HW4.
Series: previous RAG, Part 1: hallucination and retrievers | next LLM API lab | Series overview
Where this picks up, and what the recordings cover
The previous post finished the retriever story, from BM25 to DPR. All of it was about finding the right passage. Slide 60, "From Retrievers to QA," asks a new question: once you have the passage, who reads it, how, and what happens if you train the whole pipeline together? Its example asks what year Oppenheimer was born. The LLM without retrieval says 1967, which is actually the year he died. With a Wikipedia passage attached, it says 1904. The slide also notes that the generator is called the "reader," because QA is a reading-comprehension task.
The recordings need a word first. In the 2025 schedule, this slide deck sits in the W10 row. The W11 row has no slides, only two recordings. I read the caption tracks of both W11 recordings:
| Recording | Length | Content (from captions) |
|---|---|---|
| W11 Tuesday (2025-11-11) | 1:42:39 | A term-project topic change, then retrievers: BoW, TF-IDF, BM25, whether CLS is enough, Siamese networks, SimCSE, DPR, GTR. It stops at the retriever/reader boundary |
| W11 Thursday (2025-11-13) | 48:53 | The first 19 minutes cover term-project checkpoints and peer-review rules, and announce that HW4 is pushed back a week. From about minute 20: ORQA, ICT, REALM. The last few minutes start the RAG paper, with the rest promised for next week |
So the first three sections below, from BERT reading comprehension to the RAG paper, can be checked against a recording. Slides 76–125, from REPLUG onward, are not covered in either W11 recording. The W12 Tuesday recording (XGWuVpVTwTQ) is the likeliest continuation, but it has no caption track and I could not confirm its content. The second half of this post relies on the slides alone.
When the reader was BERT: ORQA and REALM
The slides split open-domain QA into two branches: the reader is an encoder (BERT), or the reader is a generator (such as BART). The first branch comes first.
BERT for reading comprehension (slide 63): concatenate the question and the passage with [SEP] and feed them to BERT. Each passage position's hidden state goes through a linear layer that outputs start logits and end logits. The argmax of each gives the answer's start and end. The answer is always a span copied from the passage.
ORQA (Lee et al., ACL 2019, slides 64–69) extends this to open-domain QA with three BERTs:
- BERT_Q encodes the query and BERT_B encodes evidence blocks. A dot product between them picks the top-k blocks
- BERT_R is the reader and extracts the answer from those blocks. The slide's example asks what the ZIP in ZIP code stands for: "Zone Improvement Plan"
The key is how the retriever gets trained in the first place. ORQA uses the Inverse Cloze Task. A normal cloze task masks a word and asks the model to guess it. ICT flips that: take a random sentence out of a passage as the query and ask the model to find the passage it came from among many blocks. It is continued pre-training starting from pre-trained BERT, and it needs no labels. Pre-training trains BERT_Q and BERT_B; fine-tuning trains BERT_Q and BERT_R.
In the recording, Kao connects ICT to RAG today. The data you want to retrieve is usually new to both the embedding model and the LLM, and it is full of domain terms. A task like ICT lets the model get used to what your data looks like. He also warns that fine-tuning an encoder yourself is hard to get right. Done badly, it can score worse than using a pre-trained model with mean pooling, so with ordinary resources his group does not fine-tune its own.
REALM (Guu et al., ICML 2020, slides 70–73) replaces ICT with MLM pre-training. The input sentence contains [MASK]; the retriever finds the 5 most similar documents and appends them; a knowledge-augmented encoder fills in the blank. The MLM loss updates the language model and also flows back into the retriever's encoder, so the retriever gradually learns which documents help most with the prediction.
The slides list three difficulties: the corpus is all of Wikipedia, so the index lives at the embedding level; the input itself contains [MASK]; and the whole pipeline trains end to end. Kao's take in the recording: you can wire it up and hit train, but getting it to train well is another matter.
The reader becomes a generator: RAG and REPLUG
RAG (Lewis et al., NeurIPS 2020, slide 75) is the first paper to use the name "RAG." The slide makes three points: no pre-training, only fine-tuning on open-domain QA; the retriever and generator are fine-tuned jointly; and MIPS (Maximum Inner-Product Search) speeds up vector search, supported by packages like FAISS. The same slide recommends reading FiD (Fusion-in-Decoder) alongside it.
This is where the W11 Thursday lecture ends. Kao's verdict: the retrieve-then-generate flow already existed in the earlier papers. What the RAG paper did was swap in a better generator and a better retriever.
REPLUG (Shi et al., NAACL 2024, slides 76–79) handles the case where the LLM is too large to train:
- At inference: each retrieved document is concatenated with the input and sent to the LLM separately. The resulting output distributions are then combined
- REPLUG LSR: the LLM stays frozen and only the retriever trains. The retriever's distribution over documents, P_R(d|x), is pushed toward Q(d|x), the distribution of which documents best help the LLM generate a good answer
Slide 79, comparing REPLUG with REPLUG LSR, leaves a question: why evaluate REPLUG with BPB (bits per byte)? The slides give no answer. It is worth coming back to after reading the paper.
Seven recent fixes in three groups
Slide 80 maps out the rest of the deck in one table:
| Group | Method | Venue (per the slides) |
|---|---|---|
| Enhancing retrieval | Query Rewriting | EMNLP 2023 |
| Enhancing retrieval | HyDE | ACL 2023 |
| Enhancing RAG | RetRobust | ICLR 2024 |
| Enhancing RAG | RAFT | COLM 2024 |
| Enhancing RAG | Self-RAG | ICLR 2024 |
| Enhancing RAG | RAAT | ACL 2024 |
| Continual retrieval | FLARE | EMNLP 2023 |
Better retrieval: make the query look like a document
Query Rewriting (Ma et al., slides 81–84) starts from the gap between the input text and the knowledge you actually need to look up. The example is "What profession does Nicholas Ray and Elia Kazan have in common?" Sent as is, it retrieves poorly. Rewritten as two queries, "Nicholas Ray profession" and "Elia Kazan profession," it finds each person's bio.
HyDE (Gao et al., slides 85–86) goes further. An LLM (InstructGPT, per the slide) writes a fake answer passage first, and an unsupervised contrastively trained encoder uses that passage to find real documents. The slide has a side note in Chinese asking whether this suits every task and query. The next slide cites Wang et al.'s 2024 best-practices study, which found HyDE can beat Query Rewriting on retrieval tasks. The site's HyDE post covers the engineering side.
A sturdier generator: don't get led astray by noise
RetRobust (Yoran et al., slides 87–91) starts with the problem. Retrieval can improve results, but it hurts on StrategyQA and Fermi, and random passages make things much worse. One slide explains StrategyQA: each question needs several unstated reasoning steps. "Could a crocodile run a marathon?" requires knowing a marathon is about 42 km and crocodiles are semi-aquatic. RetRobust keeps the retriever fixed and trains the generator on QA data that mixes relevant and irrelevant documents.
RAFT (Zhang et al., slides 92–93) also fixes the retriever and trains the LLM on both correct and incorrect documents. The slides point out the difference from RetRobust: RAFT's answers include the reasoning (a CoT answer).
RAAT (Fang et al., slides 102–105) splits noise into three kinds: relevant noise that lacks the answer, irrelevant noise, and counterfactual noise. It uses adversarial training, fine-tuning on the noise that hurts the model most, plus an auxiliary task that detects the noise type. The slides leave two questions here: how do you get these annotations, and what has the model learned in any physical sense?
Deciding when to retrieve: FLARE and Self-RAG
FLARE (Jiang et al., slides 94–101, titled "Active Retrieval Augmented Generation") starts from an observation. Traditional RAG retrieves once from the query before generating, and extra information can confuse the model. Asked to summarize Joe Biden, the retrieval results include his first wife's birthday and a quote from a speech. FLARE argues the model should retrieve only when it lacks knowledge, and the query should reflect what it is about to write. To detect a knowledge gap, it relies on language models being fairly well calibrated: when the next token's probability falls below a threshold θ, retrieval fires. In the slide's example, the model writes "Joe Biden attended," loses confidence, searches "Joe Biden University," and then writes "the University of Pennsylvania."
Self-RAG (Asai et al., slides 106–110) has the model critique and reflect on itself. For training, GPT-4 first labels reflection tokens; those labels are distilled into a critic model, which then produces training data for the generator. At inference the model emits [Retrieve], [IsRel], [IsSup] (is the answer supported by the passage), and [IsUse] (is it useful), keeping the top-B candidate segments at each step. The slides stress that these token settings can be adjusted at inference to trade precision against fluency, or accuracy against retrieval frequency. The site's Self-RAG post makes a good companion.
Challenges in modern RAG: noise and four abilities
Slides 111–116 narrow down to two challenges: the web is full of noise and even fake news, and we still don't understand well how much each model gains from retrieval. Four noise types are listed: semantically similar but missing the answer, counterfactual information, irrelevant information, and "black box digestion" (you can't see how the model absorbs the documents).
Four examples then show the abilities an LLM needs inside RAG:
| Ability | The slide's example |
|---|---|
| Noise Robustness | Asked about the 2022 Nobel Prize in Literature, with both the 2022 and 2021 winners in the documents, answer Annie Ernaux |
| Negative Rejection | With only the 2021 and 2020 winners retrieved, say there isn't enough information |
| Information Integration | Asked when the ChatGPT iOS app and the API launched, combine answers from two documents |
| Counterfactual Robustness | The documents wrongly say the 2004 Olympics were in New York; when warned they may contain errors, the model should flag the mistake and answer Athens |
Last stop: generative retrieval and RRG
Slides 117–125 introduce Generative Information Retrieval, citing a 2025 GenIR survey and an EMNLP 2023 paper on scaling generative retrieval to millions of passages. Several figures were generated with NotebookLM. The core is two new processes:
- Generative Document Retrieval (GR): instead of computing vector similarity, train an LLM to map a query directly to a DocID. The slide asks how GR differs from traditional retrieval
- Reliable Response Generation (RRG): the slides split approaches into strengthening internal knowledge and augmenting external knowledge
One slide title sums it up: "We need solution, not just documents." Users want answers, not just documents. The last slide, on RRG evaluation, is a figure with no text.
How to self-study this unit
- Read slides 60–75 first, with the W11 Thursday recording from about minute 20. The ORQA → REALM → RAG line is a story about which component gets trained. Once it clicks, every later fix has a place to go.
- Use the slide 80 table as a table of contents. For each method, ask: does it change the retriever, the generator, or the interface between them?
- The open questions in the slides (BPB, where HyDE applies, where RAAT's labels come from, how GR differs) have no official answers. They make good study-group prompts.
One thing to try tonight: take a RAG system or ChatGPT conversation you already use and write one test for each of the four abilities on slides 113–116. For example, give it only outdated documents and see whether it refuses to answer. After four tests you'll know more about where your system breaks than seven papers would tell you.
Further reading
- The same topic from another course: CS224N Lecture 10: six components of RAG and language agents
- Retrieval evaluation and neural IR: CS224U: information retrieval
- The engineering landscape: RAG patterns: a complete guide
References
- W11_RAG.pdf (Fall 2025) — slides 60–125: BERT reading comprehension, ORQA/ICT, REALM, RAG, REPLUG, the slide 80 method table, seven fixes, noise and four abilities, GenIR
- NTHU NLP 2025 schedule — W11_RAG.pdf is in the W10 row; the W11 row has only two recordings
- W11 Tuesday recording (Fall 2025, in Mandarin) — the retriever part, ending at the retriever/reader boundary
- W11 Thursday recording (Fall 2025, in Mandarin) — term-project and HW4-delay announcements, ORQA, ICT, REALM, start of the RAG paper
- W12 Tuesday recording (Fall 2025, in Mandarin) — no caption track; content not confirmed here
- Lee et al. 2019, ORQA
- Guu et al. 2020, REALM
- Lewis et al. 2020, RAG
- Izacard & Grave, FiD
- Shi et al. 2024, REPLUG
- Ma et al. 2023, Query Rewriting
- Gao et al. 2023, HyDE
- Wang et al. 2024, Searching for Best Practices in RAG
- Yoran et al. 2024, RetRobust
- Geva et al. 2021, StrategyQA
- Zhang et al. 2024, RAFT
- Jiang et al. 2023, FLARE
- Fang et al. 2024, RAAT
- Asai et al. 2024, Self-RAG
- From Matching to Generation: A Survey on Generative Information Retrieval
- How Does Generative Retrieval Scale to Millions of Passages?
Loading...