🌏 中文版
Version note: This post is based on materials from Hung-Yu Kao's Natural Language Processing course at National Tsing Hua University (NTHU), Fall 2025: rag_tutorial_1.pdf (cover dated 2024/11/28), rag_tutorial_2.pdf (cover dated 2024/12/05), RAG_tutorial_1.ipynb, RAG_lab_2, and the handout PDF,
main.ipynb,cat-facts.txt, andquestions_answers.txtin Assignment4. Both TA slide decks are reused from 2024. The recordings are W13's RAG1 and RAG2 plus the HW4 walkthrough, all in Mandarin. None has a caption track, so I did not check them section by section. I checked every fact against the official materials on 2026-09-30. Access level A3: the handout, starter code, and data are public. What's missing is NTU COOL for submission, the grading script, and solutions.
Series: previous LLM API lab | next Course summary and LLM reasoning notes | Series overview
Timeline: the assignment came before the labs
The 2025 schedule lists HW4 and its walkthrough in the W12 row; the walkthrough was uploaded on 2025-11-20. The two RAG labs are in W13 (2025-11-24 and 11-26). At the start of the W11 Thursday lecture, Kao said HW4 was meant to go out that week but was pushed back one week because the RAG material and the lab videos weren't ready. The handout gives three weeks to finish.
So the real order was: finish the RAG lectures, get the assignment, then attend the labs. For self-study, follow this post's order: lab 1, lab 2, HW4.
Lab 1: a minimal RAG with LangChain + Ollama
Running a local model on Colab
The slides introduce two tools. LangChain is a framework for building LLM apps. Ollama runs LLMs locally; the slides list its strengths as easy install, many supported models, automatic GPU use, LangChain integration, and data that never leaves for a third party.
Setup steps:
- On Colab, choose Python 3 with a T4 GPU (the slides warn free GPU time is limited)
- Request access to Llama-3.2-1B-Instruct on Hugging Face and log in with an access token. The notebook notes that Ollama alone needs no login
pip install colab-xterm,%load_ext colabxterm, then%xtermto open a terminal in a cell- In the terminal, run
curl -fsSL https://ollama.com/install.sh | sh,ollama serve, andollama pull llama3.2:1b. Idle too long and the connection drops, so rerunollama serve
Five components
The notebook uses llama3.2:1b for generation and jinaai/jina-embeddings-v2-base-en for embeddings. The knowledge base is just six sentences about Florida and Miami Dade College. The components, in order:
- LLM and embeddings:
Ollama(model=MODEL)andHuggingFaceEmbeddings(...)withnormalize_embeddingsset to False. A slide note explains that cosine similarity already normalizes during computation, so normalizing beforehand is redundant - Prompt: a
ChatPromptTemplatewhose system message says to answer from the given context, say you don't know if you don't, and use at most three sentences; the human message is{input} - Vector store: the slides compare Chroma (a vector database with metadata filtering) and FAISS (Meta's similarity-search library, using IVF or HNSW for approximate nearest-neighbor search). The notebook uses Chroma, then
.as_retriever(search_type="mmr", search_kwargs={"k": 3, "fetch_k": 5}): fetch 5 documents and let MMR pick 3 - Chain:
create_stuff_documents_chainstuffs documents into the prompt, andcreate_retrieval_chainputs retrieval in front - Run:
chain.invoke({"input": query}), passing only the question
The site has a dedicated post on MMR.
Lab 2: taking the retriever apart without LangChain
Slide 4 states the goal: apart from data prep, skip LangChain and see what the database, the retriever, and generation each do.
Data prep: WebBaseLoader scrapes a LinkedIn article about RAG by class name. The text is cleaned of newlines, repeated punctuation, URLs, HTML tags, and special characters, then TokenTextSplitter cuts it with the embedding model's tokenizer into 100-token chunks with 20 tokens of overlap. The slides give two reasons to chunk: the top-k chunks plus the query must fit the length limit, and long passages drag irrelevant content into the answer.
Building the stores: each chunk gets an id and is saved to text_db.json (text) and vector_db.json (text plus vector). For BM25, the text is lowercased and tokenized, indexed with rank_bm25's BM25Okapi, and pickled. The slides stress that real systems hold millions of chunks, so saving and loading efficiently matters.
Three retrievers:
| Retriever | How it works |
|---|---|
| Dense | Cosine similarity between the query vector and each chunk vector, sorted by score |
| Sparse | BM25 scores for each chunk, sorted |
| Hybrid | RRF (Reciprocal Rank Fusion) merges the two rankings: each document scores 1/(k + dense rank) + 1/(k + sparse rank), with k defaulting to 60. A document missing from one list gets that list's last rank plus one |
personal_retriever() wraps all three into one function that returns the top of the hybrid ranking. The site's hybrid search post goes deeper on the engineering.
Generation: transformers loads meta-llama/Llama-3.2-1B-Instruct in float16, the question and retrieved chunks fill a prompt (noted as based on LangChain hub's rlm/rag-prompt), and it generates with max_new_tokens=300. The slides note that inference time grows in proportion to max_new_tokens. The notebook also runs once without context, so you can compare with and without RAG.
Retrieval evaluation: switching to the cat-facts data HW4 uses, it checks for each question whether the gold sentence appears in the top 3, and reports Recall@3 plus a score it calls precision. Slide 29 explains with an example: when the single correct answer ranks first, recall at top 1, 2, and 3 is 1/1 each time, while precision is 1/1, 1/2, and 1/3.
Two things to watch in the code
personal_retriever()takes atopkparameter but hard-codestopk = 3inside, so passing a different value does nothing. HW4 asks for recall@5, so fix this if you reuse it- The evaluation loop's "precision" adds 1/(j+1) when the hit is at rank j, then averages. That is mean reciprocal rank (MRR@3), not precision@k as usually defined
HW4: a RAG for cat facts
The task
From the handout:
- Task: QA with RAG, short answers
- Database: cat-facts from ngxson/demo_simple_rag_py, 150 facts, one sentence each. For example: "The technical term for a cat's hairball is a "bezoar.""
- Test set: 150 QA pairs, which the handout says were generated by GPT-5, in the same order as the facts.
questions_answers.txthas one line of question and one line of answer per pair, with short answers like "Two thirds" or "Taste mutation" - Constraint: the generator is Llama-3.2-1B, frozen; no fine-tuning
- This time you may modify the code template
TODOs and weights
| Item | What | Weight |
|---|---|---|
| TODO1 | Set up Ollama: colab-xterm, install Ollama, pull llama3.2:1b, ollama serve | 5% |
| TODO2 | Load cat-facts and build a Chroma retrieval store from Document objects | 10% |
| TODO3 | Write the system prompt | 10% |
| TODO4 | Build and run the stuff-documents chain and retrieval chain. Not using jina-embeddings-v2-base-en with Llama3.2-1b halves this score | 10% |
| TODO5 | Improve the system so the LLM answers the 150 questions correctly (still Llama3.2-1b) | 10% |
| Report | See below | 55% |
For the code part, you submit a JSON file where each entry has Query, Ground_Truth, and Prediction, and the report includes a screenshot of the test log and accuracy. Report recall@1 and recall@5 for retrieval and exact match for generation. Missing any of these halves TODO5.
Report questions:
- (5%) Describe your RAG system: its components, your prompt, and what you added beyond the lab code
- (10%) How different prompts change performance
- (10%) How different input data formats for the retriever change performance
- (10%) How the order of retrieved documents fed to the generator changes performance
- (10%) How generation holds up when counterfactual information is added to the generator's input
- (10%) Anything else that strengthens the report
The last three map directly onto the noise types and four abilities in RAG, Part 2. The counterfactual question is the lecture's counterfactual robustness, now yours to measure.
Submission rules
Submit one zip: code (.py or .ipynb), predictions (.json), requirements.txt, and the report (.docx or .pdf), named NLP_HW4_school_studentID. The report must state the environment and Python version. If you use generative AI, say so in both code comments and the report, and link any code taken from the web. Highly similar submissions lose 100 points each. Submission goes through NTU COOL, which outside readers can't access.
Inconsistencies in the materials
Before you start, know where these don't line up:
- Evaluation: the handout says exact matching, but a comment in
main.ipynbsays a response counts as correct if the answer shows up in it, which is far looser. A 1B model rarely outputs just "Two thirds," so the choice swings the score a lot. State in your report which one you used - Question count: two handout slides and the notebook's TODO5 comment say "ten questions," while the grading table says 150 and the data file has 150 pairs. Go with 150
- Penalty table: that slide's filename examples say
NLP_HW3_..., carried over from the previous assignment. Follow the filenames on the HW4 slide - Package versions:
main.ipynbpins LangChain to the 0.2 series (langchain>=0.2.0,<0.3.0and friends), while the lab notebooks don't pin at all. Code written from the labs may hit import-path differences in the HW4 environment
How to self-study this unit
- Run the lab 1 notebook first and make sure Ollama starts on Colab. That's the step HW4 most often gets stuck on.
- Read lab 2's
helper_functions.py, which holds the functions HW4's retrieval evaluation uses, and fix the hard-codedtopkwhile you're there. - For HW4, change nothing at first. Get baseline recall@1, recall@5, and EM, then change one variable at a time (prompt, data format, document order). That is exactly the analysis the report asks for.
One thing to try tonight: download cat-facts.txt and questions_answers.txt and, with no model at all, run BM25 over the 150 questions and compute recall@1. That number is your retrieval floor. Whatever embedding or hybrid setup you try later has to beat it to count as progress.
Further reading
- Chunking: Chunking strategies decide whether RAG finds the answer
- Evaluating RAG: RAG evaluation frameworks and tool selection
- Another course's RAG assignment: CS224U Assignment 2: OpenQA and DSPy
References
- rag_tutorial_1.pdf (cover dated 2024/11/28) — LangChain, Ollama, Colab setup, Chroma/FAISS, MMR, retrieval chain
- rag_tutorial_2.pdf (cover dated 2024/12/05) — data prep, chunking, dense/sparse/hybrid retrieval, RRF, generation, retrieval evaluation
- RAG_tutorial_1.ipynb — the LangChain RAG
- RAG_lab_2 (RAG_tutorial_2.ipynb, helper_functions.py) — the hand-built RAG and cat-facts retrieval evaluation
- Assignment4 — NTHU_NLP_HW4_RAG.pdf, main.ipynb, cat-facts.txt, questions_answers.txt, report template
- NTHU NLP 2025 schedule — HW4 in the W12 row, RAG labs in the W13 row
- HW4 walkthrough (in Mandarin) — uploaded 2025-11-20
- W13 Tuesday: RAG1 and W13 Thursday: RAG2 (in Mandarin)
- W11 Thursday recording (in Mandarin) — opens with the one-week HW4 delay
- ngxson/demo_simple_rag_py (Hugging Face) — original source of cat-facts
- meta-llama/Llama-3.2-1B-Instruct — gated; requires an access request
- Ollama llama3.2 — the generator HW4 requires
Loading...