Skip to content

NTHU NLP RAG Labs + HW4: Building a Cat-Facts RAG Two Ways, with LangChain and by Hand

Sep 30, 20261 min
TL;DREach of the two RAG TA sessions builds one version. The first installs Ollama on Colab to run llama3.2:1b and wires up a minimal RAG with LangChain's Chroma, MMR, and retrieval chain. The second uses LangChain only for data prep and writes the rest by hand: chunking, text and vector stores, hybrid BM25 + cosine retrieval merged with RRF, then generation with Llama-3.2-1B-Instruct. HW4 applies the first session's skeleton to 150 cat facts and 150 GPT-5-generated QA pairs. The generator must be Llama3.2-1b and the embedding model jina-embeddings-v2-base-en, and you report recall@1, recall@5, and exact match. Code is 45% of the grade and the report 55%; the report analyzes how prompts, data format, document order, and counterfactual information change the results.

🌏 中文版

Version note: This post is based on materials from Hung-Yu Kao's Natural Language Processing course at National Tsing Hua University (NTHU), Fall 2025: rag_tutorial_1.pdf (cover dated 2024/11/28), rag_tutorial_2.pdf (cover dated 2024/12/05), RAG_tutorial_1.ipynb, RAG_lab_2, and the handout PDF, main.ipynb, cat-facts.txt, and questions_answers.txt in Assignment4. Both TA slide decks are reused from 2024. The recordings are W13's RAG1 and RAG2 plus the HW4 walkthrough, all in Mandarin. None has a caption track, so I did not check them section by section. I checked every fact against the official materials on 2026-09-30. Access level A3: the handout, starter code, and data are public. What's missing is NTU COOL for submission, the grading script, and solutions.

Series: previous LLM API lab | next Course summary and LLM reasoning notes | Series overview

Timeline: the assignment came before the labs

The 2025 schedule lists HW4 and its walkthrough in the W12 row; the walkthrough was uploaded on 2025-11-20. The two RAG labs are in W13 (2025-11-24 and 11-26). At the start of the W11 Thursday lecture, Kao said HW4 was meant to go out that week but was pushed back one week because the RAG material and the lab videos weren't ready. The handout gives three weeks to finish.

So the real order was: finish the RAG lectures, get the assignment, then attend the labs. For self-study, follow this post's order: lab 1, lab 2, HW4.

Lab 1: a minimal RAG with LangChain + Ollama

Running a local model on Colab

The slides introduce two tools. LangChain is a framework for building LLM apps. Ollama runs LLMs locally; the slides list its strengths as easy install, many supported models, automatic GPU use, LangChain integration, and data that never leaves for a third party.

Setup steps:

  1. On Colab, choose Python 3 with a T4 GPU (the slides warn free GPU time is limited)
  2. Request access to Llama-3.2-1B-Instruct on Hugging Face and log in with an access token. The notebook notes that Ollama alone needs no login
  3. pip install colab-xterm, %load_ext colabxterm, then %xterm to open a terminal in a cell
  4. In the terminal, run curl -fsSL https://ollama.com/install.sh | sh, ollama serve, and ollama pull llama3.2:1b. Idle too long and the connection drops, so rerun ollama serve

Five components

The notebook uses llama3.2:1b for generation and jinaai/jina-embeddings-v2-base-en for embeddings. The knowledge base is just six sentences about Florida and Miami Dade College. The components, in order:

  • LLM and embeddings: Ollama(model=MODEL) and HuggingFaceEmbeddings(...) with normalize_embeddings set to False. A slide note explains that cosine similarity already normalizes during computation, so normalizing beforehand is redundant
  • Prompt: a ChatPromptTemplate whose system message says to answer from the given context, say you don't know if you don't, and use at most three sentences; the human message is {input}
  • Vector store: the slides compare Chroma (a vector database with metadata filtering) and FAISS (Meta's similarity-search library, using IVF or HNSW for approximate nearest-neighbor search). The notebook uses Chroma, then .as_retriever(search_type="mmr", search_kwargs={"k": 3, "fetch_k": 5}): fetch 5 documents and let MMR pick 3
  • Chain: create_stuff_documents_chain stuffs documents into the prompt, and create_retrieval_chain puts retrieval in front
  • Run: chain.invoke({"input": query}), passing only the question

The site has a dedicated post on MMR.

Lab 2: taking the retriever apart without LangChain

Slide 4 states the goal: apart from data prep, skip LangChain and see what the database, the retriever, and generation each do.

Data prep: WebBaseLoader scrapes a LinkedIn article about RAG by class name. The text is cleaned of newlines, repeated punctuation, URLs, HTML tags, and special characters, then TokenTextSplitter cuts it with the embedding model's tokenizer into 100-token chunks with 20 tokens of overlap. The slides give two reasons to chunk: the top-k chunks plus the query must fit the length limit, and long passages drag irrelevant content into the answer.

Building the stores: each chunk gets an id and is saved to text_db.json (text) and vector_db.json (text plus vector). For BM25, the text is lowercased and tokenized, indexed with rank_bm25's BM25Okapi, and pickled. The slides stress that real systems hold millions of chunks, so saving and loading efficiently matters.

Three retrievers:

RetrieverHow it works
DenseCosine similarity between the query vector and each chunk vector, sorted by score
SparseBM25 scores for each chunk, sorted
HybridRRF (Reciprocal Rank Fusion) merges the two rankings: each document scores 1/(k + dense rank) + 1/(k + sparse rank), with k defaulting to 60. A document missing from one list gets that list's last rank plus one

personal_retriever() wraps all three into one function that returns the top of the hybrid ranking. The site's hybrid search post goes deeper on the engineering.

Generation: transformers loads meta-llama/Llama-3.2-1B-Instruct in float16, the question and retrieved chunks fill a prompt (noted as based on LangChain hub's rlm/rag-prompt), and it generates with max_new_tokens=300. The slides note that inference time grows in proportion to max_new_tokens. The notebook also runs once without context, so you can compare with and without RAG.

Retrieval evaluation: switching to the cat-facts data HW4 uses, it checks for each question whether the gold sentence appears in the top 3, and reports Recall@3 plus a score it calls precision. Slide 29 explains with an example: when the single correct answer ranks first, recall at top 1, 2, and 3 is 1/1 each time, while precision is 1/1, 1/2, and 1/3.

Two things to watch in the code

  • personal_retriever() takes a topk parameter but hard-codes topk = 3 inside, so passing a different value does nothing. HW4 asks for recall@5, so fix this if you reuse it
  • The evaluation loop's "precision" adds 1/(j+1) when the hit is at rank j, then averages. That is mean reciprocal rank (MRR@3), not precision@k as usually defined

HW4: a RAG for cat facts

The task

From the handout:

  • Task: QA with RAG, short answers
  • Database: cat-facts from ngxson/demo_simple_rag_py, 150 facts, one sentence each. For example: "The technical term for a cat's hairball is a "bezoar.""
  • Test set: 150 QA pairs, which the handout says were generated by GPT-5, in the same order as the facts. questions_answers.txt has one line of question and one line of answer per pair, with short answers like "Two thirds" or "Taste mutation"
  • Constraint: the generator is Llama-3.2-1B, frozen; no fine-tuning
  • This time you may modify the code template

TODOs and weights

ItemWhatWeight
TODO1Set up Ollama: colab-xterm, install Ollama, pull llama3.2:1b, ollama serve5%
TODO2Load cat-facts and build a Chroma retrieval store from Document objects10%
TODO3Write the system prompt10%
TODO4Build and run the stuff-documents chain and retrieval chain. Not using jina-embeddings-v2-base-en with Llama3.2-1b halves this score10%
TODO5Improve the system so the LLM answers the 150 questions correctly (still Llama3.2-1b)10%
ReportSee below55%

For the code part, you submit a JSON file where each entry has Query, Ground_Truth, and Prediction, and the report includes a screenshot of the test log and accuracy. Report recall@1 and recall@5 for retrieval and exact match for generation. Missing any of these halves TODO5.

Report questions:

  • (5%) Describe your RAG system: its components, your prompt, and what you added beyond the lab code
  • (10%) How different prompts change performance
  • (10%) How different input data formats for the retriever change performance
  • (10%) How the order of retrieved documents fed to the generator changes performance
  • (10%) How generation holds up when counterfactual information is added to the generator's input
  • (10%) Anything else that strengthens the report

The last three map directly onto the noise types and four abilities in RAG, Part 2. The counterfactual question is the lecture's counterfactual robustness, now yours to measure.

Submission rules

Submit one zip: code (.py or .ipynb), predictions (.json), requirements.txt, and the report (.docx or .pdf), named NLP_HW4_school_studentID. The report must state the environment and Python version. If you use generative AI, say so in both code comments and the report, and link any code taken from the web. Highly similar submissions lose 100 points each. Submission goes through NTU COOL, which outside readers can't access.

Inconsistencies in the materials

Before you start, know where these don't line up:

  1. Evaluation: the handout says exact matching, but a comment in main.ipynb says a response counts as correct if the answer shows up in it, which is far looser. A 1B model rarely outputs just "Two thirds," so the choice swings the score a lot. State in your report which one you used
  2. Question count: two handout slides and the notebook's TODO5 comment say "ten questions," while the grading table says 150 and the data file has 150 pairs. Go with 150
  3. Penalty table: that slide's filename examples say NLP_HW3_..., carried over from the previous assignment. Follow the filenames on the HW4 slide
  4. Package versions: main.ipynb pins LangChain to the 0.2 series (langchain>=0.2.0,<0.3.0 and friends), while the lab notebooks don't pin at all. Code written from the labs may hit import-path differences in the HW4 environment

How to self-study this unit

  1. Run the lab 1 notebook first and make sure Ollama starts on Colab. That's the step HW4 most often gets stuck on.
  2. Read lab 2's helper_functions.py, which holds the functions HW4's retrieval evaluation uses, and fix the hard-coded topk while you're there.
  3. For HW4, change nothing at first. Get baseline recall@1, recall@5, and EM, then change one variable at a time (prompt, data format, document order). That is exactly the analysis the report asks for.

One thing to try tonight: download cat-facts.txt and questions_answers.txt and, with no model at all, run BM25 over the 150 questions and compute recall@1. That number is your retrieval floor. Whatever embedding or hybrid setup you try later has to beat it to count as progress.

Further reading

References