🌏 中文版
This post is based on the 1132 semester (Spring 2025) of NCCU Yen-Lung Tsai's Generative AI: Text and Image Synthesis Principles and Practice. It is part 8 of the Reading NCCU Yen-Lung Tsai Generative AI series and follows L07, Building Your Own Chatbot. The L07 chatbot only knows what the model saw in training. This lecture makes it read your documents before answering.
Official sources: video 08 (2025-04-08, about 3 h 4 min, in Mandarin), the 25-page GenAI08 slides (in the instructor's slide folder), the notebooks 【Demo06a】RAG01_打造向量資料庫, 【Demo06b】RAG02_打造_RAG_系統, and 【Demo06】用_RAG_打造心靈處方籤機器人 in the AI-Demo repo, and the week-8 assignment on the Chang Gung satellite course page. Access level: A3.
The notebook drift is especially visible this week, and I flag each case below: the current repo version no longer matches what's on screen in the video.
Where this week sits
Session one of video 08 opens with an aside on Llama 4 (18:44), then covers what RAG is, how it works, memory, and finance applications (28:08–57:40), and starts on program A, the vector database. Session two builds and saves the database, introduces the embedding model, uploads it to the cloud with a direct download link, then writes program B, designs the prompt, builds the Gradio app, and explains the assignment at 2:01. Session three has lightning talks from NTHU and NCCU students and a TA segment.
The slides have three parts, "RAG", "Vector databases", and "Applications of RAG", followed by the assignment.
The Llama 4 aside is slide 2: Behemoth is a 2TB "super teacher" model, Maverick has 128 experts, and Scout has a 10-million-token context.
Concept 1: can the computer fetch the information itself?
Slide 4 brings back the "prompts are simple" slide from L06: give it the correct information, plus clear instructions. This lecture's question: can the computer find the information part in a database automatically?
Slide 5's answer is RAG (Retrieval-Augmented Generation), a way to reduce hallucination. Recall the Karpathy quote in L06: hallucination is an LLM's nature, and what we want is an assistant that doesn't hallucinate. RAG doesn't change the model. Before each answer it puts the relevant correct material into the prompt, so the model answers from the material.
Concept 2: RAG's two phases
Preparation: every document gets a feature vector
Slide 7's diagram: documents 1, 2, and 3 each pass through the same function fθ to get feature vectors k₁, k₂, k₃. "Every important document finds its own representative feature vector." fθ is the embedding model, the same idea as L02's "a neural network is a function".
Slide 8 says to drop your files into an uploaded_docs folder. Plain text works best, but PDF and Word files are fine too.
Slide 9 handles a practical problem: a long document can't become a single vector, so you cut it into chunks. The slide's illustration uses 1,000-character chunks with 200 characters of overlap between neighbors, and adds "there's actually a lot more to it". The overlap keeps a sentence that straddles a boundary from becoming unreadable on both sides.
Use: the question becomes a vector too, and you find the nearest
Slides 10–11: when the user asks a question, it goes through the same fθ to become a vector q. You compare q with k₁, k₂, k₃, find the closest and most relevant material, and put it into the prompt. The slide's aside: "This is how AI search works these days!"
Slides 12–13 split the new prompt into two variables and give a template:
question: the user's original questionretrieved_chunks: what RAG found
請根據 {retrieved_chunks} 裡的資訊, 來回應使用者的問題: {question} (Answer the user's question {question} based on the information in {retrieved_chunks}.)
That's all RAG is. Retrieval fills in retrieved_chunks; generation is still the LLM's job.
How is "closest" computed?
The slides just say "compare which is closest" without a formula. The current Demo06a/06b set normalize_embeddings=True, scaling every vector to length 1. Then the dot product of two vectors equals their cosine similarity cos θ: the closer to 1, the more aligned the direction and the closer the meaning. FAISS finds the top k nearest vectors quickly, even among many.
Concept 3: vector databases and "memory"
Slide 16: one system can use two or more vector databases, say one for product manuals and one for support logs.
Slides 17–20 are what Tsai jokingly calls "the part where I show off". You often hear that LLMs have two kinds of memory:
- Short-term memory is the current conversation. Come back next time and everything is forgotten; worse, if the conversation runs too long, the early parts are forgotten too. That's the limit of resending the whole messages list every turn in L07.
- Long-term memory: one way to build it is RAG, turning full past conversations into a vector database. The slide's example is a custom bot greeting a regular: "Welcome back! Just so you know, the dish you used to order isn't on the menu anymore."
Slide 22 lists possible RAG applications in finance: customer service and Q&A, internal knowledge management, education and training, personalized financial advice, investment research report generation, anti-money-laundering and fraud detection, regulatory and compliance advice, and financial news and trend summaries.
This week's demo notebooks
Slide 23 splits the implementation into two programs: program A reads your data and builds the vector database (yenlung.me/AI06a); program B implements the RAG system (yenlung.me/AI06b). The split is practical: building the database is slow and happens once, while Q&A runs over and over.
Program A: Demo06a builds the vector database (current repo version)
- Create an
uploaded_docsfolder and upload.txt/.pdf/.docxfiles by hand. The notes suggest practicing on NCCU's student rewards and discipline regulations (in Mandarin). - Install
langchain,langchain-community,pypdf,python-docx,faiss-cpu, andsentence-transformers. - Load each file with
TextLoader,PyPDFLoader, orUnstructuredWordDocumentLoaderdepending on its extension. - Chunk with
RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=100). - Embed with
google/embeddinggemma-300mfrom Hugging Face, which needs aHuggingFacetoken in Colab Secrets. Following Google's recommendation, the notebook prefixes documents withtitle: none | text:and queries withtask: search result | query:. - Build with
FAISS.from_documents(...),save_local("faiss_db"), thenzipit intofaiss_db.zip.
Version drift: slide 9 illustrates 1000/200 chunking; the current notebook uses 500/100. Demo06b's section heading still says "custom E5 embedding class", while the code already uses EmbeddingGemma. For which embedding model the video introduces at 1:21, go by what's on screen. The concept doesn't change: documents and questions must go through the same fθ.
Program B: Demo06b builds the RAG system (current repo version)
- Download
faiss_db.zipfrom a public Google Drive link withgdownand unzip it. This is the "download the database from the cloud" and Google Drive direct-link part of the video (1:33–1:36). - Load it with
FAISS.load_local(...)using the sameEmbeddingGemmaEmbeddingsclass, thenas_retriever(search_kwargs={"k": 4})to fetch 4 chunks per query. - Call the LLM through AISuite on Groq; currently
model = "groq:openai/gpt-oss-120b". - The prompt has two layers. The system message makes the model "NCCU's AI self-directed learning advisor". The user message is the template "Based on the following material: {retrieved_chunks}, answer the user's question: {question}", plus "if the material isn't enough, tell the student to ask the student affairs office."
chat_with_rag()retrieves, fills the template, and calls the model;gr.Blockswraps it as an "AI school discipline advisor".
That line about what to do when the material falls short, in step 4, is worth copying. It gives the model a legitimate way out instead of inventing an answer.
Note that chat_with_rag() saves each exchange to chat_history, but the messages it sends contain only the system message and the current user turn. So it is single-turn RAG and won't remember the previous question. For multi-turn, bring in the L07 approach.
Another example: Demo06, the "spiritual prescription" bot
Demo06 is a single-notebook version. Its data is the book True Happiness by Master Sheng Yen (the notes state the copyright belongs to Dharma Drum Publishing and it is used only as an example). It builds the database with OpenAIEmbeddings and FAISS, chunks at 1000/200, and uses a RetrievalQA chain with gpt-4o. The clever part is the prompt: it first draws a random "spiritual prescription" (a short aphorism), then hands the prescription, the retrieved text, and the user's question to the model and asks for a reply in a similar voice. It shows that a prompt can carry other material you want to control besides retrieved_chunks.
The assignment: week 8 (Chang Gung satellite version)
From the Chang Gung satellite course page; NCCU's own grading differs. Slide 24's assignment: you may use the simulated NCCU club data (yenlung.me/uploaded_docs); program A builds a vector database from your data and saves it as faiss_db.zip; program B must read and unzip your faiss_db.zip; design your own prompt and build your own RAG-powered chatbot.
The Chang Gung page's task: implement a RAG system
- Prepare your own practice data.
- Adapt the instructor's code to make it your own.
- Add code that reads your zip file from a URL so TAs can run it (convert the cloud file to a downloadable link).
- Demo in Gradio.
Submit: a Colab link (only the second program), a description of your data, key screenshots, the persona/background, and Gradio conversation results. The 1132 deadline was 4/21.
Rubric (out of 10): identical to the example, 1; GPT-level or off-topic, 2; a theme close to the example (campus club data), 6; mostly meets the requirements, 7–8; meets the requirements, 9; +1 for an interesting theme. Minus 1 for not importing the instructor's fixed packages. For weeks 1–4 the Chang Gung page calls these "the fixed 4 lines of packages" without listing them. Demo04, Demo04c, and Demo06 all open with the same four lines, %matplotlib inline plus imports of numpy, pandas, and matplotlib, so that's most likely what it means. The current Demo06a/06b don't have them; add them before submitting.
Choosing data: pick something you look up yourself and that general-purpose models answer badly, such as your department's degree requirements or a club's bylaws. Prepare five questions you know the answers to and test each one once the database is built. When an answer is wrong, check the 4 retrieved chunks first. If the answer isn't there, it's a retrieval problem (chunking, embedding, k). If it is there and the answer is still wrong, it's the prompt or the model.
Self-check
- Can you use L06's "information + instructions" to explain which part of the prompt RAG changes?
- Can you explain why documents are chunked and why neighboring chunks overlap?
- Can you explain why documents and questions must use the same embedding model?
- Can you put Demo06a's
faiss_db.zipin the cloud, turn it into a direct download link, and have Demo06b fetch it withgdown? - Given a wrong answer, can you tell whether retrieval or generation is at fault?
Further reading
- Every RAG component compared: The RAG Techniques Compendium; start with Naive, Advanced, and Modular RAG
- Choosing a vector database and embedding model: Vector database comparison, Choosing BGE-M3
- Letting the model decide whether to search again: CMU 11-768 AI Agents guide
Previous: L07 Building Your Own Chatbot | Next: L09 Why 2025 Is the Year of AI Agents | Series overview
References
- Generative AI 08: Retrieval-Augmented Generation (RAG), principles and practice (YouTube recording, in Mandarin)
- Yen-Lung Tsai, 1132 Generative AI slide folder (GenAI08, in Mandarin)
- 1132 video playlist (in Mandarin)
- Chang Gung satellite course page: Generative AI 2025 (in Mandarin)
- yenlung/AI-Demo: Demo06a, building the vector database
- yenlung/AI-Demo: Demo06b, building the RAG system
- yenlung/AI-Demo: Demo06, the spiritual-prescription bot with RAG
- Simulated NCCU club data, uploaded_docs.zip (in Mandarin)
- FAISS (facebookresearch/faiss)
- LangChain
- google/embeddinggemma-300m (Hugging Face)
Loading...