🌏 中文版
This post is based on the 1132 semester (Spring 2025) of NCCU Yen-Lung Tsai's Generative AI: Text and Image Synthesis Principles and Practice. It is part 7 of the Reading NCCU Yen-Lung Tsai Generative AI series and follows L06, LLM Applications and Ethical Challenges. Last lecture produced a one-shot Lucky Vicky generator. This one makes it hold a conversation, without necessarily sending your data to the cloud.
Official sources: video 07 (2025-04-01, about 3 h 3 min, in Mandarin), the 31-page GenAI07 slides (in the instructor's slide folder), the notebooks 【Demo04c】用OpenAI_API打造自己的對話機器人 and 用_Ollama_打造自己的對話機器人 in the AI-Demo repo, and the week-7 assignment on the Chang Gung satellite course page. Access level: A3. The notebooks live in a shared repo, so everything below refers to the current repo version, which may have changed since the semester ended.
Where this week sits
Video 07's chapters are clear. The first 30 minutes cover OpenAI keys, Groq, and Ollama. From 0:39 Tsai installs Ollama in Colab and builds a comforting chatbot. After the break (from 1:15) he builds a version that "keeps talking" and a Gradio web app, introduces the AISuite package at 1:41, and explains the assignment at 1:49. Session three is lightning talks, including one on Ollama applications, and a TA segment on LM Studio.
The slides have three parts: "Getting OpenAI / Groq API keys", "Ollama", and "Assignment: build your own chatbot with Ollama".
Concept 1: getting and storing keys
Slides 3–9 walk through it: create an account on the OpenAI Platform (a Google account works), find API keys, click Create new secret key. Slide 6 says in large type: make sure you record your key, since it's shown only once.
You can write OpenAI(api_key="your API key") directly, but slide 9 asks students to read it from Colab Secrets "the way we agreed":
import os
from google.colab import userdata
api_key = userdata.get('OpenAI')
os.environ['OPENAI_API_KEY'] = api_key
Keeping the key in Secrets rather than hard-coding it means sharing your Colab link doesn't share your key. Standard names let TAs run your notebook as-is.
What Groq is
Slide 10 introduces Groq: a US AI company founded in 2016 by former Google engineers. Its core product is an in-house Language Processing Unit (LPU), built to speed up LLMs at inference time and not well suited to training. It serves many open models and has a free tier. Signing up works much like OpenAI (console.groq.com).
Concept 2: the model has no memory; you resend everything
This is the lecture's most important slide. Slide 12 reviews the three roles, and slide 13 draws the structure for sending back the conversation history:
[{"role": "system", "content": "ChatGPT 的「人設」"}, # the persona
{"role": "user", "content": "使用者輸入"}, # user input
{"role": "assistant", "content": "ChatGPT 回覆"}, # model reply
{"role": "user", "content": "使用者再輸入"}] # next user input
"Send this and it replies!" The API itself keeps nothing from the previous turn. When ChatGPT's website seems to remember you, the app is resending the whole conversation. So a chatbot that keeps talking only has to do two things each turn:
- Append the user's message as a
userentry. - Once the reply comes back, append it as an
assistantentry.
Look back at Demo04 from L06: it only does step 1. The model sees everything you said but none of its own replies. Demo04c adds step 2:
def mychatbot(prompt, history):
history = history or []
global messages
messages.append({"role": "user", "content": prompt})
chat_completion = client.chat.completions.create(messages=messages, model=model)
reply = chat_completion.choices[0].message.content
messages.append({"role": "assistant", "content": reply})
history = history + [[prompt, reply]]
return history, history
The cost is just as direct: the longer the chat, the more you send each time. That is the "short-term memory" problem the L08 slides describe — when a conversation gets too long, the early parts get forgotten.
Do this: find the function that calls the API in your L06 assignment and check whether the reply is appended back to messages.
Concept 3: Ollama, running models on your own machine
L06 said to keep personal and confidential data off online services, and that local models avoid the problem. Slides 17–23 show how:
| Step | Command | Slide note |
|---|---|---|
| Install | Download from ollama.com | It starts running once installed |
| Download a model | ollama pull gemma3 | The model page gives you run; change it to pull |
| Start the server | ollama serve | Usually already running in the background |
| List installed models | ollama list | |
| See commands | ollama | There aren't many |
Slide 23 gives Ollama's standard API address, http://localhost:11434. Slide 24 says the main event today is running Ollama in Colab (yenlung.me/ollama). Slides 25–27 introduce the GUI Open WebUI (pip install open-webui, then open-webui serve) and point to more GUIs listed on Ollama's GitHub.
This week's demo notebook
用_Ollama_打造自己的對話機器人 ("Build your own chatbot with Ollama"; yenlung.me/ollama currently points here) has seven sections matching the video from 0:43:
- Install Ollama in Colab:
curlthe official install script,nohup ollama serve &to run it in the background,ollama pull gemma3:1b. - Call it with the OpenAI package: the key can be anything (
api_key = "ollama"); setbase_url="http://localhost:11434/v1". - One test message: a system message plus "你好!" ("Hello!").
- A comforting bot: the system prompt asks for a warm, best-friend tone in under twenty characters. The user says "I'm in a bad mood today", the reply is appended as assistant, then "I feel like nobody likes me." This section does concept 2's two steps by hand.
- Keep talking: a
while Trueloop that ends when the input containsbye, storing both user and assistant each turn. This is the "Pat-Pat bot" (拍拍機器人) from 1:15 in the video. - Gradio web app:
gr.Blocks,gr.Chatbot(type="messages"), andgr.Stateto hold the messages. A comment insists oncopy()(務必用 copy()); otherwise every user shares the same initial conversation. - AISuite:
model = "ollama:gemma3:1b". Change the prefix to switch providers.
One thing to watch: the text in section 1 says it demonstrates Llama 3.2, but the code actually pulls gemma3:1b. The older notebook 在_Colab_上用_Ollama ("Using Ollama in Colab") is the one that uses llama3.2. It's a trace of the repo being updated over time; trust the code.
On free Colab, a 1B model is the realistic choice. To run a larger one on your own computer, check the Gemma 3 hardware table on slide 5 of L06 first.
The assignment: week 7 (Chang Gung satellite version)
From the Chang Gung satellite course page; NCCU's own grading differs. Slides 29–31 of GenAI07 title the assignment "Build your own chatbot with Ollama": it should keep the conversation going and be fun or useful. The example persona is a virtual friend named Zhiqing (芷晴), an applied-math student who loves birdwatching and is in the coffee club. The Chang Gung page states the task as:
Build your own chatbot, advanced version. Pick one:
- Option 1: extend last week's assignment, following the instructor's example, into a version that holds an ongoing conversation. Demo in Gradio.
- Option 2: build a bot where two different models talk to each other. Demo in Gradio.
Submit: a Colab link (key points annotated in Markdown), key screenshots, the persona/background, the models used, and Gradio conversation results. The 1132 deadline was 4/14.
Rubric (out of 10): identical to the example, 1; GPT-level or off-topic, 2; a theme close to the examples (a warm chatbot, a Lucky Vicky bot, a math-recommendation bot), 6; mostly meets the requirements, 7–8; meets the requirements, 9; +1 for an interesting theme. Minus 1 for not importing the instructor's fixed packages.
How to think about option 2: each model keeps its own messages. A's reply is a user message for B, and B's reply becomes a user message for A. Each side stores its own words as assistant. Draw that table before writing code and you'll save a lot of debugging. The two models can run on different backends, say one on Groq and one on local Ollama, as long as their clients have different base_urls.
Self-check
- Can you explain why the API doesn't remember the previous turn, and how a chatbot fakes remembering?
- Can you read a key from Colab Secrets and give two reasons to do it that way?
- Can you run Ollama in Colab and point the
openaipackage at it withbase_url="http://localhost:11434/v1"? - Can you write a loop that stores both user and assistant messages and exits on
bye? - Can you say what Gradio's
gr.Stateholds here, and why its initial value needscopy()?
Further reading
- What to keep and what to drop as the history grows: Context Engineering guide, NTU ML 2026: Context Engineering
- From chatbots to tool-using agents: CMU 11-768 AI Agents guide
- How language models themselves work: Stanford CS224N guide
Previous: L06 LLM Applications and Ethical Challenges | Next: L08 Retrieval-Augmented Generation (RAG) | Series overview
References
- Generative AI 07: Build your own chatbot (YouTube recording, in Mandarin)
- Yen-Lung Tsai, 1132 Generative AI slide folder (GenAI07, in Mandarin)
- 1132 video playlist (in Mandarin)
- Chang Gung satellite course page: Generative AI 2025 (in Mandarin)
- yenlung/AI-Demo: Demo04c, build your own chatbot with the OpenAI API
- yenlung/AI-Demo: build your own chatbot with Ollama
- yenlung/AI-Demo: using Ollama in Colab
- Ollama, Ollama GitHub
- Open WebUI
- Groq Console
- andrewyng/aisuite
Loading...