🌏 中文版
Version note: This post is based on llm_api_tutorial.pdf, listed in the W12 row of the 2025 schedule for Hung-Yu Kao's Natural Language Processing course at National Tsing Hua University (NTHU), Fall 2025, and on
llm_api.ipynbandutils.pyin LLM_API_lab. The slide cover is dated 2024/11/21, so this is the 2024 session reused. The recording is the W12 Thursday one (labeled "Video2(LLM_API)" in the schedule, 2025-11-19, 56:35, in Mandarin). It has no caption track, so I did not check it section by section. I checked every fact against the official materials on 2026-09-30. Access level A3: slides and notebook are public, but theprompts.yamlthe notebook loads is not in the repo (see below).
Series: previous RAG, Part 2: from ODQA to Self-RAG | next RAG labs 1/2 + HW4 | Series overview
Why use an API
Slide 3 gives two blunt reasons:
- Using the ChatGPT web page for NLP tasks means copying and pasting by hand, which is slow
- Testing data on the web page runs into "Too many requests in 1 hour. Try again later."
Once your data runs past a dozen rows, and you want accuracy numbers or a prompt comparison, you need code. For picking a model, the slides point to the Chatbot Arena leaderboard.
The 2024 price table is history now
Slide 5 compares three APIs and Hugging Face (USD per 1M tokens):
| Gemini (gemini-1.5-pro) | OpenAI (gpt-4o) | Claude (Claude 3.5 Sonnet) | Hugging Face | |
|---|---|---|---|---|
| Free quota | 2 RPM, 32,000 TPM, 50 RPD | No | No | Free |
| Input | 1.25 (2.50 above 128k tokens) | 2.50 | 3 | — |
| Output | 5.00 (10.00 above 128k) | 10.00 | 15 | — |
| Prompt caching | 0.3125 / 0.625, plus 4.5 per hour of storage | 1.25 | 3.75 write, 0.30 read | — |
This is a November 2024 snapshot. None of the three models is current, and the prices are no basis for a budget today. What the table does teach is which dimensions to compare: is there a free tier, are input and output priced separately, does long context cost more, and does caching come with a discount.
How the notebook is laid out
llm_api.ipynb has three parts, Gemini, Claude, then OpenAI, and each runs on its own (slide 12). Every part starts the same way:
from utils import load_prompts
prompts = load_prompts("prompts.yaml")
utils.py holds a single function that loads prompts.yaml into a Python dict with the yaml package. The design is worth copying. Prompts live apart from code, so changing a prompt doesn't touch the code, and when you compare prompts you know exactly what changed.
The gap: the repo's LLM_API_lab folder contains only llm_api.ipynb and utils.py. There is no prompts.yaml. Slides 10, 15, and 23 show screenshots of it. From how the notebook uses it, it has at least these keys: system.general, user.general, user.json_mode, user.few, and mutual.few_hint. user.general carries two placeholders, {PREMISE_HERE} and {HYPOTHESIS_HERE}, and user.few carries placeholders from PREMISE_1 through HYPOTHESIS_3. To self-study, you have to rebuild the file from the screenshots.
The notebook link in the slides also points to Reference/LLM_API_lab at the repo root, which now returns 404. Use 2025/Reference/LLM_API_lab.
How system and user prompts split the work
Slide 11 breaks the prompt into three pieces:
| Where | What | Example |
|---|---|---|
| System prompt | Role (persona) | You are an expert at Natural Language Inference (NLI). |
| System prompt | Task description | Analyze a premise and a hypothesis and classify them as NEUTRAL, ENTAILMENT, or CONTRADICTION |
| User prompt | This row's input | premise: {PREMISE_HERE}, hypothesis: {HYPOTHESIS_HERE}. |
The example data is three-way entailment from SemEval 2014 Task 1, the same dataset as HW3. The sample pair is "A group of kids is playing in a yard and an old man is standing in the background" versus "A group of boys in a yard is playing and a man is standing in the background." Every example sets TEMPERATURE = 0.
Gemini: JSON output, few-shot, token counts, summarization
The Gemini part is the most complete:
- Basic call:
genai.GenerativeModel(MODEL_NAME, generation_config=..., system_instruction=system_prompt), thengenerate_content(user_prompt) - Token counts:
model.count_tokens()counts the system prompt and user prompt separately and adds them for the input total. It also counts the reply - JSON output: slide 19 asks how to get structured output for evaluating the model. The answer: append the
json_modetext to the user prompt and setresponse_mime_type: "application/json", so the reply goes straight intojson.loads() - Few-shot: two labeled examples (NEUTRAL, ENTAILMENT), then the third pair as the question. The slides note that few-shot already nudges the model into the output format, so JSON mode may not be needed
- Summarization: a demo on LCSTS (Chinese abstractive summarization), with a Chinese system prompt saying "you are an expert in Chinese text summarization"
OpenAI and Claude: few-shot format and JSON extraction differ
Slide 27 says the three providers are mostly similar. The differences come down to two things.
Few-shot format (slides 29–30): the OpenAI part writes examples as a list of messages, one user message plus one assistant message per example, with the real question last. The Claude part, like Gemini, packs all examples into one string in the user prompt.
Getting JSON: OpenAI uses response_format={"type": "json_object"}. A notebook comment warns that with this setting, you must ask for JSON in the user prompt. The Claude part has no JSON mode. It uses the regex \{.*?\} to pull the first JSON object out of the reply text and parses that.
Token usage (slide 31): both read it off the response object. OpenAI has usage.prompt_tokens and usage.completion_tokens; Claude has usage.input_tokens and usage.output_tokens.
When prompt caching helps
Slide 26 explains the idea using Gemini's docs only; the notebook has no matching code. It lists three use cases: chatbots with long system instructions, repeated queries against large document sets, and analysis of a long video. You cache the system instruction and the large file, and cached tokens cost less. The further-learning slide (33) links prompt caching and Batch API docs for the three providers.
What to change before you run it
This material ran in a late-2024 environment. To follow it in 2026, deal with these first:
- Replace every model name. Anthropic's deprecation page says
claude-3-5-sonnet-20241022was retired on 2025-10-28, before the Fall 2025 session itself (2025-11-19). Google's current Gemini deprecations table no longer lists the 1.5 family at all. Check each provider's model page before picking replacements. - Switch the Gemini SDK. The notebook uses
google.generativeai. The old SDK's repo says support ended permanently on 2025-11-30 and recommends the new Google Gen AI SDK. The install command on slide 9 isgoogle-ai-generativelanguage==0.8.3, while the notebook's comment saysgoogle-generativeai==0.8.3, so those don't match either. - The Claude few-shot cell has a bug. It builds
cur_fs_user_promptbut sendscur_user_prompt, so the few-shot examples never go out. - The token-counting claim is out of date. Slide 32 says only OpenAI offers a way to count tokens before sending. Yet the notebook's own Gemini part calls
count_tokens()before sending, and Anthropic now has a token counting endpoint.
How to self-study this lab
- Write
prompts.yamlfrom the screenshots on slides 10, 15, and 23. As long as the keys match, the notebook runs. - Pick one API with a free tier and run the whole flow: basic call, JSON output, few-shot, token counts. Slides 29–31 cover the differences between providers well enough.
- Compare zero-shot and few-shot accuracy on 10 SemEval rows. That comparison is exactly where the API beats the web page.
One thing to try tonight: move your most-used prompt out of your code into a YAML file, split into system and user keys, with {} placeholders in the user part. You'll use the habit in the next post when writing RAG prompts.
Further reading
- Iterating on prompts: Prompt engineering iteration guide
- From APIs to agents: CME295 Lecture 7: Agentic LLMs
References
- llm_api_tutorial.pdf (cover dated 2024/11/21) — why use an API, price table, install commands, prompt structure, provider differences, prompt caching
- LLM_API_lab (llm_api.ipynb, utils.py) — the three example sections and model names used
- NTHU NLP 2025 schedule — the W12 row lists llm_api_tutorial.pdf and "Video2(LLM_API)"
- W12 Thursday recording (Fall 2025, in Mandarin) — the LLM API TA session
- Anthropic model deprecations — claude-3-5-sonnet-20241022 retired on 2025-10-28
- Gemini deprecations — current Gemini models and shutdown dates
- google-gemini/deprecated-generative-ai-python — old SDK support ended 2025-11-30
- Anthropic token counting — endpoint for counting input tokens before sending
Loading...