A chatbot 'remembers' you not because the model has memory, but because your code resends the whole messages list (system, then alternating user and assistant) every turn. L07 starts with getting OpenAI and Groq keys, spells out that structure, then runs Gemma 3 locally or in Colab with Ollama, where the same openai package works after changing only base_url. The week-7 assignment offers two options: a version that keeps the conversation going, or two models talking to each other, both demoed in Gradio.
For side projects, toy demos, and RAG prototypes, nobody wants to swipe a credit card on day one. This is a verified roundup of 40+ LLM inference providers still operating as of 2026/05, tiered by whether free resources auto-replenish or are one-time grants. Each entry notes credit-card requirements, supported models, paid starting prices, and catches. Chinese-origin providers including Zhipu GLM (permanently free), Doubao (2M tokens/day), Kimi, DashScope, and the Ollama local option are all included.
Groq Console is the developer portal for Groq's in-house LPU chip, offering an OpenAI-compatible API, Playground, and free tier credits. Its selling point is running open-source models like Llama, Qwen, and DeepSeek at the fastest tokens/second on the market.
The official provider directory now lists 60 entries. The most common failure when attaching a local model is writing Ollama's base URL with /v1 — that breaks tool calling, and the model starts emitting raw tool-call JSON as plain text.