Skip to content
All tags

#nccu-generative-ai

15 posts

NCCU Generative AI L01: Why Study Generative AI, Course Intro and Colab

The first half of lecture 1 covers course rules and lightning talks. The middle answers "why learn the principles?": Yen-Lung Tsai splits the anxiety of learning AI into three kinds and argues that knowing the principles tells you a model's limits, so you stop chasing every new tool. The second half is a Colab primer, from magic commands and the four standard import lines to plt.plot, Markdown, and ipywidgets. Homework 1 is to plot a function in Colab. On the Chang Gung satellite rubric, a tweaked copy of the demo earns 6 points; a function not taught in class, with well-written Markdown notes, earns 10.

NCCU Generative AI L02: Neural Network Concepts

Lecture 2 opens up last week's "dopey AI robot." Inputs and outputs must become numbers (tensors). Classification uses one-hot labels and softmax to turn scores into probabilities. A neural network is neurons stacked layer by layer, and training means pushing the loss down with gradient descent. The lecture ends by building a first fully connected network on MNIST in Keras and wiring it to a Gradio sketchpad. Homework 2 asks you to design your own DNN, with one hard rule: it can't have three layers. The Chang Gung satellite rubric wants a screenshot of the parameters with the best validation accuracy and encourages keeping failed attempts.

NCCU Yen-Lung Tsai Generative AI L03: GANs — How Do Two Competing Networks End Up Drawing Pictures?

The prompt "a cute girl" has countless correct pictures, so training it as a function only teaches the model the average of all of them. GANs sidestep this by training two networks: a generator G turns a random latent vector into an image, a discriminator D judges real versus fake, and the two compete. L03 walks from the 2014 paper through WGAN, Progressive GAN, StyleGAN's 512-dimensional latent and AdaIN, then Pix2Pix and CycleGAN. An appendix explains cross entropy and KL divergence as a 'surprise index'. Week 3 homework: run a GAN yourself, or explain CE and KL in your own words.

NCCU Yen-Lung Tsai Generative AI L04: LLMs Are Simpler Than You Think — Next-Word Prediction, Temperature, and Your Own Benchmark

L04 reduces a large language model to one sentence: look at the preceding words, score every word in the vocabulary, turn the scores into probabilities with softmax, and sample the next word. To give the model a memory of what came before, the lecture covers RNNs and then gives a first look at Transformer Q/K/V. GPT-2's 1.5 billion and GPT-3's 175 billion parameters illustrate scale; temperature and top-p explain why every answer comes out different. The second half covers running open models locally and estimating VRAM. Week 4 homework: write test prompts on a topic you know well and compare at least two LLMs.

NCCU Yen-Lung Tsai Generative AI L05: Transformers, Explained — Reading Q/K/V, Positional Encoding, and Residuals Through Linear Algebra

L05 reads the whole Transformer with two linear-algebra rules: matrix multiplication is row-times-column dot products, and a row vector times a matrix is a linear combination of the matrix's rows. With those, attention is 'dot the query with every key, softmax into weights, take a weighted average of the values,' or softmax(QKᵀ/√d_k)V in batch form. Dividing by √d_k just pulls the numbers back toward 0 so softmax doesn't become winner-take-all. Then come multi-head attention, encoder versus decoder, masking, positional encoding as a set of sin/cos clocks, and ResNet-style residuals with layer normalization. No homework this week.

Reading NCCU Yen-Lung Tsai Generative AI, L06: LLM Applications and Ethical Challenges — Hallucination, Privacy, DeepSeek, and a One-Paragraph System Prompt Called the Lucky Vicky Generator

The first half of L06 is about ethics. Yen-Lung Tsai quotes Karpathy's line that hallucination is a feature of LLMs, then works through plagiarism, whether your data gets used for training, and DeepSeek's censorship and corpus skew, and closes with seven principles of responsible use. The second half is about applications: give the model the right information and clear instructions, and one system prompt becomes a Lucky Vicky positivity generator, a social-media copywriter, or a biased college-major counselor. The week-6 assignment moves that prompt into an OpenAI-compatible API with a Gradio front end: a chatbot with a persona.

Reading NCCU Yen-Lung Tsai Generative AI, L07: Building Your Own Chatbot — API Keys, Three Roles, Sending the History Back, and Running Models Locally with Ollama

A chatbot 'remembers' you not because the model has memory, but because your code resends the whole messages list (system, then alternating user and assistant) every turn. L07 starts with getting OpenAI and Groq keys, spells out that structure, then runs Gemma 3 locally or in Colab with Ollama, where the same openai package works after changing only base_url. The week-7 assignment offers two options: a version that keeps the conversation going, or two models talking to each other, both demoed in Gradio.

Reading NCCU Yen-Lung Tsai Generative AI, L08: Retrieval-Augmented Generation (RAG) — Chunk the Text, Embed It, Put the Closest Pieces Back in the Prompt

L06 said a prompt is two things: correct information and clear instructions. RAG lets the computer fetch the information part on its own. Split your documents into chunks, turn chunks and questions into feature vectors with the same model fθ, find the closest few chunks, and drop them into a template: 'Answer {question} based on {retrieved_chunks}.' The code comes in two notebooks: Demo06a builds a vector database with LangChain and FAISS and zips it as faiss_db.zip; Demo06b loads it back, connects an LLM, and wraps it in Gradio. The week-8 assignment is to do the same with your own data.

Reading NCCU Yen-Lung Tsai Generative AI, L09: Why 2025 Was Called the Year of AI Agents — Andrew Ng's Four Design Patterns, with Reflection and Two-Stage CoT Built in AISuite

L09 defines an AI agent in one line: the AI finishes the work you would otherwise do yourself. Yen-Lung Tsai follows Andrew Ng's four design patterns (Reflection, Tool Use, Planning, Multiagent Collaboration) but builds only the two easiest. Demo07a hands a draft between a "writer" and a "reviewer" LLM call. Demo07c splits the Lucky Vicky post generator into "think of five reasons, then write the post", a two-stage CoT. Both use AISuite with Groq and a Gradio front end. LangChain, AutoGen and CrewAI appear only on a further-learning list. The week 9 homework asks you to pick one of the two patterns.

Reading NCCU Yen-Lung Tsai Generative AI, L10: The Adventure That Starts with the VAE — Feature Vectors, Autoencoders, Diffusion, and "Without the VAE, Stable Diffusion Doesn't Run"

L10 starts from one question: how do you find a good feature vector? Word2Vec learns embeddings through a pretext task. An autoencoder squeezes out a latent vector by being forced to reproduce its input. A VAE then asks the latent to follow a normal distribution, so nearby points produce similar images. Yen-Lung Tsai then recasts diffusion as "an autoencoder whose encoder is computed and whose decoder is learned", and ends on latent diffusion: a VAE shrinks a 512×512 image to 64×64, and diffusion runs only in that small space. The week 10 homework involves no code: make several style-consistent image sets with Bing.

Reading NCCU Yen-Lung Tsai Generative AI, L11: Text-to-Image AI, Principles and Practice — CLIP Reads the Prompt, Schedulers Decide Whether It Converges, LoRA Learns Only ΔW, and You Build a Web App with diffusers

L11 fills in the rest of the Stable Diffusion diagram. CLIP is trained so that matching text and images get similar vectors, which turns a prompt into a 77×768 embedding. Schedulers compress 1,000 noising steps into twenty or thirty denoising steps, but ancestral samplers such as Euler a never settle: push to 100 steps and the subject changes jackets and seats. LoRA freezes the original W and learns only a ΔW factored into A·B. The hands-on part loads an SD 1.5-family model with diffusers, and the week 11 homework is your own image-generation web app.

NCCU Yen-Lung Tsai Generative AI L12: ControlNet and Fooocus, or How to Make an Image Model Follow Your Composition

The Stable Diffusion setup from L11 listens only to the prompt, so composition and pose are left to luck. L12 adds a steering wheel. ControlNet copies a block of SD and wires the copy back in through zero convolutions, so extra conditions such as edge maps, poses, and depth maps can steer generation. The standard example is Canny edges. The second half covers Fooocus, an SD interface that aims to be 'as simple as Midjourney': Presets, Styles, and the five Input Image features, where Image Prompt is ControlNet with a friendly wrapper. Week 12 homework: pick a use case, make at least 3 image sets in Fooocus, and write up your creative process.

NCCU Yen-Lung Tsai Generative AI L13: Reinforcement Learning, from AlphaGo to the RLHF That Makes LLMs Bluff Less

Every model in the first 12 lectures learned from training data that people prepared. L13 asks a different question: when there is no right answer, only a signal of how well you did, how does a computer learn? Tsai starts from AlphaGo and Breakout and splits the field in two. Value-based methods learn a Q function that scores each action (Deep Q-Learning, TD, experience replay, ε-greedy); policy-based methods learn the action directly (policy gradient, actor-critic). The second half returns to LLMs: ChatGPT trains a reward model from human rankings and then runs RLHF with PPO, while DeepSeek has the computer check math answers automatically and uses that as the reward. Week 13 homework is the final project proposal.

NCCU Yen-Lung Tsai Generative AI L14: Text and Image Models Invade Each Other's Territory, Plus the Final Project

The last lecture looks at two lines of technology crossing into each other. LLMs such as ChatGPT have started drawing, and the slides use early fusion plus VQ-VAE/VQGAN to explain how an image can be cut into tokens. Going the other way, Inception Labs' Mercury generates text with diffusion, noising a sentence into a row of [MASK] tokens and then restoring it. Next come a few papers anyone can use: evaluating RAG automatically, reasoning models being easier to hijack, and DeepMind's four kinds of AI risk. The lecture ends with vibe coding and a list of application tools, and the final project runs as an online conference in Gather Town.

Reading NCCU Yen-Lung Tsai's Generative AI: Overview and Self-Study Route

Generative AI: Text and Image Synthesis Principles and Practice is an introductory course taught by Yen-Lung Tsai (蔡炎龍) of NCCU's Department of Mathematical Sciences and opened to other schools as a TAICA satellite course. The most complete course page online actually belongs to the Chang Gung University satellite section, where Chih-Yuan Yang is the co-teacher. This series follows Spring 2025 (semester 1132): 14 recordings, 14 slide decks, and 12 homework specs with rubrics are public, and the demo notebooks are on GitHub, so the access grade is A3. The gaps: the notebooks keep changing, submission and grading run through each school's LMS, and final projects were never published.