Skip to content
All tags

#image-generation

10 posts

CMU 10-423 L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and Q-Former

Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.

Reading NCCU Yen-Lung Tsai Generative AI, L10: The Adventure That Starts with the VAE — Feature Vectors, Autoencoders, Diffusion, and "Without the VAE, Stable Diffusion Doesn't Run"

L10 starts from one question: how do you find a good feature vector? Word2Vec learns embeddings through a pretext task. An autoencoder squeezes out a latent vector by being forced to reproduce its input. A VAE then asks the latent to follow a normal distribution, so nearby points produce similar images. Yen-Lung Tsai then recasts diffusion as "an autoencoder whose encoder is computed and whose decoder is learned", and ends on latent diffusion: a VAE shrinks a 512×512 image to 64×64, and diffusion runs only in that small space. The week 10 homework involves no code: make several style-consistent image sets with Bing.

Reading NCCU Yen-Lung Tsai Generative AI, L11: Text-to-Image AI, Principles and Practice — CLIP Reads the Prompt, Schedulers Decide Whether It Converges, LoRA Learns Only ΔW, and You Build a Web App with diffusers

L11 fills in the rest of the Stable Diffusion diagram. CLIP is trained so that matching text and images get similar vectors, which turns a prompt into a 77×768 embedding. Schedulers compress 1,000 noising steps into twenty or thirty denoising steps, but ancestral samplers such as Euler a never settle: push to 100 steps and the subject changes jackets and seats. LoRA freezes the original W and learns only a ΔW factored into A·B. The hands-on part loads an SD 1.5-family model with diffusers, and the week 11 homework is your own image-generation web app.

NCCU Yen-Lung Tsai Generative AI L12: ControlNet and Fooocus, or How to Make an Image Model Follow Your Composition

The Stable Diffusion setup from L11 listens only to the prompt, so composition and pose are left to luck. L12 adds a steering wheel. ControlNet copies a block of SD and wires the copy back in through zero convolutions, so extra conditions such as edge maps, poses, and depth maps can steer generation. The standard example is Canny edges. The second half covers Fooocus, an SD interface that aims to be 'as simple as Midjourney': Presets, Styles, and the five Input Image features, where Image Prompt is ControlNet with a friendly wrapper. Week 12 homework: pick a use case, make at least 3 image sets in Fooocus, and write up your creative process.

NCCU Yen-Lung Tsai Generative AI L14: Text and Image Models Invade Each Other's Territory, Plus the Final Project

The last lecture looks at two lines of technology crossing into each other. LLMs such as ChatGPT have started drawing, and the slides use early fusion plus VQ-VAE/VQGAN to explain how an image can be cut into tokens. Going the other way, Inception Labs' Mercury generates text with diffusion, noising a sentence into a row of [MASK] tokens and then restoring it. Next come a few papers anyone can use: evaluating RAG automatically, reasoning models being easier to hijack, and DeepMind's four kinds of AI risk. The lecture ends with vibe coding and a list of application tools, and the final project runs as an online conference in Gather Town.

NTU ML 2026 HW4: Drawing Pokémon with Next-Token Prediction on a Decoder-Only Transformer

HW4 moves next-token prediction from text to images unchanged: 792 Pokémon sprites at 20×20, each pixel one of 167 color tokens, so one image is a 400-token sequence. Training is next-token prediction; at test time you get the first 60% of an image and the model draws the rest. Grading checks FID and a Pokémon Detection Rate (PDR) together, and the three baseline hints go from "run the sample code" to "tune hyperparameters" to "switch to Llama or Mistral". The spec, Colab, Kaggle notebook and dataset are public, but JudgeBoi returned 502 on 2026-09-30, so outside readers cannot get official FID or PDR scores.

Funding Brief|Stability AI Series B $76M

Stability AI closed a $76M Series B led jointly by Universal Music, Warner Music, Sony Music, and EA, bringing total funding to $232M. It is the first AI company to secure direct equity investment from all three major record labels simultaneously — signaling that copyright holders are shifting from 'sue AI' to 'invest in AI.'

FLUX: The Image Model Family Built by Stable Diffusion's Original Team, from 12B to a Self-Flow World Model

FLUX is Black Forest Labs' image-model family. The Stable Diffusion team launched it in August 2024 with a 12B rectified-flow transformer. Two years later it spans klein 4B ($0.014 and the only current Apache-2.0 model) / 9B, pro ($0.03), flex ($0.05), max ($0.07 with live web grounding), and open-weight 32B dev. FLUX 3 extends Self-Flow to video, synchronized audio, and robot actions. This guide covers the FLUX.1-to-FLUX 3 evolution, three-tier licensing, and model selection.

The Nous Tool Gateway: One Subscription Instead of Four Accounts, at the Cost of Concentrating Your Tool Supply Chain

The Tool Gateway routes four tool categories — web search (Firecrawl), image generation (nine FAL models), TTS (OpenAI), and cloud browser (Browser Use) — through Nous infrastructure, replacing four signups with one OAuth. It's per-tool rather than all-or-nothing, and `use_gateway: true` overrides any direct key in your `.env` — the precedence rule people most often get wrong.

aiproject

2026 Q1 Open-Source LLM Landscape: From Frontier Models to On-Device, a Complete Survey

2026 Q1 saw a full-blown open-source model explosion: on the LLM front, GLM-5, Kimi K2.5, and Qwen3.5 caught up with closed-source models; Embedding and Reranker are dominated by Qwen3 and BGE; speech has Voxtral TTS and Whisper V3; image has FLUX.2; and video has Wan 2.2 rivaling Sora. This is the complete navigation map.