Skip to content
All tags

#multimodal

22 posts

CMU 10-423 L24–L26: Audio, Video Generation, and Interactive World Models — Taking Generative Models from Images to Sound, Time, and Worlds You Can Act In

The last three lectures of CMU 10-423 carry the Transformers, tokenizers, and latent diffusion from earlier in the course over to new kinds of data. L24 covers audio: turn sound into a mel-spectrogram or discrete tokens, then transcribe with Whisper, generate with AudioLM and MusicGen, and diffuse with AudioLDM. L25 covers video: 3D UNets with spatio-temporal attention, latent video diffusion, DiT and Sora, and finally the interactive NeuralOS. The first half of L26 covers world models, which predict the next state from a state and an action, along three routes: generate a 3D scene, interactive video (Genie), and latent representations (V-JEPA, PAN).

CMU 10-423 L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and Q-Former

Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.

CMU 10-423 HW4: Text-to-Image with a Q-Former Between a Frozen GPT-2 and a Frozen DiT — Structure, Files to Edit, and Compute

HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.

CMU 10-423 L12–L13: Text-to-Image, Latent Diffusion, and Vision-Language Models

CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.

CS231N L10: Video Understanding — What Changes When You Add a Time Axis

CS231N Lecture 10 treats video as 2D plus time, a T×3×H×W tensor, and follows one thread: efficiency. Train on short clips and average several clips at test time. Architectures run from per-frame 2D CNNs and late fusion to 3D CNNs, then two-stream networks that isolate motion with optical flow, and I3D, which inflates 2D weights into 3D. After 2021 the field moved to Transformers, where token counts explode, which led to divided space-time attention, Video Swin, MViT, and tubelets. The last part covers temporal localization, audio-visual models, VideoLLMs, and long-form video, where HourVideo shows how far the field still has to go.

CS231N L16: Vision and Language — From CLIP's Contrastive Learning to Multimodal Foundation Models That Talk About Images

The CS231N Spring 2026 vision-and-language lecture replaces the "one model per task" approach of the first half of the course with foundation models: pre-train one model on a large, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use. Three threads carry the lecture. First, CLIP: contrastive learning in both directions over 400 million image-text pairs scraped from the web, then writing class names as sentences to classify without any fine-tuning; it also has weak spots, such as failing to tell "a mug in some grass" from "some grass in a mug". Second, vision-language models from LLaVA and Flamingo to Qwen3-VL and Molmo, which feed image features into an LLM so it can look at an image and output text. Third, chaining: letting an LLM write descriptions or programs that string existing vision models together.

MIT 6.5940 L14 LLM Post-Training: From SFT and RLHF to Fine-Tuning That Touches 1% of the Weights

Lecture 14 has three parts. Fine-tuning: SFT runs next-token prediction on desired answers, RLHF trains a reward model and then fine-tunes with KL-penalized RL, and DPO collapses both stages into one supervised step. Then comes a chain of PEFT methods: BitFit tunes only biases, Adapters add small layers but slow inference, Prompt/Prefix-Tuning eat input length, LoRA fixes latency with a low-rank branch you can merge back, QLoRA stores the backbone in NF4, and BitDelta compresses the fine-tune delta to 1 bit. Multimodal LLMs: Flamingo uses cross-attention, PaLM-E and VILA feed images in as tokens, and VILA-U can also output images. Prompt engineering: zero/few-shot, CoT, and RAG.

Reading NTU ML 2026: HW10 Spoken Language Model — Three Architectures, Mimi's 32 Token Layers, and How Moshi Listens While It Talks

HW10 is 12 multiple-choice questions answered only on NTU COOL. Section 1 compares three spoken language model architectures: Cascade (ASR → LLM → TTS, with text in the middle), End-to-End (a language model over discrete speech tokens), and Thinker-Talker (an LLM thinks, a separate decoder speaks). In the Colab, two models listen to three clips and guess the speaker's gender, and you work out which one is the cascade. Section 2 takes Mimi apart: tokenize an emotion corpus into 32 RVQ layers, plot UMAP for layers 0, 6, 16, and 31, then encode and decode speech, laughter, and music to hear what breaks. The rest are paper questions on TWIST, AudioLM, LLaMA-Omni 2, Moshi, and GLM-4-Voice, covering initialization, pretraining, interleaving, and realtime/full-duplex behavior. The Colab needs Llama-3.2-3B-Instruct access and an HF token. Questions and Colab are public; outside readers miss only the COOL grading and answers.

Multimodal and Vision: WebWatcher Redefines Deep Research

All deep research agents are 'text-first'—but the real world isn't just text. WebWatcher (NeurIPS 2025) is the first system to integrate visual reasoning into deep research, using OCR, image search, code execution, and other tools to handle charts, screenshots, videos, and other diverse information.

How to Spend Every Parameter: OpenELM's Layer-wise Scaling and MiniCPM's Three-stage Unfreezing

OpenELM uses layer-wise scaling to shift parameters toward layers near the output; with 1.08B parameters and 1.5T tokens it beats OLMo 1.2B (+2.36% on the LLM360 average) despite OLMo training on 3T tokens. MiniCPM trains multimodal small models from scratch with a three-stage unfreezing recipe (Resampler first, vision encoder next, everything unfrozen last); MiniCPM-V 4.5 reaches sub-30B SOTA on VideoMME with only 8B parameters, and 4-bit quantization squeezes fp16's 16–17GB memory footprint down to about 5GB for phones.

Inkling: From an OpenAI Exodus Team to a 975B Open Flagship, and Tinker's Fine-Tuning Bet

Thinking Machines Lab (founded 2025 by Mira Murati, $2B seed at a $12B valuation) released Inkling in July 2026 under Apache 2.0 (975B total / 41B active params, 1M context, native multimodality, controllable thinking effort) plus a smaller Inkling-Small (276B / 12B), paired with the Tinker fine-tuning platform—turning customizability itself into the product.

What Top AI Conferences Accepted in 2024: The Year of Agents and the Scaling Debate

The defining conference keywords of 2024 were agents, alignment, multimodal LLMs, and inference-time compute. The LLM share at five major conferences doubled again after its sharp 2023 rise; agent-related terms grew 4.3 times; and diffusion models graduated from an emerging topic to a second generative-AI pillar alongside LLMs. Traditional task-oriented NLP continued to contract, while GANs almost disappeared from top venues.

Gemini——Google's Native Multimodal Flagship: 1M Context and Scientific Reasoning Champion

Gemini is Google DeepMind's native multimodal LLM family, famed for a 1M-token context window and native video/speech input plus scientific reasoning. 3.1 Pro tops GPQA Diamond 94.1% and ARC-AGI-2 77.1% to claim science-reasoning dual crowns, at $2/$12—1/6 of Claude. 3.7 Flash delivers near-Pro agent capability for $0.75/$3.75.

Llama——From Open-Source Experiment to the Most Deployed Open LLM, and Meta's Closed-Source Pivot

Llama is Meta's open-source LLM family, with the largest enterprise deployment footprint and the most mature ecosystem. Llama 4 Scout (10M context) and Maverick (17B active / 400B total MoE) are the current open multimodal benchmarks, but Meta pivoted to closed-source Muse Spark in April 2026—Llama 4 is likely the last major open Llama, and its license is not truly open (Llama 4 Community License, separate license required above 700M MAU).

Mistral——Europe's Open AI Challenger: Smaller Models and European Sovereignty as a Different Bet

Mistral is Europe's most successful AI startup, cutting through the market with a 'smaller, faster, cheaper' strategy and European data-sovereignty positioning. Mistral Large 3 is Europe's strongest commercial LLM, Small 4 is the 24B efficiency king, and Medium 3.5 is the open Modified-MIT model optimized for agentic coding. Its moat is not technical scale but the 'European compliance' card.

AI Model Landscape: The 2026 Map You Need

In 2026, AI models span seven major categories and more than 20 subcategories. This introduction to the AI Model Families series maps use cases to models and models to families, with current rankings and selection advice for each use case.

Stanford CS224V Lecture 13: ReactGenie Gives Voice and Native GUIs Shared State

ReactGenie annotates React components to expose data, actions, and views, parses composite voice commands into a DSL, and renders native graphical output against shared UI context.

CS336 Lecture 17: Multimodal Models Turn Images into Tokens, Then Reconcile Semantics with Detail

Lecture 17 organizes CLIP/SigLIP, LLaVA, Qwen-VL, and Chameleon into three paths: contrastive encoders learn semantics, vision-encoder/projector/LM stacks provide understanding, and discrete image tokens enable generation. Resolution, token budgets, and modality balance constrain them all.

aideep-dive

Stanford CS25 V6: A Course Called Transformers United Whose First Two Talks Weren't About Transformers

CS25 is Stanford's 1-unit seminar where attendance is the only homework and anyone can audit. Of the nine talks in the Spring 2026 season, the three worth your time are Albert Gu on the inductive biases of SSMs vs Transformers, Charles Frye on serving inference across thousands of GPUs, and Victoria Lin on what native multimodality still hasn't solved.

aideep-dive

Multimodal Models, First Half of 2026: Native Fusion vs. Bolted-On Vision, and Why Leaderboards Contradict Each Other

Pure image understanding has flattened out — four frontier models all clear 80% on MMMU-Pro within 3 points of each other. The real differentiation is video, long-document OCR, and realtime speech, each with a different leader. But the most useful lesson from assembling these rankings is that two credible sources named different Video-MME leaders more than 10 points apart — and that July and August each turned the field over again.

NVIDIA NCA-GENM: The Multimodal One, With Two Required Courses Only Sold as $500 Workshops

NCA-GENM matches NCA-GENL on price, length, and level but not on emphasis: Experimentation rises to 25% (the heaviest), Core ML drops from 30% to 20%, and two new areas appear — Multimodal Data 15% and Performance Optimization 10%. The content covers U-Net, CLIP, diffusion models, multimodal loss functions, attention maps, and NVIDIA's Riva / NeMo / Triton / ACE SDKs. Watch the cost structure: two of the five recommended courses exist only as $500 workshops with no self-paced option, so a self-study path cannot cover the official set. Official specs: $125, 1 hour, 50–60 items, two-year validity, English only.

Multimodal RAG: Bringing Images into the Knowledge Base

Climbing routes carry a ton of visual information (topos, wall photos) that text-only RAG misses entirely. Multimodal RAG makes images searchable and understandable.