Skip to content

CMU 10-423 L24–L26: Audio, Video Generation, and Interactive World Models — Taking Generative Models from Images to Sound, Time, and Worlds You Can Act In

Sep 30, 20261 min
TL;DRThe last three lectures of CMU 10-423 carry the Transformers, tokenizers, and latent diffusion from earlier in the course over to new kinds of data. L24 covers audio: turn sound into a mel-spectrogram or discrete tokens, then transcribe with Whisper, generate with AudioLM and MusicGen, and diffuse with AudioLDM. L25 covers video: 3D UNets with spatio-temporal attention, latent video diffusion, DiT and Sora, and finally the interactive NeuralOS. The first half of L26 covers world models, which predict the next state from a state and an action, along three routes: generate a 3D scene, interactive video (Genie), and latent representations (V-JEPA, PAN).

🌏 中文版

This post is based on the Spring 2026 edition of CMU 10-423/623/723 Generative AI. It is part 22 of the Reading CMU 10-423 series and follows L23: code generation and autonomous agents. It covers the course's last three lectures:

LectureDateTitleSpeakers (title slide)
Lecture 24April 13Audio understanding and synthesisMatt Gormley
Lecture 25April 15Generative Models for VideosAran Nayebi & Matt Gormley
Lecture 26 (first half)April 20Interactive World ModelsMatt Gormley & Aran Nayebi

Official materials used: the schedule, the L24 slides (38 pages) and inked version (38 pages), the L25 slides (34 pages), and the L26 world models slides (40 pages). L26's other deck, "Science of Alignment," is covered in part 20. None of the three lectures lists readings on the schedule. The course's access grade is A3 (definitions in the global AI/CS course map), but the recordings sit behind a CMU Panopto login and the audio and video demos survive only as links, so this post relies entirely on slide text and figure captions.

All three lectures share one question: how do the tokenizers, Transformers, and latent diffusion you learned on text and images extend to sound, to the time axis, and to worlds that respond to a user's actions? The answer keeps the same shape. First find a good representation (a spectrogram, discrete tokens, a latent space), then apply a generative model you already know.

L24: understanding and synthesizing audio

Turning sound into something a model can eat

The deck starts with raw audio. A recording device keeps sampling air pressure (amplitude), which gives the PCM representation. A raw audio file has three parameters: number of channels, bit depth (bits per amplitude), and sampling rate (samples per second). The example is a 44.1 kHz, 16-bit stereo recording: 44,100 samples per second, each sample two 16-bit integers, one per channel.

If a sound never changed, a single Fast Fourier Transform (FFT) would recover the sine waves behind it. Real sound changes over time, so you run many FFTs over overlapping windows and get a spectrogram you can view as an image. Pass the frequencies through a frequency-to-mel map and you have a mel-spectrogram.

Speech recognition: Whisper, Gemini, Parakeet

Whisper is an encoder-decoder Transformer that the slides call almost identical to Vaswani et al. (2017). The concrete settings:

  • The input is a log-mel spectrogram of a 30-second chunk (16 kHz, 80 channels, 25 ms window, 10 ms stride).
  • The encoder starts with two convolution layers plus sinusoidal positional embeddings; the decoder uses learned positional embeddings.
  • Encoder and decoder have the same number of Transformer blocks (32).
  • One model is trained on many tasks, with special tokens marking the task and the language; timestamp tokens let it produce time-aligned transcripts.
  • The Large model has 1.5B parameters, and the dataset is so large that only 2–3 epochs are used.

The results slides say Whisper closes the gap to human-level word error rate on English LibriSpeech, does well on high-resource languages and less well on low-resource ones, and translates speech across many languages.

Two other routes:

  • Gemini: a multimodal LLM that, unlike Whisper, is decoder-only. Audio is converted to 16 kHz and each second becomes 25 tokens. Transcription falls out naturally, and text prompts can be interleaved with the audio.
  • Parakeet (NVIDIA): the slide says it tends to be about 10x faster than Whisper while keeping very high recognition quality. The family is built on FastConformer, paired with three output schemes: TDT (Token and Duration Transducer), RNN-Transducer, and CTC. The figure comes from Hugging Face's Open ASR Leaderboard.

Audio generation: sound as a token sequence

The generation section opens with a figure from an overview of audio language modeling: many models first use an audio codec to turn the signal into discrete tokens, then train a language model on those tokens to generate speech, music, or natural sounds.

AudioLM is built from:

ComponentSetting on the slides
InputSingle-channel audio of length T = 16000
Acoustic tokenizerSoundStream neural audio codec, vocabulary N = 1024
Semantic tokenizerw2v-BERT, vocabulary K = 1024, sequence length T′ = T/640
GeneratorDecoder-only Transformer in a three-stage hierarchy; each stage conditions on the previous one, with a separate model per stage to keep sequences short
OutputSoundStream decoder

Music generation uses two representations. Music Transformer generates MIDI (the slides compare its continuations with a baseline Transformer and an LSTM). MusicGen generates audio tokens directly:

  • The audio tokenizer is EnCodec, a convolutional autoencoder
  • Codebook interleaving: several token streams predicted in parallel
  • Text prompts are encoded with T5, Flan-T5, or CLAP; melody prompts go through an information bottleneck
  • The decoder is a Transformer LM of up to 3.3B parameters

The diffusion route: AudioLDM handles text to audio, text plus audio to audio (style transfer or completion), and audio to audio (inpainting). It has three parts: a VAE that compresses the mel-spectrogram into a latent space (compression ratio r = 4), a DDPM with a DDIM schedule and classifier-free guidance, and encoders trained with contrastive language-audio pre-training (CLAP, like CLIP for audio). StableAudio is architecturally similar but also conditions on start and end times, so the model learns where a snippet sits within a song and can generate audio of variable length.

This section maps almost one to one onto part 13 on text-to-image: the VAE plays the role of latent diffusion's compressor, and CLAP plays the role of CLIP.

Audio language models: SpeechVerse

SpeechVerse represents a common architecture for audio language models. Audio passes through a pre-trained audio encoder and an adapter to become vectors, text goes through the LLM's embedding matrix, and the two are concatenated and fed to a pre-trained LLM. The slide points out that this mirrors the pattern from part 13: attach an image encoder to a pre-trained LLM to get a VLM.

The last slide previews the next lecture: Veo 3 is a video diffusion model that diffuses audio at the same time, so its clips come with synced sound.

L25: generating and understanding video

Data: most captions are written by models

The deck notes that the largest captioned video datasets mostly use a model to write the captions, for example Panda-70M, and that caption quality and style vary a lot depending on which model was used.

From 3D UNet to the Video Diffusion Model

The Video Diffusion Model (2022) is a 3D UNet plus spatio-temporal attention, with relative positional embeddings along the time axis. The slides explain it through two predecessors:

  • 3D UNet: originally for segmenting 3D biomedical images (the example is a 3D scan of a Xenopus kidney). It is almost identical to the standard UNet, except that 2D convolution (height, width, channel) becomes 3D convolution (height, width, depth, channel).
  • ViViT (written VViT on the slides): an image ViT treats every frame as independent and only attends over space; the video version alternates spatial and temporal attention.
Factorized spatio-temporal attention as shown on the slides

Spatial attention:

  1. reshape: b t h w c -> (b t) (h w) c
  2. multi-headed attention
  3. reshape back to b t h w c

Temporal attention:

  1. reshape: b t h w c -> (b h w) t c
  2. multi-headed attention
  3. reshape back to the original shape

The trick is to fold whichever axes are not part of this attention into the batch dimension.

The Video Diffusion Model adds two more techniques:

  • Joint image and video training: when training on images, the temporal attention is masked so that all attention mass stays on the current image. The slides say this improves video generation.
  • Reconstruction guided sampling: used for sampling from a conditional distribution, for example generating the next 16 frames given the first 16, or filling in frames to raise a low frame rate. The key idea is to guide the sample based on the model's reconstruction of the conditioning data.

Latent video diffusion, DiT, and Sora

  • Video LDM: two parts, an encoder/decoder that compresses video to a latent representation and back, and a video diffusion model trained in that latent space. Same idea as latent diffusion in part 13.
  • DiT: a Transformer replaces the UNet as the diffusion backbone, covered in part 14.
  • Sora: the slide says only that Sora uses a DiT backbone trained on images and videos.
  • Open-Sora 2.0: training costs are falling as models and datasets improve; it uses a three-stage training recipe and is evaluated with human evaluation plus benchmarks and metrics.

Video understanding and "interactive video"

Video understanding gets a single example: the Large World Model, which both generates and understands text, images, and video.

The final section, "agentic video generation + understanding," is the newest material in the lecture:

  • NeuralOS (2025): interactive video diffusion for GUIs. User actions plus a hidden state go in, the next screen comes out.
  • Neural Computers (2026): the slide shows only a demo link and the paper source, with no explanatory text.

NeuralOS already takes a state and an action and predicts the next frame, which leads straight into L26's definition of a world model.

L26, first half: interactive world models

Definition and uses

The deck sets the tone with psychology: a world model is "hypothetical thinking," informally a thought experiment. Then it gives the course's definition:

A world model takes a previous world state s and an action a, and samples (or predicts) the next world state s′ through a distribution (or function): s′ ∼ p(s′ | s, a).

What the state holds depends on what you want to model. The slide contrasts "all geopolitical entities, their leaders, and the decisions they make" with "physical interactions in a 3D world," and the course focuses on the latter.

The slides list six uses: simulated training data for robotics, simulated training data for autonomous driving, interactive worlds (i.e. computer games) created by prompting, automatic generation and editing of 3D animation or graphics, letting a thinking model hypothesize about physical interactions in the real world, and faster creation of VR/AR environments. Data sources include large image and video datasets, smaller 3D scene datasets, text, audio, and multimodal data.

Three routes

RouteApproachExamples on the slides
1. Generate a 3D scene, then render/simulateProduce a representation an existing renderer can drawMarble (World Labs), GSGen, DreamFusion, NeRF-VAE
2. Interactive videosThe feel of interacting with a game engineGenie 1–3, GameNGen, Muse, Oasis, GAIA 1–2
3. Latent world representationsRepresent the world state in a low-dimensional latent space (continuous, discrete, or both)PAN, V-JEPA 1–2

Route 1: generating 3D scenes

NeRF and Gaussian Splatting share a goal: synthesize a 3D scene from 2D images and define a differentiable renderer.

  • NeRF: a continuous neural field mapping 3D position and view direction to density and emitted radiance. Its drawback is slow volumetric rendering, which is a poor fit for real-time interaction.
  • Gaussian Splatting: represent the scene as a cloud of 3D Gaussians (splats), each with position, shape, opacity, and color, and render in real time or near real time by rasterizing them efficiently.

Generative models sit on top of these representations. DreamFusion trains a NeRF from scratch for each caption, which works because NeRF is differentiable. GSGen does the same but generates Gaussian splats. The section ends with World Labs' Marble.

Route 2: interactive video (Genie)

Genie-1 gets more slides than any other model in the lecture. At test time it takes an image as a prompt, accepts one user action per time step, and generates the next frame from the current frame and the action, mimicking how you'd interact with a video game.

ItemWhat the slides say
Training data6.8M 16-second videos (30k hours) of 2D platformer games
Size11B parameters
Video tokenizerA VQ-VAE that encodes T frames into discrete tokens; ST-Transformer backbone
Latent action model (LAM)Another VQ-VAE that learns latent actions from unlabeled video: the encoder infers actions from consecutive frames, the decoder reconstructs the next frame from past frames plus actions. Discarded at test time, since a human supplies the actions
Dynamics modelA decoder-only MaskGIT that predicts the next frame's tokens from previous video tokens and actions, trained with cross-entropy; action embeddings are added rather than concatenated

The LAM is the part to notice. The training data contains no key presses at all; the model learns the actions from how frames change. The qualitative results show that images from a text-to-image model can serve as prompts, and that the learned latent actions map onto real keys so users can discover what each one does.

The slides cover two follow-up experiments:

  • Robotics: training Genie on 130k robot demonstrations, a simulation dataset, and 209k episodes of real robot data (all video, no actions) yields a model that lets you play with a robot arm "game" style.
  • Do latent actions transfer? On CoinRun, a procedurally generated platformer, three policies are compared: behavior cloning on expert actions (the skyline), random actions (the baseline), and a LAM-based policy trained on Genie's latent actions, with a mapping from latent to expert actions learned from a few examples.

The slides call Genie-2's technical description vague and guess that the main change is scale (more parameters, more data). For Genie-3 they list six limitations:

  • The interaction window lasts only a few minutes
  • Actions are (likely) latent and come from a fixed set; you can't define new ones
  • Prompting for world changes is an open set, creating an asymmetry between closed actions and open world changes
  • Seemingly limited to one character
  • Fundamentally still a video model, with no game-engine-readable world representation (point cloud, mesh, Gaussian splats)
  • Very high compute demands

Route 3: latent world representations

  • V-JEPA (1–2): a non-generative world model. It always works in latent space, minimizing the gap between the latent representation of the true masked regions of a raw video and the latent representation predicted from the masked video, with the loss defined directly between the two latents. The goal is to transfer the learned representation to other tasks.
  • PAN: aims at long-horizon, conditioned interactive simulation. Three parts: a vision encoder h maps observations to structured latent states; an autoregressive world model f predicts the next latent state from actions and history (long horizon); a video diffusion decoder g reconstructs high-quality frames from latent states (short horizon). Unlike JEPA, PAN keeps its training objective in observation space. The slides' reasoning: the JEPA objective is prone to collapse and needs careful regularization, while an observation-space objective forces the model to stay grounded in the real observations in the training data.

How the course tests these lectures

  • Quiz 6: in class on April 20 (the day of L26) per the schedule, covering L21–L24. So L24 audio is in scope and L25 and L26 are not. The questions are not public.
  • Homework and exam: there are no programming assignments after L15, and the practice exam, released before the March 30 exam, doesn't cover these lectures either.
  • HW623: papers on the list that relate to these lectures include NeRF, Video Diffusion Models, Video-LLaVA, and A Recipe for Generating 3D Worlds From a Single Image.
  • Project: both the L24 and L25 decks open with project reminders (midway report April 13, poster upload April 26, presentations April 28, final report and code April 30). Details are in part 23.

Try this tonight: take a 30-second recording and plot its log-mel spectrogram with librosa or torchaudio using 80 channels, a 25 ms window, and a 10 ms stride, which are Whisper's input settings. Looking at that image, you'll see why the diffusion models in the second half of L24 can treat sound as a picture.

What this post can and cannot confirm

Confirmed: schedule dates and titles, the text and figure captions of the three decks, and the paper URLs the slides cite. Not confirmed: spoken explanations and demo content (Panopto requires a login and the slides keep only demo links), technical details of Genie-2 and Genie-3 (the slides themselves call them vague), the content of Neural Computers (the slide has only a link), and the Quiz 6 questions. The reminder slide at the start of the L26 deck lists HW623 due April 21, poster upload April 27, and the final report May 1. The schedule says April 20, 26, and 30. The reminder slide looks carried over from an earlier term, so this post follows the schedule.

Further reading: for the math of diffusion and flow matching, see Reading MIT 6.S184; for the role of world models in reinforcement learning, see CS234 guest lecture: World of World Modeling.

Series navigation: previous L23: code generation and autonomous agents | next Wrap-up: practice exam, HW623, and the final project | series overview

References