CS231N Lecture 15 runs on one question: what data structure should a 3D shape use so a neural network can read it and produce it? The slides walk through five representations (depth map/surface normals, voxels, point clouds, triangle meshes, implicit surfaces), each with a signature architecture (fully convolutional depth prediction, 3D convolution, PointNet, Pixel2Mesh and Mesh R-CNN, DeepSDF). Then comes the speed trade-off between NeRF and 3D Gaussian Splatting, and a closing roll call of 2025–2026 models: VGGT, TRELLIS, Marble. The 2025 recording uses a different slide deck, with a different order and emphasis.
2022 marked computer vision’s turn from recognition toward generation. Latent Diffusion Models appeared at CVPR and led to Stable Diffusion; NeRF research jumped from 25 papers in 2021 to more than 50 at CVPR alone; ConvNeXt mounted a compelling counterattack for CNNs; and ECCV in Tel Aviv set a record with 157 oral papers.
The dominant paradigm in 3D generation in 2026 is video diffusion feeding feed-forward 3D reconstruction, and Lyra 2.0 is the flagship of that line. But three Best Papers at CVPR 2026 point at what comes next: SAM 3D brings foundation-model-scale object reconstruction, D4RT rebuilds dynamic 4D scenes in seconds from a unified transformer, and O-Voxel replaces Gaussians with structured latents. 3DGS still rules, but surface primitives are challenging it, and pixel-space diffusion is pushing back against latent space.