Skip to content

CS231N L15: 3D Vision — One Shape, Five Ways to Store It

Sep 30, 20261 min
TL;DRCS231N Lecture 15 runs on one question: what data structure should a 3D shape use so a neural network can read it and produce it? The slides walk through five representations (depth map/surface normals, voxels, point clouds, triangle meshes, implicit surfaces), each with a signature architecture (fully convolutional depth prediction, 3D convolution, PointNet, Pixel2Mesh and Mesh R-CNN, DeepSDF). Then comes the speed trade-off between NeRF and 3D Gaussian Splatting, and a closing roll call of 2025–2026 models: VGGT, TRELLIS, Marble. The 2025 recording uses a different slide deck, with a different order and emphasis.

🌏 中文版

Which year: The slides are the CS231N Spring 2026 Lecture 15 deck (87 pages, cover dated 2026-05-21). The recording is Spring 2025 Lecture 15 (YouTube, about 1h11m, lecturer Jiajun Wu). Note: the 2025 lecture used Jiajun Wu's own 105-page deck, organized differently from the 2026 one, so you can't follow the recording page by page against the 2026 slides. The 2026 recordings are on Canvas for enrolled students only.

This is post 19 in the Reading Stanford CS231N series.

Series: previous A3: Transformer Captioning, SSL, DDPM, CLIP & DINO | next Wrap-up: World Modeling / Robot Learning, Human-Centered AI and the Final Project | series overview

On the schedule, L15 sits between L14 and L16. This series moves it after A3. It has nothing to do with A3's four questions, and moving it keeps "generation → multimodal → assignment" in one run.

Every lecture so far has worked with 2D images, plus a time axis for video in L10. This one adds a spatial dimension. The slides open with how many subfields 3D vision has: 3D representations, correspondences, multi-view stereo, structure from motion, pose estimation, SLAM, differentiable graphics, 3D sensors. One lecture can't cover them, so the 2026 version picks the first one: how to represent a 3D shape.

The schedule's three keywords are 3D shape representations, shape reconstruction, and neural implicit representations. One sentence ties the lecture together: each way of storing a shape dictates what the network looks like and how the loss is written.

Sidestepping 3D: Multi-View CNN

The first example is the laziest and one of the most instructive. Don't handle 3D at all. Render 2D images from several viewpoints, run each through the same CNN, max-pool element-wise across views, and feed a second CNN that outputs a shape descriptor (Su et al., ICCV 2015).

On ModelNet40 classification and retrieval, it beats the non-deep methods and 3D ShapeNets in the slide's table. The lesson is plain: 2D CNNs are already strong, so borrow them when you can. But it also dodges the question. The five representations that follow are the real subject.

Five representations, one table

RepresentationWhat it storesUpsideTroubleSignature work in the slides
Depth map (plus surface normals)Distance from camera per pixel (or surface normal)It's a 2D image; fully convolutional nets just workFront view only, so it's 2.5DEigen & Fergus depth prediction; DepthAnything
Voxel gridV×V×V occupancy gridSimplest idea, like a 3D segmentation maskMemory explodes at high resolution3D ShapeNets, 3D-R2N2, octrees
Point cloudA set of P points in 3DFine detail from few pointsNo surface; rendering needs post-processingPointNet, Point Transformer
Triangle meshVertices and triangular facesGraphics standard; few faces on flat areas, more where detail isAwkward for neural netsPixel2Mesh, Mesh R-CNN, MeshAnything
Implicit surfaceA function that says inside or outsideResolution not tied to a gridNeeds sampling and many queriesDeepSDF, NeRF

Each section below covers how the "Trouble" column gets handled.

Depth maps: one image can't tell size

Predicting a depth map from an RGB image is a fully convolutional network with a per-pixel L2 loss. The catch is physics: a small, close object and a large, far object look exactly the same in a single image. Absolute scale and depth are ambiguous from one view.

The slides answer with a scale-invariant loss that doesn't penalize a global scale factor. Surface normals use the same architecture with a cosine loss between predicted and true vectors (the slide recalls x·y = |x||y|cos θ). The section ends with the modern Depth Anything series.

Voxels: simple, but cubic

A voxel grid stores the shape as V×V×V occupancies. You process it with 3D convolutions and still train with a classification loss. The slides state the drawback bluntly: detail needs resolution, and memory grows with the cube of V. Their example: a 1024³ voxel grid takes 4 GB.

One fix is an octree, which subdivides only where detail is needed (Tatarchenko et al., ICCV 2017).

Point clouds: order shouldn't matter

A point cloud is a set of points. The difficulty is that it is a set: reorder the points and it should still be the same shape.

PointNet runs the same MLP on every point, then max-pools all point features into one vector. Max doesn't care about order, so the network is permutation-invariant by construction. The slides list its uses (classification, part segmentation, semantic segmentation of scenes) and DenseFusion, which fuses point clouds with RGB per point.

To generate point clouds you need a loss that compares two sets. The slides use Chamfer distance: each point finds its nearest neighbor in the other set, squared distances are summed, in both directions. It needs no matching and is differentiable. The modern representative is the Point Transformer series.

Meshes: topology is fixed

Triangle meshes are the graphics standard, and vertices can carry colors, texture coordinates and normals. The price is that neural nets handle them poorly.

Pixel2Mesh outputs a mesh from a single RGB image. The slides list its four key ideas: iterative refinement from an initial mesh, graph convolution, image features aligned to vertices, and a Chamfer loss.

Deformation has a hard limit: topology comes from the initial mesh. No amount of deforming a sphere produces a mug handle's hole. Mesh R-CNN (Gkioxari, Malik, Johnson, ICCV 2019) answers with a hybrid. It adds a mesh head to Mask R-CNN that first predicts voxels to get a coarse mesh with the right topology, then refines it by deformation. For each detected object it outputs a box, a label, an instance mask, and a 3D mesh.

This section extends L9's detection and segmentation straight into 3D. The modern representative is MeshAnything, which generates meshes with an autoregressive Transformer.

Implicit surfaces: turn the shape into a function

The first four representations list where the shape is. An implicit function flips that. Learn o(x): give it any 3D point and it returns the probability the point is inside, and the surface is the level set o(x) = ½. The slides borrow constructive solid geometry slides from Ren Ng's Berkeley CS184 to show the idea is old, then move to DeepSDF, which learns a signed distance function.

NeRF's inputs and outputs

NeRF (Mildenhall et al., ECCV 2020) takes a position (x, y, z) and a viewing direction (θ, φ) and returns a color (r, g, b) and a density σ. To render a pixel, it queries the MLP at many points along the camera ray and composites the results. The task is novel view synthesis: after seeing several photos of a scene, render angles nobody photographed.

After NeRF, the slides quickly show three extensions: Nerfies for deformable scenes, RawNeRF for high dynamic range, and Block-NeRF for stitching together a San Francisco neighborhood.

Then the main problem: it's very slow. Per the slides, one scene takes 1–2 days to train on a V100. Rendering one 256×256 image at 224 samples per pixel takes 14.6 million MLP forward passes.

3D Gaussian Splatting goes back to an explicit representation: the scene is a set of 3D Gaussians, and rendering blends discrete Gaussians along each ray instead of querying a continuous MLP. The slides' comparison: NeRF fitting takes several hours even on the best GPUs (the 1–2 days figure on the previous slide is for a V100), and rendering takes about 10 seconds per frame at moderate resolution; 3DGS fits in a few minutes and renders in real time. Dynamic 3D Gaussians extend it to tracking dynamic scenes.

The new models the slides name at the end

After the summary slide, a few "advanced use cases" pages each show one name and one figure:

  • SLAM: the VGGT series
  • 3D object generation: the TRELLIS series
  • 3D world generation: Marble from World Labs

The slides don't explain how these work, and this post won't fill that in. Their job is to show that the five representations aren't history. Today's large models still pick among them.

What the 2025 recording does differently

If you watch the 2025 recording alongside, the order is completely different. Jiajun Wu's 2025 deck starts from graphics. It splits representations into explicit (point clouds, polygon meshes, parametric surfaces, Bézier, subdivision) and implicit (algebraic surfaces, CSG, distance functions, level sets) and compares what each makes easy, sampling versus inside/outside tests. Then it covers datasets like ShapeNet, and moves through voxels, octrees, PointNet, DeepSDF, Occupancy Networks, NeRF, Gaussian splatting, and generating tree- and graph-structured shapes.

The core vocabulary overlaps (voxel, point cloud, mesh, implicit, NeRF). The 2026 deck adds Mesh R-CNN, depth and normal prediction, and the 2025–2026 models. Use the 2026 slides as the spine and treat the recording as another teacher covering the same topic.

How to study this lecture

  1. Memorize the table first. For every later model, ask which representation it uses and how it handles that representation's trouble.
  2. Stop at the PointNet slide. Check on paper why a per-point MLP followed by max-pool ignores point order.
  3. Write Chamfer distance yourself. With two tensors of shape (N, 3) and (M, 3), torch.cdist gets you there in a few lines, and then it's obvious why no matching is needed.
  4. NeRF versus 3DGS isn't "newer is better." It's the implicit-versus-explicit trade-off; compare it with the voxel and implicit rows of the table.

One thing to do tonight: pick any 3D object on your desk (a mug works) and write down how much data each of the five representations would need to store it, and which one best captures the hole in the handle.

Further reading

References