Table of Contents
🌏 中文版
2025 was a two-conference year for computer vision: CVPR in Nashville in June and ICCV in Honolulu in October. Both set submission records. Together they received more than 24,000 papers and accepted 5,500, nearly twice the combined totals from 2021. Four themes stood out: 3D Gaussian Splatting fully displaced NeRF, video generation moved from academia toward commerce, flow models challenged diffusion's dominance, and embodied AI drew visual understanding closer to robotic control.
CVPR 2025
CVPR received 13,008 submissions, up 13% from 11,532 in 2024, and accepted 2,878; 2,872 were ultimately presented. Its 22.1% acceptance rate was a record low, and only 3.3% of papers received oral presentations. The conference drew 9,375 attendees from 75 countries, with 118 workshops, 25 tutorials, and 69 demos—33% more demos than in 2024.
Best Paper
VGGT: Visual Geometry Grounded Transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, David Novotny (University of Oxford / Meta AI)
VGGT is a feed-forward network that takes anywhere from one to hundreds of images and estimates camera pose, scene depth, point correspondences, and other 3D scene properties within seconds. Traditional reconstruction requires a multistage pipeline—SfM, then MVS, then a mesh. VGGT compresses that process into one forward pass fast enough for real-time use. It exemplifies a shift from treating 3D vision as an optimization problem to treating it as direct inference. Compared with 2021-era NeRF, which required half an hour of per-scene training, the difference is conceptual as well as quantitative.
Best Student Paper
Neural Inverse Rendering from Propagating Light
Anagh Malik, Benjamin Attal, Andrew Xie, Matthew O'Toole, David B. Lindell (University of Toronto / Vector Institute / Carnegie Mellon University)
The paper used multiview time-resolved LiDAR measurements for physically grounded inverse rendering. It reconstructed not only geometry but also the propagation of light through a scene, combining computational photography with neural rendering.
Best Paper Honorable Mention
- MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos — Zhengqi Li, Richard Tucker et al. (Google Research / UC Berkeley) recovered structure and motion from casual phone video without assuming a static scene.
- Navigation World Models — Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, Yann LeCun (UC Berkeley / Meta AI / NYU) predicted future views to navigate robots in unfamiliar environments.
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models — Matt Deitke and 47 coauthors (AI2) released the fully open Molmo VLM family and a new PixMo dataset not generated by an external VLM, reaching the open-model state of the art.
- 3D Student Splatting and Scooping — Jialin Zhu, Jiangbei Yue, Feixiang He, He Wang (UCL / Peking University).
Best Student Paper Honorable Mention
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens — Kaihang Pan, Wang Lin, Zhongqi Yue et al. (NTU / Zhejiang University).
Best Demo (joint)
DynaMem and Robot Utility Models — Haritheja Etukuru et al. (NYU / Meta AI) demonstrated robots using dynamic memory to manipulate objects in real environments.
Influential Papers Outside the Awards
- DepthCrafter (Wenbo Hu et al., Tencent / HKUST) generated temporally consistent, long-sequence depth maps for open-world video without camera pose or optical-flow input.
- Video Depth Anything (Sili Chen et al., ByteDance / HKU) extended Depth Anything to consistent depth estimation over very long videos.
- GEN3C (NVIDIA) used 3D information to guide world-consistent video generation with precise camera control.
ICCV 2025
ICCV received 11,239 submissions, up 39% from 8,088 in 2023, and accepted 2,698 for a 24.0% rate. It took place in Honolulu from October 19 to 23.
Its review-integrity policy was especially notable. Continuing CVPR 2025's approach, Program Chairs proactively identified 25 seriously irresponsible reviewers and desk-rejected 29 associated submissions—12 of which otherwise would have been accepted. Review integrity moved from punishment after the fact toward proactive removal.
Marr Prize (Best Paper)
BrickGPT: Generating Physically Stable and Buildable Brick Structures from Text
Ava Pun, Kangle Deng, Ruixuan Liu, Deva Ramanan, Changliu Liu, Jun-Yan Zhu (Carnegie Mellon University)
BrickGPT was the first method to generate physically stable LEGO models from text that could actually be assembled. Its key innovation was physics-aware validation during inference: physical laws and assembly constraints pruned infeasible token predictions as generation proceeded. The team released StableText2Lego, with more than 47,000 stable structures and 28,000 unique 3D objects, along with code and models.
The contribution went beyond bricks. It combined next-token prediction with hard physical constraints, validating during generation rather than filtering results afterward. That pattern can apply to any generation task with hard constraints.
Best Student Paper
FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, Tomer Michaeli (Technion)
FlowEdit is a model-agnostic text-based image editor built on pretrained flow models, requiring neither inversion nor optimization. Its importance extends beyond editing: it showed flow models replacing diffusion at the application layer and offering a cleaner editing pipeline.
Honorable Mention
- Spatially-Varying Autofocus — Yingsi Qin, Aswin C. Sankaranarayanan, Matthew O'Toole (CMU) introduced spatially varying autofocus.
- RayZer: A Self-supervised Large View Synthesis Model — Hanwen Jiang et al. (UT Austin / Adobe / Cornell) learned 3D scene representations from 2D images without 3D labels.
Influential Papers Outside the Awards
- EVER (Exact Volumetric Ellipsoid Rendering) replaced Gaussian splatting with volumetric ellipsoids, eliminating 3DGS popping artifacts and even surpassing Zip-NeRF on its dataset.
- SceneSplat introduced the first large-scale indoor 3DGS dataset, SceneSplat-7K, with 6,868 scenes, addressing generalization for semantic reasoning over 3DGS.
- GeometryCrafter (Tencent / Tsinghua) extended DepthCrafter from depth to full geometry estimation.
- MINDCUBE benchmarked the spatial mental models of VLMs and found that leading systems performed only slightly above random guessing on spatial reasoning.
Four Major CV Trends in 2025
1. 3D Gaussian Splatting Fully Replaced NeRF
ICCV 2021 had more than 25 NeRF papers. 3DGS appeared in 2023; by 2025 it was the dominant representation. Paper Digest's ICCV 2023–2025 comparison described explosive 3DGS growth from a topic that had barely existed two years earlier.
NeRF survived mainly in specialized settings such as precise optical simulation. 3DGS won on real-time rendering and editability. EVER addressed original 3DGS artifacts, while SceneSplat laid groundwork for large-scale dataset training.
2. Video Generation Moved from Research Toward Products
Sora, Kling, Runway Gen-3, and other commercial video generators launched in 2024–2025. Academic work moved beyond whether video could be generated:
- Depth consistency: DepthCrafter and Video Depth Anything made depth stable across long videos.
- 3D consistency: GEN3C guided video with 3D information for multiview consistency.
- 4D scene understanding: FICTION, Uni4D, and related work added time to 3D and attempted to understand dynamic 4D scenes directly from video.
3. Flow Models Began Replacing Diffusion Models
The signal was particularly strong at ICCV. FlowEdit won Best Student Paper, and Paper Digest identified the move from diffusion to flow as a major trend. Flow models learn a continuous mapping from noise to data; relative to iterative diffusion denoising, they offer structural advantages in inference efficiency and controllability.
4. Embodied AI Integrated with World Models
Navigation World Models received a CVPR Honorable Mention and DynaMem won Best Demo. Robot-vision integration was no longer confined to workshops. CVPR's official trend summary highlighted autonomous driving's move from modular pipelines toward end-to-end systems and world models. Work such as Genesis, a general-purpose differentiable physics simulator, advanced the complementary route of training embodied AI entirely in simulation rather than on real data.
Compared with 2024
| Dimension | 2024 | 2025 |
|---|---|---|
| Submissions | CVPR 11,532 / ECCV 8,585 | CVPR 13,008 / ICCV 11,239 |
| Acceptance | CVPR 23.6% / ECCV 27.8% | CVPR 22.1% / ICCV 24.0% |
| Mainstream 3D representation | 3DGS growing fast; NeRF still visible | 3DGS dominant; NeRF became niche |
| Generative models | Diffusion dominant | Flow models began replacing diffusion across tasks |
| Video research | Primarily generation quality | Depth and 3D consistency; 4D understanding |
| Embodied AI | Mostly workshops | Best Paper Honorable Mention and Best Demo |
| Review integrity | Concern about AI-generated reviews began | Papers tied to irresponsible reviewers were proactively desk-rejected |
Looking Back from 2026: Which Papers Mattered Most?
- VGGT. If feed-forward 3D reconstruction continues to improve, it could make per-scene optimization obsolete.
- FlowEdit. As a marker of flow models replacing diffusion at the application layer, its influence can extend far beyond image editing.
- Navigation World Models. A milestone for embodied AI that moved world models from a proof of concept into practical navigation.
- Molmo / PixMo. A fully open VLM ecosystem with substantial value to academic researchers and independent developers.
- DepthCrafter / Video Depth Anything. Infrastructure-level work in video depth estimation, useful to AR, editing, autonomous driving, and other downstream applications.
References
- CVPR 2025 Awards Press Release
- CVPR 2025 Best Papers and Best Demos
- CVPR 2025 Conference Wrap-Up
- ICCV 2025 Paper Awards — IEEE TCPAMI
- Computer Vision Awards — The Computer Vision Foundation
- BrickGPT — CMU Intelligent Control Lab
- BrickGPT GitHub Repository
- Top CVPR 2025 Papers — GitHub
- Key Computer Vision Trends: ICCV 2023 vs 2025 — Paper Digest
- Top Computer Vision Research Topics: CVPR 2025 — Paper Digest
- ICCV Acceptance Rate and Submission Statistics
- ICCV 2025 Accepted Papers
- CVPR 2025 Technical Program
- ICCV 2025 Best Paper Awards
Loading...