A collection of 95 oral papers from CVPR 2025, organized by topic, with links to papers, project pages, and code.
Official resources: Program · Proceedings · Awards
Award labels follow the official CVPR 2025 results.
| Award | Paper |
|---|---|
| Best Paper | VGGT |
| Best Student Paper | Neural Inverse Rendering from Propagating Light |
| Best Paper Honorable Mention | MegaSaM |
| Best Paper Honorable Mention | Navigation World Models |
| Best Paper Honorable Mention | Molmo and PixMo |
| Best Paper Honorable Mention | 3D Student Splatting and Scooping |
| Best Student Paper Honorable Mention | Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens |
| Paper | Links |
|---|---|
| VGGT: Visual Geometry Grounded Transformer | Paper · Project · Code |
| CUT3R — Continuous 3D Perception Model with Persistent State | Paper · Project · Code |
| MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision | Paper · Project · Code |
| FoundationStereo: Zero-Shot Stereo Matching | Paper · Project · Code |
| Multi-view Reconstruction via SfM-guided Monocular Depth Estimation | Paper · Project · Code |
| MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | Paper · Project · Code |
| MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos | Paper · Project · Code |
| Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos | Paper · Project · Code |
| Zero-Shot Monocular Scene Flow Estimation in the Wild | Paper · Project · Code |
| DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models | Paper · Project · Code |
| TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion | Paper · Code |
| Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World | Paper · Code |
| Camera Resection from Known Line Pencils and a Radially Distorted Scanline | Paper · Code |
| Paper | Links |
|---|---|
| Neural Inverse Rendering from Propagating Light | Paper · Project · Code |
| Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models | Paper · Project · Code |
| 3D Student Splatting and Scooping | Paper · Code |
| 3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splatting | Paper · Code |
| Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance Fields | Paper · Project · Code |
| FluidNexus: 3D Fluid Reconstruction and Prediction from a Single Video | Paper · Project · Code |
| CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner | Paper · Project · Code |
| CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models | Paper · Project |
| DNF: Unconditional 4D Generation with Dictionary-based Neural Fields | Paper · Project · Code |
| Birth and Death of a Rose | Paper · Project |
| Paper | Links |
|---|---|
| CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models | Paper · Project · Code |
| Reconstructing Humans with a Biomechanically Accurate Skeleton | Paper · Project · Code |
| MEGA: Masked Generative Autoencoder for Human Mesh Recovery | Paper |
| TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization | Paper · Project · Code |
| EgoLM: Multi-Modal Language Model of Egocentric Motions | Paper · Project |
| Paper | Links |
|---|---|
| Infinity∞: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis | Paper · Project · Code |
| RandAR: Decoder-only Autoregressive Visual Generation in Random Orders | Paper · Project · Code |
| Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models | Paper · Code |
| Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space | Paper · Project · Code |
| Autoregressive Distillation of Diffusion Transformers | Paper · Code |
| Language-Guided Image Tokenization for Generation | Paper · Project |
| AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea | Paper · Project · Code |
| CustAny: Customizing Anything from A Single Example | Paper · Project · Code |
| DreamRelation: Bridging Customization and Relation Generation | Paper · Project · Code |
| Minority-Focused Text-to-Image Generation via Prompt Optimization | Paper · Code |
| DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models | Paper |
| Motion Prompting: Controlling Video Generation with Motion Trajectories | Paper · Project |
| Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise | Paper · Project · Code |
| LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping | Paper · Project |
| Reanimating Images using Neural Representations of Dynamic Stimuli | Paper |
| Paper | Links |
|---|---|
| Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models | Paper · Blog · Code |
| Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens | Paper · Project · Code |
| LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models | Paper |
| OPA-DPO — Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key | Paper · Project · Code |
| Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding | Paper · Project · Code |
| Identifying and Mitigating Position Bias of Multi-image Vision-Language Models | Paper |
| Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces | Paper · Project · Code |
| Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding | Paper · Code |
| VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection | Paper · Code |
| OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation | Paper · Project · Code |
| Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content | Paper · Code |
| SEAL: Semantic Attention Learning for Long Video Representation | Paper · Code |
| Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval | Paper |
| Temporal Alignment-Free Video Matching for Few-shot Action Recognition | Paper · Code |
| Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning | Paper · Project |
| Temporally Consistent Object-Centric Learning by Contrasting Slots | Paper · Project · Code |
| The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recognition | Paper · Project |
| Paper | Links |
|---|---|
| Navigation World Models | Paper · Project · Code |
| RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics | Paper · Project · Code |
| From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons | Paper |
| PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation | Paper |
| GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill | Paper · Project · Code |
| Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models | Paper · Project · Code |
| Paper | Links |
|---|---|
| SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images | Paper · Project · Code |
| Effective SAM Combination for Open-Vocabulary Semantic Segmentation | Paper |
| Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation | Paper |
| Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weather | Paper |
| Efficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruning | Paper |
| Paper | Links |
|---|---|
| Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus Cues | Paper · Project · Code |
| Opportunistic Single-Photon Time of Flight | Paper |
| Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor Arrays | Paper · Project |
| Removing Reflections from RAW Photos | Paper · Project |
| DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-Resolution | Paper · Code |
| Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing | Paper · Project · Code |
| DiffFNO: Diffusion Fourier Neural Operator | Paper · Project |
| Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining | Paper |
| Paper | Links |
|---|---|
| OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | Paper · Code |
| CleanDIFT: Diffusion Features without Noise | Paper · Project · Code |
| LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions | Paper · Code |
| Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild | Paper |
| Rethinking Spiking Self-Attention Mechanism: Implementing α-XNOR Similarity Calculation in Spiking Transformers | Paper |
| Towards Universal Dataset Distillation via Task-Driven Diffusion | Paper |
| Gromov–Wasserstein Problem with Cyclic Symmetry | Paper |
| UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming | Paper |
| Enhancing Diversity for Data-free Quantization | Paper |
| Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learning | Paper · Code |
| Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models | Paper · Code |
| Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks | Paper · Code |
| Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector | Paper · Code |
| Paper | Links |
|---|---|
| TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model | Paper · Code |
| Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation | Paper |
| IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion Prior | Paper |
These entries are retained for reference and are not counted among the 95 papers in the official oral program.
| Paper | Links |
|---|---|
| Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation | Paper · Code |
| FedSPA: Generalizable Federated Graph Learning under Homophily Heterogeneity | Paper |
| One Category One Prompt: Dataset Distillation using Diffusion Models | Paper · Code |
| Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding | Paper · Code |
Missing a link or spotted a mistake? Open an issue or submit a pull request with the paper title and updated links.
Thanks to the paper authors for sharing their work, and to cvpr25_oral_gpu_info for the original collection reference.
A collection of 95 oral papers from CVPR 2025, organized by topic, with links to papers, project pages, and code.
Official resources: Program · Proceedings · Awards
Award labels follow the official CVPR 2025 results.
| Award | Paper |
|---|---|
| Best Paper | VGGT |
| Best Student Paper | Neural Inverse Rendering from Propagating Light |
| Best Paper Honorable Mention | MegaSaM |
| Best Paper Honorable Mention | Navigation World Models |
| Best Paper Honorable Mention | Molmo and PixMo |
| Best Paper Honorable Mention | 3D Student Splatting and Scooping |
| Best Student Paper Honorable Mention | Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens |
| Paper | Links |
|---|---|
| VGGT: Visual Geometry Grounded Transformer | Paper · Project · Code |
| CUT3R — Continuous 3D Perception Model with Persistent State | Paper · Project · Code |
| MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision | Paper · Project · Code |
| FoundationStereo: Zero-Shot Stereo Matching | Paper · Project · Code |
| Multi-view Reconstruction via SfM-guided Monocular Depth Estimation | Paper · Project · Code |
| MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | Paper · Project · Code |
| MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos | Paper · Project · Code |
| Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos | Paper · Project · Code |
| Zero-Shot Monocular Scene Flow Estimation in the Wild | Paper · Project · Code |
| DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models | Paper · Project · Code |
| TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion | Paper · Code |
| Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World | Paper · Code |
| Camera Resection from Known Line Pencils and a Radially Distorted Scanline | Paper · Code |
| Paper | Links |
|---|---|
| Neural Inverse Rendering from Propagating Light | Paper · Project · Code |
| Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models | Paper · Project · Code |
| 3D Student Splatting and Scooping | Paper · Code |
| 3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splatting | Paper · Code |
| Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance Fields | Paper · Project · Code |
| FluidNexus: 3D Fluid Reconstruction and Prediction from a Single Video | Paper · Project · Code |
| CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner | Paper · Project · Code |
| CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models | Paper · Project |
| DNF: Unconditional 4D Generation with Dictionary-based Neural Fields | Paper · Project · Code |
| Birth and Death of a Rose | Paper · Project |
| Paper | Links |
|---|---|
| CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models | Paper · Project · Code |
| Reconstructing Humans with a Biomechanically Accurate Skeleton | Paper · Project · Code |
| MEGA: Masked Generative Autoencoder for Human Mesh Recovery | Paper |
| TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization | Paper · Project · Code |
| EgoLM: Multi-Modal Language Model of Egocentric Motions | Paper · Project |
| Paper | Links |
|---|---|
| Infinity∞: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis | Paper · Project · Code |
| RandAR: Decoder-only Autoregressive Visual Generation in Random Orders | Paper · Project · Code |
| Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models | Paper · Code |
| Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space | Paper · Project · Code |
| Autoregressive Distillation of Diffusion Transformers | Paper · Code |
| Language-Guided Image Tokenization for Generation | Paper · Project |
| AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea | Paper · Project · Code |
| CustAny: Customizing Anything from A Single Example | Paper · Project · Code |
| DreamRelation: Bridging Customization and Relation Generation | Paper · Project · Code |
| Minority-Focused Text-to-Image Generation via Prompt Optimization | Paper · Code |
| DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models | Paper |
| Motion Prompting: Controlling Video Generation with Motion Trajectories | Paper · Project |
| Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise | Paper · Project · Code |
| LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping | Paper · Project |
| Reanimating Images using Neural Representations of Dynamic Stimuli | Paper |
| Paper | Links |
|---|---|
| Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models | Paper · Blog · Code |
| Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens | Paper · Project · Code |
| LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models | Paper |
| OPA-DPO — Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key | Paper · Project · Code |
| Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding | Paper · Project · Code |
| Identifying and Mitigating Position Bias of Multi-image Vision-Language Models | Paper |
| Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces | Paper · Project · Code |
| Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding | Paper · Code |
| VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection | Paper · Code |
| OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation | Paper · Project · Code |
| Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content | Paper · Code |
| SEAL: Semantic Attention Learning for Long Video Representation | Paper · Code |
| Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval | Paper |
| Temporal Alignment-Free Video Matching for Few-shot Action Recognition | Paper · Code |
| Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning | Paper · Project |
| Temporally Consistent Object-Centric Learning by Contrasting Slots | Paper · Project · Code |
| The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recognition | Paper · Project |
| Paper | Links |
|---|---|
| Navigation World Models | Paper · Project · Code |
| RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics | Paper · Project · Code |
| From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons | Paper |
| PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation | Paper |
| GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill | Paper · Project · Code |
| Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models | Paper · Project · Code |
| Paper | Links |
|---|---|
| SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images | Paper · Project · Code |
| Effective SAM Combination for Open-Vocabulary Semantic Segmentation | Paper |
| Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation | Paper |
| Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weather | Paper |
| Efficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruning | Paper |
| Paper | Links |
|---|---|
| Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus Cues | Paper · Project · Code |
| Opportunistic Single-Photon Time of Flight | Paper |
| Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor Arrays | Paper · Project |
| Removing Reflections from RAW Photos | Paper · Project |
| DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-Resolution | Paper · Code |
| Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing | Paper · Project · Code |
| DiffFNO: Diffusion Fourier Neural Operator | Paper · Project |
| Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining | Paper |
| Paper | Links |
|---|---|
| OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | Paper · Code |
| CleanDIFT: Diffusion Features without Noise | Paper · Project · Code |
| LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions | Paper · Code |
| Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild | Paper |
| Rethinking Spiking Self-Attention Mechanism: Implementing α-XNOR Similarity Calculation in Spiking Transformers | Paper |
| Towards Universal Dataset Distillation via Task-Driven Diffusion | Paper |
| Gromov–Wasserstein Problem with Cyclic Symmetry | Paper |
| UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming | Paper |
| Enhancing Diversity for Data-free Quantization | Paper |
| Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learning | Paper · Code |
| Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models | Paper · Code |
| Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks | Paper · Code |
| Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector | Paper · Code |
| Paper | Links |
|---|---|
| TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model | Paper · Code |
| Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation | Paper |
| IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion Prior | Paper |
These entries are retained for reference and are not counted among the 95 papers in the official oral program.
| Paper | Links |
|---|---|
| Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation | Paper · Code |
| FedSPA: Generalizable Federated Graph Learning under Homophily Heterogeneity | Paper |
| One Category One Prompt: Dataset Distillation using Diffusion Models | Paper · Code |
| Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding | Paper · Code |
Missing a link or spotted a mistake? Open an issue or submit a pull request with the paper title and updated links.
Thanks to the paper authors for sharing their work, and to cvpr25_oral_gpu_info for the original collection reference.