yejun688/CVPR_2025_Oral_Paper_List

😎 A curated list of CVPR 2025 Oral paper. Total 96

67

50 commits

updated Sep 6, 2026

See the code

README

🎉 CVPR 2025 Oral Paper List

A collection of 95 oral papers from CVPR 2025, organized by topic, with links to papers, project pages, and code.

CVPR 2025 Oral Papers

Official resources: Program · Proceedings · Awards

📚 Browse by Topic

🏆 Awards

Award labels follow the official CVPR 2025 results.

AwardPaper
Best PaperVGGT
Best Student PaperNeural Inverse Rendering from Propagating Light
Best Paper Honorable MentionMegaSaM
Best Paper Honorable MentionNavigation World Models
Best Paper Honorable MentionMolmo and PixMo
Best Paper Honorable Mention3D Student Splatting and Scooping
Best Student Paper Honorable MentionGenerative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

📐 3D Geometry & Reconstruction

PaperLinks
VGGT: Visual Geometry Grounded TransformerPaper · Project · Code
CUT3R — Continuous 3D Perception Model with Persistent StatePaper · Project · Code
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionPaper · Project · Code
FoundationStereo: Zero-Shot Stereo MatchingPaper · Project · Code
Multi-view Reconstruction via SfM-guided Monocular Depth EstimationPaper · Project · Code
MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsPaper · Project · Code
MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosPaper · Project · Code
Stereo4D: Learning How Things Move in 3D from Internet Stereo VideosPaper · Project · Code
Zero-Shot Monocular Scene Flow Estimation in the WildPaper · Project · Code
DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion ModelsPaper · Project · Code
TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage FusionPaper · Code
Convex Relaxation for Robust Vanishing Point Estimation in Manhattan WorldPaper · Code
Camera Resection from Known Line Pencils and a Radially Distorted ScanlinePaper · Code

✨ Rendering & 3D/4D Generation

PaperLinks
Neural Inverse Rendering from Propagating LightPaper · Project · Code
Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion ModelsPaper · Project · Code
3D Student Splatting and ScoopingPaper · Code
3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian SplattingPaper · Code
Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance FieldsPaper · Project · Code
FluidNexus: 3D Fluid Reconstruction and Prediction from a Single VideoPaper · Project · Code
CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry RefinerPaper · Project · Code
CAT4D: Create Anything in 4D with Multi-View Video Diffusion ModelsPaper · Project
DNF: Unconditional 4D Generation with Dictionary-based Neural FieldsPaper · Project · Code
Birth and Death of a RosePaper · Project

🧍 Humans, Avatars & Motion

PaperLinks
CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion ModelsPaper · Project · Code
Reconstructing Humans with a Biomechanically Accurate SkeletonPaper · Project · Code
MEGA: Masked Generative Autoencoder for Human Mesh RecoveryPaper
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task TokenizationPaper · Project · Code
EgoLM: Multi-Modal Language Model of Egocentric MotionsPaper · Project

🎨 Image & Video Generation

PaperLinks
Infinity∞: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image SynthesisPaper · Project · Code
RandAR: Decoder-only Autoregressive Visual Generation in Random OrdersPaper · Project · Code
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion ModelsPaper · Code
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent SpacePaper · Project · Code
Autoregressive Distillation of Diffusion TransformersPaper · Code
Language-Guided Image Tokenization for GenerationPaper · Project
AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaPaper · Project · Code
CustAny: Customizing Anything from A Single ExamplePaper · Project · Code
DreamRelation: Bridging Customization and Relation GenerationPaper · Project · Code
Minority-Focused Text-to-Image Generation via Prompt OptimizationPaper · Code
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion ModelsPaper
Motion Prompting: Controlling Video Generation with Motion TrajectoriesPaper · Project
Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped NoisePaper · Project · Code
LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPaper · Project
Reanimating Images using Neural Representations of Dynamic StimuliPaper

💬 Vision-Language & Video Understanding

PaperLinks
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsPaper · Blog · Code
Generative Multimodal Pretraining with Discrete Diffusion Timestep TokensPaper · Project · Code
LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language ModelsPaper
OPA-DPO — Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the KeyPaper · Project · Code
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingPaper · Project · Code
Identifying and Mitigating Position Bias of Multi-image Vision-Language ModelsPaper
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall SpacesPaper · Project · Code
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video UnderstandingPaper · Code
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame SelectionPaper · Code
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPaper · Project · Code
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision ContentPaper · Code
SEAL: Semantic Attention Learning for Long Video RepresentationPaper · Code
Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalPaper
Temporal Alignment-Free Video Matching for Few-shot Action RecognitionPaper · Code
Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningPaper · Project
Temporally Consistent Object-Centric Learning by Contrasting SlotsPaper · Project · Code
The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionPaper · Project

🤖 Embodied AI & Autonomous Driving

PaperLinks
Navigation World ModelsPaper · Project · Code
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for RoboticsPaper · Project · Code
From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsPaper
PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationPaper
GROVE: A Generalized Reward for Learning Open-Vocabulary Physical SkillPaper · Project · Code
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic ModelsPaper · Project · Code

🎯 Segmentation & Detection

PaperLinks
SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesPaper · Project · Code
Effective SAM Combination for Open-Vocabulary Semantic SegmentationPaper
Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationPaper
Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse WeatherPaper
Efficient Test-time Adaptive Object Detection via Sensitivity-Guided PruningPaper

📷 Computational Imaging & Restoration

PaperLinks
Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus CuesPaper · Project · Code
Opportunistic Single-Photon Time of FlightPaper
Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor ArraysPaper · Project
Removing Reflections from RAW PhotosPaper · Project
DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionPaper · Code
Improving Diffusion Inverse Problem Solving with Decoupled Noise AnnealingPaper · Project · Code
DiffFNO: Diffusion Fourier Neural OperatorPaper · Project
Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingPaper

🧠 Learning, Efficiency & Trustworthiness

PaperLinks
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic KernelsPaper · Code
CleanDIFT: Diffusion Features without NoisePaper · Project · Code
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer AttributionsPaper · Code
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the WildPaper
Rethinking Spiking Self-Attention Mechanism: Implementing α-XNOR Similarity Calculation in Spiking TransformersPaper
Towards Universal Dataset Distillation via Task-Driven DiffusionPaper
Gromov–Wasserstein Problem with Cyclic SymmetryPaper
UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic ProgrammingPaper
Enhancing Diversity for Data-free QuantizationPaper
Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated LearningPaper · Code
Black-Box Forgery Attacks on Semantic Watermarks for Diffusion ModelsPaper · Code
Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial AttacksPaper · Code
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorPaper · Code

🔬 Medical & Scientific Vision

PaperLinks
TopoCellGen: Generating Histopathology Cell Topology with a Diffusion ModelPaper · Code
Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image SegmentationPaper
IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion PriorPaper
4 additional papers from the original collection

These entries are retained for reference and are not counted among the 95 papers in the official oral program.

PaperLinks
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationPaper · Code
FedSPA: Generalizable Federated Graph Learning under Homophily HeterogeneityPaper
One Category One Prompt: Dataset Distillation using Diffusion ModelsPaper · Code
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video UnderstandingPaper · Code

🌷 Contributing & Acknowledgments

Missing a link or spotted a mistake? Open an issue or submit a pull request with the paper title and updated links.

Thanks to the paper authors for sharing their work, and to cvpr25_oral_gpu_info for the original collection reference.

autoregressive
backbone
computer-graphics
computer-vision
cvpr2025
diffusion-models
dust3r
foundation-models
gaussian-splatting
inverse-rendering
sam
vggt

yejun688/CVPR_2025_Oral_Paper_List

😎 A curated list of CVPR 2025 Oral paper. Total 96

67

50 commits

updated Sep 6, 2026

See the code

README

🎉 CVPR 2025 Oral Paper List

A collection of 95 oral papers from CVPR 2025, organized by topic, with links to papers, project pages, and code.

CVPR 2025 Oral Papers

Official resources: Program · Proceedings · Awards

📚 Browse by Topic

🏆 Awards

Award labels follow the official CVPR 2025 results.

AwardPaper
Best PaperVGGT
Best Student PaperNeural Inverse Rendering from Propagating Light
Best Paper Honorable MentionMegaSaM
Best Paper Honorable MentionNavigation World Models
Best Paper Honorable MentionMolmo and PixMo
Best Paper Honorable Mention3D Student Splatting and Scooping
Best Student Paper Honorable MentionGenerative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

📐 3D Geometry & Reconstruction

PaperLinks
VGGT: Visual Geometry Grounded TransformerPaper · Project · Code
CUT3R — Continuous 3D Perception Model with Persistent StatePaper · Project · Code
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionPaper · Project · Code
FoundationStereo: Zero-Shot Stereo MatchingPaper · Project · Code
Multi-view Reconstruction via SfM-guided Monocular Depth EstimationPaper · Project · Code
MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsPaper · Project · Code
MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic VideosPaper · Project · Code
Stereo4D: Learning How Things Move in 3D from Internet Stereo VideosPaper · Project · Code
Zero-Shot Monocular Scene Flow Estimation in the WildPaper · Project · Code
DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion ModelsPaper · Project · Code
TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage FusionPaper · Code
Convex Relaxation for Robust Vanishing Point Estimation in Manhattan WorldPaper · Code
Camera Resection from Known Line Pencils and a Radially Distorted ScanlinePaper · Code

✨ Rendering & 3D/4D Generation

PaperLinks
Neural Inverse Rendering from Propagating LightPaper · Project · Code
Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion ModelsPaper · Project · Code
3D Student Splatting and ScoopingPaper · Code
3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian SplattingPaper · Code
Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance FieldsPaper · Project · Code
FluidNexus: 3D Fluid Reconstruction and Prediction from a Single VideoPaper · Project · Code
CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry RefinerPaper · Project · Code
CAT4D: Create Anything in 4D with Multi-View Video Diffusion ModelsPaper · Project
DNF: Unconditional 4D Generation with Dictionary-based Neural FieldsPaper · Project · Code
Birth and Death of a RosePaper · Project

🧍 Humans, Avatars & Motion

PaperLinks
CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion ModelsPaper · Project · Code
Reconstructing Humans with a Biomechanically Accurate SkeletonPaper · Project · Code
MEGA: Masked Generative Autoencoder for Human Mesh RecoveryPaper
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task TokenizationPaper · Project · Code
EgoLM: Multi-Modal Language Model of Egocentric MotionsPaper · Project

🎨 Image & Video Generation

PaperLinks
Infinity∞: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image SynthesisPaper · Project · Code
RandAR: Decoder-only Autoregressive Visual Generation in Random OrdersPaper · Project · Code
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion ModelsPaper · Code
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent SpacePaper · Project · Code
Autoregressive Distillation of Diffusion TransformersPaper · Code
Language-Guided Image Tokenization for GenerationPaper · Project
AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaPaper · Project · Code
CustAny: Customizing Anything from A Single ExamplePaper · Project · Code
DreamRelation: Bridging Customization and Relation GenerationPaper · Project · Code
Minority-Focused Text-to-Image Generation via Prompt OptimizationPaper · Code
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion ModelsPaper
Motion Prompting: Controlling Video Generation with Motion TrajectoriesPaper · Project
Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped NoisePaper · Project · Code
LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPaper · Project
Reanimating Images using Neural Representations of Dynamic StimuliPaper

💬 Vision-Language & Video Understanding

PaperLinks
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsPaper · Blog · Code
Generative Multimodal Pretraining with Discrete Diffusion Timestep TokensPaper · Project · Code
LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language ModelsPaper
OPA-DPO — Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the KeyPaper · Project · Code
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingPaper · Project · Code
Identifying and Mitigating Position Bias of Multi-image Vision-Language ModelsPaper
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall SpacesPaper · Project · Code
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video UnderstandingPaper · Code
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame SelectionPaper · Code
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPaper · Project · Code
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision ContentPaper · Code
SEAL: Semantic Attention Learning for Long Video RepresentationPaper · Code
Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalPaper
Temporal Alignment-Free Video Matching for Few-shot Action RecognitionPaper · Code
Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningPaper · Project
Temporally Consistent Object-Centric Learning by Contrasting SlotsPaper · Project · Code
The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionPaper · Project

🤖 Embodied AI & Autonomous Driving

PaperLinks
Navigation World ModelsPaper · Project · Code
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for RoboticsPaper · Project · Code
From Multimodal LLMs to Generalist Embodied Agents: Methods and LessonsPaper
PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationPaper
GROVE: A Generalized Reward for Learning Open-Vocabulary Physical SkillPaper · Project · Code
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic ModelsPaper · Project · Code

🎯 Segmentation & Detection

PaperLinks
SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing ImagesPaper · Project · Code
Effective SAM Combination for Open-Vocabulary Semantic SegmentationPaper
Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationPaper
Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse WeatherPaper
Efficient Test-time Adaptive Object Detection via Sensitivity-Guided PruningPaper

📷 Computational Imaging & Restoration

PaperLinks
Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus CuesPaper · Project · Code
Opportunistic Single-Photon Time of FlightPaper
Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor ArraysPaper · Project
Removing Reflections from RAW PhotosPaper · Project
DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionPaper · Code
Improving Diffusion Inverse Problem Solving with Decoupled Noise AnnealingPaper · Project · Code
DiffFNO: Diffusion Fourier Neural OperatorPaper · Project
Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingPaper

🧠 Learning, Efficiency & Trustworthiness

PaperLinks
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic KernelsPaper · Code
CleanDIFT: Diffusion Features without NoisePaper · Project · Code
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer AttributionsPaper · Code
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the WildPaper
Rethinking Spiking Self-Attention Mechanism: Implementing α-XNOR Similarity Calculation in Spiking TransformersPaper
Towards Universal Dataset Distillation via Task-Driven DiffusionPaper
Gromov–Wasserstein Problem with Cyclic SymmetryPaper
UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic ProgrammingPaper
Enhancing Diversity for Data-free QuantizationPaper
Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated LearningPaper · Code
Black-Box Forgery Attacks on Semantic Watermarks for Diffusion ModelsPaper · Code
Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial AttacksPaper · Code
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorPaper · Code

🔬 Medical & Scientific Vision

PaperLinks
TopoCellGen: Generating Histopathology Cell Topology with a Diffusion ModelPaper · Code
Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image SegmentationPaper
IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion PriorPaper
4 additional papers from the original collection

These entries are retained for reference and are not counted among the 95 papers in the official oral program.

PaperLinks
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationPaper · Code
FedSPA: Generalizable Federated Graph Learning under Homophily HeterogeneityPaper
One Category One Prompt: Dataset Distillation using Diffusion ModelsPaper · Code
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video UnderstandingPaper · Code

🌷 Contributing & Acknowledgments

Missing a link or spotted a mistake? Open an issue or submit a pull request with the paper title and updated links.

Thanks to the paper authors for sharing their work, and to cvpr25_oral_gpu_info for the original collection reference.

autoregressive
backbone
computer-graphics
computer-vision
cvpr2025
diffusion-models
dust3r
foundation-models
gaussian-splatting
inverse-rendering
sam
vggt