ICCV 2025 Accepance Rate of 24% = 2699 / 11239
注1:欢迎各位大佬提交issue,分享ICCV 2025论文和开源项目!
注2:关于往年CV顶会论文以及其他优质CV论文和大盘点,详见: https://github.com/amusi/daily-paper-computer-vision
欢迎扫码加入【CVer学术交流群】,可以获取ICCV 2025等最前沿工作!这是最大的计算机视觉AI知识星球!每日更新,第一时间分享最新最前沿的计算机视觉、AIGC、扩散模型、多模态、深度学习、自动驾驶、医疗影像和遥感等方向的学习资料,快加入学起来!

TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
Where, What, Why: Towards Explainable Driver Attention Prediction
ROADWork Dataset: Learning to Recognize, Observe, Analyze and Drive Through Work Zones
DriveMM: All-in-One Large Multimodal Model for Autonomous Driving
EAMamba: Efficient All-Around Vision State Space Model for Image Restoration
#3D Visual Grounding(3D视觉定位)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
ROADWork Dataset: Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Music Grounding by Short Video
ICCV 2025 Accepance Rate of 24% = 2699 / 11239
注1:欢迎各位大佬提交issue,分享ICCV 2025论文和开源项目!
注2:关于往年CV顶会论文以及其他优质CV论文和大盘点,详见: https://github.com/amusi/daily-paper-computer-vision
欢迎扫码加入【CVer学术交流群】,可以获取ICCV 2025等最前沿工作!这是最大的计算机视觉AI知识星球!每日更新,第一时间分享最新最前沿的计算机视觉、AIGC、扩散模型、多模态、深度学习、自动驾驶、医疗影像和遥感等方向的学习资料,快加入学起来!

TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
Where, What, Why: Towards Explainable Driver Attention Prediction
ROADWork Dataset: Learning to Recognize, Observe, Analyze and Drive Through Work Zones
DriveMM: All-in-One Large Multimodal Model for Autonomous Driving
EAMamba: Efficient All-Around Vision State Space Model for Image Restoration
#3D Visual Grounding(3D视觉定位)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
ROADWork Dataset: Learning to Recognize, Observe, Analyze and Drive Through Work Zones
Music Grounding by Short Video