Paper2Chinese/CVPR-2026-reading-papers-with-code

218

24 commits

updated Mar 22, 2026

See the code

README

History

CVPR-2025-reading-papers-with-code

CVPR-2026-reading-papers-with-code

收集CVPR 2026论文&源码

收集全网对CVPR 2026论文的优质讲解


注1:欢迎各位作者大佬提交issue,分享CVPR 2026论文和开源项目!

注2:关于CV领域顶级期刊(TPAMI、IJCV等)论文解读大盘点,详见: https://github.com/Paper2Chinese/Paper2Chinese

注3:关于人工智能领域NeurIPS顶会论文解读大盘点,详见: https://github.com/Paper2Chinese/NeurIPS2024-Reading-Paper-With-Code

【CVPR 2026 论文开源目录】

3DGS(Gaussian Splatting)Mamba / (SSM)AvatarsBackboneCLIPMAE联邦学习(Federated Learning)
多模态大语言模型(MLLM)大语言模型(LLM)视觉语言模型(VLM)多模态(Multi-modal)NASOCRNeRF
视觉问答(Visual Question Answering)强化学习(Reinforcement Learning)扩散模型(Diffusion Models)ReID(重识别)长尾分布(Long-Tail)视频压缩(Video Compression)
增量学习(Incremental Learning)数据增强(Data Augmentation)目标检测(Object Detection)异常检测(Anomaly Detection)目标跟踪(Visual Tracking)语义分割(Semantic Segmentation)实例分割(Instance Segmentation)
医学图像(Medical Image)医学图像分割(Medical Image Segmentation)视频目标分割(Video Object Segmentation)视频实例分割(Video Instance Segmentation)参考图像分割(Referring Image Segmentation)图像抠图(Image Matting)图像编辑(Image Editing)
具身智能Prompt自监督学习(Self-supervised Learning)生物工程(bioengineering)Low-level Vision超分辨率(Super-Resolution)去模糊(Deblur)
生成对抗网络(GAN)3D点云(3D Point Cloud)3D目标检测(3D Object Detection)3D语义分割(3D Semantic Segmentation)3D目标跟踪(3D Object Tracking)3D语义场景补全(3D Semantic Scene Completion)视频理解(Video Understanding)
3D人体姿态估计(3D Human Pose Estimation)3D人体Mesh估计(3D Human Mesh Estimation)少样本学习(Few-Shot Learning)图像生成(Image Generation)视频生成(Video Generation)3D生成(3D Generation)图像压缩(Image Compression)
持续学习(Continual Learning)行为识别(Action Recognition)行为检测(Action Detection)人脸识别(Face Recognition)文本检测(Text Detection)知识蒸馏(Knowledge Distillation)三维重建(3D Reconstruction)
GNNDETRVision Transformer全景分割(Panoptic Segmentation)去噪(Denoising)自动驾驶(Autonomous Driving)3D配准(3D Registration)
模型剪枝(Model Pruning)深度估计(Depth Estimation)轨迹预测(Trajectory Prediction)车道线检测(Lane Detection)图像描述(Image Captioning)手语识别(Sign Language Recognition)视频预测(Video Prediction)
新视点合成(Novel View Synthesis)Zero-Shot Learning(零样本学习)立体匹配(Stereo Matching)特征匹配(Feature Matching)场景图生成(Scene Graph Generation)计数(Counting)隐式神经表示(Implicit Neural Representations)
图像质量评价(Image Quality Assessment)视频质量评价(Video Quality Assessment)数据集(Datasets)反学习(Machine Unlearning)新任务(New Tasks)模型加速(Improving Reasoning)时间序列(Time Series)
其他(Others)脉冲网络图像检索图像去雾(Dehazing)

图像去雾(Dehazing)

Bilevel Layer-Positioning LoRA for Real Image Dehazing

UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

具身智能(Embodied AI)

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation

RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual Reasoning

Structural Action Transformer for 3D Dexterous Manipulation

Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation

GeCo-SRT: Geometry-aware Continual Adaptation for Robotic Cross-Task Sim-to-Real Transfer

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

3DGS(Gaussian Splatting)

BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting

CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Image

ReLaGS: Relational Language Gaussian Splatting

Motion-Aware Animatable Gaussian Avatars Deblurring

LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates

EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopy

PhyGaP: Physically-Grounded Gaussians with Polarization Cues

Speeding Up the Learning of 3D Gaussians with Much Shorter Gaussian Lists

REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM

E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction

ProgressiveAvatars: Progressive Animatable 3D Gaussian Avatars

Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation

OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution

STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction

Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting

RAP: Fast Feedforward Rendering-Free Attribute-Guided Primitive Importance Score Prediction for Efficient 3D Gaussian Splatting Processing

三维重建(3D Reconstruction)

GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

Parallelised Differentiable Straightest Geodesics for 3D Meshes

FoV-Net: Rotation-Invariant CAD B-rep Learning via Field-of-View Ray Casting

A2Z-10M+: Geometric Deep Learning with A-to-Z BRep Annotations for AI-Assisted CAD Modeling and Reverse Engineering

Order Matters: 3D Shape Generation from Sequential VR Sketches

CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customization

PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery

SimRecon: SimReady Compositional Scene Reconstruction from Real Videos

MoRe: Motion-aware Feed-forward 4D Reconstruction Transformer

RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

FoV-Net: Rotation-Invariant CAD B-rep Learning via Field-of-View Ray Casting

Global-Aware Edge Prioritization for Pose Graph Initialization

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

模型剪枝(Model Pruning)

When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs

Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving

Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

深度估计(Depth Estimation)

SpiderCam: Low-Power Snapshot Depth from Differential Defocus

轨迹预测(Trajectory Prediction)

FoSS: Modeling Long Range Dependencies and Multimodal Uncertainty in Trajectory Prediction via Fourier State Space Integration

Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction

Mamba / SSM

DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection

Avatars

自动驾驶

AdaRadar: Rate Adaptive Spectral Compression for Radar-based Perception

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

SimScale: Learning to Drive via Real-World Simulation at Scale

VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation

All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception

HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation

KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System

SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance

CoLC: Communication-Efficient Collaborative Perception with LiDAR Completion

Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

LiREC-Net: A Target-Free and Learning-Based Network for LiDAR, RGB, and Event Calibration

RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space

HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning

Perception Characteristics Distance: Measuring Stability and Robustness of Perception System in Dynamic Conditions under a Certain Decision Rule

SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World

Backbone

Hyperbolic Busemann Neural Networks

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

CLIP

FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data Selection

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy

MAE

Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification

SARMAE: Masked Autoencoder for SAR Representation Learning

OCR

What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

Efficient Document Parsing via Parallel Token Prediction

D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping

Occupancy

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

NeRF

Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesis

Node-RF: Learning Generalized Continuous Space-Time Scene Dynamics with Neural ODE-based NeRFs

Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Events

NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code

DETR

EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer

GNN

Prompt

Towards Calibrating Prompt Tuning of Vision-Language Models

PHAC: Promptable Human Amodal Completion

FOZO: Forward-Only Zeroth-Order Prompt Optimization for Test-Time Adaptation

FOZO: Forward-Only Zeroth-Order Prompt Optimization for Test-Time Adaptation

大语言模型(LLM)

VecGlypher: Unified Vector Glyph Generation with Language Models

LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference

视觉语言模型(LLM)

Draft and Refine with Visual Experts

Interpretable Debiasing of Vision-Language Models for Social Fairness

Ego: Embedding-Guided Personalization of Vision-Language Models

GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training

V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs

It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles

Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks

Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization

多模态大语言模型(MLLM)

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients

How to Take a Memorable Picture? Empowering Users with Actionable Feedback

Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token

MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction

HIFICL: High-Fidelity In-Context Learning for Multimodal Tasks

Parallel In-context Learning for Large Vision Language Models

Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans

WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence

LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models

Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents

Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought

ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking

See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition

Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation

Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs

MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs

MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping

多模态

ConsistCompose: Unified Multimodal Layout Control for Image Composition

Mixture of States: Routing Token-Level Dynamics for Multimodal Generation

Intrinsic Concept Extraction Based on Compositional Interpretability

ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis

x²-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space

EI: Early Intervention for Multimodal Imaging based Disease Recognition

Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation

Linking Modality Isolation in Heterogeneous Collaborative Perception

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

Beyond Global Similarity: Towards Fine-Grained, Multi-Condition Multimodal Retrieval

Adaptive Confidence Regularization for Multimodal Failure Detection

Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning

Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery

CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion

NAS

视觉问答(Visual Question Answering)

Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models

强化学习(Reinforcement Learning)

Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial Odometry

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360°Image Quality Assessment

Specificity-aware reinforcement learning for fine-grained open-world classification

From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation

ReID(重识别)

长尾分布(Long-Tail)

Hier-COS: Making Deep Features Hierarchy-aware via Composition of Orthogonal Subspaces

Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning

视频压缩(Video Compression)

扩散模型(Diffusion Models)

Prototype-Guided Concept Erasure in Diffusion Models

Pixel Motion Diffusion is What We Need for Robot Control

CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

TAUE: Training-free Noise Transplant and Cultivation Diffusion Model

3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal Diffusion

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

CoD: A Diffusion Foundation Model for Image Compression

SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras

All-in-One Slider for Attribute Manipulation in Diffusion Models

LESA: Learnable Stage-Aware Predictors for Diffusion Model Acceleration

When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters

SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion Transformer

Guiding Diffusion Models with Semantically Degraded Conditions

COT-FM: Cluster-wise Optimal Transport Flow Matching

Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning

Making Training-Free Diffusion Segmentors Scale with the Generative Power

Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens

Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generation

ADAPT: Attention Driven Adaptive Prompt Scheduling and InTerpolating Orthogonal Complements for Rare Concepts Generation

All-in-One Slider for Attribute Manipulation in Diffusion Models

TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models

Elucidating the Design Space of Arbitrary-Noise-Based Diffusion Models

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization

SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models

Vision Transformer

Make it SING: Analyzing Semantic Invariants in Classifiers

Revisiting Model Stitching In the Foundation Model Era

BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers

MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy

全景分割(Panoptic Segmentation)

Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentation

视觉和语言(Vision-Language)

VL-RouterBench: A Benchmark for Vision-Language Model Routing

CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network

HoneyBee: Data Recipes for Vision-Language Reasoners

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

目标检测(Object Detection)

Prompt-Free Universal Region Proposal Network

Does YOLO Really Need to See Every Training Image in Every Epoch?

Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection

数据增强(Data Augmentation)

Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

异常检测(Anomaly Detection)

MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection

EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection

RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer

Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning

Diversity over Uniformity: Rethinking Representation in Generated Image Detection

The Invisible Gorilla Effect in Out-of-distribution Detection

GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

目标跟踪(Object Tracking)

UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking

Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking

Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking

Changes in Real Time: Online Scene Change Detection with Multi-View Fusion

Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction

语义分割(Semantic Segmentation)

ACPV-Net: All-Class Polygonal Vectorization for Seamless Vector Map Generation from Aerial Imagery

Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels

MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention

Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wild

Discriminative Perception via Anchored Description for Reasoning Segmentation

实例分割(Instance Segmentation)

少样本学习(Few-Shot Learning)

Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection

MAGIC: Few-Shot Mask-Guided Anomaly Inpainting with Prompt Perturbation, Spatially Adaptive Guidance, and Context Awareness

SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation

MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification

Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection

生物医学

医学图像(Medical Image)

Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosis

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Code

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data

Every Error has Its Magnitude: Asymmetric Mistake Severity Training for Multiclass Multiple Instance Learning

X-WIN: Building Chest Radiograph World Model via Predictive Sensing

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Cell-Type Prototype-Informed Neural Network for Gene Expression Estimation from Pathology Images

Benchmarking Endoscopic Surgical Image Restoration and Beyond

MRI Contrast Enhancement Kinetics World Model

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priors

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

医学图像分割(Medical Image Segmentation)

Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model

SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation

视频目标分割(Video Object Segmentation)

MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator

Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos

行为检测(Action Detection)

SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion

人脸识别(Face Recognition)

RecoverMark: Robust Watermarking for Localization and Recovery of Manipulated Faces

IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbation

3D点云(3D-Point-Cloud)

Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis

Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors

Universal 3D Shape Matching via Coarse-to-Fine Language Guidance

QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

自监督学习(Self-supervised Learning)

BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images

Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

生物工程(bioengineering)

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

联邦学习(Federated Learning)

Federated Active Learning Under Extreme Non-IID and Global Class Imbalance

Domain-Skewed Federated Learning with Feature Decoupling and Calibration

Fed-ADE: Adaptive Learning Rate for Federated Post-adaptation under Distribution Shift

HiLoRA: Hierarchical Low-Rank Adaptation for Personalized Federated Learning

FedVG: Gradient-Guided Aggregation for Enhanced Federated Learning

FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation

增量学习(Incremental Learning)

3D目标检测(3D Object Detection)

Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection

VirPro: Visual-referred Probabilistic Prompt Learning for Weakly-Supervised Monocular 3D Detection

CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection

R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection

SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection

Learning Mutual View Information Graph for Adaptive Adversarial Collaborative Perception

SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

3D语义分割(3D Semantic Segmentation)

图像编辑(Image Editing)

Precise Object and Effect Removal with Adaptive Target-Aware Attention

CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing

Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing

Cycle-Consistent Tuning for Layered Image Decomposition

BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modeling

owards Source-Aware Object Swapping with Initial Noise Perturbation

Cycle-Consistent Tuning for Layered Image Decomposition

ChordEdit: One-Step Low-Energy Transport for Image Editing

Towards Source-Aware Object Swapping with Initial Noise Perturbation

图像补全/图像修复(Image Inpainting)

1. HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

生成对抗网络(GAN)

视频编辑(Video Editing)

Object-WIPER: Training-Free Object and Associated Effect Removal in Videos

HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles

Low-level Vision

ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration

F²HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

BluRef: Unsupervised Image Deblurring with Dense-Matching References

Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration

Reparameterized Tensor Ring Functional Decomposition for Multi-Dimensional Data Recovery

Lumosaic: Hyperspectral Video via Active Illumination and Coded-Exposure Pixels

MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision

Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis

超分辨率(Super-Resolution)

RAW-Domain Degradation Models for Realistic Smartphone Super-Resolution

UCAN: Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-Resolution

FiDeSR: High-Fidelity and Detail-Preserving One-Step Diffusion Super-Resolution

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution

Spectral Super-Resolution via Adversarial Unfolding and Data-Driven Spectrum Regularization: From Multispectral Satellite Data to NASA Hyperspectral Image

AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution

去噪(Denoising)

Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imaging

图像生成(Image Generation)

coDrawAgents: A Multi-Agent Dialogue Framework for Compositional Image Generation

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards

VeCoR -- Velocity Contrastive Regularization for Flow Matching

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards

Enhancing Spatial Understanding in Image Generation via Reward Modeling

AutoDebias: Automated Framework for Debiasing Text-to-Image Models

视频生成(Video Generation)

Anchoring and Rescaling Attention for Semantically Coherent Inbetweening

Training-free Motion Factorization for Compositional Video Generation

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters

EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing

CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video

FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning

UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation

ExpPortrait: Expressive Portrait Generation via Personalized Representation

LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation

FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

3D生成

Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editing

NI-Tex: Non-isometric Image-based Garment Texture Generation

ForgeDreamer: Industrial Text-to-3D Generation with Multi-Expert LoRA and Cross-View Hypergraph

Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single Image

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation

MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing

Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

视频理解(Video Understanding)

F2HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling

Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding

StreamingTOM: Streaming Token Compression for Efficient Video Understanding

Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding

StreamReady: Learning What to Answer and When in Long Streaming Videos

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding

ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark

Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding

FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

3D人体姿态估计(3D Human Pose Estimation)

GazeOnce360: Fisheye-Based 360° Multi-Person Gaze Estimation with Global-Local Feature Fusion

OnlineHMR: Video-based Online World-Grounded Human Mesh Recovery

Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation

Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation

Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator

CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation

SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking

VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery

持续学习(Continual Learning)

Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment

Elastic Weight Consolidation Done Right for Continual Learning

行为识别(Action Recognition)

BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment

Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency

知识蒸馏(Knowledge Distillation)

Momentum Memory for Knowledge Distillation in Computational Pathology

WaDi: Weight Direction-aware Distillation for One-step Image Synthesis

Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation

Distilling Balanced Knowledge from a Biased Teacher

Momentum Memory for Knowledge Distillation in Computational Pathology

图像压缩(Image Compression)

SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representation

Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression

Zero-Shot Learning(零样本学习)

Learning through Creation: A Hash-Free Framework for On-the-Fly Category Discovery

TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery

Learning through Creation: A Hash-Free Framework for On-the-Fly Category Discovery

Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

立体匹配(Stereo Matching)

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo Matching

场景图生成(Scene Graph Generation)

DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime

计数(Counting)

UNICBench: UNIfied Counting Benchmark for MLLM

隐式神经表示(Implicit Neural Representations)

图像质量评价(Image Quality Assessment)

How to Take a Memorable Picture? Empowering Users with Actionable Feedback

Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks

视频质量评价(Video Quality Assessment)

数据集(Datasets)

E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought

Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis

LenghuSky-8: An 8-Year All-Sky Cloud Dataset with Star-Aware Masks and Alt-Az Calibration for Segmentation and Nowcasting

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

反学习(Machine Unlearning)

SineProject: Machine Unlearning for Stable Vision Language Alignment

RAZOR: Ratio-Aware Layer Editing for Targeted Unlearning in Vision Transformers and Diffusion Models

Stake the Points: Structure-Faithful Instance Unlearning

新任务(New Tasks)

模型加速(Improving Reasoning)

Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models

Variation-aware Vision Token Dropping for Faster Large Vision-Language Models

Model Merging in the Essential Subspace

时间序列(Time Series)

STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting

脉冲网络

Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networks

Rethinking SNN Online Training and Deployment: Gradient-Coherent Learning via Hybrid-Driven LIF Model

图像检索

TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking

PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing

NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries

其他(Others)

NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training

Defending Unauthorized Model Merging via Dual-Stage Weight Protection

BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning

ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimation

GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement

DC-Merge: Improving Model Merging with Directional Consistency

Defending Unauthorized Model Merging via Dual-Stage Weight Protection

BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning

Bridging Domains through Subspace-Aware Model Merging

Significant stargazers

Matiur Rahman Minar

155 followers · starred Mar 2026

Paper2Chinese/CVPR-2026-reading-papers-with-code

218

24 commits

updated Mar 22, 2026

See the code

README

History

CVPR-2025-reading-papers-with-code

CVPR-2026-reading-papers-with-code

收集CVPR 2026论文&源码

收集全网对CVPR 2026论文的优质讲解


注1:欢迎各位作者大佬提交issue,分享CVPR 2026论文和开源项目!

注2:关于CV领域顶级期刊(TPAMI、IJCV等)论文解读大盘点,详见: https://github.com/Paper2Chinese/Paper2Chinese

注3:关于人工智能领域NeurIPS顶会论文解读大盘点,详见: https://github.com/Paper2Chinese/NeurIPS2024-Reading-Paper-With-Code

【CVPR 2026 论文开源目录】

3DGS(Gaussian Splatting)Mamba / (SSM)AvatarsBackboneCLIPMAE联邦学习(Federated Learning)
多模态大语言模型(MLLM)大语言模型(LLM)视觉语言模型(VLM)多模态(Multi-modal)NASOCRNeRF
视觉问答(Visual Question Answering)强化学习(Reinforcement Learning)扩散模型(Diffusion Models)ReID(重识别)长尾分布(Long-Tail)视频压缩(Video Compression)
增量学习(Incremental Learning)数据增强(Data Augmentation)目标检测(Object Detection)异常检测(Anomaly Detection)目标跟踪(Visual Tracking)语义分割(Semantic Segmentation)实例分割(Instance Segmentation)
医学图像(Medical Image)医学图像分割(Medical Image Segmentation)视频目标分割(Video Object Segmentation)视频实例分割(Video Instance Segmentation)参考图像分割(Referring Image Segmentation)图像抠图(Image Matting)图像编辑(Image Editing)
具身智能Prompt自监督学习(Self-supervised Learning)生物工程(bioengineering)Low-level Vision超分辨率(Super-Resolution)去模糊(Deblur)
生成对抗网络(GAN)3D点云(3D Point Cloud)3D目标检测(3D Object Detection)3D语义分割(3D Semantic Segmentation)3D目标跟踪(3D Object Tracking)3D语义场景补全(3D Semantic Scene Completion)视频理解(Video Understanding)
3D人体姿态估计(3D Human Pose Estimation)3D人体Mesh估计(3D Human Mesh Estimation)少样本学习(Few-Shot Learning)图像生成(Image Generation)视频生成(Video Generation)3D生成(3D Generation)图像压缩(Image Compression)
持续学习(Continual Learning)行为识别(Action Recognition)行为检测(Action Detection)人脸识别(Face Recognition)文本检测(Text Detection)知识蒸馏(Knowledge Distillation)三维重建(3D Reconstruction)
GNNDETRVision Transformer全景分割(Panoptic Segmentation)去噪(Denoising)自动驾驶(Autonomous Driving)3D配准(3D Registration)
模型剪枝(Model Pruning)深度估计(Depth Estimation)轨迹预测(Trajectory Prediction)车道线检测(Lane Detection)图像描述(Image Captioning)手语识别(Sign Language Recognition)视频预测(Video Prediction)
新视点合成(Novel View Synthesis)Zero-Shot Learning(零样本学习)立体匹配(Stereo Matching)特征匹配(Feature Matching)场景图生成(Scene Graph Generation)计数(Counting)隐式神经表示(Implicit Neural Representations)
图像质量评价(Image Quality Assessment)视频质量评价(Video Quality Assessment)数据集(Datasets)反学习(Machine Unlearning)新任务(New Tasks)模型加速(Improving Reasoning)时间序列(Time Series)
其他(Others)脉冲网络图像检索图像去雾(Dehazing)

图像去雾(Dehazing)

Bilevel Layer-Positioning LoRA for Real Image Dehazing

UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

具身智能(Embodied AI)

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation

RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual Reasoning

Structural Action Transformer for 3D Dexterous Manipulation

Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation

GeCo-SRT: Geometry-aware Continual Adaptation for Robotic Cross-Task Sim-to-Real Transfer

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

3DGS(Gaussian Splatting)

BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting

CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Image

ReLaGS: Relational Language Gaussian Splatting

Motion-Aware Animatable Gaussian Avatars Deblurring

LTGS: Long-Term Gaussian Scene Chronology From Sparse View Updates

EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopy

PhyGaP: Physically-Grounded Gaussians with Polarization Cues

Speeding Up the Learning of 3D Gaussians with Much Shorter Gaussian Lists

REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM

E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction

ProgressiveAvatars: Progressive Animatable 3D Gaussian Avatars

Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation

OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution

STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction

Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting

RAP: Fast Feedforward Rendering-Free Attribute-Guided Primitive Importance Score Prediction for Efficient 3D Gaussian Splatting Processing

三维重建(3D Reconstruction)

GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

Parallelised Differentiable Straightest Geodesics for 3D Meshes

FoV-Net: Rotation-Invariant CAD B-rep Learning via Field-of-View Ray Casting

A2Z-10M+: Geometric Deep Learning with A-to-Z BRep Annotations for AI-Assisted CAD Modeling and Reverse Engineering

Order Matters: 3D Shape Generation from Sequential VR Sketches

CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customization

PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery

SimRecon: SimReady Compositional Scene Reconstruction from Real Videos

MoRe: Motion-aware Feed-forward 4D Reconstruction Transformer

RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

FoV-Net: Rotation-Invariant CAD B-rep Learning via Field-of-View Ray Casting

Global-Aware Edge Prioritization for Pose Graph Initialization

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

模型剪枝(Model Pruning)

When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs

Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving

Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

深度估计(Depth Estimation)

SpiderCam: Low-Power Snapshot Depth from Differential Defocus

轨迹预测(Trajectory Prediction)

FoSS: Modeling Long Range Dependencies and Multimodal Uncertainty in Trajectory Prediction via Fourier State Space Integration

Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory Prediction

Mamba / SSM

DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection

Avatars

自动驾驶

AdaRadar: Rate Adaptive Spectral Compression for Radar-based Perception

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

SimScale: Learning to Drive via Real-World Simulation at Scale

VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation

All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception

HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation

KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System

SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance

CoLC: Communication-Efficient Collaborative Perception with LiDAR Completion

Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

LiREC-Net: A Target-Free and Learning-Based Network for LiDAR, RGB, and Event Calibration

RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space

HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles

NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning

Perception Characteristics Distance: Measuring Stability and Robustness of Perception System in Dynamic Conditions under a Certain Decision Rule

SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World

Backbone

Hyperbolic Busemann Neural Networks

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

CLIP

FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data Selection

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

FluoCLIP: Stain-Aware Focus Quality Assessment in Fluorescence Microscopy

MAE

Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification

SARMAE: Masked Autoencoder for SAR Representation Learning

OCR

What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

Efficient Document Parsing via Parallel Token Prediction

D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping

Occupancy

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

NeRF

Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesis

Node-RF: Learning Generalized Continuous Space-Time Scene Dynamics with Neural ODE-based NeRFs

Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Events

NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code

DETR

EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer

GNN

Prompt

Towards Calibrating Prompt Tuning of Vision-Language Models

PHAC: Promptable Human Amodal Completion

FOZO: Forward-Only Zeroth-Order Prompt Optimization for Test-Time Adaptation

FOZO: Forward-Only Zeroth-Order Prompt Optimization for Test-Time Adaptation

大语言模型(LLM)

VecGlypher: Unified Vector Glyph Generation with Language Models

LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic Inference

视觉语言模型(LLM)

Draft and Refine with Visual Experts

Interpretable Debiasing of Vision-Language Models for Social Fairness

Ego: Embedding-Guided Personalization of Vision-Language Models

GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training

V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs

It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles

Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks

Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness

Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization

多模态大语言模型(MLLM)

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients

How to Take a Memorable Picture? Empowering Users with Actionable Feedback

Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token

MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction

HIFICL: High-Fidelity In-Context Learning for Multimodal Tasks

Parallel In-context Learning for Large Vision Language Models

Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans

WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence

LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models

Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents

Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought

ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking

See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles

Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition

Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation

Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs

MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs

MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping

多模态

ConsistCompose: Unified Multimodal Layout Control for Image Composition

Mixture of States: Routing Token-Level Dynamics for Multimodal Generation

Intrinsic Concept Extraction Based on Compositional Interpretability

ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis

x²-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Space

EI: Early Intervention for Multimodal Imaging based Disease Recognition

Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation

Linking Modality Isolation in Heterogeneous Collaborative Perception

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

Beyond Global Similarity: Towards Fine-Grained, Multi-Condition Multimodal Retrieval

Adaptive Confidence Regularization for Multimodal Failure Detection

Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

MultiModalPFN: Extending Prior-Data Fitted Networks for Multimodal Tabular Learning

Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery

CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion

NAS

视觉问答(Visual Question Answering)

Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models

强化学习(Reinforcement Learning)

Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial Odometry

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360°Image Quality Assessment

Specificity-aware reinforcement learning for fine-grained open-world classification

From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation

ReID(重识别)

长尾分布(Long-Tail)

Hier-COS: Making Deep Features Hierarchy-aware via Composition of Orthogonal Subspaces

Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning

视频压缩(Video Compression)

扩散模型(Diffusion Models)

Prototype-Guided Concept Erasure in Diffusion Models

Pixel Motion Diffusion is What We Need for Robot Control

CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

TAUE: Training-free Noise Transplant and Cultivation Diffusion Model

3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal Diffusion

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

CoD: A Diffusion Foundation Model for Image Compression

SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras

All-in-One Slider for Attribute Manipulation in Diffusion Models

LESA: Learnable Stage-Aware Predictors for Diffusion Model Acceleration

When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters

SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion Transformer

Guiding Diffusion Models with Semantically Degraded Conditions

COT-FM: Cluster-wise Optimal Transport Flow Matching

Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning

Making Training-Free Diffusion Segmentors Scale with the Generative Power

Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens

Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generation

ADAPT: Attention Driven Adaptive Prompt Scheduling and InTerpolating Orthogonal Complements for Rare Concepts Generation

All-in-One Slider for Attribute Manipulation in Diffusion Models

TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models

Elucidating the Design Space of Arbitrary-Noise-Based Diffusion Models

TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration

CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization

SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Models

Vision Transformer

Make it SING: Analyzing Semantic Invariants in Classifiers

Revisiting Model Stitching In the Foundation Model Era

BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers

MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy

全景分割(Panoptic Segmentation)

Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentation

视觉和语言(Vision-Language)

VL-RouterBench: A Benchmark for Vision-Language Model Routing

CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network

HoneyBee: Data Recipes for Vision-Language Reasoners

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

目标检测(Object Detection)

Prompt-Free Universal Region Proposal Network

Does YOLO Really Need to See Every Training Image in Every Epoch?

Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection

数据增强(Data Augmentation)

Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

异常检测(Anomaly Detection)

MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection

EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection

RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer

Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning

Diversity over Uniformity: Rethinking Representation in Generated Image Detection

The Invisible Gorilla Effect in Out-of-distribution Detection

GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

目标跟踪(Object Tracking)

UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking

Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking

Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking

Changes in Real Time: Online Scene Change Detection with Multi-View Fusion

Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction

语义分割(Semantic Segmentation)

ACPV-Net: All-Class Polygonal Vectorization for Seamless Vector Map Generation from Aerial Imagery

Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels

MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention

Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wild

Discriminative Perception via Anchored Description for Reasoning Segmentation

实例分割(Instance Segmentation)

少样本学习(Few-Shot Learning)

Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection

MAGIC: Few-Shot Mask-Guided Anomaly Inpainting with Prompt Perturbation, Spatially Adaptive Guidance, and Context Awareness

SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation

MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification

Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection

生物医学

医学图像(Medical Image)

Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosis

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Code

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data

Every Error has Its Magnitude: Asymmetric Mistake Severity Training for Multiclass Multiple Instance Learning

X-WIN: Building Chest Radiograph World Model via Predictive Sensing

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Cell-Type Prototype-Informed Neural Network for Gene Expression Estimation from Pathology Images

Benchmarking Endoscopic Surgical Image Restoration and Beyond

MRI Contrast Enhancement Kinetics World Model

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priors

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

医学图像分割(Medical Image Segmentation)

Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Model

SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation

视频目标分割(Video Object Segmentation)

MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator

Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos

行为检测(Action Detection)

SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion

人脸识别(Face Recognition)

RecoverMark: Robust Watermarking for Localization and Recovery of Manipulated Faces

IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbation

3D点云(3D-Point-Cloud)

Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis

Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors

Universal 3D Shape Matching via Coarse-to-Fine Language Guidance

QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

自监督学习(Self-supervised Learning)

BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images

Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

生物工程(bioengineering)

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

联邦学习(Federated Learning)

Federated Active Learning Under Extreme Non-IID and Global Class Imbalance

Domain-Skewed Federated Learning with Feature Decoupling and Calibration

Fed-ADE: Adaptive Learning Rate for Federated Post-adaptation under Distribution Shift

HiLoRA: Hierarchical Low-Rank Adaptation for Personalized Federated Learning

FedVG: Gradient-Guided Aggregation for Enhanced Federated Learning

FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation

增量学习(Incremental Learning)

3D目标检测(3D Object Detection)

Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection

VirPro: Visual-referred Probabilistic Prompt Learning for Weakly-Supervised Monocular 3D Detection

CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection

R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection

SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection

Learning Mutual View Information Graph for Adaptive Adversarial Collaborative Perception

SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

3D语义分割(3D Semantic Segmentation)

图像编辑(Image Editing)

Precise Object and Effect Removal with Adaptive Target-Aware Attention

CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing

Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing

Cycle-Consistent Tuning for Layered Image Decomposition

BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modeling

owards Source-Aware Object Swapping with Initial Noise Perturbation

Cycle-Consistent Tuning for Layered Image Decomposition

ChordEdit: One-Step Low-Energy Transport for Image Editing

Towards Source-Aware Object Swapping with Initial Noise Perturbation

图像补全/图像修复(Image Inpainting)

1. HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

生成对抗网络(GAN)

视频编辑(Video Editing)

Object-WIPER: Training-Free Object and Associated Effect Removal in Videos

HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles

Low-level Vision

ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration

F²HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling

Missing No More: Dictionary-Guided Cross-Modal Image Fusion under Missing Infrared

BluRef: Unsupervised Image Deblurring with Dense-Matching References

Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration

Reparameterized Tensor Ring Functional Decomposition for Multi-Dimensional Data Recovery

Lumosaic: Hyperspectral Video via Active Illumination and Coded-Exposure Pixels

MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision

Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis

超分辨率(Super-Resolution)

RAW-Domain Degradation Models for Realistic Smartphone Super-Resolution

UCAN: Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-Resolution

FiDeSR: High-Fidelity and Detail-Preserving One-Step Diffusion Super-Resolution

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution

Spectral Super-Resolution via Adversarial Unfolding and Data-Driven Spectrum Regularization: From Multispectral Satellite Data to NASA Hyperspectral Image

AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution

去噪(Denoising)

Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imaging

图像生成(Image Generation)

coDrawAgents: A Multi-Agent Dialogue Framework for Compositional Image Generation

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards

VeCoR -- Velocity Contrastive Regularization for Flow Matching

Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards

Enhancing Spatial Understanding in Image Generation via Reward Modeling

AutoDebias: Automated Framework for Debiasing Text-to-Image Models

视频生成(Video Generation)

Anchoring and Rescaling Attention for Semantically Coherent Inbetweening

Training-free Motion Factorization for Compositional Video Generation

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters

EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing

CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video

FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning

UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation

ExpPortrait: Expressive Portrait Generation via Personalized Representation

LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation

FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection

3D生成

Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editing

NI-Tex: Non-isometric Image-based Garment Texture Generation

ForgeDreamer: Industrial Text-to-3D Generation with Multi-Expert LoRA and Cross-View Hypergraph

Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single Image

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation

MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing

Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow

BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation

视频理解(Video Understanding)

F2HDR: Two-Stage HDR Video Reconstruction via Flow Adapter and Physical Motion Modeling

Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding

StreamingTOM: Streaming Token Compression for Efficient Video Understanding

Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding

StreamReady: Learning What to Answer and When in Long Streaming Videos

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding

ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videos

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark

Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding

FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

3D人体姿态估计(3D Human Pose Estimation)

GazeOnce360: Fisheye-Based 360° Multi-Person Gaze Estimation with Global-Local Feature Fusion

OnlineHMR: Video-based Online World-Grounded Human Mesh Recovery

Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation

Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation

Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator

CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillation

SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking

VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery

持续学习(Continual Learning)

Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment

Elastic Weight Consolidation Done Right for Continual Learning

行为识别(Action Recognition)

BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment

Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency

知识蒸馏(Knowledge Distillation)

Momentum Memory for Knowledge Distillation in Computational Pathology

WaDi: Weight Direction-aware Distillation for One-step Image Synthesis

Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation

Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation

Distilling Balanced Knowledge from a Biased Teacher

Momentum Memory for Knowledge Distillation in Computational Pathology

图像压缩(Image Compression)

SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representation

Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression

Zero-Shot Learning(零样本学习)

Learning through Creation: A Hash-Free Framework for On-the-Fly Category Discovery

TALON: Test-time Adaptive Learning for On-the-Fly Category Discovery

Learning through Creation: A Hash-Free Framework for On-the-Fly Category Discovery

Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

立体匹配(Stereo Matching)

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo Matching

场景图生成(Scene Graph Generation)

DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime

计数(Counting)

UNICBench: UNIfied Counting Benchmark for MLLM

隐式神经表示(Implicit Neural Representations)

图像质量评价(Image Quality Assessment)

How to Take a Memorable Picture? Empowering Users with Actionable Feedback

Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks

视频质量评价(Video Quality Assessment)

数据集(Datasets)

E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought

Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis

LenghuSky-8: An 8-Year All-Sky Cloud Dataset with Star-Aware Masks and Alt-Az Calibration for Segmentation and Nowcasting

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

反学习(Machine Unlearning)

SineProject: Machine Unlearning for Stable Vision Language Alignment

RAZOR: Ratio-Aware Layer Editing for Targeted Unlearning in Vision Transformers and Diffusion Models

Stake the Points: Structure-Faithful Instance Unlearning

新任务(New Tasks)

模型加速(Improving Reasoning)

Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models

Variation-aware Vision Token Dropping for Faster Large Vision-Language Models

Model Merging in the Essential Subspace

时间序列(Time Series)

STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting

脉冲网络

Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networks

Rethinking SNN Online Training and Deployment: Gradient-Coherent Learning via Hybrid-Driven LIF Model

图像检索

TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking

PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing

NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries

其他(Others)

NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training

Defending Unauthorized Model Merging via Dual-Stage Weight Protection

BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning

ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimation

GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement

DC-Merge: Improving Model Merging with Directional Consistency

Defending Unauthorized Model Merging via Dual-Stage Weight Protection

BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learning

Bridging Domains through Subspace-Aware Model Merging

Significant stargazers

Matiur Rahman Minar

155 followers · starred Mar 2026