ashishpatel26/CVPR2024

CVPR 2024 Research Paper with Code

47

3 commits

updated Jun 28, 2024

See the code

README

CVPR 2024

Research Paper with Code


Table of Contents

Domain-wise Table

3DGS (Gaussian Splatting)

IndexPaper TitlePaper LinkCodeOfficial Repo
1Scaffold-GS: Structured 3D Gaussians for View-Adaptive RenderingPaperCodeHomepage
2GPS-Gaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View SynthesisPaperCodeHomepage
3GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansPaperCodeN/A
4GaussianEditor: Swift and Controllable 3D Editing with Gaussian SplattingPaperCodeN/A
5Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene ReconstructionPaperCodeHomepage

Avatars

IndexPaper TitlePaper LinkCodeOfficial Repo
6GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansPaperCodeN/A
7Real-Time Simulated Avatar from Head-Mounted SensorsPaperN/AHomepage

Backbone

IndexPaper TitlePaper LinkCodeOfficial Repo
8RepViT: Revisiting Mobile CNN From ViT PerspectivePaperCodeN/A
9TransNeXt: Robust Foveal Visual Perception for Vision TransformersPaperCodeN/A

CLIP

IndexPaper TitlePaper LinkCodeOfficial Repo
10Alpha-CLIP: A CLIP Model Focusing on Wherever You WantPaperCodeN/A
11FairCLIP: Harnessing Fairness in Vision-Language LearningPaperCodeN/A

Embodied AI

IndexPaper TitlePaper LinkCodeOfficial Repo
12EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AIPaperCodeHomepage
13MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionPaperCodeHomepage

OCR

IndexPaper TitlePaper LinkCodeOfficial Repo
14An Empirical Study of Scaling Law for OCRPaperCodeN/A
15ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingPaperCodeN/A

NeRF

IndexPaper TitlePaper LinkCodeOfficial Repo
16PIE-NeRF: Physics-based Interactive Elastodynamics with NeRFPaperCodeN/A

DETR

IndexPaper TitlePaper LinkCodeOfficial Repo
17DETRs Beat YOLOs on Real-time Object DetectionPaperCodeN/A
18Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementPaperCodeN/A

ReID

IndexPaper TitlePaper LinkCodeOfficial Repo
19Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPaperCodeN/A
20Noisy-Correspondence Learning for Text-to-Image Person Re-identificationPaperCodeN/A

Long-Tail

IndexPaper TitlePaper LinkCodeOfficial Repo
1Delving into the Trajectory Long-tail Distribution for Multi-object TrackingPaperCodeN/A

Vision Transformer

IndexPaper TitlePaper LinkCodeOfficial Repo
2TransNeXt: Robust Foveal Visual Perception for Vision TransformersPaperCodeN/A
3RepViT: Revisiting Mobile CNN From ViT PerspectivePaperCodeN/A

Vision-Language

IndexPaper TitlePaper LinkCodeOfficial Repo
4PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsPaperCodeN/A
5FairCLIP: Harnessing Fairness in Vision-Language LearningPaperCodeN/A

Self-supervised Learning

IndexPaper TitlePaper LinkCodeOfficial Repo
6N/AN/AN/AN/A

Data Augmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
7N/AN/AN/AN/A

Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
8DETRs Beat YOLOs on Real-time Object DetectionPaperCodeN/A
9Boosting Object Detection with Zero-Shot Day-Night Domain AdaptationPaperCodeN/A
10YOLO-World: Real-Time Open-Vocabulary Object DetectionPaperCodeN/A
11Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementPaperCodeN/A

Anomaly Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
12Anomaly Heterogeneity Learning for Open-set Supervised Anomaly DetectionPaperCodeN/A

Visual Tracking

IndexPaper TitlePaper LinkCodeOfficial Repo
13N/AN/AN/AN/A

Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
14Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationPaperCodeN/A
15SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationPaperCodeN/A

Instance Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
16N/AN/AN/AN/A

Panoptic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
17N/AN/AN/AN/A

Medical Image

IndexPaper TitlePaper LinkCodeOfficial Repo
18Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyPaperCodeN/A
19VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisPaperCodeN/A
20ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImagesPaperCodeN/A

Medical Image Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
21N/AN/AN/AN/A

Video Object Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
22N/AN/AN/AN/A

Video Instance Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
23N/AN/AN/AN/A

Referring Image Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
24N/AN/AN/AN/A

Image Matting

IndexPaper TitlePaper LinkCodeOfficial Repo
25N/AN/AN/AN/A

Image Editing

IndexPaper TitlePaper LinkCodeOfficial Repo
26Edit One for All: Interactive Batch Image EditingPaperCodeHomepage

Low-level Vision

IndexPaper TitlePaper LinkCodeOfficial Repo
27Residual Denoising Diffusion ModelsPaperCodeN/A
28Boosting Image Restoration via Priors from Pre-trained ModelsPaperN/AN/A

Super-Resolution)

IndexPaper TitlePaper LinkCodeOfficial Repo
29SeD: Semantic-Aware Discriminator for Image Super-ResolutionPaperCodeN/A
30APISR: Anime Production Inspired Real-World Anime Super-ResolutionPaper[Code](https://github.com/Kiter### Domain-wise Table

Denoising

IndexPaper TitlePaper LinkCodeOfficial Repo
31Residual Denoising Diffusion ModelsPaperCodeN/A

Deblur

IndexPaper TitlePaper LinkCodeOfficial Repo
32N/AN/AN/AN/A

Autonomous Driving

IndexPaper TitlePaper LinkCodeOfficial Repo
33UniPAD: A Universal Pre-training Paradigm for Autonomous DrivingPaperCodeN/A
34Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsPaperCodeN/A
35Memory-based Adapters for Online 3D Scene PerceptionPaperCodeN/A
36Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesPaperCodeN/A
37A Real-world Large-scale Dataset for Roadside Cooperative PerceptionPaperCodeN/A
38Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingPaperCodeN/A

3D Point Cloud

IndexPaper TitlePaper LinkCodeOfficial Repo
40N/AN/AN/AN/A

3D Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
41PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionPaperCodeN/A
42UniMODE: Unified Monocular 3D Object DetectionPaperN/AN/A

3D Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
43N/AN/AN/AN/A

3D Object Tracking

IndexPaper TitlePaper LinkCodeOfficial Repo
44N/AN/AN/AN/A

3D Semantic Scene Completion

IndexPaper TitlePaper LinkCodeOfficial Repo
45Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesPaperCodeN/A

3D Registration

IndexPaper TitlePaper LinkCodeOfficial Repo
46N/AN/AN/AN/A

3D Human Pose Estimation

IndexPaper TitlePaper LinkCodeOfficial Repo
47Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationPaperCodeN/A

3D Human Mesh Estimation

IndexPaper TitlePaper LinkCodeOfficial Repo
48N/AN/AN/AN/A

Medical Image

IndexPaper TitlePaper LinkCodeOfficial Repo
49Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyPaperCodeN/A
50VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisPaperCodeN/A
51ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImagesPaperCodeN/A

Image Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
52InstanceDiffusion: Instance-level Control for Image GenerationPaperCodeHomepage
53ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image GenerationsPaperCodeHomepage
54Instruct-Imagen: Image Generation with Multi-modal InstructionPaperN/AN/A
55UniGS: Unified Representation for Image Generation and SegmentationPaperN/AN/A
56Multi-Instance Generation Controller for Text-to-Image SynthesisPaperCodeN/A
57SVGDreamer: Text Guided SVG Generation with Diffusion ModelPaperCodeN/A
58InteractDiffusion: Interaction-Control for Text-to-Image Diffusion ModelPaperCodeN/A
59Ranni: Taming Text-to-Image Diffusion for Accurate Prompt FollowingPaperCodeN/A

Video Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
60Vlogger: Make Your Dream A VlogPaperCodeN/A
61VBench: Comprehensive Benchmark Suite for Video Generative ModelsPaperCodeHomepage
62VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion ModelsPaperCodeHomepage

Vision Transformer

IndexPaper TitlePaper LinkCodeOfficial Repo
63TransNeXt: Robust Foveal Visual Perception for Vision TransformersPaperCodeN/A
64RepViT: Revisiting Mobile CNN From ViT PerspectivePaperCodeN/A
65A General and Efficient Training for Transformer via Token ExpansionPaperCodeN/A

Vision-Language

IndexPaper TitlePaper LinkCodeOfficial Repo
66PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsPaperCodeN/A
67FairCLIP: Harnessing Fairness in Vision-Language LearningPaperCodeN/A

Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
68DETRs Beat YOLOs on Real-time Object DetectionPaperCodeN/A
69Boosting Object Detection with Zero-Shot Day-Night Domain AdaptationPaperCodeN/A
70YOLO-World: Real-Time Open-Vocabulary Object DetectionPaperCodeN/A
71Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementPaperCodeN/A

Anomaly Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
72Anomaly Heterogeneity Learning for Open-set Supervised Anomaly DetectionPaperCodeN/A

Object Tracking

IndexPaper TitlePaper LinkCodeOfficial Repo
73Delving into the Trajectory Long-tail Distribution for Multi-object TrackingPaperCodeN/A

Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
74Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationPaperCodeN/A
75SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationPaperCodeN/A

Medical Image

IndexPaper TitlePaper LinkCodeOfficial Repo
76Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyPaperCodeN/A
77VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisPaperCodeN/A
78ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImagesPaperCodeN/A

Medical Image Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
76N/AN/AN/AN/A

Autonomous Driving

IndexPaper TitlePaper LinkCodeOfficial Repo
77UniPAD: A Universal Pre-training Paradigm for Autonomous DrivingPaperCodeN/A
78Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsPaperCodeN/A
79Memory-based Adapters for Online 3D Scene PerceptionPaperCodeN/A
80Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesPaperCodeN/A
81A Real-world Large-scale Dataset for Roadside Cooperative PerceptionPaperCodeN/A
82Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingPaperCodeN/A
83Traffic Scene Parsing through the TSP6K DatasetPaperCodeN/A

3D Point Cloud

IndexPaper TitlePaper LinkCodeOfficial Repo
84N/AN/AN/AN/A

3D Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
85PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionPaperCodeN/A
86UniMODE: Unified Monocular 3D Object DetectionPaperN/AN/A

3D Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
87N/AN/AN/AN/A

Image Editing

IndexPaper TitlePaper LinkCodeOfficial Repo
88Edit One for All: Interactive Batch Image EditingPaperCodeHomepage

Video Editing

IndexPaper TitlePaper LinkCodeOfficial Repo
89MaskINT: Video Editing via Interpolative Non-autoregressive Masked TransformersPaperN/AHomepage

Low-level Vision

IndexPaper TitlePaper LinkCodeOfficial Repo
90Residual Denoising Diffusion ModelsPaperCodeN/A
91Boosting Image Restoration via Priors from Pre-trained ModelsPaperN/AN/A

Super-Resolution

IndexPaper TitlePaper LinkCodeOfficial Repo
92SeD: Semantic-Aware Discriminator for Image Super-ResolutionPaperCodeN/A
93APISR: Anime Production Inspired Real-World Anime Super-ResolutionPaperCodeN/A

Denoising

IndexPaper TitlePaper LinkCodeOfficial Repo
94N/AN/AN/AN/A

3D Human Pose Estimation

IndexPaper TitlePaper LinkCodeOfficial Repo
95Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationPaperCodeN/A

Image Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
96InstanceDiffusion: Instance-level Control for Image GenerationPaperCodeHomepage
97ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image GenerationsPaperCodeHomepage
98Instruct-Imagen: Image Generation with Multi-modal InstructionPaperN/AN/A
99Residual Denoising Diffusion ModelsPaperCodeN/A
100UniGS: Unified Representation for Image Generation and SegmentationPaperN/AN/A
101Multi-Instance Generation Controller for Text-to-Image SynthesisPaperCodeN/A
102SVGDreamer: Text Guided SVG Generation with Diffusion ModelPaperCodeN/A
103InteractDiffusion: Interaction-Control for Text-to-Image Diffusion ModelPaperCodeN/A
104Ranni: Taming Text-to-Image Diffusion for Accurate Prompt FollowingPaperCodeN/A

Video Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
105Vlogger: Make Your Dream A VlogPaperCodeN/A
106VBench: Comprehensive Benchmark Suite for Video Generative ModelsPaperCodeHomepage
107VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion ModelsPaperCodeHomepage

3D Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
108CityDreamer: Compositional Generative Model of Unbounded 3D CitiesPaperCodeHomepage
109LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score MatchingPaperCodeN/A

Video Understanding

IndexPaper TitlePaper LinkCodeOfficial Repo
110MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkPaperCodeN/A

Knowledge Distillation

IndexPaper TitlePaper LinkCodeOfficial Repo
111Logit Standardization in Knowledge DistillationPaperCodeN/A
112Efficient Dataset Distillation via Minimax DiffusionPaperCodeN/A

Stereo Matching

IndexPaper TitlePaper LinkCodeOfficial Repo
113Neural Markov Random Field for Stereo MatchingPaperCodeN/A

Scene Graph Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
114HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph GenerationPaperCodeHomepage

Video Quality Assessment

IndexPaper TitlePaper LinkCodeOfficial Repo
115KVQ: Kaleidoscope Video Quality Assessment for Short-form VideosPaperCodeHomepage

Datasets

IndexPaper TitlePaper LinkCodeOfficial Repo
116A Real-world Large-scale Dataset for Roadside Cooperative PerceptionPaperCodeN/A
117Traffic Scene Parsing through the TSP6K DatasetPaperCodeN/A

Others

IndexPaper TitlePaper LinkCodeOfficial Repo
118Object Recognition as Next Token PredictionPaperCodeN/A
119ParameterNet: Parameters Are All You Need for Large-scale Visual Pretraining of Mobile NetworksPaperCodeN/A
120Seamless Human Motion Composition with Blended Positional EncodingsPaperCodeN/A
121LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and PlanningPaperCodeHomepage
122CLOVA: A Closed-LOop Visual Assistant with Tool Usage and UpdatePaperN/AHomepage
123MoMask: Generative Masked Modeling of 3D Human MotionsPaperCodeN/A
124Amodal Ground Truth and Completion in the WildPaperCodeHomepage
125Improved Visual Grounding through Self-Consistent ExplanationsPaperCodeN/A
126ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectPaperCodeHomepage
127Learning from Synthetic Human Group ActivitiesPaperCodeHomepage
128A Cross-Subject Brain Decoding FrameworkPaperCodeHomepage
129Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsPaperCodeN/A
130Contrastive Mean-Shift Learning for Generalized Category DiscoveryPaperCodeHomepage

Thank you for Reading

computervision
cvpr
cvpr2024

ashishpatel26/CVPR2024

CVPR 2024 Research Paper with Code

47

3 commits

updated Jun 28, 2024

See the code

README

CVPR 2024

Research Paper with Code


Table of Contents

Domain-wise Table

3DGS (Gaussian Splatting)

IndexPaper TitlePaper LinkCodeOfficial Repo
1Scaffold-GS: Structured 3D Gaussians for View-Adaptive RenderingPaperCodeHomepage
2GPS-Gaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View SynthesisPaperCodeHomepage
3GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansPaperCodeN/A
4GaussianEditor: Swift and Controllable 3D Editing with Gaussian SplattingPaperCodeN/A
5Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene ReconstructionPaperCodeHomepage

Avatars

IndexPaper TitlePaper LinkCodeOfficial Repo
6GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D GaussiansPaperCodeN/A
7Real-Time Simulated Avatar from Head-Mounted SensorsPaperN/AHomepage

Backbone

IndexPaper TitlePaper LinkCodeOfficial Repo
8RepViT: Revisiting Mobile CNN From ViT PerspectivePaperCodeN/A
9TransNeXt: Robust Foveal Visual Perception for Vision TransformersPaperCodeN/A

CLIP

IndexPaper TitlePaper LinkCodeOfficial Repo
10Alpha-CLIP: A CLIP Model Focusing on Wherever You WantPaperCodeN/A
11FairCLIP: Harnessing Fairness in Vision-Language LearningPaperCodeN/A

Embodied AI

IndexPaper TitlePaper LinkCodeOfficial Repo
12EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AIPaperCodeHomepage
13MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionPaperCodeHomepage

OCR

IndexPaper TitlePaper LinkCodeOfficial Repo
14An Empirical Study of Scaling Law for OCRPaperCodeN/A
15ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and SpottingPaperCodeN/A

NeRF

IndexPaper TitlePaper LinkCodeOfficial Repo
16PIE-NeRF: Physics-based Interactive Elastodynamics with NeRFPaperCodeN/A

DETR

IndexPaper TitlePaper LinkCodeOfficial Repo
17DETRs Beat YOLOs on Real-time Object DetectionPaperCodeN/A
18Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementPaperCodeN/A

ReID

IndexPaper TitlePaper LinkCodeOfficial Repo
19Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-IdentificationPaperCodeN/A
20Noisy-Correspondence Learning for Text-to-Image Person Re-identificationPaperCodeN/A

Long-Tail

IndexPaper TitlePaper LinkCodeOfficial Repo
1Delving into the Trajectory Long-tail Distribution for Multi-object TrackingPaperCodeN/A

Vision Transformer

IndexPaper TitlePaper LinkCodeOfficial Repo
2TransNeXt: Robust Foveal Visual Perception for Vision TransformersPaperCodeN/A
3RepViT: Revisiting Mobile CNN From ViT PerspectivePaperCodeN/A

Vision-Language

IndexPaper TitlePaper LinkCodeOfficial Repo
4PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsPaperCodeN/A
5FairCLIP: Harnessing Fairness in Vision-Language LearningPaperCodeN/A

Self-supervised Learning

IndexPaper TitlePaper LinkCodeOfficial Repo
6N/AN/AN/AN/A

Data Augmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
7N/AN/AN/AN/A

Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
8DETRs Beat YOLOs on Real-time Object DetectionPaperCodeN/A
9Boosting Object Detection with Zero-Shot Day-Night Domain AdaptationPaperCodeN/A
10YOLO-World: Real-Time Open-Vocabulary Object DetectionPaperCodeN/A
11Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementPaperCodeN/A

Anomaly Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
12Anomaly Heterogeneity Learning for Open-set Supervised Anomaly DetectionPaperCodeN/A

Visual Tracking

IndexPaper TitlePaper LinkCodeOfficial Repo
13N/AN/AN/AN/A

Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
14Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationPaperCodeN/A
15SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationPaperCodeN/A

Instance Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
16N/AN/AN/AN/A

Panoptic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
17N/AN/AN/AN/A

Medical Image

IndexPaper TitlePaper LinkCodeOfficial Repo
18Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyPaperCodeN/A
19VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisPaperCodeN/A
20ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImagesPaperCodeN/A

Medical Image Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
21N/AN/AN/AN/A

Video Object Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
22N/AN/AN/AN/A

Video Instance Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
23N/AN/AN/AN/A

Referring Image Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
24N/AN/AN/AN/A

Image Matting

IndexPaper TitlePaper LinkCodeOfficial Repo
25N/AN/AN/AN/A

Image Editing

IndexPaper TitlePaper LinkCodeOfficial Repo
26Edit One for All: Interactive Batch Image EditingPaperCodeHomepage

Low-level Vision

IndexPaper TitlePaper LinkCodeOfficial Repo
27Residual Denoising Diffusion ModelsPaperCodeN/A
28Boosting Image Restoration via Priors from Pre-trained ModelsPaperN/AN/A

Super-Resolution)

IndexPaper TitlePaper LinkCodeOfficial Repo
29SeD: Semantic-Aware Discriminator for Image Super-ResolutionPaperCodeN/A
30APISR: Anime Production Inspired Real-World Anime Super-ResolutionPaper[Code](https://github.com/Kiter### Domain-wise Table

Denoising

IndexPaper TitlePaper LinkCodeOfficial Repo
31Residual Denoising Diffusion ModelsPaperCodeN/A

Deblur

IndexPaper TitlePaper LinkCodeOfficial Repo
32N/AN/AN/AN/A

Autonomous Driving

IndexPaper TitlePaper LinkCodeOfficial Repo
33UniPAD: A Universal Pre-training Paradigm for Autonomous DrivingPaperCodeN/A
34Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsPaperCodeN/A
35Memory-based Adapters for Online 3D Scene PerceptionPaperCodeN/A
36Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesPaperCodeN/A
37A Real-world Large-scale Dataset for Roadside Cooperative PerceptionPaperCodeN/A
38Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingPaperCodeN/A

3D Point Cloud

IndexPaper TitlePaper LinkCodeOfficial Repo
40N/AN/AN/AN/A

3D Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
41PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionPaperCodeN/A
42UniMODE: Unified Monocular 3D Object DetectionPaperN/AN/A

3D Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
43N/AN/AN/AN/A

3D Object Tracking

IndexPaper TitlePaper LinkCodeOfficial Repo
44N/AN/AN/AN/A

3D Semantic Scene Completion

IndexPaper TitlePaper LinkCodeOfficial Repo
45Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesPaperCodeN/A

3D Registration

IndexPaper TitlePaper LinkCodeOfficial Repo
46N/AN/AN/AN/A

3D Human Pose Estimation

IndexPaper TitlePaper LinkCodeOfficial Repo
47Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationPaperCodeN/A

3D Human Mesh Estimation

IndexPaper TitlePaper LinkCodeOfficial Repo
48N/AN/AN/AN/A

Medical Image

IndexPaper TitlePaper LinkCodeOfficial Repo
49Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyPaperCodeN/A
50VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisPaperCodeN/A
51ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImagesPaperCodeN/A

Image Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
52InstanceDiffusion: Instance-level Control for Image GenerationPaperCodeHomepage
53ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image GenerationsPaperCodeHomepage
54Instruct-Imagen: Image Generation with Multi-modal InstructionPaperN/AN/A
55UniGS: Unified Representation for Image Generation and SegmentationPaperN/AN/A
56Multi-Instance Generation Controller for Text-to-Image SynthesisPaperCodeN/A
57SVGDreamer: Text Guided SVG Generation with Diffusion ModelPaperCodeN/A
58InteractDiffusion: Interaction-Control for Text-to-Image Diffusion ModelPaperCodeN/A
59Ranni: Taming Text-to-Image Diffusion for Accurate Prompt FollowingPaperCodeN/A

Video Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
60Vlogger: Make Your Dream A VlogPaperCodeN/A
61VBench: Comprehensive Benchmark Suite for Video Generative ModelsPaperCodeHomepage
62VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion ModelsPaperCodeHomepage

Vision Transformer

IndexPaper TitlePaper LinkCodeOfficial Repo
63TransNeXt: Robust Foveal Visual Perception for Vision TransformersPaperCodeN/A
64RepViT: Revisiting Mobile CNN From ViT PerspectivePaperCodeN/A
65A General and Efficient Training for Transformer via Token ExpansionPaperCodeN/A

Vision-Language

IndexPaper TitlePaper LinkCodeOfficial Repo
66PromptKD: Unsupervised Prompt Distillation for Vision-Language ModelsPaperCodeN/A
67FairCLIP: Harnessing Fairness in Vision-Language LearningPaperCodeN/A

Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
68DETRs Beat YOLOs on Real-time Object DetectionPaperCodeN/A
69Boosting Object Detection with Zero-Shot Day-Night Domain AdaptationPaperCodeN/A
70YOLO-World: Real-Time Open-Vocabulary Object DetectionPaperCodeN/A
71Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementPaperCodeN/A

Anomaly Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
72Anomaly Heterogeneity Learning for Open-set Supervised Anomaly DetectionPaperCodeN/A

Object Tracking

IndexPaper TitlePaper LinkCodeOfficial Repo
73Delving into the Trajectory Long-tail Distribution for Multi-object TrackingPaperCodeN/A

Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
74Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationPaperCodeN/A
75SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic SegmentationPaperCodeN/A

Medical Image

IndexPaper TitlePaper LinkCodeOfficial Repo
76Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyPaperCodeN/A
77VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisPaperCodeN/A
78ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy ImagesPaperCodeN/A

Medical Image Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
76N/AN/AN/AN/A

Autonomous Driving

IndexPaper TitlePaper LinkCodeOfficial Repo
77UniPAD: A Universal Pre-training Paradigm for Autonomous DrivingPaperCodeN/A
78Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsPaperCodeN/A
79Memory-based Adapters for Online 3D Scene PerceptionPaperCodeN/A
80Symphonize 3D Semantic Scene Completion with Contextual Instance QueriesPaperCodeN/A
81A Real-world Large-scale Dataset for Roadside Cooperative PerceptionPaperCodeN/A
82Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingPaperCodeN/A
83Traffic Scene Parsing through the TSP6K DatasetPaperCodeN/A

3D Point Cloud

IndexPaper TitlePaper LinkCodeOfficial Repo
84N/AN/AN/AN/A

3D Object Detection

IndexPaper TitlePaper LinkCodeOfficial Repo
85PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object DetectionPaperCodeN/A
86UniMODE: Unified Monocular 3D Object DetectionPaperN/AN/A

3D Semantic Segmentation

IndexPaper TitlePaper LinkCodeOfficial Repo
87N/AN/AN/AN/A

Image Editing

IndexPaper TitlePaper LinkCodeOfficial Repo
88Edit One for All: Interactive Batch Image EditingPaperCodeHomepage

Video Editing

IndexPaper TitlePaper LinkCodeOfficial Repo
89MaskINT: Video Editing via Interpolative Non-autoregressive Masked TransformersPaperN/AHomepage

Low-level Vision

IndexPaper TitlePaper LinkCodeOfficial Repo
90Residual Denoising Diffusion ModelsPaperCodeN/A
91Boosting Image Restoration via Priors from Pre-trained ModelsPaperN/AN/A

Super-Resolution

IndexPaper TitlePaper LinkCodeOfficial Repo
92SeD: Semantic-Aware Discriminator for Image Super-ResolutionPaperCodeN/A
93APISR: Anime Production Inspired Real-World Anime Super-ResolutionPaperCodeN/A

Denoising

IndexPaper TitlePaper LinkCodeOfficial Repo
94N/AN/AN/AN/A

3D Human Pose Estimation

IndexPaper TitlePaper LinkCodeOfficial Repo
95Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationPaperCodeN/A

Image Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
96InstanceDiffusion: Instance-level Control for Image GenerationPaperCodeHomepage
97ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image GenerationsPaperCodeHomepage
98Instruct-Imagen: Image Generation with Multi-modal InstructionPaperN/AN/A
99Residual Denoising Diffusion ModelsPaperCodeN/A
100UniGS: Unified Representation for Image Generation and SegmentationPaperN/AN/A
101Multi-Instance Generation Controller for Text-to-Image SynthesisPaperCodeN/A
102SVGDreamer: Text Guided SVG Generation with Diffusion ModelPaperCodeN/A
103InteractDiffusion: Interaction-Control for Text-to-Image Diffusion ModelPaperCodeN/A
104Ranni: Taming Text-to-Image Diffusion for Accurate Prompt FollowingPaperCodeN/A

Video Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
105Vlogger: Make Your Dream A VlogPaperCodeN/A
106VBench: Comprehensive Benchmark Suite for Video Generative ModelsPaperCodeHomepage
107VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion ModelsPaperCodeHomepage

3D Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
108CityDreamer: Compositional Generative Model of Unbounded 3D CitiesPaperCodeHomepage
109LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score MatchingPaperCodeN/A

Video Understanding

IndexPaper TitlePaper LinkCodeOfficial Repo
110MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkPaperCodeN/A

Knowledge Distillation

IndexPaper TitlePaper LinkCodeOfficial Repo
111Logit Standardization in Knowledge DistillationPaperCodeN/A
112Efficient Dataset Distillation via Minimax DiffusionPaperCodeN/A

Stereo Matching

IndexPaper TitlePaper LinkCodeOfficial Repo
113Neural Markov Random Field for Stereo MatchingPaperCodeN/A

Scene Graph Generation

IndexPaper TitlePaper LinkCodeOfficial Repo
114HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph GenerationPaperCodeHomepage

Video Quality Assessment

IndexPaper TitlePaper LinkCodeOfficial Repo
115KVQ: Kaleidoscope Video Quality Assessment for Short-form VideosPaperCodeHomepage

Datasets

IndexPaper TitlePaper LinkCodeOfficial Repo
116A Real-world Large-scale Dataset for Roadside Cooperative PerceptionPaperCodeN/A
117Traffic Scene Parsing through the TSP6K DatasetPaperCodeN/A

Others

IndexPaper TitlePaper LinkCodeOfficial Repo
118Object Recognition as Next Token PredictionPaperCodeN/A
119ParameterNet: Parameters Are All You Need for Large-scale Visual Pretraining of Mobile NetworksPaperCodeN/A
120Seamless Human Motion Composition with Blended Positional EncodingsPaperCodeN/A
121LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and PlanningPaperCodeHomepage
122CLOVA: A Closed-LOop Visual Assistant with Tool Usage and UpdatePaperN/AHomepage
123MoMask: Generative Masked Modeling of 3D Human MotionsPaperCodeN/A
124Amodal Ground Truth and Completion in the WildPaperCodeHomepage
125Improved Visual Grounding through Self-Consistent ExplanationsPaperCodeN/A
126ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectPaperCodeHomepage
127Learning from Synthetic Human Group ActivitiesPaperCodeHomepage
128A Cross-Subject Brain Decoding FrameworkPaperCodeHomepage
129Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsPaperCodeN/A
130Contrastive Mean-Shift Learning for Generalized Category DiscoveryPaperCodeHomepage

Thank you for Reading

computervision
cvpr
cvpr2024