lliai/Awesome-Vision-Knowledge-Distillation

Awesome Knowledge-Distillation for CV

95

15 commits

updated Apr 30, 2024

See the code

README

Awesome Knowledge Distillation in Computer vision

[TOC]

Diffusion Knowledge Distillation

TitleVenueNote
A Comprehensive Survey on Knowledge Distillation of Diffusion Models2023Weijian Luo. [pdf]
Knowledge distillation in iterative generative models for improved sampling speed2021Eric Luhman, Troy Luhman. [pdf]
Progressive Distillation for Fast Sampling of Diffusion ModelsICLR 2022Tim Salimans and Jonathan Ho. [pdf]
On Distillation of Guided Diffusion ModelsCVPR 2023Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, Tim Salimans. [pdf]
TRACT: Denoising Diffusion Models with Transitive Closure Time-Distillation2023Berthelot, David, Autef, Arnaud, Lin, Jierui, Yap, Dian Ang, Zhai, Shuangfei, Hu, Siyuan, Zheng, Daniel, Talbott, Walter, Gu, Eric. [pdf]
BK-SDM: Architecturally Compressed Stable Diffusion for Efficient Text-to-Image GenerationICML 2023Kim, Bo-Kyeong, Song, Hyoung-Kyu, Castells, Thibault, Choi, Shinkook. [pdf]
On Architectural Compression of Text-to-Image Diffusion Models2023Kim, Bo-Kyeong, Song, Hyoung-Kyu, Castells, Thibault, Choi, Shinkook. [pdf]
Knowledge Diffusion for Distillation2023Tao Huang, Yuan Zhang, Mingkai Zheng, Shan You, Fei Wang, Chen Qian, Chang Xu. [pdf]
SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds2023Yanyu Li, Huan Wang, Qing Jin, Ju Hu, Pavlo Chemerys, Yun Fu, Yanzhi Wang, Sergey Tulyakov, Jian Ren1. [pdf]
BOOT: Data-free Distillation of Denoising Diffusion Models with Bootstrapping2023Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Lingjie Liu, Josh Susskind. [pdf]
Consistency models2023Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. [pdf]

Knowledge Distillation for Semantic Segmentation

TitleVenueNote
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial NetworksarXiv:1709.00513
Knowledge Distillation for Semantic Segmentation
Structured knowledge distillation for semantic segmentationCVPR-2019
Intra-class feature variation distillation for semantic segmentationECCV-2020
Channel-wise knowledge distillation for dense predictionICCV-2021
Double Similarity Distillation for Semantic Image SegmentationTIP-2021
Cross-Image Relational Knowledge Distillation for Semantic SegmentationCVPR-2022

Knowledge Distillation for Object Detection

TitleVenueNote
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial NetworksarXiv:1709.00513
Mimicking very efficient network for object detectionCVPR 2017pdf
Distilling object detectors with fine-grained feature imitationCVPR 2019pdf
General instance distillation for object detectionCVPR 2021pdf
Distilling object detectors via decoupled featuresCVPR 2021pdf
Distilling object detectors with feature richnessNeurIPS 2021pdf
Focal and global knowledge distillation for detectorsCVPR 2022pdf
Rank Mimicking and Prediction-guided Feature ImitationAAAI 2022pdf
Prediction-Guided DistillationECCV 2022pdf
Masked Distillation with Receptive TokensICLR 2023pdf
Structural Knowledge Distillation for Object DetectionNeurIPS 2022OpenReview
Dual Relation Knowledge Distillation for Object DetectionIJCAI 2023pdf
GLAMD: Global and Local Attention Mask Distillation for Object DetectorsECCV 2022ECVA
G-DetKD: Towards General Distillation Framework for Object Detectors via Contrastive and Semantic-guided Feature ImitationICCV 2021CVF
PKD: General Distillation Framework for Object Detectors via Pearson Correlation CoefficientNeurIPS 2022OpenReview
MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object DetectionECCV 2020ECVA
LabelEnc: A New Intermediate Supervision Method for Object DetectionECCV 2020ECVA
TitleVenueNote
HEtero-Assists Distillation for Heterogeneous Object DetectorsECCV 2022HEAD
LGD: Label-Guided Self-Distillation for Object DetectionAAAI 2022LGD
When Object Detection Meets Knowledge Distillation: A SurveyTPAMI
ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorCVPR 2023ScaleKD
CrossKD: Cross-Head Knowledge Distillation for Dense Object DetectionarXiv:2306.11369CrossKD

Knowledge Distillation in Vision Transformers

TitleVenueNote
Training data-efficient image transformers & distillation through attentionICML2021
Co-advise: Cross inductive bias distillationCVPR2022
Tinyvit: Fast pretraining distillation for small vision transformersarXiv:2207.10666
Attention Probe: Vision Transformer Distillation in the WildICASSP2022
Dear KD: Data-Efficient Early Knowledge Distillation for Vision TransformersCVPR2022
Efficient vision transformers via fine-grained manifold distillationNIPS2022
Cross-Architecture Knowledge DistillationarXiv:2207.05273
MiniViT: Compressing Vision Transformers with Weight MultiplexingCVPR2022
ViTKD: Practical Guidelines for ViT feature knowledge distillationarXiv 2022code

Knowledge Distillation for Teacher-Student Gaps

TitleVenueNote
Improved Knowledge Distillation via Teacher Assistant: Bridging the Gap Between Student and TeacherAAAI2020
Search to Distill: Pearls are Everywhere but not the EyesCVPR 2020
Reducing the Teacher-Student Gap via Spherical Knowledge DisitllationarXiv:2020
Knowledge Distillation via the Target-aware TransformerCVPR2022
Decoupled Knowledge DistillationCVPR 2022code
Prune Your Model Before Distill ItECCV 2022code
Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainNeurIPS 2022
Weighted Distillation with Unlabeled ExamplesNeurIPS 2022
Respecting Transfer Gap in Knowledge DistillationNeurIPS 2022
Knowledge Distillation from A Stronger TeacherarXiv:2205.10536
Masked Generative DistillationECCV 2022code
Curriculum Temperature for Knowledge DistillationAAAI 2023code
Knowledge distillation: A good teacher is patient and consistentCVPR 2022
Knowledge Distillation with the Reused Teacher ClassifierCVPR 2022
Scaffolding a Student to Instill KnowledgeICLR2023
Function-Consistent Feature DistillationICLR2023
Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge DistillationICLR2023
Supervision Complexity and its Role in Knowledge DistillationICLR2023

Logits Knowledge Distillation

TitleVenueNote
Distilling the knowledge in a neural networkarXiv:1503.2531
Deep Model Compression: Distilling Knowledge from Noisy TeachersarXiv:161009650
Semi-Supervised Knowledge Transfer for Deep Learning from Private Training DataICLR 2017
Knowledge Adaptation: Teaching to AdaptArxiv:17022052
Learning from Multiple Teacher NetworksKDD 2017
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning resultsNIPS 2017
Training Deep Neural Networks in Generations:A More Tolerant Teacher Educates Better StudentsarXiv:1805.551
Moonshine:Distilling with Cheap ConvolutionsNIPS 2018
Learning from Multiple Teacher NetworksKDD 2017
Positive-Unlabeled Compression on the CloudNIPS 2019
Variational Student: Learning Compact and Sparser Networks in Knowledge Distillation FrameworkarXiv:1910.12061
Preparing Lessons: Improve Knowledge Distillation with Better SupervisionarXiv:1911.7471
Adaptive Regularization of LabelsarXiv:1908.5474
Learning Metrics from Teachers: Compact Networks for Image EmbeddingCVPR 2019
Diversity with Cooperation: Ensemble Methods for Few-Shot ClassificationICCV 2019
Improved Knowledge Distillation via Teacher Assistant: Bridging the Gap Between Student and TeacherarXiv:1902.3393
MEAL: Multi-Model Ensemble via Adversarial LearningAAAI 2019
Revisit Knowledge Distillation: a Teacher-free FrameworkCVPR 2020 [code]
Ensemble Distribution DistillationICLR 2020
Noisy Collaboration in Knowledge DistillationICLR 2020
Self-training with Noisy Student improves ImageNet classificationCVPR 2020
QUEST: Quantized embedding space for transferring knowledgeCVPR 2020(pre)
Meta Pseudo LabelsICML 2020
Subclass DistillationICML2020
Boosting Self-Supervised Learning via Knowledge TransferCVPR 2018
Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox ModelCVPR 2020 [code]
Regularizing Class-wise Predictions via Self-knowledge DistillationCVPR 2020 [code]
Rethinking Data Augmentation: Self-Supervision and Self-DistillationICLR 2020
What it Thinks is Important is Important: Robustness Transfers through Input GradientsCVPR 2020
Role-Wise Data Augmentation for Knowledge DistillationICLR 2020 [code]
Distilling Effective Supervision from Severe Label NoiseCVPR 2020
Learning with Noisy Class Labels for Instance SegmentationECCV 2020
Self-Distillation Amplifies Regularization in Hilbert SpacearXiv:2002.5715
MINILM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersarXiv:200210957
Hydra: Preserving Ensemble Diversity for Model DistillationarXiv:20014694
Teacher-Class Network: A Neural Network Compression MechanismarXiv:2004.3281
Learning from a Lightweight Teacher for Efficient Knowledge DistillationarXiv:2005.9163
Self-Distillation as Instance-Specific Label SmoothingarXiv:2006.5065
Self-supervised Knowledge Distillation for Few-shot LearningarXiv:2006.09785
Improving Weakly Supervised Visual Grounding by Contrastive Knowledge DistillationarXiv:2007.1951
Few Sample Knowledge Distillation for Efficient Network CompressionCVPR 2020
Learning What and Where to TransferICML 2019
Transferring Knowledge across Learning ProcessesICLR 2019
Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalICCV 2019
Diversity with Cooperation: Ensemble Methods for Few-Shot ClassificationICCV 2019
Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge DistillationarXiv:191105329v1
Progressive Knowledge Distillation For Generative ModelingICLR 2020
Few Shot Network Compression via Cross DistillationAAAI 2020

Intermediate Knowledge Distillation

TitleVenueNote
Fitnets: Hints for thin deep netsarXiv:1412.6550
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transferICLR 2017
Knowledge Projection for Effective Design of Thinner and Faster Deep Neural NetworksarXiv:1710.9505
A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer LearningCVPR 2017
Paraphrasing complex network: Network compression via factor transferNIPS 2018
Knowledge transfer with jacobian matchingICML 2018
Like What You Like: Knowledge Distill via Neuron Selectivity TransferCVPR2018
An Embarrassingly Simple Approach for Knowledge DistillationMLR 2018
Self-supervised knowledge distillation using singular value decompositionECCV 2018
Learning Deep Representations with Probabilistic Knowledge TransferECCV 2018
Correlation Congruence for Knowledge DistillationICCV 2019
Similarity-Preserving Knowledge DistillationICCV 2019
Variational Information Distillation for Knowledge TransferCVPR 2019
Knowledge Distillation via Instance Relationship GraphCVPR 2019
Knowledge Distillation via Instance Relationship GraphCVPR 2019
Knowledge Distillation via Route Constrained OptimizationICCV 2019
Similarity-Preserving Knowledge DistillationICCV 2019
Stagewise Knowledge DistillationarXiv: 1911.6786
Distilling Object Detectors with Fine-grained Feature ImitationICLR 2020
Knowledge Squeezed Adversarial Network CompressionAAAI 2020
Knowledge Distillation from Internal RepresentationsAAAI 2020
Knowledge Flow:Improve Upon Your TeachersICLR 2019
LIT: Learned Intermediate Representation Training for Model CompressionICML 2019
A Comprehensive Overhaul of Feature DistillationICCV 2019
Residual Knowledge DistillationarXiv:2002.9168
Knowledge distillation via adaptive instance normalizationarXiv:2003.4289
Channel Distillation: Channel-Wise Attention for Knowledge DistillationarXiv:2006.01683
Matching Guided DistillationECCV 2020
Differentiable Feature Aggregation Search for Knowledge DistillationECCV 2020
Local Correlation Consistency for Knowledge DistillationECCV 2020

Oneline Knowledge Distillation

TitleVenueNote
Deep Mutual LearningCVPR 2018
Born-Again Neural NetworksICML 2018
Knowledge distillation by on-the-fly native ensembleNIPS 2018
Collaborative learning for deep neural networksNIPS 2018
Unifying Heterogeneous Classifiers with DistillationCVPR 2019
Snapshot Distillation: Teacher-Student Optimization in One GenerationCVPR 2019
Deeply-supervised knowledge synergyCVPR 2019
Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationICCV 2019
Distillation-Based Training for Multi-Exit ArchitecturesICCV 2019
MSD: Multi-Self-Distillation Learning via Multi-classifiers within Deep Neural NetworksarXiv:1911.9418
FEED: Feature-level Ensemble for Knowledge DistillationAAAI 2020
Stochasticity and Skip Connection Improve Knowledge TransferICLR 2020
Online Knowledge Distillation with Diverse PeersAAAI 2020
Online Knowledge Distillation via Collaborative LearningCVPR 2020
Collaborative Learning for Faster StyleGAN EmbeddingarXiv:20071758
Online Knowledge Distillation via Collaborative LearningCVPR 2020
Feature-map-level Online Adversarial Knowledge DistillationICML 2020
Knowledge Transfer via Dense Cross-layer Mutual-distillationECCV 2020
MetaDistiller: Network Self-boosting via Meta-learned Top-down DistillationECCV 2020
ResKD: Residual-Guided Knowledge DistillationarXiv:2006.4719
Interactive Knowledge DistillationarXiv:2007.1476

Understanding Knowledge Distillation

TitleVenueNote
Do deep nets really need to be deep?NIPS 2014
When Does Label Smoothing Help?NIPS 2019
Towards Understanding Knowledge DistillationAAAI 2019
Harnessing deep neural networks with logical rulesACL 2016
Adaptive Regularization of LabelsarXiv:1908
Knowledge Isomorphism between Neural NetworksarXiv:1908
Understanding and Improving Knowledge DistillationarXiv:2002.3532
The State of Knowledge Distillation for ClassificationarXiv:1912.10850
Explaining Knowledge Distillation by Quantifying the KnowledgeCVPR 2020
DeepVID: deep visual interpretation and diagnosis for image classifiers via knowledge distillationIEEE Trans, 2019
On the Unreasonable Effectiveness of Knowledge Distillation: Analysis in the Kernel RegimearXiv:2003.13438
Why distillation helps: a statistical perspectivearXiv:2005.10419
Transferring Inductive Biases through Knowledge DistillationarXiv:2006.555
Does label smoothing mitigate label noise? Lukasik, Michal et alICML 2020
An Empirical Analysis of the Impact of Data Augmentation on Knowledge DistillationarXiv:2006.3810
Does Adversarial Transferability Indicate Knowledge Transferability?arXiv:2006.14512
On the Demystification of Knowledge Distillation: A Residual Network PerspectivearXiv:2006.16589
Teaching To Teach By Structured Dark KnowledgeICLR 2020
Inter-Region Affinity Distillation for Road Marking SegmentationCVPR 2020 [code]
Heterogeneous Knowledge Distillation using Information Flow ModelingCVPR 2020 [code]
Local Correlation Consistency for Knowledge DistillationECCV2020
Few-Shot Class-Incremental LearningCVPR 2020
Unifying distillation and privileged informationICLR 2016

Knowledge Distillation with Pruning , Quantization, NAS

TitleVenueNote
Accelerating Convolutional Neural Networks with Dominant Convolutional Kernel and Knowledge Pre-regressionECCV 2016
N2N Learning: Network to Network Compression via Policy Gradient Reinforcement LearningICLR 2018
Slimmable Neural NetworksICLR 2018
Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network AccuracyNIPS 2018
MetaPruning: Meta Learning for Automatic Neural Network Channel PruningICCV 2019
LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuningICLR 2020
Pruning with hints: an efficient framework for model accelerationICLR 2020
Knapsack Pruning with Inner DistillationarXiv:2002.8258
Training convolutional neural networks with cheap convolutions and online distillationarXiv:190913063
Cooperative Pruning in Cross-Domain Deep Neural Network CompressionIJCAI 2019
QKD: Quantization-aware Knowledge DistillationarXiv:191112491v1
Neural Network Pruning with Residual-Connections and Limited-DataCVPR 2020
Training Quantized Neural Networks with a Full-precision Auxiliary ModuleCVPR 2020
Towards Effective Low-bitwidth Convolutional Neural NetworksCVPR 2018
Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and ActivationsarXiv:19084680
Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble DistillationarXiv:200611487
Knowledge Distillation Beyond Model Compressionarxiv:20071493
Teacher Guided Architecture SearchICCV 2019
Distillation Guided Residual Learning for Binary Convolutional Neural NetworksECCV 2020
MutualNet: Adaptive ConvNet via Mutual Learning from Network Width and ResolutionECCV 2020
Improving Neural Architecture Search Image Classifiers via Ensemble LearningarXiv:19036236
Blockwisely Supervised Neural Architecture Search with Knowledge DistillationarXiv:191113053v1
Towards Oracle Knowledge Distillation with Neural Architecture SearchAAAI 2020
Search for Better Students to Learn Distilled KnowledgearXiv:200111612
Circumventing Outliers of AutoAugment with Knowledge DistillationarXiv:200311342
Network Pruning via Transformable Architecture SearchNIPS 2019
Search to Distill: Pearls are Everywhere but not the EyesCVPR 2020
AutoGAN-Distiller: Searching to Compress Generative Adversarial NetworksICML 2020 [code]

Application of Knowledge Distillation

SubTitleVenue
GraphGraph-based Knowledge Distillation by Multi-head Attention NetworkarXiv:19072226
Graph Representation Learning via Multi-task Knowledge DistillationarXiv:19115700
Deep geometric knowledge distillation with graphsarXiv:19113080
Better and faster: Knowledge transfer from multiple self-supervised learning tasks via graph distillation for video classificationIJCAI 2018
Distillating Knowledge from Graph Convolutional NetworksCVPR 2020
FaceFace model compression by distilling knowledge from neuronsAAAI 2016
MarginDistillation: distillation for margin-based softmaxarXiv:2003.2586
ReIDDistilled Person Re-Identification: Towards a More Scalable SystemCVPR 2019
Robust Re-Identification by Multiple Views Knowledge DistillationECCV 2020 [code]
DetectionLearning efficient object detection models with knowledge distillationNIPS 2017
Distilling Object Detectors with Fine-grained Feature ImitationCVPR 2019
Relation Distillation Networks for Video Object DetectionICCV 2019
Learning Lightweight Face Detector with Knowledge DistillationIEEE 2019
Teacher Supervises Students How to Learn From Partially Labeled Images for Facial Landmark DetectionICCV 2019
Learning Lightweight Lane Detection CNNs by Self Attention DistillationICCV 2019
A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionCVPR 2020 [code]
Boosting Weakly Supervised Object Detection with Progressive Knowledge TransferECCV 2020
A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionCVPR 2020 [code]
Temporal Self-Ensembling Teacher for Semi-Supervised Object DetectionIEEE 2020 [code]
Uninformed Students: Student-Teacher Anomaly Detection with Discriminative Latent EmbeddingsCVPR 2020
Distilling Knowledge from Refinement in Multiple Instance Detection NetworksarXiv:2004.10943
Enabling Incremental Knowledge Transfer for Object Detection at the EdgearXiv:2004.5746
PoseDOPE: Distillation Of Part Experts for whole-body 3D pose estimation in the wildECCV 2020
Fast Human Pose EstimationCVPR 2019
Distill Knowledge From NRSfM for Weakly Supervised 3D Pose LearningICCV 2019
SegmentationROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban ScenesCVPR 2018
Knowledge Distillation for Incremental Learning in Semantic SegmentationarXiv:1911.3462
Geometry-Aware Distillation for Indoor Semantic SegmentationCVPR 2019
Structured Knowledge Distillation for Semantic SegmentationCVPR 2019
Self-similarity Student for Partial Label Histopathology Image SegmentationECCV 2020
Knowledge Distillation for Brain Tumor SegmentationarXiv:2002.3688
ROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban ScenesCVPR 2018
Low-VisionLightweight Image Super-Resolution with Information Multi-distillation NetworkICCVW 2019
Collaborative Distillation for Ultra-Resolution Universal Style TransferCVPR 2020 [code]
VideoEfficient Video Classification Using Fewer FramesCVPR 2019
Relation Distillation Networks for Video Object DetectionICCV 2019
Teacher Supervises Students How to Learn From Partially Labeled Images for Facial Landmark DetectionICCV 2019
Progressive Teacher-student Learning for Early Action PredictionCVPR 2019
MOD: A Deep Mixture Model with Online Knowledge Distillation for Large Scale Video Temporal Concept LocalizationarXiv:1910.12295
AWSD:Adaptive Weighted Spatiotemporal Distillation for Video RepresentationICCV 2019
Dynamic Kernel Distillation for Efficient Pose Estimation in VideosICCV 2019
Online Model Distillation for Efficient Video InferenceICCV 2019
Optical Flow Distillation: Towards Efficient and Stable Video Style TransferECCV 2020
Adversarial Self-Supervised Learning for Semi-Supervised 3D Action RecognitionECCV 2020
Object Relational Graph with Teacher-Recommended Learning for Video CaptioningCVPR 2020
Spatio-Temporal Graph for Video Captioning with Knowledge distillationCVPR 2020 [code]
TA-Student VQA: Multi-Agents Training by Self-QuestioningCVPR 2020

Data-free Knowledge Distillation

TitleVenueNote
Data-Free Knowledge Distillation for Deep Neural NetworksNIPS 2017
Zero-Shot Knowledge Distillation in Deep NetworksICML 2019
DAFL:Data-Free Learning of Student NetworksICCV 2019
Zero-shot Knowledge Transfer via Adversarial Belief MatchingNIPS 2019
Dream Distillation: A Data-Independent Model Compression FrameworkICML 2019
Dreaming to Distill: Data-free Knowledge Transfer via DeepInversionCVPR 2020
Data-Free Adversarial DistillationCVPR 2020
The Knowledge Within: Methods for Data-Free Model CompressionCVPR 2020
Knowledge Extraction with No Observable DataNIPS 2019
Data-Free Knowledge Amalgamation via Group-Stack Dual-GANCVPR 2020
DeGAN : Data-Enriching GAN for Retrieving Representative Samples from a Trained ClassifierarXiv:1912.11960
Generative Low-bitwidth Data Free QuantizationarXiv:2003.3603
This dataset does not exist: training models from generated imagesarXiv:1911.2888
MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient EstimationarXiv:2005.3161
Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training DataECCV 2020
Billion-scale semi-supervised learning for image classificationarXiv:1905.00546
Data-free Parameter Pruning for Deep Neural NetworksarXiv:1507.6149
Data-Free Quantization Through Weight Equalization and Bias CorrectionICCV 2019
DAC: Data-free Automatic Acceleration of Convolutional NetworksWACV 2019

Cross-modal Knowledge Distillation

TitleVenueNote
SoundNet: Learning Sound Representations from Unlabeled Video SoundNet ArchitectureECCV 2016
Cross Modal Distillation for Supervision TransferCVPR 2016
Emotion recognition in speech using cross-modal transfer in the wildACM MM 2018
Through-Wall Human Pose Estimation Using Radio SignalsCVPR 2018
Compact Trilinear Interaction for Visual Question AnsweringICCV 2019
Cross-Modal Knowledge Distillation for Action RecognitionICIP 2019
Learning to Map Nearly AnythingarXiv:1909.6928
Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalICCV 2019
UM-Adapt: Unsupervised Multi-Task Adaptation Using Adversarial Cross-Task DistillationICCV 2019
CrDoCo: Pixel-level Domain Transfer with Cross-Domain ConsistencyCVPR 2019
XD:Cross lingual Knowledge Distillation for Polyglot Sentence Embeddings
Effective Domain Knowledge Transfer with Soft Fine-tuningarXiv:1909.2236
ASR is all you need: cross-modal distillation for lip readingarXiv:1911.12747
Knowledge distillation for semi-supervised domain adaptationarXiv:1908.7355
Domain Adaptation via Teacher-Student Learning for End-to-End Speech RecognitionarXiv:2001.1798
Cluster Alignment with a Teacher for Unsupervised Domain AdaptationICCV 2019.
Attention Bridging Network for Knowledge TransferICCV 2019
Unpaired Multi-modal Segmentation via Knowledge DistillationarXiv:2001.3111
Multi-source Distilling Domain AdaptationarXiv:1911.11554
Creating Something from Nothing: Unsupervised Knowledge Distillation for Cross-Modal HashingCVPR 2020
Improving Semantic Segmentation via Self-TrainingarXiv:2004.14960
Speech to Text Adaptation: Towards an Efficient Cross-Modal DistillationarXiv:2005.8213
Joint Progressive Knowledge Distillation and Unsupervised Domain AdaptationarXiv:2005.7839
Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior KnowledgeCVPR 2020
Large-Scale Domain Adaptation via Teacher-Student LearningarXiv:1708.5466
Large Scale Audiovisual Learning of Sounds with Weakly Labeled DataIJCAI 2020
Distilling Cross-Task Knowledge via Relationship MatchingCVPR 2020 [code]
Modality distillation with multiple stream networks for action recognitionECCV 2018
Domain Adaptation through Task DistillationECCV 2020

Adversarial Knowledge Distillation

TitleVenueNote
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial NetworksarXiv:1709.00513
KTAN: Knowledge Transfer Adversarial NetworkarXiv:1810.08126
KDGAN:Knowledge Distillation with Generative Adversarial Networks.NIPS 2018
Adversarial Learning of Portable Student NetworksAAAI 2018
Adversarial Network CompressionECCV 2018
Cross-Modality Distillation: A case for Conditional Generative Adversarial NetworksICASSP 2018
Adversarial Distillation for Efficient Recommendation with External KnowledgeTOIS 2018
Training student networks for acceleration with conditional adversarial networksBMVC 2018
Adversarial network compressionECCV 2018
KDGAN:Knowledge Distillation with Generative Adversarial NetworksNIPS 2018
DAFL:Data-Free Learning of Student NetworksICCV 2019
MEAL: Multi-Model Ensemble via Adversarial LearningAAAI 2019
Exploiting the Ground-Truth: An Adversarial Imitation Based Knowledge Distillation Approach for Event DetectionAAAI 2019
Adversarially Robust DistillationAAAI 2020
GAN-Knowledge Distillation for one-stage Object DetectionarXiv:1906.08467
Lifelong GAN: Continual Learning for Conditional Image GenerationarXiv:1908.03884
Compressing GANs using Knowledge DistillationarXiv:1902.00159
Feature-map-level Online Adversarial Knowledge DistillationICML 2020
MineGAN: effective knowledge transfer from GANs to target domains with few imagesCVPR 2020
Distilling portable Generative Adversarial Networks for Image TranslationAAAI 2020
GAN Compression: Efficient Architectures for Interactive Conditional GANsCVPR 2020

lliai/Awesome-Vision-Knowledge-Distillation

Awesome Knowledge-Distillation for CV

95

15 commits

updated Apr 30, 2024

See the code

README

Awesome Knowledge Distillation in Computer vision

[TOC]

Diffusion Knowledge Distillation

TitleVenueNote
A Comprehensive Survey on Knowledge Distillation of Diffusion Models2023Weijian Luo. [pdf]
Knowledge distillation in iterative generative models for improved sampling speed2021Eric Luhman, Troy Luhman. [pdf]
Progressive Distillation for Fast Sampling of Diffusion ModelsICLR 2022Tim Salimans and Jonathan Ho. [pdf]
On Distillation of Guided Diffusion ModelsCVPR 2023Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma, Stefano Ermon, Jonathan Ho, Tim Salimans. [pdf]
TRACT: Denoising Diffusion Models with Transitive Closure Time-Distillation2023Berthelot, David, Autef, Arnaud, Lin, Jierui, Yap, Dian Ang, Zhai, Shuangfei, Hu, Siyuan, Zheng, Daniel, Talbott, Walter, Gu, Eric. [pdf]
BK-SDM: Architecturally Compressed Stable Diffusion for Efficient Text-to-Image GenerationICML 2023Kim, Bo-Kyeong, Song, Hyoung-Kyu, Castells, Thibault, Choi, Shinkook. [pdf]
On Architectural Compression of Text-to-Image Diffusion Models2023Kim, Bo-Kyeong, Song, Hyoung-Kyu, Castells, Thibault, Choi, Shinkook. [pdf]
Knowledge Diffusion for Distillation2023Tao Huang, Yuan Zhang, Mingkai Zheng, Shan You, Fei Wang, Chen Qian, Chang Xu. [pdf]
SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds2023Yanyu Li, Huan Wang, Qing Jin, Ju Hu, Pavlo Chemerys, Yun Fu, Yanzhi Wang, Sergey Tulyakov, Jian Ren1. [pdf]
BOOT: Data-free Distillation of Denoising Diffusion Models with Bootstrapping2023Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Lingjie Liu, Josh Susskind. [pdf]
Consistency models2023Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. [pdf]

Knowledge Distillation for Semantic Segmentation

TitleVenueNote
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial NetworksarXiv:1709.00513
Knowledge Distillation for Semantic Segmentation
Structured knowledge distillation for semantic segmentationCVPR-2019
Intra-class feature variation distillation for semantic segmentationECCV-2020
Channel-wise knowledge distillation for dense predictionICCV-2021
Double Similarity Distillation for Semantic Image SegmentationTIP-2021
Cross-Image Relational Knowledge Distillation for Semantic SegmentationCVPR-2022

Knowledge Distillation for Object Detection

TitleVenueNote
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial NetworksarXiv:1709.00513
Mimicking very efficient network for object detectionCVPR 2017pdf
Distilling object detectors with fine-grained feature imitationCVPR 2019pdf
General instance distillation for object detectionCVPR 2021pdf
Distilling object detectors via decoupled featuresCVPR 2021pdf
Distilling object detectors with feature richnessNeurIPS 2021pdf
Focal and global knowledge distillation for detectorsCVPR 2022pdf
Rank Mimicking and Prediction-guided Feature ImitationAAAI 2022pdf
Prediction-Guided DistillationECCV 2022pdf
Masked Distillation with Receptive TokensICLR 2023pdf
Structural Knowledge Distillation for Object DetectionNeurIPS 2022OpenReview
Dual Relation Knowledge Distillation for Object DetectionIJCAI 2023pdf
GLAMD: Global and Local Attention Mask Distillation for Object DetectorsECCV 2022ECVA
G-DetKD: Towards General Distillation Framework for Object Detectors via Contrastive and Semantic-guided Feature ImitationICCV 2021CVF
PKD: General Distillation Framework for Object Detectors via Pearson Correlation CoefficientNeurIPS 2022OpenReview
MimicDet: Bridging the Gap Between One-Stage and Two-Stage Object DetectionECCV 2020ECVA
LabelEnc: A New Intermediate Supervision Method for Object DetectionECCV 2020ECVA
TitleVenueNote
HEtero-Assists Distillation for Heterogeneous Object DetectorsECCV 2022HEAD
LGD: Label-Guided Self-Distillation for Object DetectionAAAI 2022LGD
When Object Detection Meets Knowledge Distillation: A SurveyTPAMI
ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorCVPR 2023ScaleKD
CrossKD: Cross-Head Knowledge Distillation for Dense Object DetectionarXiv:2306.11369CrossKD

Knowledge Distillation in Vision Transformers

TitleVenueNote
Training data-efficient image transformers & distillation through attentionICML2021
Co-advise: Cross inductive bias distillationCVPR2022
Tinyvit: Fast pretraining distillation for small vision transformersarXiv:2207.10666
Attention Probe: Vision Transformer Distillation in the WildICASSP2022
Dear KD: Data-Efficient Early Knowledge Distillation for Vision TransformersCVPR2022
Efficient vision transformers via fine-grained manifold distillationNIPS2022
Cross-Architecture Knowledge DistillationarXiv:2207.05273
MiniViT: Compressing Vision Transformers with Weight MultiplexingCVPR2022
ViTKD: Practical Guidelines for ViT feature knowledge distillationarXiv 2022code

Knowledge Distillation for Teacher-Student Gaps

TitleVenueNote
Improved Knowledge Distillation via Teacher Assistant: Bridging the Gap Between Student and TeacherAAAI2020
Search to Distill: Pearls are Everywhere but not the EyesCVPR 2020
Reducing the Teacher-Student Gap via Spherical Knowledge DisitllationarXiv:2020
Knowledge Distillation via the Target-aware TransformerCVPR2022
Decoupled Knowledge DistillationCVPR 2022code
Prune Your Model Before Distill ItECCV 2022code
Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainNeurIPS 2022
Weighted Distillation with Unlabeled ExamplesNeurIPS 2022
Respecting Transfer Gap in Knowledge DistillationNeurIPS 2022
Knowledge Distillation from A Stronger TeacherarXiv:2205.10536
Masked Generative DistillationECCV 2022code
Curriculum Temperature for Knowledge DistillationAAAI 2023code
Knowledge distillation: A good teacher is patient and consistentCVPR 2022
Knowledge Distillation with the Reused Teacher ClassifierCVPR 2022
Scaffolding a Student to Instill KnowledgeICLR2023
Function-Consistent Feature DistillationICLR2023
Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge DistillationICLR2023
Supervision Complexity and its Role in Knowledge DistillationICLR2023

Logits Knowledge Distillation

TitleVenueNote
Distilling the knowledge in a neural networkarXiv:1503.2531
Deep Model Compression: Distilling Knowledge from Noisy TeachersarXiv:161009650
Semi-Supervised Knowledge Transfer for Deep Learning from Private Training DataICLR 2017
Knowledge Adaptation: Teaching to AdaptArxiv:17022052
Learning from Multiple Teacher NetworksKDD 2017
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning resultsNIPS 2017
Training Deep Neural Networks in Generations:A More Tolerant Teacher Educates Better StudentsarXiv:1805.551
Moonshine:Distilling with Cheap ConvolutionsNIPS 2018
Learning from Multiple Teacher NetworksKDD 2017
Positive-Unlabeled Compression on the CloudNIPS 2019
Variational Student: Learning Compact and Sparser Networks in Knowledge Distillation FrameworkarXiv:1910.12061
Preparing Lessons: Improve Knowledge Distillation with Better SupervisionarXiv:1911.7471
Adaptive Regularization of LabelsarXiv:1908.5474
Learning Metrics from Teachers: Compact Networks for Image EmbeddingCVPR 2019
Diversity with Cooperation: Ensemble Methods for Few-Shot ClassificationICCV 2019
Improved Knowledge Distillation via Teacher Assistant: Bridging the Gap Between Student and TeacherarXiv:1902.3393
MEAL: Multi-Model Ensemble via Adversarial LearningAAAI 2019
Revisit Knowledge Distillation: a Teacher-free FrameworkCVPR 2020 [code]
Ensemble Distribution DistillationICLR 2020
Noisy Collaboration in Knowledge DistillationICLR 2020
Self-training with Noisy Student improves ImageNet classificationCVPR 2020
QUEST: Quantized embedding space for transferring knowledgeCVPR 2020(pre)
Meta Pseudo LabelsICML 2020
Subclass DistillationICML2020
Boosting Self-Supervised Learning via Knowledge TransferCVPR 2018
Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox ModelCVPR 2020 [code]
Regularizing Class-wise Predictions via Self-knowledge DistillationCVPR 2020 [code]
Rethinking Data Augmentation: Self-Supervision and Self-DistillationICLR 2020
What it Thinks is Important is Important: Robustness Transfers through Input GradientsCVPR 2020
Role-Wise Data Augmentation for Knowledge DistillationICLR 2020 [code]
Distilling Effective Supervision from Severe Label NoiseCVPR 2020
Learning with Noisy Class Labels for Instance SegmentationECCV 2020
Self-Distillation Amplifies Regularization in Hilbert SpacearXiv:2002.5715
MINILM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersarXiv:200210957
Hydra: Preserving Ensemble Diversity for Model DistillationarXiv:20014694
Teacher-Class Network: A Neural Network Compression MechanismarXiv:2004.3281
Learning from a Lightweight Teacher for Efficient Knowledge DistillationarXiv:2005.9163
Self-Distillation as Instance-Specific Label SmoothingarXiv:2006.5065
Self-supervised Knowledge Distillation for Few-shot LearningarXiv:2006.09785
Improving Weakly Supervised Visual Grounding by Contrastive Knowledge DistillationarXiv:2007.1951
Few Sample Knowledge Distillation for Efficient Network CompressionCVPR 2020
Learning What and Where to TransferICML 2019
Transferring Knowledge across Learning ProcessesICLR 2019
Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalICCV 2019
Diversity with Cooperation: Ensemble Methods for Few-Shot ClassificationICCV 2019
Knowledge Representing: Efficient, Sparse Representation of Prior Knowledge for Knowledge DistillationarXiv:191105329v1
Progressive Knowledge Distillation For Generative ModelingICLR 2020
Few Shot Network Compression via Cross DistillationAAAI 2020

Intermediate Knowledge Distillation

TitleVenueNote
Fitnets: Hints for thin deep netsarXiv:1412.6550
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transferICLR 2017
Knowledge Projection for Effective Design of Thinner and Faster Deep Neural NetworksarXiv:1710.9505
A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer LearningCVPR 2017
Paraphrasing complex network: Network compression via factor transferNIPS 2018
Knowledge transfer with jacobian matchingICML 2018
Like What You Like: Knowledge Distill via Neuron Selectivity TransferCVPR2018
An Embarrassingly Simple Approach for Knowledge DistillationMLR 2018
Self-supervised knowledge distillation using singular value decompositionECCV 2018
Learning Deep Representations with Probabilistic Knowledge TransferECCV 2018
Correlation Congruence for Knowledge DistillationICCV 2019
Similarity-Preserving Knowledge DistillationICCV 2019
Variational Information Distillation for Knowledge TransferCVPR 2019
Knowledge Distillation via Instance Relationship GraphCVPR 2019
Knowledge Distillation via Instance Relationship GraphCVPR 2019
Knowledge Distillation via Route Constrained OptimizationICCV 2019
Similarity-Preserving Knowledge DistillationICCV 2019
Stagewise Knowledge DistillationarXiv: 1911.6786
Distilling Object Detectors with Fine-grained Feature ImitationICLR 2020
Knowledge Squeezed Adversarial Network CompressionAAAI 2020
Knowledge Distillation from Internal RepresentationsAAAI 2020
Knowledge Flow:Improve Upon Your TeachersICLR 2019
LIT: Learned Intermediate Representation Training for Model CompressionICML 2019
A Comprehensive Overhaul of Feature DistillationICCV 2019
Residual Knowledge DistillationarXiv:2002.9168
Knowledge distillation via adaptive instance normalizationarXiv:2003.4289
Channel Distillation: Channel-Wise Attention for Knowledge DistillationarXiv:2006.01683
Matching Guided DistillationECCV 2020
Differentiable Feature Aggregation Search for Knowledge DistillationECCV 2020
Local Correlation Consistency for Knowledge DistillationECCV 2020

Oneline Knowledge Distillation

TitleVenueNote
Deep Mutual LearningCVPR 2018
Born-Again Neural NetworksICML 2018
Knowledge distillation by on-the-fly native ensembleNIPS 2018
Collaborative learning for deep neural networksNIPS 2018
Unifying Heterogeneous Classifiers with DistillationCVPR 2019
Snapshot Distillation: Teacher-Student Optimization in One GenerationCVPR 2019
Deeply-supervised knowledge synergyCVPR 2019
Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationICCV 2019
Distillation-Based Training for Multi-Exit ArchitecturesICCV 2019
MSD: Multi-Self-Distillation Learning via Multi-classifiers within Deep Neural NetworksarXiv:1911.9418
FEED: Feature-level Ensemble for Knowledge DistillationAAAI 2020
Stochasticity and Skip Connection Improve Knowledge TransferICLR 2020
Online Knowledge Distillation with Diverse PeersAAAI 2020
Online Knowledge Distillation via Collaborative LearningCVPR 2020
Collaborative Learning for Faster StyleGAN EmbeddingarXiv:20071758
Online Knowledge Distillation via Collaborative LearningCVPR 2020
Feature-map-level Online Adversarial Knowledge DistillationICML 2020
Knowledge Transfer via Dense Cross-layer Mutual-distillationECCV 2020
MetaDistiller: Network Self-boosting via Meta-learned Top-down DistillationECCV 2020
ResKD: Residual-Guided Knowledge DistillationarXiv:2006.4719
Interactive Knowledge DistillationarXiv:2007.1476

Understanding Knowledge Distillation

TitleVenueNote
Do deep nets really need to be deep?NIPS 2014
When Does Label Smoothing Help?NIPS 2019
Towards Understanding Knowledge DistillationAAAI 2019
Harnessing deep neural networks with logical rulesACL 2016
Adaptive Regularization of LabelsarXiv:1908
Knowledge Isomorphism between Neural NetworksarXiv:1908
Understanding and Improving Knowledge DistillationarXiv:2002.3532
The State of Knowledge Distillation for ClassificationarXiv:1912.10850
Explaining Knowledge Distillation by Quantifying the KnowledgeCVPR 2020
DeepVID: deep visual interpretation and diagnosis for image classifiers via knowledge distillationIEEE Trans, 2019
On the Unreasonable Effectiveness of Knowledge Distillation: Analysis in the Kernel RegimearXiv:2003.13438
Why distillation helps: a statistical perspectivearXiv:2005.10419
Transferring Inductive Biases through Knowledge DistillationarXiv:2006.555
Does label smoothing mitigate label noise? Lukasik, Michal et alICML 2020
An Empirical Analysis of the Impact of Data Augmentation on Knowledge DistillationarXiv:2006.3810
Does Adversarial Transferability Indicate Knowledge Transferability?arXiv:2006.14512
On the Demystification of Knowledge Distillation: A Residual Network PerspectivearXiv:2006.16589
Teaching To Teach By Structured Dark KnowledgeICLR 2020
Inter-Region Affinity Distillation for Road Marking SegmentationCVPR 2020 [code]
Heterogeneous Knowledge Distillation using Information Flow ModelingCVPR 2020 [code]
Local Correlation Consistency for Knowledge DistillationECCV2020
Few-Shot Class-Incremental LearningCVPR 2020
Unifying distillation and privileged informationICLR 2016

Knowledge Distillation with Pruning , Quantization, NAS

TitleVenueNote
Accelerating Convolutional Neural Networks with Dominant Convolutional Kernel and Knowledge Pre-regressionECCV 2016
N2N Learning: Network to Network Compression via Policy Gradient Reinforcement LearningICLR 2018
Slimmable Neural NetworksICLR 2018
Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network AccuracyNIPS 2018
MetaPruning: Meta Learning for Automatic Neural Network Channel PruningICCV 2019
LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuningICLR 2020
Pruning with hints: an efficient framework for model accelerationICLR 2020
Knapsack Pruning with Inner DistillationarXiv:2002.8258
Training convolutional neural networks with cheap convolutions and online distillationarXiv:190913063
Cooperative Pruning in Cross-Domain Deep Neural Network CompressionIJCAI 2019
QKD: Quantization-aware Knowledge DistillationarXiv:191112491v1
Neural Network Pruning with Residual-Connections and Limited-DataCVPR 2020
Training Quantized Neural Networks with a Full-precision Auxiliary ModuleCVPR 2020
Towards Effective Low-bitwidth Convolutional Neural NetworksCVPR 2018
Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and ActivationsarXiv:19084680
Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble DistillationarXiv:200611487
Knowledge Distillation Beyond Model Compressionarxiv:20071493
Teacher Guided Architecture SearchICCV 2019
Distillation Guided Residual Learning for Binary Convolutional Neural NetworksECCV 2020
MutualNet: Adaptive ConvNet via Mutual Learning from Network Width and ResolutionECCV 2020
Improving Neural Architecture Search Image Classifiers via Ensemble LearningarXiv:19036236
Blockwisely Supervised Neural Architecture Search with Knowledge DistillationarXiv:191113053v1
Towards Oracle Knowledge Distillation with Neural Architecture SearchAAAI 2020
Search for Better Students to Learn Distilled KnowledgearXiv:200111612
Circumventing Outliers of AutoAugment with Knowledge DistillationarXiv:200311342
Network Pruning via Transformable Architecture SearchNIPS 2019
Search to Distill: Pearls are Everywhere but not the EyesCVPR 2020
AutoGAN-Distiller: Searching to Compress Generative Adversarial NetworksICML 2020 [code]

Application of Knowledge Distillation

SubTitleVenue
GraphGraph-based Knowledge Distillation by Multi-head Attention NetworkarXiv:19072226
Graph Representation Learning via Multi-task Knowledge DistillationarXiv:19115700
Deep geometric knowledge distillation with graphsarXiv:19113080
Better and faster: Knowledge transfer from multiple self-supervised learning tasks via graph distillation for video classificationIJCAI 2018
Distillating Knowledge from Graph Convolutional NetworksCVPR 2020
FaceFace model compression by distilling knowledge from neuronsAAAI 2016
MarginDistillation: distillation for margin-based softmaxarXiv:2003.2586
ReIDDistilled Person Re-Identification: Towards a More Scalable SystemCVPR 2019
Robust Re-Identification by Multiple Views Knowledge DistillationECCV 2020 [code]
DetectionLearning efficient object detection models with knowledge distillationNIPS 2017
Distilling Object Detectors with Fine-grained Feature ImitationCVPR 2019
Relation Distillation Networks for Video Object DetectionICCV 2019
Learning Lightweight Face Detector with Knowledge DistillationIEEE 2019
Teacher Supervises Students How to Learn From Partially Labeled Images for Facial Landmark DetectionICCV 2019
Learning Lightweight Lane Detection CNNs by Self Attention DistillationICCV 2019
A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionCVPR 2020 [code]
Boosting Weakly Supervised Object Detection with Progressive Knowledge TransferECCV 2020
A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionCVPR 2020 [code]
Temporal Self-Ensembling Teacher for Semi-Supervised Object DetectionIEEE 2020 [code]
Uninformed Students: Student-Teacher Anomaly Detection with Discriminative Latent EmbeddingsCVPR 2020
Distilling Knowledge from Refinement in Multiple Instance Detection NetworksarXiv:2004.10943
Enabling Incremental Knowledge Transfer for Object Detection at the EdgearXiv:2004.5746
PoseDOPE: Distillation Of Part Experts for whole-body 3D pose estimation in the wildECCV 2020
Fast Human Pose EstimationCVPR 2019
Distill Knowledge From NRSfM for Weakly Supervised 3D Pose LearningICCV 2019
SegmentationROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban ScenesCVPR 2018
Knowledge Distillation for Incremental Learning in Semantic SegmentationarXiv:1911.3462
Geometry-Aware Distillation for Indoor Semantic SegmentationCVPR 2019
Structured Knowledge Distillation for Semantic SegmentationCVPR 2019
Self-similarity Student for Partial Label Histopathology Image SegmentationECCV 2020
Knowledge Distillation for Brain Tumor SegmentationarXiv:2002.3688
ROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban ScenesCVPR 2018
Low-VisionLightweight Image Super-Resolution with Information Multi-distillation NetworkICCVW 2019
Collaborative Distillation for Ultra-Resolution Universal Style TransferCVPR 2020 [code]
VideoEfficient Video Classification Using Fewer FramesCVPR 2019
Relation Distillation Networks for Video Object DetectionICCV 2019
Teacher Supervises Students How to Learn From Partially Labeled Images for Facial Landmark DetectionICCV 2019
Progressive Teacher-student Learning for Early Action PredictionCVPR 2019
MOD: A Deep Mixture Model with Online Knowledge Distillation for Large Scale Video Temporal Concept LocalizationarXiv:1910.12295
AWSD:Adaptive Weighted Spatiotemporal Distillation for Video RepresentationICCV 2019
Dynamic Kernel Distillation for Efficient Pose Estimation in VideosICCV 2019
Online Model Distillation for Efficient Video InferenceICCV 2019
Optical Flow Distillation: Towards Efficient and Stable Video Style TransferECCV 2020
Adversarial Self-Supervised Learning for Semi-Supervised 3D Action RecognitionECCV 2020
Object Relational Graph with Teacher-Recommended Learning for Video CaptioningCVPR 2020
Spatio-Temporal Graph for Video Captioning with Knowledge distillationCVPR 2020 [code]
TA-Student VQA: Multi-Agents Training by Self-QuestioningCVPR 2020

Data-free Knowledge Distillation

TitleVenueNote
Data-Free Knowledge Distillation for Deep Neural NetworksNIPS 2017
Zero-Shot Knowledge Distillation in Deep NetworksICML 2019
DAFL:Data-Free Learning of Student NetworksICCV 2019
Zero-shot Knowledge Transfer via Adversarial Belief MatchingNIPS 2019
Dream Distillation: A Data-Independent Model Compression FrameworkICML 2019
Dreaming to Distill: Data-free Knowledge Transfer via DeepInversionCVPR 2020
Data-Free Adversarial DistillationCVPR 2020
The Knowledge Within: Methods for Data-Free Model CompressionCVPR 2020
Knowledge Extraction with No Observable DataNIPS 2019
Data-Free Knowledge Amalgamation via Group-Stack Dual-GANCVPR 2020
DeGAN : Data-Enriching GAN for Retrieving Representative Samples from a Trained ClassifierarXiv:1912.11960
Generative Low-bitwidth Data Free QuantizationarXiv:2003.3603
This dataset does not exist: training models from generated imagesarXiv:1911.2888
MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient EstimationarXiv:2005.3161
Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training DataECCV 2020
Billion-scale semi-supervised learning for image classificationarXiv:1905.00546
Data-free Parameter Pruning for Deep Neural NetworksarXiv:1507.6149
Data-Free Quantization Through Weight Equalization and Bias CorrectionICCV 2019
DAC: Data-free Automatic Acceleration of Convolutional NetworksWACV 2019

Cross-modal Knowledge Distillation

TitleVenueNote
SoundNet: Learning Sound Representations from Unlabeled Video SoundNet ArchitectureECCV 2016
Cross Modal Distillation for Supervision TransferCVPR 2016
Emotion recognition in speech using cross-modal transfer in the wildACM MM 2018
Through-Wall Human Pose Estimation Using Radio SignalsCVPR 2018
Compact Trilinear Interaction for Visual Question AnsweringICCV 2019
Cross-Modal Knowledge Distillation for Action RecognitionICIP 2019
Learning to Map Nearly AnythingarXiv:1909.6928
Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalICCV 2019
UM-Adapt: Unsupervised Multi-Task Adaptation Using Adversarial Cross-Task DistillationICCV 2019
CrDoCo: Pixel-level Domain Transfer with Cross-Domain ConsistencyCVPR 2019
XD:Cross lingual Knowledge Distillation for Polyglot Sentence Embeddings
Effective Domain Knowledge Transfer with Soft Fine-tuningarXiv:1909.2236
ASR is all you need: cross-modal distillation for lip readingarXiv:1911.12747
Knowledge distillation for semi-supervised domain adaptationarXiv:1908.7355
Domain Adaptation via Teacher-Student Learning for End-to-End Speech RecognitionarXiv:2001.1798
Cluster Alignment with a Teacher for Unsupervised Domain AdaptationICCV 2019.
Attention Bridging Network for Knowledge TransferICCV 2019
Unpaired Multi-modal Segmentation via Knowledge DistillationarXiv:2001.3111
Multi-source Distilling Domain AdaptationarXiv:1911.11554
Creating Something from Nothing: Unsupervised Knowledge Distillation for Cross-Modal HashingCVPR 2020
Improving Semantic Segmentation via Self-TrainingarXiv:2004.14960
Speech to Text Adaptation: Towards an Efficient Cross-Modal DistillationarXiv:2005.8213
Joint Progressive Knowledge Distillation and Unsupervised Domain AdaptationarXiv:2005.7839
Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior KnowledgeCVPR 2020
Large-Scale Domain Adaptation via Teacher-Student LearningarXiv:1708.5466
Large Scale Audiovisual Learning of Sounds with Weakly Labeled DataIJCAI 2020
Distilling Cross-Task Knowledge via Relationship MatchingCVPR 2020 [code]
Modality distillation with multiple stream networks for action recognitionECCV 2018
Domain Adaptation through Task DistillationECCV 2020

Adversarial Knowledge Distillation

TitleVenueNote
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial NetworksarXiv:1709.00513
KTAN: Knowledge Transfer Adversarial NetworkarXiv:1810.08126
KDGAN:Knowledge Distillation with Generative Adversarial Networks.NIPS 2018
Adversarial Learning of Portable Student NetworksAAAI 2018
Adversarial Network CompressionECCV 2018
Cross-Modality Distillation: A case for Conditional Generative Adversarial NetworksICASSP 2018
Adversarial Distillation for Efficient Recommendation with External KnowledgeTOIS 2018
Training student networks for acceleration with conditional adversarial networksBMVC 2018
Adversarial network compressionECCV 2018
KDGAN:Knowledge Distillation with Generative Adversarial NetworksNIPS 2018
DAFL:Data-Free Learning of Student NetworksICCV 2019
MEAL: Multi-Model Ensemble via Adversarial LearningAAAI 2019
Exploiting the Ground-Truth: An Adversarial Imitation Based Knowledge Distillation Approach for Event DetectionAAAI 2019
Adversarially Robust DistillationAAAI 2020
GAN-Knowledge Distillation for one-stage Object DetectionarXiv:1906.08467
Lifelong GAN: Continual Learning for Conditional Image GenerationarXiv:1908.03884
Compressing GANs using Knowledge DistillationarXiv:1902.00159
Feature-map-level Online Adversarial Knowledge DistillationICML 2020
MineGAN: effective knowledge transfer from GANs to target domains with few imagesCVPR 2020
Distilling portable Generative Adversarial Networks for Image TranslationAAAI 2020
GAN Compression: Efficient Architectures for Interactive Conditional GANsCVPR 2020

Significant stargazers

Siyuan Li

83 followers · starred Apr 2024