This is an on-going attempt to consolidate interesting efforts in the area of understanding / interpreting / explaining / visualizing a pre-trained ML model.
DeepVis: Deep Visualization Toolbox. Yosinski et al. ICML 2015 code | pdfSWAP: Generate adversarial poses of objects in a 3D space. Alcorn et al. CVPR 2019 code | pdfAllenNLP: Query online NLP models with user-provided inputs and observe explanations (Gradient, Integrated Gradient, SmoothGrad). Last accessed 03/2020 demo3DB: A framework for analyzing computer vision models with simulated data codeAM: Visualizing higher-layer features of a deep network. Erhan et al. 2009 pdfDeepVis: Understanding Neural Networks through Deep Visualization. Yosinski et al. ICML workshop 2015 pdf | urlMFV: Multifaceted Feature Visualization: Uncovering the different types of features learned by each neuron in deep neural networks. Nguyen et al. ICML workshop 2016 pdf | codeDGN-AM: Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. Nguyen et al. NIPS 2016 pdf | codePPGN: Plug and Play Generative Networks. Nguyen et al. CVPR 2017 pdf | codeBigGAN-AM: A cost-effective method for improving and re-purposing large, pre-trained GANs by fine-tuning their class-embeddings. Li et al. ACCV 2020 pdf | codeTCAV: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors. Kim et al. 2018 pdf | code
DTCAV: Automating Interpretability: Discovering and Testing Visual Concepts Learned by Neural Networks. Ghorbani et al. 2019 pdfSVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability. Raghu et al. 2017 pdf | codeNetwork Dissection: Quantifying Interpretability of Deep Visual Representations. Bau et al. CVPR 2017 url | pdf
GAN Dissection: Visualizing and Understanding Generative Adversarial Networks. Bau et al. ICLR 2019 pdfNet2Vec: Quantifying and Explaining how Concepts are Encoded by Filters in Deep Neural Networks. Fong & Vedaldi CVPR 2018 pdfNLIZE: A Perturbation-Driven Visual Interrogation Tool for Analyzing and Interpreting Natural Language Inference Models. Liu et al. 2018 pdfGradient: Deep inside convolutional networks: Visualising image classification models and saliency maps. Simonyan et al. 2013 pdfDeconvnet: Visualizing and understanding convolutional networks. Zeiler et al. 2014 pdfGuided-backprop: Striving for simplicity: The all convolutional net. Springenberg et al. 2015 pdfSmoothGrad: removing noise by adding noise. Smilkov et al. 2017 pdfDeepLIFT: Learning important features through propagating activation differences. Shrikumar et al. 2017 pdfIG: Axiomatic Attribution for Deep Networks. Sundararajan et al. 2018 pdf | code
EG: Learning Explainable Models Using Attribution Priors. Erion et al. 2019 pdf | codeI-GOR: Visualizing Deep Networks by Optimizing with Integrated Gradients. Qi et al. 2019 pdfBlurIG: Attribution in Scale and Space. Xu et al. CVPR 2020 pdf | codeXRAI: Better Attributions Through Regions. Kapishnikov et al. ICCV 2019 pdf | codeLRP: Beyond saliency: understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation pdf
DTD: Explaining NonLinear Classification Decisions With Deep Tayor Decomposition pdfCAM: Learning Deep Features for Discriminative Localization. Zhou et al. 2016 code | web
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization. Selvaraju et al. 2017 pdf
Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks. Chattopadhyay et al. 2017 pdf | code
Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models. Omeiza et al. 2019 pdf
NormGrad: There and Back Again: Revisiting Backpropagation Saliency Methods. Rebuffi et al. CVPR 2020 pdf | code
Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks. Wang et al. CVPR 2020 workshop pdf | code
Relevance-CAM: Your Model Already Knows Where to Look. Lee et al. CVPR 2021 pdf | code
LIFT-CAM: Towards Better Explanations of Class Activation Mapping. Jung & Oh ICCV 2021 pdf
MP: Interpretable Explanations of Black Boxes by Meaningful Perturbation. Fong et al. 2017 pdf
FIDO: Explaining image classifiers by counterfactual generation. Chang et al. ICLR 2019 pdfFG-Vis: Interpretable and Fine-Grained Visual Explanations for Convolutional Neural Networks. Wagner et al. CVPR 2019 pdfCEM: Explanations based on the Missing: Towards Contrastive Explanations with Pertinent Negatives. Dhurandhar & Chen et al. NeurIPS 2018 pdf | codeFullGrad: Full-Gradient Representation for Neural Network Visualization. Srinivas et al. NeurIPS 2019 pdfIA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers. Pan et al. NeurIPS 2021 pdfSliding-Patch: Visualizing and understanding convolutional networks. Zeiler et al. 2014 pdfPDA: Visualizing deep neural network decisions: Prediction difference analysis. Zintgraf et al. ICLR 2017 pdfRISE: Randomized Input Sampling for Explanation of Black-box Models. Petsiuk et al. BMVC 2018 pdfLIME: Why should i trust you?: Explaining the predictions of any classifier. Ribeiro et al. 2016 pdf | blog
SHAP: A Unified Approach to Interpreting Model Predictions. Lundberg et al. 2017 pdf | codeOSFT: Interpreting Black Box Models via Hypothesis Testing. Burns et al. 2019 pdfIM: Interpretation of NLP models through input marginalization. Kim et al. EMNLP 2020 pdf
Deletion & Insertion: Randomized Input Sampling for Explanation of Black-box Models. Petsiuk et al. BMVC 2018 pdfROAD: A Consistent and Efficient Evaluation Strategy for Attribution Methods. Rong & Leemann, et al. ICML 2022 pdf | codeROAR: A Benchmark for Interpretability Methods in Deep Neural Networks. Hooker et al. NeurIPS 2019 pdf | code
Sanity Checks for Saliency Maps. Adebayo et al. 2018 pdfBIM: Towards Quantitative Evaluation of Attribution Methods with Ground Truth. Yang et al. 2019 pdfSAM: The Sensitivity of Attribution Methods to Hyperparameters. Bansal, Agarwal, Nguyen. CVPR 2020 pdf | codeDeletion_BERT: Double Trouble: How to not explain a text classifier’s decisions using counterfactuals synthesized by masked language models. Pham et al. 2022 pdf | code
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior? Hase & Bansal ACL 2020 pdf | code
Teach Me to Explain: A Review of Datasets for Explainable NLP. Wiegreffe & Marasović 2021 pdf | web
BiLRP: Building and Interpreting Deep Similarity Models. Jie Zhou et al. TPAMI 2020 pdfSANE: Why do These Match? Explaining the Behavior of Image Similarity Models. Plummer et al. ECCV 2020 pdfDISE: Explainable Face Recognition. Williford et al. ECCV 2020 pdf | codexCos: An explainable cosine metric for face verification task. Lin et al. 2021 pdf | codeDeepFace-EMD: Re-ranking Using Patch-wise Earth Movers Distance Improves Out-Of-Distribution Face Identification. Phan & Nguyen. CVPR 2022 (pdf | code)L2E: Learning to Explain: Generating Stable Explanations Fast. Situ et al. ACL 2021 pdf | codeProtoPNet This Looks Like That: Deep Learning for Interpretable Image Recognition. Chen et al. NeurIPS 2019 pdf | code
ProtoTree Neural Prototype Trees for Interpretable Fine-grained Image Recognition. Nauta et al. CVPR 2021 pdf | codeEMD-Corr & CHM-Corr: Visual correspondence-based explanations improve AI robustness and human-AI team accuracy. Nguyen, Taesiri, Nguyen 2022. pdf | code141 commits
This is an on-going attempt to consolidate interesting efforts in the area of understanding / interpreting / explaining / visualizing a pre-trained ML model.
DeepVis: Deep Visualization Toolbox. Yosinski et al. ICML 2015 code | pdfSWAP: Generate adversarial poses of objects in a 3D space. Alcorn et al. CVPR 2019 code | pdfAllenNLP: Query online NLP models with user-provided inputs and observe explanations (Gradient, Integrated Gradient, SmoothGrad). Last accessed 03/2020 demo3DB: A framework for analyzing computer vision models with simulated data codeAM: Visualizing higher-layer features of a deep network. Erhan et al. 2009 pdfDeepVis: Understanding Neural Networks through Deep Visualization. Yosinski et al. ICML workshop 2015 pdf | urlMFV: Multifaceted Feature Visualization: Uncovering the different types of features learned by each neuron in deep neural networks. Nguyen et al. ICML workshop 2016 pdf | codeDGN-AM: Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. Nguyen et al. NIPS 2016 pdf | codePPGN: Plug and Play Generative Networks. Nguyen et al. CVPR 2017 pdf | codeBigGAN-AM: A cost-effective method for improving and re-purposing large, pre-trained GANs by fine-tuning their class-embeddings. Li et al. ACCV 2020 pdf | codeTCAV: Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors. Kim et al. 2018 pdf | code
DTCAV: Automating Interpretability: Discovering and Testing Visual Concepts Learned by Neural Networks. Ghorbani et al. 2019 pdfSVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability. Raghu et al. 2017 pdf | codeNetwork Dissection: Quantifying Interpretability of Deep Visual Representations. Bau et al. CVPR 2017 url | pdf
GAN Dissection: Visualizing and Understanding Generative Adversarial Networks. Bau et al. ICLR 2019 pdfNet2Vec: Quantifying and Explaining how Concepts are Encoded by Filters in Deep Neural Networks. Fong & Vedaldi CVPR 2018 pdfNLIZE: A Perturbation-Driven Visual Interrogation Tool for Analyzing and Interpreting Natural Language Inference Models. Liu et al. 2018 pdfGradient: Deep inside convolutional networks: Visualising image classification models and saliency maps. Simonyan et al. 2013 pdfDeconvnet: Visualizing and understanding convolutional networks. Zeiler et al. 2014 pdfGuided-backprop: Striving for simplicity: The all convolutional net. Springenberg et al. 2015 pdfSmoothGrad: removing noise by adding noise. Smilkov et al. 2017 pdfDeepLIFT: Learning important features through propagating activation differences. Shrikumar et al. 2017 pdfIG: Axiomatic Attribution for Deep Networks. Sundararajan et al. 2018 pdf | code
EG: Learning Explainable Models Using Attribution Priors. Erion et al. 2019 pdf | codeI-GOR: Visualizing Deep Networks by Optimizing with Integrated Gradients. Qi et al. 2019 pdfBlurIG: Attribution in Scale and Space. Xu et al. CVPR 2020 pdf | codeXRAI: Better Attributions Through Regions. Kapishnikov et al. ICCV 2019 pdf | codeLRP: Beyond saliency: understanding convolutional neural networks from saliency prediction on layer-wise relevance propagation pdf
DTD: Explaining NonLinear Classification Decisions With Deep Tayor Decomposition pdfCAM: Learning Deep Features for Discriminative Localization. Zhou et al. 2016 code | web
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization. Selvaraju et al. 2017 pdf
Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks. Chattopadhyay et al. 2017 pdf | code
Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models. Omeiza et al. 2019 pdf
NormGrad: There and Back Again: Revisiting Backpropagation Saliency Methods. Rebuffi et al. CVPR 2020 pdf | code
Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks. Wang et al. CVPR 2020 workshop pdf | code
Relevance-CAM: Your Model Already Knows Where to Look. Lee et al. CVPR 2021 pdf | code
LIFT-CAM: Towards Better Explanations of Class Activation Mapping. Jung & Oh ICCV 2021 pdf
MP: Interpretable Explanations of Black Boxes by Meaningful Perturbation. Fong et al. 2017 pdf
FIDO: Explaining image classifiers by counterfactual generation. Chang et al. ICLR 2019 pdfFG-Vis: Interpretable and Fine-Grained Visual Explanations for Convolutional Neural Networks. Wagner et al. CVPR 2019 pdfCEM: Explanations based on the Missing: Towards Contrastive Explanations with Pertinent Negatives. Dhurandhar & Chen et al. NeurIPS 2018 pdf | codeFullGrad: Full-Gradient Representation for Neural Network Visualization. Srinivas et al. NeurIPS 2019 pdfIA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers. Pan et al. NeurIPS 2021 pdfSliding-Patch: Visualizing and understanding convolutional networks. Zeiler et al. 2014 pdfPDA: Visualizing deep neural network decisions: Prediction difference analysis. Zintgraf et al. ICLR 2017 pdfRISE: Randomized Input Sampling for Explanation of Black-box Models. Petsiuk et al. BMVC 2018 pdfLIME: Why should i trust you?: Explaining the predictions of any classifier. Ribeiro et al. 2016 pdf | blog
SHAP: A Unified Approach to Interpreting Model Predictions. Lundberg et al. 2017 pdf | codeOSFT: Interpreting Black Box Models via Hypothesis Testing. Burns et al. 2019 pdfIM: Interpretation of NLP models through input marginalization. Kim et al. EMNLP 2020 pdf
Deletion & Insertion: Randomized Input Sampling for Explanation of Black-box Models. Petsiuk et al. BMVC 2018 pdfROAD: A Consistent and Efficient Evaluation Strategy for Attribution Methods. Rong & Leemann, et al. ICML 2022 pdf | codeROAR: A Benchmark for Interpretability Methods in Deep Neural Networks. Hooker et al. NeurIPS 2019 pdf | code
Sanity Checks for Saliency Maps. Adebayo et al. 2018 pdfBIM: Towards Quantitative Evaluation of Attribution Methods with Ground Truth. Yang et al. 2019 pdfSAM: The Sensitivity of Attribution Methods to Hyperparameters. Bansal, Agarwal, Nguyen. CVPR 2020 pdf | codeDeletion_BERT: Double Trouble: How to not explain a text classifier’s decisions using counterfactuals synthesized by masked language models. Pham et al. 2022 pdf | code
Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior? Hase & Bansal ACL 2020 pdf | code
Teach Me to Explain: A Review of Datasets for Explainable NLP. Wiegreffe & Marasović 2021 pdf | web
BiLRP: Building and Interpreting Deep Similarity Models. Jie Zhou et al. TPAMI 2020 pdfSANE: Why do These Match? Explaining the Behavior of Image Similarity Models. Plummer et al. ECCV 2020 pdfDISE: Explainable Face Recognition. Williford et al. ECCV 2020 pdf | codexCos: An explainable cosine metric for face verification task. Lin et al. 2021 pdf | codeDeepFace-EMD: Re-ranking Using Patch-wise Earth Movers Distance Improves Out-Of-Distribution Face Identification. Phan & Nguyen. CVPR 2022 (pdf | code)L2E: Learning to Explain: Generating Stable Explanations Fast. Situ et al. ACL 2021 pdf | codeProtoPNet This Looks Like That: Deep Learning for Interpretable Image Recognition. Chen et al. NeurIPS 2019 pdf | code
ProtoTree Neural Prototype Trees for Interpretable Fine-grained Image Recognition. Nauta et al. CVPR 2021 pdf | codeEMD-Corr & CHM-Corr: Visual correspondence-based explanations improve AI robustness and human-AI team accuracy. Nguyen, Taesiri, Nguyen 2022. pdf | code141 commits