A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.
220
134 commits
updated Mar 4, 2026
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs).
📬 Have a new paper or collaboration idea? Reach out to me at: qyhhere@gmail.com.
🤝 Seeking Opportunities: I'm eager to discuss and collaborate, especially for industry internships in Large Multimodal Models!
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models (Jun. 06, 2025)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models (Feb. 22, 2025)
Open Problems in Mechanistic Interpretability (Jan. 27, 2025)
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future (Dec. 18, 2024)
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey (Dec. 3, 2024)
Mechanistic Interpretability Meets Vision Language Models: Insights and Limitations (Apr. 28, 2025)
Are SAE features from the Base Model still meaningful to LLaVA? (Dec. 6, 2024)
Bridging the VLM and mech interp communities for multimodal interpretability (Oct. 28, 2024)
Case Study: Interpreting, Manipulating, and Controlling CLIP With Sparse Autoencoders (Aug. 2, 2025)
Interpreting and Steering Features in Images (Apr. 28, 2025)
Case Study: Interpreting, Manipulating, and Controlling CLIP With Sparse Autoencoders (Apr. 28, 2025)
Laying the Foundations for Vision and Multimodal Mechanistic Interpretability & Open Problems (May. 14, 2024)
Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers (Apr. 30, 2024)
Token Activation Map for VLLM Token Activation Map to Visually Explain Multimodal LLMs (Jun. 29, 2025)
Attributing Decision Region for Object-level Foundation Model Interpreting Object-level Foundation Models via Visual Precision Search (Apr. 4, 2025)
PixelSHAP PixelSHAP: Model-agnostic Visual Attribution (Mar. 9, 2025)
Understanding the decision basis of MLLMs based on integrated gradient Where do Large Vision-Language Models Look at when Answering Questions? (Mar. 18, 2025)
Understanding the decision basis of MLLMs based on gradient and attention methods From redundancy to relevance: Enhancing explainability in multimodal large language models (Oct. 17, 2024)
Tool to Visualize LVLM Decision LVLM-Interpret: an interpretability tool for large vision-language models (Jun. 24, 2024)
Attributing LLM's Chain-of-Thought Analyzing chain-of-thought prompting in large language models via gradient-based feature attributions (Jul. 25, 2023)
Causal Interpretation of SAE Features in Vision Causal Interpretation of Sparse Autoencoder Features in Vision (Aug. 31, 2025)
L0’s Impact on SAE Features Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders (Sep. 26, 2025)
Understanding Feature Mappings in MLLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs (Jun. 13, 2025)
SAE Learning Monosemantic Features in VLMs Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models (Jun. 6, 2025)
SAE Attributing CLIP's Latent Components From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance (May. 26, 2025)
SAE Analyzing Hierarchical Structure in VMs Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders (May. 21, 2025)
DiffLens: Dissecting and Mitigating Diffusion Bias Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability (Mar. 26, 2025)
SAE Reveal Selective Remapping of Visual Concepts Sparse autoencoders reveal selective remapping of visual concepts during adaptation (Mar. 21, 2025)
SAE-V SAE-V: Interpreting Multimodal Models for Enhanced Alignment (Feb. 22, 2025)
SAE 4 Scientifically Rigorous Interpretation Sparse Autoencoders for Scientifically Rigorous Interpretation of Vision Models (Feb. 10, 2025)
Universal SAE Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment (Feb. 6, 2025)
SAE 4 LMM Large Multi-modal Models Can Interpret Features in Large Multi-modal Models (Nov. 22, 2024)
Probing MLLMs: Visual Grounding & Reasoning How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding (Aug. 27, 2025)
Probing MLLMs Probing Multimodal Large Language Models for Global and Local Semantic Representations (Jan. 6, 2025)
Probing the Role of Positional Information in VLMs Probing the Role of Positional Information in Vision-Language Models (May. 17, 2023)
Probing Hallucination in VIT Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training (Feb. 10, 2023)
Probing Representations of Numbers in VLMs Probing Representations of Numbers in Vision and Language Models (May. 06, 2023)
Probing VIT Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective (Jun. 28, 2022)
Probing VLMs 4 Visio-Linguistic Compositionality Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality (Apr. 22, 2022)
A Probing Perspective of VITs Learning Multimodal Representations Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective (Jun. 28, 2022)
Diffusion Steering Lens Decoding ViTs Decoding Vision Transformers: the Diffusion Steering Lens (Apr. 23, 2025)
Beyond Logit Lens Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs (Feb. 19, 2025)
Diffusion Lens Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines (Oct. 21, 2024)
Attention Lens Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens (Nov. 23, 2024)
Causal Interpretation of SAE Features in Vision Causal Interpretation of Sparse Autoencoder Features in Vision (Aug. 31, 2025)
Mechanisms for In-Context Vision Language Binding Investigating Mechanisms for In-Context Vision Language Binding (May. 17, 2025)
LLaVA in VQA Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering (Nov. 17, 2024)
Information Storage and Transfer Understanding Information Storage and Transfer in Multi-modal Large Language Models (June. 06, 2024)
Causal Tracing Tool 4 BLIP Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP (Aug. 27, 2023)
Mechanistic Interpretability 4 VLA Mechanistic Interpretability for Steering Vision-Language-Action Models (Aug. 30, 2025)
Textual Steering Vectors 4 MLLMs Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models (May. 20, 2025)
Diffusion Steering Lens Decoding ViTs Decoding Vision Transformers: the Diffusion Steering Lens (Apr. 23, 2025)
DiffLens: Dissecting and Mitigating Diffusion Bias Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability (Mar. 26, 2025)
LMM Concept Explainability A Concept-Based Explainability Framework for Large Multimodal Models (Nov. 30, 2024)
Concept Sliders Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models (Nov. 10, 2024)
Latent Space Steering Reducing Hallucinations in VLMs Reducing Hallucinations in Vision-Language Models via Latent Space Steering (Oct. 22, 2024)
Attribute Control Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions (Mar. 24, 2024)
Understanding Feature Mappings in VLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs (Jun. 13, 2025)
Embedding Shift Dissection on CLIP Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning (Apr. 10, 2025)
Fine-tuning Representation Shift Analyzing Fine-tuning Representation Shift for Multimodal LLMs Steering (Jan. 6, 2025)
Decomposing and Interpreting Image Representations Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP (Oct. 21, 2024)
Implicit Multimodal Alignment Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs (Oct. 5, 2024)
Interp and Editing VLR to Mitigate Hallucinations Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations (Oct. 5, 2024)
Text-Based Decomposition Interpreting CLIP's Image Representation via Text-Based Decomposition (Mar. 29, 2024)
Interpreting CLIP with Sparse Linear Concept Embeddings Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE) (Feb. 16, 2024)
A Probing Perspective of VITs Learning Multimodal Representations Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective (Jun. 28, 2022)
Bridging Modality-Specific Circuits Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs (Jun. 10, 2025)
Do VLMs Have Bad Eyes? Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability (Aug. 26, 2025)
ViTs Don't Need Trained Registers Vision Transformers Don't Need Trained Registers (Jun. 10, 2025)
Interpreting Neurons in VLMs Deciphering Functions of Neurons in Vision-Language Models (Feb. 10, 2025)
Interpreting the Second-Order Effects of Neurons in CLIP Interpreting the Second-Order Effects of Neurons in CLIP (Feb. 12, 2025)
Multi-Modal Neurons Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers (Jun. 11, 2024)
Multimodal In-Context Learning What Makes Multimodal In-Context Learning Work? (April. 25, 2024)
Modality-Specific Neurons in MLLMs MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models (Oct. 7, 2024)
What Do VLMs NOTICE? What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation (Oct. 18, 2024)
Multimodal Neurons in Pretrained Text-Only Transformers Multimodal Neurons in Pretrained Text-Only Transformers (Aug. 18, 2023)
Map the Flow: Information Flow in VideoLLMs Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs (Oct. 15, 2025)
V-SEAM: Visual Semantic Editing & Attention Modulating V-SEAM: Visual Semantic Editing & Attention Modulating (Sep. 18, 2025)
Head Attribution for Image-to-Text Information Flow Head Attribution for Image-to-Text Information Flow (Sep. 22, 2025)
Spatial Reasoning 4 VLMs Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas (Mar. 4, 2025)
Attention Lens Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens (Nov. 23, 2024)
MAIA: Multimodal Automated Interpretability Agent MAIA: Multimodal Automated Interpretability Agent (Feb. 11, 2025)
Prisma Prisma : An Open Source Toolkit for Mechanistic Interpretability in Vision and Video (Apr. 28, 2025)
LLaVA-Intrepret Towards Interpreting Visual Information Processing in Vision-Language Models (Oct. 09, 2024)
LVLM-Intrepret LVLM-Intrepret: An Interpretability Tool for Large Vision-Language Models (June. 24, 2024)
ViT Prisma ViT Prisma: A Mechanistic Interpretability Library for Vision Transformers (2023)
Causal Tracing Tool 4 BLIP Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP (Aug. 27, 2023)
The main maintainer is Yihao Quan (@Yihao Quan).
Future contributors are welcome, and feel free to send pull requests in hopes that Awesome-LMMs-Mechanistic-Interpretability can become a more mature repo in the Large Multimodal Models' community.
70 followers · starred Nov 2025
31 followers · starred Jun 2025
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.
220
134 commits
updated Mar 4, 2026
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs).
📬 Have a new paper or collaboration idea? Reach out to me at: qyhhere@gmail.com.
🤝 Seeking Opportunities: I'm eager to discuss and collaborate, especially for industry internships in Large Multimodal Models!
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models (Jun. 06, 2025)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models (Feb. 22, 2025)
Open Problems in Mechanistic Interpretability (Jan. 27, 2025)
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future (Dec. 18, 2024)
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey (Dec. 3, 2024)
Mechanistic Interpretability Meets Vision Language Models: Insights and Limitations (Apr. 28, 2025)
Are SAE features from the Base Model still meaningful to LLaVA? (Dec. 6, 2024)
Bridging the VLM and mech interp communities for multimodal interpretability (Oct. 28, 2024)
Case Study: Interpreting, Manipulating, and Controlling CLIP With Sparse Autoencoders (Aug. 2, 2025)
Interpreting and Steering Features in Images (Apr. 28, 2025)
Case Study: Interpreting, Manipulating, and Controlling CLIP With Sparse Autoencoders (Apr. 28, 2025)
Laying the Foundations for Vision and Multimodal Mechanistic Interpretability & Open Problems (May. 14, 2024)
Towards Multimodal Interpretability: Learning Sparse Interpretable Features in Vision Transformers (Apr. 30, 2024)
Token Activation Map for VLLM Token Activation Map to Visually Explain Multimodal LLMs (Jun. 29, 2025)
Attributing Decision Region for Object-level Foundation Model Interpreting Object-level Foundation Models via Visual Precision Search (Apr. 4, 2025)
PixelSHAP PixelSHAP: Model-agnostic Visual Attribution (Mar. 9, 2025)
Understanding the decision basis of MLLMs based on integrated gradient Where do Large Vision-Language Models Look at when Answering Questions? (Mar. 18, 2025)
Understanding the decision basis of MLLMs based on gradient and attention methods From redundancy to relevance: Enhancing explainability in multimodal large language models (Oct. 17, 2024)
Tool to Visualize LVLM Decision LVLM-Interpret: an interpretability tool for large vision-language models (Jun. 24, 2024)
Attributing LLM's Chain-of-Thought Analyzing chain-of-thought prompting in large language models via gradient-based feature attributions (Jul. 25, 2023)
Causal Interpretation of SAE Features in Vision Causal Interpretation of Sparse Autoencoder Features in Vision (Aug. 31, 2025)
L0’s Impact on SAE Features Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders (Sep. 26, 2025)
Understanding Feature Mappings in MLLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs (Jun. 13, 2025)
SAE Learning Monosemantic Features in VLMs Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models (Jun. 6, 2025)
SAE Attributing CLIP's Latent Components From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance (May. 26, 2025)
SAE Analyzing Hierarchical Structure in VMs Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders (May. 21, 2025)
DiffLens: Dissecting and Mitigating Diffusion Bias Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability (Mar. 26, 2025)
SAE Reveal Selective Remapping of Visual Concepts Sparse autoencoders reveal selective remapping of visual concepts during adaptation (Mar. 21, 2025)
SAE-V SAE-V: Interpreting Multimodal Models for Enhanced Alignment (Feb. 22, 2025)
SAE 4 Scientifically Rigorous Interpretation Sparse Autoencoders for Scientifically Rigorous Interpretation of Vision Models (Feb. 10, 2025)
Universal SAE Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment (Feb. 6, 2025)
SAE 4 LMM Large Multi-modal Models Can Interpret Features in Large Multi-modal Models (Nov. 22, 2024)
Probing MLLMs: Visual Grounding & Reasoning How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding (Aug. 27, 2025)
Probing MLLMs Probing Multimodal Large Language Models for Global and Local Semantic Representations (Jan. 6, 2025)
Probing the Role of Positional Information in VLMs Probing the Role of Positional Information in Vision-Language Models (May. 17, 2023)
Probing Hallucination in VIT Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-training (Feb. 10, 2023)
Probing Representations of Numbers in VLMs Probing Representations of Numbers in Vision and Language Models (May. 06, 2023)
Probing VIT Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective (Jun. 28, 2022)
Probing VLMs 4 Visio-Linguistic Compositionality Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality (Apr. 22, 2022)
A Probing Perspective of VITs Learning Multimodal Representations Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective (Jun. 28, 2022)
Diffusion Steering Lens Decoding ViTs Decoding Vision Transformers: the Diffusion Steering Lens (Apr. 23, 2025)
Beyond Logit Lens Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs (Feb. 19, 2025)
Diffusion Lens Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines (Oct. 21, 2024)
Attention Lens Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens (Nov. 23, 2024)
Causal Interpretation of SAE Features in Vision Causal Interpretation of Sparse Autoencoder Features in Vision (Aug. 31, 2025)
Mechanisms for In-Context Vision Language Binding Investigating Mechanisms for In-Context Vision Language Binding (May. 17, 2025)
LLaVA in VQA Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering (Nov. 17, 2024)
Information Storage and Transfer Understanding Information Storage and Transfer in Multi-modal Large Language Models (June. 06, 2024)
Causal Tracing Tool 4 BLIP Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP (Aug. 27, 2023)
Mechanistic Interpretability 4 VLA Mechanistic Interpretability for Steering Vision-Language-Action Models (Aug. 30, 2025)
Textual Steering Vectors 4 MLLMs Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models (May. 20, 2025)
Diffusion Steering Lens Decoding ViTs Decoding Vision Transformers: the Diffusion Steering Lens (Apr. 23, 2025)
DiffLens: Dissecting and Mitigating Diffusion Bias Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability (Mar. 26, 2025)
LMM Concept Explainability A Concept-Based Explainability Framework for Large Multimodal Models (Nov. 30, 2024)
Concept Sliders Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models (Nov. 10, 2024)
Latent Space Steering Reducing Hallucinations in VLMs Reducing Hallucinations in Vision-Language Models via Latent Space Steering (Oct. 22, 2024)
Attribute Control Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions (Mar. 24, 2024)
Understanding Feature Mappings in VLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs (Jun. 13, 2025)
Embedding Shift Dissection on CLIP Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning (Apr. 10, 2025)
Fine-tuning Representation Shift Analyzing Fine-tuning Representation Shift for Multimodal LLMs Steering (Jan. 6, 2025)
Decomposing and Interpreting Image Representations Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP (Oct. 21, 2024)
Implicit Multimodal Alignment Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs (Oct. 5, 2024)
Interp and Editing VLR to Mitigate Hallucinations Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations (Oct. 5, 2024)
Text-Based Decomposition Interpreting CLIP's Image Representation via Text-Based Decomposition (Mar. 29, 2024)
Interpreting CLIP with Sparse Linear Concept Embeddings Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE) (Feb. 16, 2024)
A Probing Perspective of VITs Learning Multimodal Representations Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective (Jun. 28, 2022)
Bridging Modality-Specific Circuits Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs (Jun. 10, 2025)
Do VLMs Have Bad Eyes? Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability (Aug. 26, 2025)
ViTs Don't Need Trained Registers Vision Transformers Don't Need Trained Registers (Jun. 10, 2025)
Interpreting Neurons in VLMs Deciphering Functions of Neurons in Vision-Language Models (Feb. 10, 2025)
Interpreting the Second-Order Effects of Neurons in CLIP Interpreting the Second-Order Effects of Neurons in CLIP (Feb. 12, 2025)
Multi-Modal Neurons Finding and Editing Multi-Modal Neurons in Pre-Trained Transformers (Jun. 11, 2024)
Multimodal In-Context Learning What Makes Multimodal In-Context Learning Work? (April. 25, 2024)
Modality-Specific Neurons in MLLMs MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models (Oct. 7, 2024)
What Do VLMs NOTICE? What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation (Oct. 18, 2024)
Multimodal Neurons in Pretrained Text-Only Transformers Multimodal Neurons in Pretrained Text-Only Transformers (Aug. 18, 2023)
Map the Flow: Information Flow in VideoLLMs Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs (Oct. 15, 2025)
V-SEAM: Visual Semantic Editing & Attention Modulating V-SEAM: Visual Semantic Editing & Attention Modulating (Sep. 18, 2025)
Head Attribution for Image-to-Text Information Flow Head Attribution for Image-to-Text Information Flow (Sep. 22, 2025)
Spatial Reasoning 4 VLMs Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas (Mar. 4, 2025)
Attention Lens Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens (Nov. 23, 2024)
MAIA: Multimodal Automated Interpretability Agent MAIA: Multimodal Automated Interpretability Agent (Feb. 11, 2025)
Prisma Prisma : An Open Source Toolkit for Mechanistic Interpretability in Vision and Video (Apr. 28, 2025)
LLaVA-Intrepret Towards Interpreting Visual Information Processing in Vision-Language Models (Oct. 09, 2024)
LVLM-Intrepret LVLM-Intrepret: An Interpretability Tool for Large Vision-Language Models (June. 24, 2024)
ViT Prisma ViT Prisma: A Mechanistic Interpretability Library for Vision Transformers (2023)
Causal Tracing Tool 4 BLIP Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP (Aug. 27, 2023)
The main maintainer is Yihao Quan (@Yihao Quan).
Future contributors are welcome, and feel free to send pull requests in hopes that Awesome-LMMs-Mechanistic-Interpretability can become a more mature repo in the Large Multimodal Models' community.
70 followers · starred Nov 2025
31 followers · starred Jun 2025