xlyu0106/Awesome-Latent-Space

A paper list of Awesome Latent Space.

972

120 commits

updated Jul 13, 2026

See the code

README

icon The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Awesome list badge GitHub stars MIT License arXiv Hugging Face PRs welcome WeChat Group Semantic Scholar Citations

This repository manually collects works in latent space, which will be continuously updated.

πŸ“– News

[2026/04/03] We release our survey: The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook!

[2025/11/30] We release the initial version!

Star History Chart

🌟 Overview

πŸ“„ Citation

If you find this survey helpful, a citation to our paper would be greatly appreciated:

@article{yu2026latent,
  title={The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook},
  author={Yu, Xinlei and Chen, Zhangquan and He, Yongbo and Fu, Tianyu and Yang, Cheng and Xu, Chengming and Ma, Yue and Hu, Xiaobin and Cao, Zhe and Xu, Jie and others},
  journal={arXiv preprint arXiv:2604.02029},
  year={2026}
}

🀝 Contributing

We warmly welcome contributions of excellent resources you find via pull request. Please follow the instruction in CONTRIBUTING.md if you want to make one. Additionally, if you want to have any other issue, please add our wechat group.

πŸ”₯ Methods

Large-Language-Model

DatePaper TitleIntroductionCode
2024/09Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decodingimage-
2024/09Uncovering Latent Chain of Thought Vectors in Language Modelsimage-
2024/10Understanding Reasoning in Chain-of-Thought from the Hopfieldian Viewimage-
2024/10ICLR'25
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
imageGithub
2024/11Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-RewardingimageGithub
2024/12COLM'25
Training Large Language Models to Reason in a Continuous Latent Space
imageGithub
2024/12Compressed Chain of Thought: Efficient Reasoning Through Dense Representationsimage-
2024/12ICML'25
Deliberation in Latent Space via Differentiable Cache Augmentation
image-
2025/01Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacksimageGithub
2025/01LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Modelsimage-
2025/02ICML'25
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
image-
2025/02ICML'25
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization
image-
2025/02NeurIPS'25
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
imageGithub
2025/02ICLR'26
LLM Pretraining with Continuous Concepts
imageGithub
2025/02ACL'25
SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
imageGithub
2025/02Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Predictionimage-
2025/02ICLR'25
Reasoning with Latent Thoughts: On the Power of Looped Transformers
image-
2025/02Beyond Words: A Latent Memory Approach to Internal Reasoning in LLMsimage-
2025/02EMNLP'25
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
imageGithub
2025/03ICLR'25
Reasoning to Learn from Latent Thoughts
imageGithub
2025/03Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendationimage-
2025/03MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Modelsimage-
2025/04Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models--
2025/04Efficient Pretraining Length Scalingimage-
2025/05SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought ReasoningimageGithub
2025/05NeurIPS'25
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
imageGithub
2025/05Enhancing Latent Computation in Transformers with Latent Tokensimage-
2025/05Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent SpaceimageGithub
2025/05Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMsimageGithub
2025/05Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Spaceimage-
2025/05NeurIPS'25
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
imageGithub
2025/05LARES: Latent Reasoning for Sequential Recommendationimage-
2025/05NeurIPS'25
Hybrid Latent Reasoning via Reinforcement Learning
imageGithub
2025/05NeurIPS'25
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
image-
2025/05ICLR'26
Reinforced Latent Reasoning for LLM-based Recommendation
imageGithub
2025/05Continuous Chain of Thought Enables Parallel Exploration and ReasoningimageGithub
2025/05ICML'25
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
imageGithub
2025/06Efficient Post-Training Refinement of Latent Reasoning in Large Language ModelsimageGithub
2025/06DART: Distilling Autoregressive Reasoning to Silent Thoughtimage-
2025/06EMNLP'25
Parallel Continuous Chain-of-Thought with Jacobi Iteration
imageGithub
2025/07Latent Chain-of-Thought? Decoding the Depth-Recurrent TransformerimageGithub
2025/07CTRLS: Chain-of-Thought Reasoning via Latent State Transitionimage-
2025/07Geometry of Knowledge Allows Extending Diversity Boundaries of Large Language Modelsimage-
2025/08Bridging Search and Recommendation through Latent Cross Reasoningimage-
2025/08LatentPrompt: Optimizing Promts in Latent Spaceimage-
2025/08Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputsimage-
2025/09Decoding in Latent Spaces for Efficient Inference in LLM-based Recommendationimage-
2025/09LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex ReasoningimageGithub
2025/09EMNLP'25
The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
image-
2025/09LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generationimage-
2025/09ICLR'26
SIM-CoT: Supervised Implicit Chain-of-Thought
imageGithub
2025/09PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous SpaceimageGithub
2025/09Fast Thinking for Large Language Modelsimage-
2025/09Learning to Ponder: Adaptive Reasoning in Latent Spaceimage-
2025/09Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memoryimage-
2025/09ICLR'26
MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
imageGithub
2025/09LatentEvolve: Self-Evolving Test-Time Scaling in Latent SpaceimageGithub
2025/09MARCOS: Deep Thinking by Markov Chain of Continuous Thoughtsimage-
2025/09A Formal Comparison Between Chain of Thought and Latent Thought--
2025/09ICLR'26
Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
image-
2025/10Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent SpaceimageGithub
2025/10Analyzing Latent Concepts in Code Language Modelsimage-
2025/10Exploring System 1 and 2 communication for latent reasoning in LLMsimage-
2025/10ICLR'26
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
image-
2025/10ICLR'26
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
imageGithub
2025/10ICLR'26
LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
imageGithub
2025/10ICLR'26
SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
imageGithub
2025/10Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughtsimage-
2025/10Parallel Test-Time Scaling for Latent Reasoning ModelsimageGithub
2025/10LatentBreak: Jailbreaking Large Language Models through Latent Space Feedbackimage-
2025/10Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk DetectionimageGithub
2025/10Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoningimage-
2025/10Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space ReasoningimageGithub
2025/10Language Models are Injective and Hence Invertibleimage-
2025/10LLM Latent Reasoning as Chain of SuperpositionimageGithub
2025/10ICLR'26
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
image-
2025/10ICLR'26
Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
image-
2025/10NeurIPS'25
SALS: Sparse Attention in Latent Space for KV cache Compression
image-
2025/10NeurIPS'25
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
imageGithub
2025/10Scaling Latent Reasoning via Looped Language Modelsimage-
2025/10ICLR'26
Cache-to-Cache: Direct Semantic Communication Between Large Language Model
imageGithub
2025/10NeurIPS'25
Thought Communication in Multiagent Collaboration
image-
2025/11SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy OptimizationimageGithub
2025/11Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thoughtimage-
2025/11Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language ModelsimageGithub
2025/11SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleavingimage-
2025/11Enabling Agents to Communicate Entirely in Latent Spaceimage-
2025/11Improving Latent Reasoning in LLMs via Soft Concept Mixingimage-
2025/11Your Latent Reasoning is Secretly Policy Improvement Operator--
2025/11CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningimageGithub
2025/11Learning When to Stop: Adaptive Latent Reasoning via Reinforcement LearningimageGithub
2025/11Visualizing LLM Latent Space Geometry Through Dimensionality ReductionimageGithub
2025/11Polarity-Aware Probing for Quantifying Latent Alignment in Language ModelsimageGithub
2025/11Latent Collaboration in Multi-Agent SystemsimageGithub
2025/11Next-Latent Prediction Transformers Learn Compact World ModelsimageGithub
2025/12Latent Debate: A Surrogate Framework for Interpreting LLM ThinkingimageGithub
2025/12Lightweight Latent Reasoning for Narrative Tasksimage-
2025/12ICLR'26
Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
image-
2025/12ReLaX: Reasoning with Latent Exploration for Large Reasoning Modelsimage-
2025/12Reinforcement Learning for Latent-Space Thinking in LLMsimageGithub
2025/12Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMsimage-
2025/12JEPA-Reasoner: Decoupling Latent Reasoning from Token Generationimage-
2025/12Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thoughtimage-
2025/12iCLP: Large Language Model Reasoning with Implicit Cognition Latent PlanningimageGithub
2025/12Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Spaceimage-
2025/12Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning--
2026/01Parallel Latent Reasoning for Sequential Recommendationimage-
2026/01Latent Space Communication via K-V Cache Alignmentimage-
2026/01Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language ModelsimageGithub
2026/01FlashMem: Distilling Intrinsic Latent Memory via Computation Reuseimage-
2026/01IIB-LPO: Latent Policy Optimization via Iterative Information BottleneckimageGithub
2026/01Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent SpaceimageGithub
2026/01Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clusteringimage-
2026/01Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Modelsimage-
2026/01RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation SteeringimageGithub
2026/01GeoSteer: Faithful Chain-of-Thought Steering via Latent Manifold Gradientsimage-
2026/01Reasoning While Recommending: Entropy-Guided Latent Reasoning in Generative Re-ranking Modelsimage-
2026/01Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoningimage-
2026/01UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space AnalysisimageGithub
2026/01S2GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendationimage-
2026/01The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoningimage-
2026/01PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Modelsimage-
2026/01Beyond Imitation: Reinforcement Learning for Active Latent PlanningimageGithub
2026/01Latent Adversarial Regularization for Offline Preference Optimizationimage-
2026/01Latent Chain-of-Thought as Planning: Decoupling Reasoning from VerbalizationimageGithub
2026/01Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deservesimage-
2026/01From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning--
2026/01ReGuLaR: Variational Latent Reasoning Guided by Rendered Chain-of-ThoughtimageGithub
2026/02G-MemLLM: Gated Latent Memory Augmentation for Long-Context Reasoning in Large Language Modelsimage-
2026/02Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning TasksimageGithub
2026/02Capabilities and Fundamental Limits of Latent Chain-of-Thought--
2026/02Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning ModelsimageGithub
2026/02No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs-Github
2026/02CoLT: Reasoning with Chain of Latent Tool Callsimage-
2026/02Internalizing LLM Reasoning via Discovery and Replay of Latent ActionsimageGithub
2026/02Inference-Time Rethinking with Latent Thought Vectors for Math Reasoningimage-
2026/02LatentChem: From Textual CoT to Latent Thinking in Chemical ReasoningimageGithub
2026/02DeltaKV: Residual-Based KV Cache Compression via Long-Range SimilarityimageGithub
2026/02Pretraining with Token-Level Adaptive Latent Chain-of-Thoughtimage-
2026/02Latent Reasoning with Supervised Thinking Statesimage-
2026/02Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structureimage-
2026/02Next Concept Prediction in Discrete Latent Space Leads to Stronger Language ModelsimageGithub
2026/02Talking with the Latents -- how to convert your LLM into an astronomerimage-
2026/02Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent TokensimageGithub
2026/02Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Modelsimage-
2026/02ICLR'26
LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
imageGithub
2026/02Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Modelsimage-
2026/02ICLR'26
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
image-
2026/02ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spacesimage-
2026/02SpiralFormer: Looped Transformers Can Learn Hierarchical Dependencies via Multi-Resolution Recursionimage-
2026/02GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Samplerimage-
2026/02Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generationimage-
2026/02Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training-Github
2026/02LatentMem: Customizing Latent Memory for Multi-Agent SystemsimageGithub
2026/02Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systemsimage-
2026/03LaSER: Internalizing Explicit Reasoning into Latent Space for Dense RetrievalimageGithub
2026/03ICLR'26
Multi-Head Low-Rank Attention
-Github
2026/03AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depthimage-
2026/03PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Maskingimage-
2026/03When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning-Github
2026/03ICLR'26
βˆ‡-REASONER: LLM REASONING VIA TEST-TIMEGRADIENT DESCENT IN LATENT SPACE
--
2026/03SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language Modelsimage-
2026/03NextMem: Towards Latent Factual Memory for LLM-based AgentsimageGithub
2026/03Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representationsimage-
2026/03LoopRPT: Reinforcement Pre-Training for Looped Language Modelsimage-

Vision-Language-Model

DatePaper TitleIntroductionCode
2024/10Reducing hallucinations in large vision-language models via latent space steeringimageGithub
2024/12CVPR'25
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
imageGithub
2025/01Efficient Reasoning with Hidden ThinkingimageGithub
2025/02NeurIPS'25
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
image-
2025/03Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsimageGithub
2025/05NeurIPS'25
Towards General Continuous Memory for Vision-Language Models
imageGithub
2025/05NeurIPS'25
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
imageGithub
2025/06Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensimageGithub
2025/08Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Modelsimage-
2025/09ICLR'26
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
imageGithub
2025/09ICLR'26
Latent Visual reasoning
imageGithub
2025/10Auto-scaling Continuous Memory for GUI AgentimageGithub
2025/10Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent SpaceimageGithub
2025/10CVPR'26
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
imageGithub
2025/10Latent Chain-of-Thought for Visual ReasoningimageGithub
2025/10Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMsimageGithub
2025/11Multimodal Reasoning via Latent Refocusingimage-
2025/11CVPR'26
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Model
imageGithub
2025/11L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Interventionimage-
2025/11Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual TokensimageGithub
2025/11Reading Between the Lines: Abstaining from VLM-Generated OCR Errors via Latent Representation Probesimage-
2025/11Monet: Reasoning in Latent Visual Space Beyond Image and LanguageimageGithub
2025/12Interleaved Latent Visual Reasoning with Selective Perceptual ModelingimageGithub
2025/12Mull-Tokens: Modality-Agnostic Latent Thinkingimage-
2025/12VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageimage-
2025/12Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent SpaceimageGithub
2025/12Sketch-in-Latents: Eliciting Unified Reasoning in MLLMsimageGithub
2025/12Latent Implicit Visual Reasoningimage-
2026/01Forest Before Trees: Latent Superposition for Efficient Visual ReasoningimageGithub
2026/01Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actionsimage-
2026/01LaViT: Aligning Latent Visual Thoughts for Multi-modal ReasoningimageGithub
2026/01PREGEN: Uncovering Latent Thoughts in Composed Video Retrievalimage-
2026/01Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent ReasoningimageGithub
2026/01CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embeddingimage-
2026/02PolarMem: A Training-Free Polarized Latent Graph Memory for Verifiable Multimodal AgentsimageGithub
2026/02LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsimageGithub
2026/02Dual Latent Memory for Visual Multi-agent SystemimageGithub
2026/02Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddingsimage-
2026/02Toward Cognitive Supersensing in Multimodal Large Language ModelimageGithub
2026/02Visual Reasoning over Time Series via Multi-Agent Systemsimage-
2026/02Vision-aligned Latent Reasoning for Multi-modal Large Language Modelimage-
2026/02Multimodal Latent Reasoning via Hierarchical Visual Cues Injectionimage-
2026/02LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigationimage-
2026/02MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Trainingimage-
2026/02Reason-IAD: Knowledge-Guided Dynamic Latent Reasoning for Explainable Industrial Anomaly DetectionimageGithub
2026/02Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language ModelsimageGithub
2026/02OneLatent: Single-Token Compression for Visual Latent Reasoningimage-
2026/02The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent SystemsimageGithub
2026/02Test-Time Computing for Referring Multimodal Large Language Modelsimage-
2026/02CrystaL: Spontaneous Emergence of Visual Latents in MLLMsimageGithub
2026/02Imagination Helps Visual Reasoning, But Not Yet in Latent Spaceimage-
2026/02Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detectionimage-
2026/03Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decodingimage-
2026/04Visual Enhanced Depth Scaling for Multimodal Latent ReasoningimageGithub
2026/04HyLaR: Hybrid Latent Reasoning with Decoupled Policy OptimizationimageGithub
2026/03MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolutionimage-
2026/05Representative Attention For Vision TransformersimageGithub
2026/06Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event PredictionimageGithub

Vision-Language-Action-Model

DatePaper TitleIntroductionCode
2024/10ICLR'25
Latent Action Pretraining from Videos
imageGithub
2025/05UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsimageGithub
2025/07NeurIPS'25
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
imageGithub
2025/09Align-Then-Steer: Adapting the Vision-Language Action Models through Unified Latent GuidanceimageGithub
2025/09OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervisionimage-
2025/09Latent Action Pretraining Through World Modelingimage-
2025/09Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLAimage-
2025/11SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsimageGithub
2025/11LatBot: Distilling Universal Latent Actions for Vision-Language-Action Modelsimage-
2025/11Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationimageGithub
2025/12SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadimageGithub
2025/12GLaD: Geometric Latent Distillation for Vision-Language-Action Modelsimage-
2025/12Latent Chain-of-Thought World Modeling for End-to-End Autonomous Drivingimage-
2025/12WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation ControlimageGithub
2025/12Motus: A Unified Latent Action World ModelimageGithub
2025/12LoLA: Long Horizon Latent Action Learning for General Robot Manipulationimage-
2025/12ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous DrivingimageGithub
2026/01Learning to Act Robustly with View-Invariant Latent Actionsimage-
2026/01CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videosimage-
2026/01LaST0: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Modelimage-
2026/01LatentVLA: Efficient Vision-Language Models for Autonomous Driving via Latent Action Predictionimage-
2026/01Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planningimage-
2026/01LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action QueriesimageGithub
2026/01CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Controlimage-
2026/01Vision-Language Models Unlock Task-Centric Latent Actionsimage-
2026/02Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action ModelsimageGithub
2026/02DriveWorld-VLA: Unified Latent-Space World Modeling for Autonomous Drivingimage-
2026/02Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoningimage-
2026/02ConLA: Contrastive Latent Action Learning from Human Videos for Robotic ManipulationimageGithub
2026/02VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelimageGithub
2026/02FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Executionimage-
2026/02UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Modelsimage-
2026/02CVPR'26
JALA: Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
image-
2026/03LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous DrivingimageGithub
2026/03Chain of World: World Model Thinking in Latent MotionimageGithub
2026/04OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanationimage-

Contributors

xlyu0106

88 commits

NIneeeeeem

7 commits

ImYangC7

6 commits

MichaelCao0

5 commits

xlyu0106/Awesome-Latent-Space

A paper list of Awesome Latent Space.

972

120 commits

updated Jul 13, 2026

See the code

README

icon The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Awesome list badge GitHub stars MIT License arXiv Hugging Face PRs welcome WeChat Group Semantic Scholar Citations

This repository manually collects works in latent space, which will be continuously updated.

πŸ“– News

[2026/04/03] We release our survey: The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook!

[2025/11/30] We release the initial version!

Star History Chart

🌟 Overview

πŸ“„ Citation

If you find this survey helpful, a citation to our paper would be greatly appreciated:

@article{yu2026latent,
  title={The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook},
  author={Yu, Xinlei and Chen, Zhangquan and He, Yongbo and Fu, Tianyu and Yang, Cheng and Xu, Chengming and Ma, Yue and Hu, Xiaobin and Cao, Zhe and Xu, Jie and others},
  journal={arXiv preprint arXiv:2604.02029},
  year={2026}
}

🀝 Contributing

We warmly welcome contributions of excellent resources you find via pull request. Please follow the instruction in CONTRIBUTING.md if you want to make one. Additionally, if you want to have any other issue, please add our wechat group.

πŸ”₯ Methods

Large-Language-Model

DatePaper TitleIntroductionCode
2024/09Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decodingimage-
2024/09Uncovering Latent Chain of Thought Vectors in Language Modelsimage-
2024/10Understanding Reasoning in Chain-of-Thought from the Hopfieldian Viewimage-
2024/10ICLR'25
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
imageGithub
2024/11Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-RewardingimageGithub
2024/12COLM'25
Training Large Language Models to Reason in a Continuous Latent Space
imageGithub
2024/12Compressed Chain of Thought: Efficient Reasoning Through Dense Representationsimage-
2024/12ICML'25
Deliberation in Latent Space via Differentiable Cache Augmentation
image-
2025/01Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacksimageGithub
2025/01LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Modelsimage-
2025/02ICML'25
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
image-
2025/02ICML'25
Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization
image-
2025/02NeurIPS'25
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
imageGithub
2025/02ICLR'26
LLM Pretraining with Continuous Concepts
imageGithub
2025/02ACL'25
SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
imageGithub
2025/02Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Predictionimage-
2025/02ICLR'25
Reasoning with Latent Thoughts: On the Power of Looped Transformers
image-
2025/02Beyond Words: A Latent Memory Approach to Internal Reasoning in LLMsimage-
2025/02EMNLP'25
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
imageGithub
2025/03ICLR'25
Reasoning to Learn from Latent Thoughts
imageGithub
2025/03Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendationimage-
2025/03MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Modelsimage-
2025/04Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models--
2025/04Efficient Pretraining Length Scalingimage-
2025/05SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought ReasoningimageGithub
2025/05NeurIPS'25
Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
imageGithub
2025/05Enhancing Latent Computation in Transformers with Latent Tokensimage-
2025/05Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent SpaceimageGithub
2025/05Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMsimageGithub
2025/05Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Spaceimage-
2025/05NeurIPS'25
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
imageGithub
2025/05LARES: Latent Reasoning for Sequential Recommendationimage-
2025/05NeurIPS'25
Hybrid Latent Reasoning via Reinforcement Learning
imageGithub
2025/05NeurIPS'25
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
image-
2025/05ICLR'26
Reinforced Latent Reasoning for LLM-based Recommendation
imageGithub
2025/05Continuous Chain of Thought Enables Parallel Exploration and ReasoningimageGithub
2025/05ICML'25
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
imageGithub
2025/06Efficient Post-Training Refinement of Latent Reasoning in Large Language ModelsimageGithub
2025/06DART: Distilling Autoregressive Reasoning to Silent Thoughtimage-
2025/06EMNLP'25
Parallel Continuous Chain-of-Thought with Jacobi Iteration
imageGithub
2025/07Latent Chain-of-Thought? Decoding the Depth-Recurrent TransformerimageGithub
2025/07CTRLS: Chain-of-Thought Reasoning via Latent State Transitionimage-
2025/07Geometry of Knowledge Allows Extending Diversity Boundaries of Large Language Modelsimage-
2025/08Bridging Search and Recommendation through Latent Cross Reasoningimage-
2025/08LatentPrompt: Optimizing Promts in Latent Spaceimage-
2025/08Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputsimage-
2025/09Decoding in Latent Spaces for Efficient Inference in LLM-based Recommendationimage-
2025/09LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex ReasoningimageGithub
2025/09EMNLP'25
The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
image-
2025/09LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generationimage-
2025/09ICLR'26
SIM-CoT: Supervised Implicit Chain-of-Thought
imageGithub
2025/09PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous SpaceimageGithub
2025/09Fast Thinking for Large Language Modelsimage-
2025/09Learning to Ponder: Adaptive Reasoning in Latent Spaceimage-
2025/09Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memoryimage-
2025/09ICLR'26
MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
imageGithub
2025/09LatentEvolve: Self-Evolving Test-Time Scaling in Latent SpaceimageGithub
2025/09MARCOS: Deep Thinking by Markov Chain of Continuous Thoughtsimage-
2025/09A Formal Comparison Between Chain of Thought and Latent Thought--
2025/09ICLR'26
Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
image-
2025/10Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent SpaceimageGithub
2025/10Analyzing Latent Concepts in Code Language Modelsimage-
2025/10Exploring System 1 and 2 communication for latent reasoning in LLMsimage-
2025/10ICLR'26
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
image-
2025/10ICLR'26
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
imageGithub
2025/10ICLR'26
LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
imageGithub
2025/10ICLR'26
SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
imageGithub
2025/10Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughtsimage-
2025/10Parallel Test-Time Scaling for Latent Reasoning ModelsimageGithub
2025/10LatentBreak: Jailbreaking Large Language Models through Latent Space Feedbackimage-
2025/10Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk DetectionimageGithub
2025/10Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoningimage-
2025/10Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space ReasoningimageGithub
2025/10Language Models are Injective and Hence Invertibleimage-
2025/10LLM Latent Reasoning as Chain of SuperpositionimageGithub
2025/10ICLR'26
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
image-
2025/10ICLR'26
Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
image-
2025/10NeurIPS'25
SALS: Sparse Attention in Latent Space for KV cache Compression
image-
2025/10NeurIPS'25
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
imageGithub
2025/10Scaling Latent Reasoning via Looped Language Modelsimage-
2025/10ICLR'26
Cache-to-Cache: Direct Semantic Communication Between Large Language Model
imageGithub
2025/10NeurIPS'25
Thought Communication in Multiagent Collaboration
image-
2025/11SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy OptimizationimageGithub
2025/11Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thoughtimage-
2025/11Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language ModelsimageGithub
2025/11SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleavingimage-
2025/11Enabling Agents to Communicate Entirely in Latent Spaceimage-
2025/11Improving Latent Reasoning in LLMs via Soft Concept Mixingimage-
2025/11Your Latent Reasoning is Secretly Policy Improvement Operator--
2025/11CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningimageGithub
2025/11Learning When to Stop: Adaptive Latent Reasoning via Reinforcement LearningimageGithub
2025/11Visualizing LLM Latent Space Geometry Through Dimensionality ReductionimageGithub
2025/11Polarity-Aware Probing for Quantifying Latent Alignment in Language ModelsimageGithub
2025/11Latent Collaboration in Multi-Agent SystemsimageGithub
2025/11Next-Latent Prediction Transformers Learn Compact World ModelsimageGithub
2025/12Latent Debate: A Surrogate Framework for Interpreting LLM ThinkingimageGithub
2025/12Lightweight Latent Reasoning for Narrative Tasksimage-
2025/12ICLR'26
Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
image-
2025/12ReLaX: Reasoning with Latent Exploration for Large Reasoning Modelsimage-
2025/12Reinforcement Learning for Latent-Space Thinking in LLMsimageGithub
2025/12Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMsimage-
2025/12JEPA-Reasoner: Decoupling Latent Reasoning from Token Generationimage-
2025/12Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thoughtimage-
2025/12iCLP: Large Language Model Reasoning with Implicit Cognition Latent PlanningimageGithub
2025/12Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Spaceimage-
2025/12Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning--
2026/01Parallel Latent Reasoning for Sequential Recommendationimage-
2026/01Latent Space Communication via K-V Cache Alignmentimage-
2026/01Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language ModelsimageGithub
2026/01FlashMem: Distilling Intrinsic Latent Memory via Computation Reuseimage-
2026/01IIB-LPO: Latent Policy Optimization via Iterative Information BottleneckimageGithub
2026/01Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent SpaceimageGithub
2026/01Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clusteringimage-
2026/01Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Modelsimage-
2026/01RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation SteeringimageGithub
2026/01GeoSteer: Faithful Chain-of-Thought Steering via Latent Manifold Gradientsimage-
2026/01Reasoning While Recommending: Entropy-Guided Latent Reasoning in Generative Re-ranking Modelsimage-
2026/01Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoningimage-
2026/01UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space AnalysisimageGithub
2026/01S2GR: Stepwise Semantic-Guided Reasoning in Latent Space for Generative Recommendationimage-
2026/01The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoningimage-
2026/01PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Modelsimage-
2026/01Beyond Imitation: Reinforcement Learning for Active Latent PlanningimageGithub
2026/01Latent Adversarial Regularization for Offline Preference Optimizationimage-
2026/01Latent Chain-of-Thought as Planning: Decoupling Reasoning from VerbalizationimageGithub
2026/01Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deservesimage-
2026/01From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning--
2026/01ReGuLaR: Variational Latent Reasoning Guided by Rendered Chain-of-ThoughtimageGithub
2026/02G-MemLLM: Gated Latent Memory Augmentation for Long-Context Reasoning in Large Language Modelsimage-
2026/02Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning TasksimageGithub
2026/02Capabilities and Fundamental Limits of Latent Chain-of-Thought--
2026/02Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning ModelsimageGithub
2026/02No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs-Github
2026/02CoLT: Reasoning with Chain of Latent Tool Callsimage-
2026/02Internalizing LLM Reasoning via Discovery and Replay of Latent ActionsimageGithub
2026/02Inference-Time Rethinking with Latent Thought Vectors for Math Reasoningimage-
2026/02LatentChem: From Textual CoT to Latent Thinking in Chemical ReasoningimageGithub
2026/02DeltaKV: Residual-Based KV Cache Compression via Long-Range SimilarityimageGithub
2026/02Pretraining with Token-Level Adaptive Latent Chain-of-Thoughtimage-
2026/02Latent Reasoning with Supervised Thinking Statesimage-
2026/02Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structureimage-
2026/02Next Concept Prediction in Discrete Latent Space Leads to Stronger Language ModelsimageGithub
2026/02Talking with the Latents -- how to convert your LLM into an astronomerimage-
2026/02Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent TokensimageGithub
2026/02Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Modelsimage-
2026/02ICLR'26
LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
imageGithub
2026/02Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Modelsimage-
2026/02ICLR'26
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
image-
2026/02ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spacesimage-
2026/02SpiralFormer: Looped Transformers Can Learn Hierarchical Dependencies via Multi-Resolution Recursionimage-
2026/02GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Samplerimage-
2026/02Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generationimage-
2026/02Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training-Github
2026/02LatentMem: Customizing Latent Memory for Multi-Agent SystemsimageGithub
2026/02Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systemsimage-
2026/03LaSER: Internalizing Explicit Reasoning into Latent Space for Dense RetrievalimageGithub
2026/03ICLR'26
Multi-Head Low-Rank Attention
-Github
2026/03AdaPonderLM: Gated Pondering Language Models with Token-Wise Adaptive Depthimage-
2026/03PonderLM-3: Adaptive Token-Wise Pondering with Differentiable Maskingimage-
2026/03When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning-Github
2026/03ICLR'26
βˆ‡-REASONER: LLM REASONING VIA TEST-TIMEGRADIENT DESCENT IN LATENT SPACE
--
2026/03SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language Modelsimage-
2026/03NextMem: Towards Latent Factual Memory for LLM-based AgentsimageGithub
2026/03Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representationsimage-
2026/03LoopRPT: Reinforcement Pre-Training for Looped Language Modelsimage-

Vision-Language-Model

DatePaper TitleIntroductionCode
2024/10Reducing hallucinations in large vision-language models via latent space steeringimageGithub
2024/12CVPR'25
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
imageGithub
2025/01Efficient Reasoning with Hidden ThinkingimageGithub
2025/02NeurIPS'25
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
image-
2025/03Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsimageGithub
2025/05NeurIPS'25
Towards General Continuous Memory for Vision-Language Models
imageGithub
2025/05NeurIPS'25
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
imageGithub
2025/06Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensimageGithub
2025/08Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Modelsimage-
2025/09ICLR'26
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
imageGithub
2025/09ICLR'26
Latent Visual reasoning
imageGithub
2025/10Auto-scaling Continuous Memory for GUI AgentimageGithub
2025/10Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent SpaceimageGithub
2025/10CVPR'26
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
imageGithub
2025/10Latent Chain-of-Thought for Visual ReasoningimageGithub
2025/10Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMsimageGithub
2025/11Multimodal Reasoning via Latent Refocusingimage-
2025/11CVPR'26
VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Model
imageGithub
2025/11L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Interventionimage-
2025/11Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual TokensimageGithub
2025/11Reading Between the Lines: Abstaining from VLM-Generated OCR Errors via Latent Representation Probesimage-
2025/11Monet: Reasoning in Latent Visual Space Beyond Image and LanguageimageGithub
2025/12Interleaved Latent Visual Reasoning with Selective Perceptual ModelingimageGithub
2025/12Mull-Tokens: Modality-Agnostic Latent Thinkingimage-
2025/12VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageimage-
2025/12Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent SpaceimageGithub
2025/12Sketch-in-Latents: Eliciting Unified Reasoning in MLLMsimageGithub
2025/12Latent Implicit Visual Reasoningimage-
2026/01Forest Before Trees: Latent Superposition for Efficient Visual ReasoningimageGithub
2026/01Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actionsimage-
2026/01LaViT: Aligning Latent Visual Thoughts for Multi-modal ReasoningimageGithub
2026/01PREGEN: Uncovering Latent Thoughts in Composed Video Retrievalimage-
2026/01Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent ReasoningimageGithub
2026/01CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embeddingimage-
2026/02PolarMem: A Training-Free Polarized Latent Graph Memory for Verifiable Multimodal AgentsimageGithub
2026/02LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsimageGithub
2026/02Dual Latent Memory for Visual Multi-agent SystemimageGithub
2026/02Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddingsimage-
2026/02Toward Cognitive Supersensing in Multimodal Large Language ModelimageGithub
2026/02Visual Reasoning over Time Series via Multi-Agent Systemsimage-
2026/02Vision-aligned Latent Reasoning for Multi-modal Large Language Modelimage-
2026/02Multimodal Latent Reasoning via Hierarchical Visual Cues Injectionimage-
2026/02LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigationimage-
2026/02MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Trainingimage-
2026/02Reason-IAD: Knowledge-Guided Dynamic Latent Reasoning for Explainable Industrial Anomaly DetectionimageGithub
2026/02Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language ModelsimageGithub
2026/02OneLatent: Single-Token Compression for Visual Latent Reasoningimage-
2026/02The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent SystemsimageGithub
2026/02Test-Time Computing for Referring Multimodal Large Language Modelsimage-
2026/02CrystaL: Spontaneous Emergence of Visual Latents in MLLMsimageGithub
2026/02Imagination Helps Visual Reasoning, But Not Yet in Latent Spaceimage-
2026/02Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detectionimage-
2026/03Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decodingimage-
2026/04Visual Enhanced Depth Scaling for Multimodal Latent ReasoningimageGithub
2026/04HyLaR: Hybrid Latent Reasoning with Decoupled Policy OptimizationimageGithub
2026/03MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolutionimage-
2026/05Representative Attention For Vision TransformersimageGithub
2026/06Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event PredictionimageGithub

Vision-Language-Action-Model

DatePaper TitleIntroductionCode
2024/10ICLR'25
Latent Action Pretraining from Videos
imageGithub
2025/05UniVLA: Learning to Act Anywhere with Task-centric Latent ActionsimageGithub
2025/07NeurIPS'25
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
imageGithub
2025/09Align-Then-Steer: Adapting the Vision-Language Action Models through Unified Latent GuidanceimageGithub
2025/09OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervisionimage-
2025/09Latent Action Pretraining Through World Modelingimage-
2025/09Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLAimage-
2025/11SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsimageGithub
2025/11LatBot: Distilling Universal Latent Actions for Vision-Language-Action Modelsimage-
2025/11Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationimageGithub
2025/12SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadimageGithub
2025/12GLaD: Geometric Latent Distillation for Vision-Language-Action Modelsimage-
2025/12Latent Chain-of-Thought World Modeling for End-to-End Autonomous Drivingimage-
2025/12WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation ControlimageGithub
2025/12Motus: A Unified Latent Action World ModelimageGithub
2025/12LoLA: Long Horizon Latent Action Learning for General Robot Manipulationimage-
2025/12ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous DrivingimageGithub
2026/01Learning to Act Robustly with View-Invariant Latent Actionsimage-
2026/01CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videosimage-
2026/01LaST0: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Modelimage-
2026/01LatentVLA: Efficient Vision-Language Models for Autonomous Driving via Latent Action Predictionimage-
2026/01Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planningimage-
2026/01LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action QueriesimageGithub
2026/01CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Controlimage-
2026/01Vision-Language Models Unlock Task-Centric Latent Actionsimage-
2026/02Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action ModelsimageGithub
2026/02DriveWorld-VLA: Unified Latent-Space World Modeling for Autonomous Drivingimage-
2026/02Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoningimage-
2026/02ConLA: Contrastive Latent Action Learning from Human Videos for Robotic ManipulationimageGithub
2026/02VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelimageGithub
2026/02FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Executionimage-
2026/02UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Modelsimage-
2026/02CVPR'26
JALA: Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
image-
2026/03LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous DrivingimageGithub
2026/03Chain of World: World Model Thinking in Latent MotionimageGithub
2026/04OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanationimage-

Contributors

xlyu0106

88 commits

NIneeeeeem

7 commits

ImYangC7

6 commits

MichaelCao0

5 commits