Oliverbansk/Awesome-Test-Time-Robot-Learning

A curated list and taxonomy of test-time robot learning: RL post-training, policy steering, test-time adaptation and training, test-time scaling, in-context learning, and policy self-improvement.

Python

25

12 commits

updated Sep 30, 2026

See the code

README

🤖 Awesome Test-Time Robot Learning

Awesome License: MIT

A curated research map of how robot policies improve after pre-training and task-level post-training, using context, test-time computation, deployment experience, or online interaction.

🎯 Scope

We use test-time robot learning as a broad deployment-stage umbrella: a post-trained robot policy encounters a concrete environment and improves within an inference call, an episode, or a sequence of online interactions. The list is centered on manipulation, while including general robot-learning methods when their test-time mechanism transfers directly.

A paper belongs here when it changes deployment context; action sampling, guidance, or selection; candidate generation and verification; a fast state, memory, or representation; or policy-side parameters using deployment data. Routine pre-training, ordinary task fine-tuning with a fixed offline dataset, and planning without a learned robot policy are outside the main scope.

🗺️ Research Map

Research map of the six test-time robot learning lines

The six lines are organized by their primary deployment-time mechanism and supervision source. Hybrid methods are expected; each paper is assigned to the line that best captures its central contribution.

Compare the six lines
LineCore mechanismDeployment signalExpected effectMain limitation
RL Post-TrainingRL optimization after task-level post-trainingRewards, success labels, interventions, online rolloutsProduces an RL-improved policy, adapter, residual controller, or latent interfaceInteraction cost, resets, safety, and reward design
Test-Time Policy SteeringGuides sampling, denoising, flow, or action selection at executionValue functions, VLM rewards, dynamics, constraintsImproves action generation or selection for the current settingCandidate coverage, guidance quality, and inference latency
Test-Time Adaptation and Training (TTA & TTT)Adapts restricted parameters, state, or representations from deployment dataSelf-supervised losses, feedback, predictive objectives, trajectory historyResponds to distribution shift, temporal context, or model mismatchObjective alignment, stability, and forgetting
In-Context Learning and PromptingConditions a policy on demonstrations, video, language, or sensorimotor contextPrompt examples supplied at deploymentUses demonstrations and instructions without a gradient updateRequires a policy trained to interpret the prompt modality
Test-Time ScalingAllocates more inference compute to reasoning, sampling, search, prediction, refinement, or verificationInternal confidence, value, predicted outcomes, self-consistency, or reasoning tracesImproves decisions within the current inference call without collecting new demonstrationsGains depend on useful candidate diversity or reasoning; latency grows with compute
Policy Self-ImprovementRepeats deployment data collection and model updatesRollouts, failures, corrections, rewards, or model-generated experienceImproves behavior across deployment iterationsRequires safe collection, reliable feedback, resets, and control of policy or model drift

📚 Contents

Papers are sorted by their first public release date, from oldest to newest. Venue denotes the latest known publication venue; otherwise it is listed as arXiv.

1. 🧪 RL Post-Training

These methods use reinforcement learning after task-level post-training to improve a policy, adapter, residual controller, or latent interface from reward-bearing interaction.

Typical mechanism: Policy-wide or restricted parameter optimization, residual control, latent-space optimization, or online actor-critic updates.

Core advantage: Can improve state-action associations using rewards and interaction beyond behavior cloning.

Main limitation: Interaction cost, resets, safety, reward design, and stability of online optimization.

Papers (15)

DatePaperVenueResourcesTags
2024-07-23From Imitation to Refinement: Residual RL for Precise AssemblyICRA 2025ProjectResidual RL, Assembly, PPO, Sim-to-Real
2024-09-01Diffusion Policy Policy Optimization (DPPO)ICLR 2025-Diffusion Policy, On-Policy RL, PPO
2024-12-18Policy Decorator: Model-Agnostic Online Refinement for Large Policy ModelICLR 2025ProjectResidual RL, Online Adaptation, SAC, Model-Agnostic
2025-05-24VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningarXiv-VLA, Online RL, PPO, Process Reward
2025-06-18Steering Your Diffusion Policy with Latent Space Reinforcement LearningCoRL 2025ProjectDiffusion Policy, Latent-Space RL, Black-Box Policy, Real-World RL
2025-08-20HIL-SERL: Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement LearningScience Robotics-Human-in-the-Loop RL, Real-World RL, Human Intervention
2025-09-18Unified Latent Steering and Residual Refinement for Online Improvement of Diffusion Policy ModelsICML 2026 Workshop / CoRL 2026 Submission-Diffusion Policy, Latent Steering, Residual RL, Online RL
2025-09-23Residual Off-Policy RL for Finetuning Behavior Cloning PoliciesICLR 2026 Workshop-Residual RL, Off-Policy RL, Behavior Cloning
2025-10-16RL-100: Performant Robotic Manipulation with Real-World Reinforcement LearningScience RoboticsProjectReal-World RL, Offline-to-Online RL, Diffusion Policy
2026-01-11On-the-Fly VLA Adaptation via Test-Time Reinforcement LearningarXiv-VLA, Test-Time RL, LoRA, Online Adaptation
2026-02-13Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA ModelsICRA 2026 Workshop-VLA, Sim-to-Real, Co-Training, RL Fine-Tuning
2026-05-12TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy FinetuningRSS 2026-Diffusion Policy, Exploration, RL Fine-Tuning, Action Coverage
2026-05-19Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement LearningarXivProjectLatent-Space RL, Flow Matching, Real-World RL, Policy Adaptation
2026-06-30Adapting Generalist Robot Policies with Semantic Reinforcement LearningarXiv-VLA, Semantic Actions, Real-World RL, Online Adaptation
2026-07-09FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent SpacearXiv-Human-in-the-Loop, DAgger, Generative Policy, Action Inversion

2. 🧭 Test-Time Policy Steering

These methods keep the base policy frozen and intervene in its action-generation process by re-ranking candidates, guiding denoising or flow trajectories, optimizing noise, or applying external value and constraint signals.

Typical mechanism: The sampling process, candidate action, noise variable, denoising or flow trajectory, or execution-time controller around a frozen policy.

Core advantage: It is modular and often data-light, making it attractive when the desired behavior already exists in the policy distribution.

Main limitation: It usually cannot repair missing support; guidance quality, repeated sampling, and external models can also add substantial latency.

Papers (12)

DatePaperVenueResourcesTags
2024-10-18Steering Your Generalists: Improving Robotic Foundation Models via Value GuidanceCoRL 2024-Value Guidance, Action Re-Ranking, Offline RL, Black-Box Steering
2025-06-16DynaGuide: Steering Diffusion Policies with Active Dynamic GuidanceNeurIPS 2025-Diffusion Guidance, Latent Dynamics, Model-Based Control
2025-08-08Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-DistributionNeurIPS 2025-Diffusion Policy, Latent Dynamics, OOD Detection, Recovery
2025-11-18Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary DiffusionRA-L 2026ProjectVLA, Evolutionary Diffusion, VLM Reward, Training-Free
2026-02-03VLS: Steering Pretrained Robot Policies via Vision-Language ModelsarXiv-VLM, Reward-Guided Denoising, Diffusion Policy, Flow Matching
2026-03-09OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot PoliciesarXiv-Guidance Field, VLA, Human Demonstrations, Flow Matching
2026-05-12Retrieve-then-SteerarXiv-Retrieval, Test-Time Memory, Frozen Policy, Flow Policy
2026-06-09Test-Time Gradient Guidance of Flow Policies in Reinforcement LearningarXivCodeFlow Matching, Q-Guidance, Offline RL, Frozen Policy
2026-06-12Improving Robotic Generalist Policies via Flow Reversal SteeringarXiv-Flow Matching, Flow Inversion, VLM Guidance, Noise-Space Policy
2026-07-02Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action PoliciesarXiv-VLA, Flow Matching, Q-Guidance, Real Robot
2026-09-08Proxy Policy SteeringCoRL 2026-Proxy Policies, Velocity-Space Guidance, Flow Matching, Frozen Policy
2026-09-18Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment TransferarXiv-Cross-Embodiment, World Model, Latent-Space Search, Frozen Policy

3. 🧠 Test-Time Adaptation and Training (TTA & TTT)

These methods adapt a policy, representation, or temporal state from deployment data. They include gradient-based self-supervised adaptation, feedback-driven test-time optimization, fast-weight updates, and adaptive memories.

Typical mechanism: Model parameters or restricted adapters, fast weights, adaptive memory, temporal state, or perception and control representations.

Core advantage: Can respond to visual changes, temporal context, feedback, or model mismatch without a dedicated offline adaptation dataset for each deployment setting.

Main limitation: The deployment objective or feedback signal may not track task success, and continual updates can drift, forget, or destabilize control.

Papers (11)

DatePaperVenueResourcesTags
2018-03-30Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement LearningICLR 2019-Meta-Reinforcement Learning, Model-Based RL, Online Adaptation, Dynamics Model, MPC
2020-07-08Self-Supervised Policy Adaptation during DeploymentICLR 2021-Self-Supervised Learning, Inverse Dynamics, Visual Shift, Policy Adaptation
2023-07-03MoVie: Visual Model-Based Policy Adaptation for View GeneralizationNeurIPS 2023-Test-Time Adaptation, View Generalization, Model-Based RL, Forward Dynamics
2023-11-22Fast-Slow Test-Time Adaptation for Online Vision-and-Language NavigationICML 2024-Test-Time Adaptation, Vision-Language Navigation, Entropy Minimization, Online Adaptation
2023-12-24ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic ManipulationCVPR 2024-Test-Time Adaptation, Robot Manipulation, Multimodal LLM, Affordance
2024-02-04Fast Peer Adaptation with Context-aware ExplorationICML 2024ProjectOnline Adaptation, Context-Aware Policy, Multi-Episode History, Active Exploration
2025-05-28Communication-Efficient Desire Alignment for Proactive Embodied Human–Agent InteractionACL 2026 (Main, Oral)-Embodied Human–Agent Interaction, Online Desire Adaptation, Reflection-Based Communication, Persistent Memory
2025-07-13Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement LearningICML 2025-Test-Time Adaptation, Vision-Language Navigation, Feedback-Based RL, REINFORCE
2026-07-01FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy ImprovementCoRL 2026-Failure Recovery, Preference Adaptation, Continual Policy Improvement, Action Perturbation
2026-07-08WAM-TTT: Steering World-Action Models by Watching Human Play at Test TimearXiv-World-Action Model, Human Video, Fast Weights, Test-Time Training
2026-07-16RoboTTT: Context Scaling for Robot PoliciesarXiv-Long Context, Fast Weights, VLA, Human Video

4. 🎬 In-Context Learning and Prompting

These methods use demonstrations, videos, language, or sensorimotor trajectories in the deployment-time context to specify robot behavior. The direct in-context imitation subset predicts actions from task examples without task-specific parameter updates; the broader prompting subset studies compatible task interfaces such as human video and multimodal instructions.

Typical mechanism: The deployment-time context sequence; policy parameters are usually unchanged.

Core advantage: A user can specify a task, behavior, or constraint through examples without a task-specific gradient update or online reward optimization.

Main limitation: The policy must be trained to read the prompt modality, and reliable prompting does not guarantee the required low-level behavior is in the policy's learned repertoire.

Papers (20)

DatePaperVenueResourcesTags
2017-03-21One-Shot Imitation LearningNeurIPS 2017-Meta-Imitation, One-Shot Learning, Demonstration Conditioning, Attention
2021-05-13Coarse-to-Fine Imitation Learning: Robot Manipulation from a Single DemonstrationICRA 2021ProjectSingle Demonstration, Human Video, Visual Servoing, Trajectory Replay
2022-02-04BC-Z: Zero-Shot Task Generalization with Robotic Imitation LearningCoRL 2021ProjectHuman Video Prompt, Task Embedding, Multi-Task Imitation, Real Robot
2022-10-06VIMA: General Robot Manipulation with Multimodal PromptsICML 2023ProjectMultimodal Prompt, One-Shot Video Imitation, Cross-Attention, Simulation Benchmark
2023-01-18Human-Timescale Adaptation in an Open-Ended Task SpaceICML 2023-Meta-RL, Attention Memory, Embodied 3D, Demonstration Prompt
2024-03-19Vid2Robot: End-to-End Video-Conditioned Policy Learning with Cross-Attention TransformersRSS 2024ProjectHuman Video Prompt, Cross-Embodiment, Cross-Attention, Paired Data
2024-08-28In-Context Imitation Learning via Next-Token Prediction (ICRT)ICRA 2025-Sensorimotor Prompt, Next-Token Prediction, Transformer, In-Context Learning
2024-11-19Instant Policy: In-Context Imitation Learning via Graph DiffusionICLR 2025-Graph Diffusion, One-Shot Imitation, 3D Representation, Pseudo-Demonstrations
2025-05-27Learning Generalizable Robot Policy with Human Demonstration Video as a PromptICRA 2026-Human Video Prompt, Cross-Embodiment, In-Context Learning
2025-05-27Hierarchical Instruction-aware Embodied Visual TrackingFCS 2026ProjectLanguage-Conditioned Policy, LLM Spatial Reasoning, RAG Goal Correction, Embodied Visual Tracking
2025-05-28Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement LearningICML 2025ProjectOffline In-Context RL, Task Inference, Context Trajectories, Distribution Shift
2025-06-18Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot ManipulationIROS 2025ProjectIn-Context Imitation, Trajectory Aggregation, Uncertainty, Real Robot
2025-12-08See Once, Then Act: VLA Task Learning from One-Shot Video Demonstrations (ViVLA)arXiv-Human Video Prompt, Cross-Embodiment, Latent Action, VLA
2026-04-22AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual TrackingRA-L 2026-Cross-Embodiment, Active Visual Tracking, In-Context Policy, Zero-Shot Adaptation
2026-06-02Instant-Fold: In-Context Imitation Learning for Deformable Object ManipulationCoRL 2026-Deformable Manipulation, Human Demonstration, Flow Matching, 3D Tokens
2026-06-06SynthICL: Scalable In-context Imitation Learning with Synthetic DataCoRL 2026ProjectSynthetic Data, RGB-Only, Flow Matching, One-Shot Imitation
2026-06-29Behavior Prompting Policy: Demonstrations as Prompts for ManipulationCoRL 2026ProjectSensorimotor Prompt, Robot Demonstration, Behavior Prompting
2026-08-18Introducing S1: In-Context Learning for RoboticsSkild AI Blog-In-Context Learning, Video Prompt, Cross-Embodiment, Long-Horizon Manipulation
2026-08-19GEN-1.5: Embodied Foundation Models are One-Shot LearnersGeneralist AI Blog-Physical Prompting, Continuous Pretraining, Long Context, Few-Step Adaptation
2026-08-26RA-VLA: Retrieval-Augmented VLA for Test-Time AdaptationarXiv-VLA, In-Context Imitation, Behavior-Aligned Retrieval, Flow Matching

5. 📈 Test-Time Scaling

These methods allocate additional inference compute to reason, sample, search, predict outcomes, refine actions, or verify candidates before execution. Verification is one mechanism within this broader line rather than a separate category.

Typical mechanism: Reasoning depth, candidate count, search depth, world-model rollouts, refinement steps, verifier calls, or an adaptively allocated inference budget.

Core advantage: Can improve a fixed policy at deployment and expose a measurable compute-performance trade-off without collecting new demonstrations.

Main limitation: More computation helps only when reasoning, candidate diversity, predictive models, or selection scores contain useful signal; latency and hardware cost increase with the budget.

Papers (12)

DatePaperVenueResourcesTags
2024-07-11Robotic Control via Embodied Chain-of-Thought ReasoningCoRL 2024Project / CodeVLA, Embodied Chain-of-Thought, Grounded Reasoning, Intermediate Computation
2025-05-27Hume: Introducing System-2 Thinking in Visual-Language-Action ModelarXivProjectVLA, Value-Guided Sampling, Best-of-N, Cascaded Denoising
2025-06-21RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action ModelsRSS 2025 OOD WorkshopProjectVLA, Test-Time Sampling, Action Verification, Inference Scaling Law
2025-08-17Improving Pre-Trained Vision-Language-Action Policies with Model-Based SearchCoRL 2025 Workshop-VLA, MCTS, Model-Based Search, Inference-Time Planning
2025-09-26VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree SearchICRA 2026ProjectVLA, MCTS, World Model, Value-Guided Search
2025-10-07Verifier-free Test-Time Sampling for Vision-Language-Action ModelsICLR 2026Project / CodeVLA, Verifier-Free Selection, Condition Masking, Best-of-N
2025-10-13RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action ModelarXiv-VLA, Process Reward Model, Candidate Refinement, Test-Time Verification
2025-12-02Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling ApproacharXiv-VLA, Flow Matching, Pseudo-Count Verifier, Best-of-N
2026-02-12Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action AlignmentarXiv-Test-Time Scaling, Action Verification, VLA Alignment, Best-of-N
2026-04-21FASTER: Value-Guided Sampling for Fast RLarXivProject / CodeDiffusion Policy, Value-Guided Sampling, Early Candidate Filtering, Denoising MDP
2026-05-31τ0-WM: A Unified Video-Action World Model for Robotic ManipulationarXivProjectWorld Model, Adaptive Test-Time Compute, Imagined Rollouts, Action Rectification
2026-08-17τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time ComputationarXivProjectHierarchical VLA, World Model, Test-Time Compute, Subtask Verification

6. 🔄 Policy Self-Improvement

These methods use deployment rollouts, failures, corrections, or model-generated experience to improve future behavior across successive data-collection and update rounds.

Typical mechanism: The base policy, residual policy, critic, world model, recovery behavior, or data-generation loop between deployment iterations.

Core advantage: Turns autonomous experience and intervention data into better future behavior, reducing dependence on repeated full demonstrations.

Main limitation: Requires reliable rewards or success labels, safe data collection, practical resets, and controls against policy drift or errors in model-generated data.

Papers (8)

DatePaperVenueResourcesTags
2023-06-20RoboCat: A Self-Improving Generalist Agent for Robotic ManipulationTMLR 2024-Generalist Policy, Self-Generated Data, Iterative Training, Multi-Embodiment
2025-05-28VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language ModelsIROS 2025ProjectEmbodied Visual Tracking, Failure Recovery, Memory-Augmented Self-Reflection, Iterative Recovery Improvement
2025-09-09RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and CorrectionarXivProjectLong-Horizon Manipulation, Human Intervention, Recovery Data, Iterative Imitation Learning
2025-10-30Self-Improving Vision-Language-Action Models with Data Generation via Residual RLICLR 2026ProjectVLA, Residual RL, Deployment-Aligned Data, Policy Distillation
2026-02-12VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelarXivProjectVLA, World Model, Synthetic Rollouts, Iterative Co-Improvement
2026-03-17DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World ModelsarXivProjectVision-Language Planner, Video World Model, Synthetic Rollouts, Reinforcement Fine-Tuning
2026-05-06When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement LearningarXivProjectOffline-to-Online RL, Q-Estimation, Q-Gating, On-Robot Learning
2026-08-21Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-PlanningCoRL 2026Project / CodeFrozen BC Policy, Off-Policy Q-Learning, Failure Rollouts, Value-Guided Planning

🤝 Contributing

Paper additions and corrections are welcome. Please read CONTRIBUTING.md, edit data/papers.json, and run:

python3 scripts/validate_data.py
python3 scripts/build_readme.py --check

You can also use the Add a paper issue template.

🙏 Acknowledgements

The repository structure draws inspiration from community-maintained collections such as Awesome World Models for Robotics, Awesome VLA Post-Training, and Awesome VLA.

📄 License

This repository is released under the MIT License.

Oliverbansk/Awesome-Test-Time-Robot-Learning

A curated list and taxonomy of test-time robot learning: RL post-training, policy steering, test-time adaptation and training, test-time scaling, in-context learning, and policy self-improvement.

Python

25

12 commits

updated Sep 30, 2026

See the code

README

🤖 Awesome Test-Time Robot Learning

Awesome License: MIT

A curated research map of how robot policies improve after pre-training and task-level post-training, using context, test-time computation, deployment experience, or online interaction.

🎯 Scope

We use test-time robot learning as a broad deployment-stage umbrella: a post-trained robot policy encounters a concrete environment and improves within an inference call, an episode, or a sequence of online interactions. The list is centered on manipulation, while including general robot-learning methods when their test-time mechanism transfers directly.

A paper belongs here when it changes deployment context; action sampling, guidance, or selection; candidate generation and verification; a fast state, memory, or representation; or policy-side parameters using deployment data. Routine pre-training, ordinary task fine-tuning with a fixed offline dataset, and planning without a learned robot policy are outside the main scope.

🗺️ Research Map

Research map of the six test-time robot learning lines

The six lines are organized by their primary deployment-time mechanism and supervision source. Hybrid methods are expected; each paper is assigned to the line that best captures its central contribution.

Compare the six lines
LineCore mechanismDeployment signalExpected effectMain limitation
RL Post-TrainingRL optimization after task-level post-trainingRewards, success labels, interventions, online rolloutsProduces an RL-improved policy, adapter, residual controller, or latent interfaceInteraction cost, resets, safety, and reward design
Test-Time Policy SteeringGuides sampling, denoising, flow, or action selection at executionValue functions, VLM rewards, dynamics, constraintsImproves action generation or selection for the current settingCandidate coverage, guidance quality, and inference latency
Test-Time Adaptation and Training (TTA & TTT)Adapts restricted parameters, state, or representations from deployment dataSelf-supervised losses, feedback, predictive objectives, trajectory historyResponds to distribution shift, temporal context, or model mismatchObjective alignment, stability, and forgetting
In-Context Learning and PromptingConditions a policy on demonstrations, video, language, or sensorimotor contextPrompt examples supplied at deploymentUses demonstrations and instructions without a gradient updateRequires a policy trained to interpret the prompt modality
Test-Time ScalingAllocates more inference compute to reasoning, sampling, search, prediction, refinement, or verificationInternal confidence, value, predicted outcomes, self-consistency, or reasoning tracesImproves decisions within the current inference call without collecting new demonstrationsGains depend on useful candidate diversity or reasoning; latency grows with compute
Policy Self-ImprovementRepeats deployment data collection and model updatesRollouts, failures, corrections, rewards, or model-generated experienceImproves behavior across deployment iterationsRequires safe collection, reliable feedback, resets, and control of policy or model drift

📚 Contents

Papers are sorted by their first public release date, from oldest to newest. Venue denotes the latest known publication venue; otherwise it is listed as arXiv.

1. 🧪 RL Post-Training

These methods use reinforcement learning after task-level post-training to improve a policy, adapter, residual controller, or latent interface from reward-bearing interaction.

Typical mechanism: Policy-wide or restricted parameter optimization, residual control, latent-space optimization, or online actor-critic updates.

Core advantage: Can improve state-action associations using rewards and interaction beyond behavior cloning.

Main limitation: Interaction cost, resets, safety, reward design, and stability of online optimization.

Papers (15)

DatePaperVenueResourcesTags
2024-07-23From Imitation to Refinement: Residual RL for Precise AssemblyICRA 2025ProjectResidual RL, Assembly, PPO, Sim-to-Real
2024-09-01Diffusion Policy Policy Optimization (DPPO)ICLR 2025-Diffusion Policy, On-Policy RL, PPO
2024-12-18Policy Decorator: Model-Agnostic Online Refinement for Large Policy ModelICLR 2025ProjectResidual RL, Online Adaptation, SAC, Model-Agnostic
2025-05-24VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningarXiv-VLA, Online RL, PPO, Process Reward
2025-06-18Steering Your Diffusion Policy with Latent Space Reinforcement LearningCoRL 2025ProjectDiffusion Policy, Latent-Space RL, Black-Box Policy, Real-World RL
2025-08-20HIL-SERL: Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement LearningScience Robotics-Human-in-the-Loop RL, Real-World RL, Human Intervention
2025-09-18Unified Latent Steering and Residual Refinement for Online Improvement of Diffusion Policy ModelsICML 2026 Workshop / CoRL 2026 Submission-Diffusion Policy, Latent Steering, Residual RL, Online RL
2025-09-23Residual Off-Policy RL for Finetuning Behavior Cloning PoliciesICLR 2026 Workshop-Residual RL, Off-Policy RL, Behavior Cloning
2025-10-16RL-100: Performant Robotic Manipulation with Real-World Reinforcement LearningScience RoboticsProjectReal-World RL, Offline-to-Online RL, Diffusion Policy
2026-01-11On-the-Fly VLA Adaptation via Test-Time Reinforcement LearningarXiv-VLA, Test-Time RL, LoRA, Online Adaptation
2026-02-13Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA ModelsICRA 2026 Workshop-VLA, Sim-to-Real, Co-Training, RL Fine-Tuning
2026-05-12TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy FinetuningRSS 2026-Diffusion Policy, Exploration, RL Fine-Tuning, Action Coverage
2026-05-19Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement LearningarXivProjectLatent-Space RL, Flow Matching, Real-World RL, Policy Adaptation
2026-06-30Adapting Generalist Robot Policies with Semantic Reinforcement LearningarXiv-VLA, Semantic Actions, Real-World RL, Online Adaptation
2026-07-09FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent SpacearXiv-Human-in-the-Loop, DAgger, Generative Policy, Action Inversion

2. 🧭 Test-Time Policy Steering

These methods keep the base policy frozen and intervene in its action-generation process by re-ranking candidates, guiding denoising or flow trajectories, optimizing noise, or applying external value and constraint signals.

Typical mechanism: The sampling process, candidate action, noise variable, denoising or flow trajectory, or execution-time controller around a frozen policy.

Core advantage: It is modular and often data-light, making it attractive when the desired behavior already exists in the policy distribution.

Main limitation: It usually cannot repair missing support; guidance quality, repeated sampling, and external models can also add substantial latency.

Papers (12)

DatePaperVenueResourcesTags
2024-10-18Steering Your Generalists: Improving Robotic Foundation Models via Value GuidanceCoRL 2024-Value Guidance, Action Re-Ranking, Offline RL, Black-Box Steering
2025-06-16DynaGuide: Steering Diffusion Policies with Active Dynamic GuidanceNeurIPS 2025-Diffusion Guidance, Latent Dynamics, Model-Based Control
2025-08-08Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-DistributionNeurIPS 2025-Diffusion Policy, Latent Dynamics, OOD Detection, Recovery
2025-11-18Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary DiffusionRA-L 2026ProjectVLA, Evolutionary Diffusion, VLM Reward, Training-Free
2026-02-03VLS: Steering Pretrained Robot Policies via Vision-Language ModelsarXiv-VLM, Reward-Guided Denoising, Diffusion Policy, Flow Matching
2026-03-09OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot PoliciesarXiv-Guidance Field, VLA, Human Demonstrations, Flow Matching
2026-05-12Retrieve-then-SteerarXiv-Retrieval, Test-Time Memory, Frozen Policy, Flow Policy
2026-06-09Test-Time Gradient Guidance of Flow Policies in Reinforcement LearningarXivCodeFlow Matching, Q-Guidance, Offline RL, Frozen Policy
2026-06-12Improving Robotic Generalist Policies via Flow Reversal SteeringarXiv-Flow Matching, Flow Inversion, VLM Guidance, Noise-Space Policy
2026-07-02Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action PoliciesarXiv-VLA, Flow Matching, Q-Guidance, Real Robot
2026-09-08Proxy Policy SteeringCoRL 2026-Proxy Policies, Velocity-Space Guidance, Flow Matching, Frozen Policy
2026-09-18Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment TransferarXiv-Cross-Embodiment, World Model, Latent-Space Search, Frozen Policy

3. 🧠 Test-Time Adaptation and Training (TTA & TTT)

These methods adapt a policy, representation, or temporal state from deployment data. They include gradient-based self-supervised adaptation, feedback-driven test-time optimization, fast-weight updates, and adaptive memories.

Typical mechanism: Model parameters or restricted adapters, fast weights, adaptive memory, temporal state, or perception and control representations.

Core advantage: Can respond to visual changes, temporal context, feedback, or model mismatch without a dedicated offline adaptation dataset for each deployment setting.

Main limitation: The deployment objective or feedback signal may not track task success, and continual updates can drift, forget, or destabilize control.

Papers (11)

DatePaperVenueResourcesTags
2018-03-30Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement LearningICLR 2019-Meta-Reinforcement Learning, Model-Based RL, Online Adaptation, Dynamics Model, MPC
2020-07-08Self-Supervised Policy Adaptation during DeploymentICLR 2021-Self-Supervised Learning, Inverse Dynamics, Visual Shift, Policy Adaptation
2023-07-03MoVie: Visual Model-Based Policy Adaptation for View GeneralizationNeurIPS 2023-Test-Time Adaptation, View Generalization, Model-Based RL, Forward Dynamics
2023-11-22Fast-Slow Test-Time Adaptation for Online Vision-and-Language NavigationICML 2024-Test-Time Adaptation, Vision-Language Navigation, Entropy Minimization, Online Adaptation
2023-12-24ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic ManipulationCVPR 2024-Test-Time Adaptation, Robot Manipulation, Multimodal LLM, Affordance
2024-02-04Fast Peer Adaptation with Context-aware ExplorationICML 2024ProjectOnline Adaptation, Context-Aware Policy, Multi-Episode History, Active Exploration
2025-05-28Communication-Efficient Desire Alignment for Proactive Embodied Human–Agent InteractionACL 2026 (Main, Oral)-Embodied Human–Agent Interaction, Online Desire Adaptation, Reflection-Based Communication, Persistent Memory
2025-07-13Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement LearningICML 2025-Test-Time Adaptation, Vision-Language Navigation, Feedback-Based RL, REINFORCE
2026-07-01FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy ImprovementCoRL 2026-Failure Recovery, Preference Adaptation, Continual Policy Improvement, Action Perturbation
2026-07-08WAM-TTT: Steering World-Action Models by Watching Human Play at Test TimearXiv-World-Action Model, Human Video, Fast Weights, Test-Time Training
2026-07-16RoboTTT: Context Scaling for Robot PoliciesarXiv-Long Context, Fast Weights, VLA, Human Video

4. 🎬 In-Context Learning and Prompting

These methods use demonstrations, videos, language, or sensorimotor trajectories in the deployment-time context to specify robot behavior. The direct in-context imitation subset predicts actions from task examples without task-specific parameter updates; the broader prompting subset studies compatible task interfaces such as human video and multimodal instructions.

Typical mechanism: The deployment-time context sequence; policy parameters are usually unchanged.

Core advantage: A user can specify a task, behavior, or constraint through examples without a task-specific gradient update or online reward optimization.

Main limitation: The policy must be trained to read the prompt modality, and reliable prompting does not guarantee the required low-level behavior is in the policy's learned repertoire.

Papers (20)

DatePaperVenueResourcesTags
2017-03-21One-Shot Imitation LearningNeurIPS 2017-Meta-Imitation, One-Shot Learning, Demonstration Conditioning, Attention
2021-05-13Coarse-to-Fine Imitation Learning: Robot Manipulation from a Single DemonstrationICRA 2021ProjectSingle Demonstration, Human Video, Visual Servoing, Trajectory Replay
2022-02-04BC-Z: Zero-Shot Task Generalization with Robotic Imitation LearningCoRL 2021ProjectHuman Video Prompt, Task Embedding, Multi-Task Imitation, Real Robot
2022-10-06VIMA: General Robot Manipulation with Multimodal PromptsICML 2023ProjectMultimodal Prompt, One-Shot Video Imitation, Cross-Attention, Simulation Benchmark
2023-01-18Human-Timescale Adaptation in an Open-Ended Task SpaceICML 2023-Meta-RL, Attention Memory, Embodied 3D, Demonstration Prompt
2024-03-19Vid2Robot: End-to-End Video-Conditioned Policy Learning with Cross-Attention TransformersRSS 2024ProjectHuman Video Prompt, Cross-Embodiment, Cross-Attention, Paired Data
2024-08-28In-Context Imitation Learning via Next-Token Prediction (ICRT)ICRA 2025-Sensorimotor Prompt, Next-Token Prediction, Transformer, In-Context Learning
2024-11-19Instant Policy: In-Context Imitation Learning via Graph DiffusionICLR 2025-Graph Diffusion, One-Shot Imitation, 3D Representation, Pseudo-Demonstrations
2025-05-27Learning Generalizable Robot Policy with Human Demonstration Video as a PromptICRA 2026-Human Video Prompt, Cross-Embodiment, In-Context Learning
2025-05-27Hierarchical Instruction-aware Embodied Visual TrackingFCS 2026ProjectLanguage-Conditioned Policy, LLM Spatial Reasoning, RAG Goal Correction, Embodied Visual Tracking
2025-05-28Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement LearningICML 2025ProjectOffline In-Context RL, Task Inference, Context Trajectories, Distribution Shift
2025-06-18Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot ManipulationIROS 2025ProjectIn-Context Imitation, Trajectory Aggregation, Uncertainty, Real Robot
2025-12-08See Once, Then Act: VLA Task Learning from One-Shot Video Demonstrations (ViVLA)arXiv-Human Video Prompt, Cross-Embodiment, Latent Action, VLA
2026-04-22AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual TrackingRA-L 2026-Cross-Embodiment, Active Visual Tracking, In-Context Policy, Zero-Shot Adaptation
2026-06-02Instant-Fold: In-Context Imitation Learning for Deformable Object ManipulationCoRL 2026-Deformable Manipulation, Human Demonstration, Flow Matching, 3D Tokens
2026-06-06SynthICL: Scalable In-context Imitation Learning with Synthetic DataCoRL 2026ProjectSynthetic Data, RGB-Only, Flow Matching, One-Shot Imitation
2026-06-29Behavior Prompting Policy: Demonstrations as Prompts for ManipulationCoRL 2026ProjectSensorimotor Prompt, Robot Demonstration, Behavior Prompting
2026-08-18Introducing S1: In-Context Learning for RoboticsSkild AI Blog-In-Context Learning, Video Prompt, Cross-Embodiment, Long-Horizon Manipulation
2026-08-19GEN-1.5: Embodied Foundation Models are One-Shot LearnersGeneralist AI Blog-Physical Prompting, Continuous Pretraining, Long Context, Few-Step Adaptation
2026-08-26RA-VLA: Retrieval-Augmented VLA for Test-Time AdaptationarXiv-VLA, In-Context Imitation, Behavior-Aligned Retrieval, Flow Matching

5. 📈 Test-Time Scaling

These methods allocate additional inference compute to reason, sample, search, predict outcomes, refine actions, or verify candidates before execution. Verification is one mechanism within this broader line rather than a separate category.

Typical mechanism: Reasoning depth, candidate count, search depth, world-model rollouts, refinement steps, verifier calls, or an adaptively allocated inference budget.

Core advantage: Can improve a fixed policy at deployment and expose a measurable compute-performance trade-off without collecting new demonstrations.

Main limitation: More computation helps only when reasoning, candidate diversity, predictive models, or selection scores contain useful signal; latency and hardware cost increase with the budget.

Papers (12)

DatePaperVenueResourcesTags
2024-07-11Robotic Control via Embodied Chain-of-Thought ReasoningCoRL 2024Project / CodeVLA, Embodied Chain-of-Thought, Grounded Reasoning, Intermediate Computation
2025-05-27Hume: Introducing System-2 Thinking in Visual-Language-Action ModelarXivProjectVLA, Value-Guided Sampling, Best-of-N, Cascaded Denoising
2025-06-21RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action ModelsRSS 2025 OOD WorkshopProjectVLA, Test-Time Sampling, Action Verification, Inference Scaling Law
2025-08-17Improving Pre-Trained Vision-Language-Action Policies with Model-Based SearchCoRL 2025 Workshop-VLA, MCTS, Model-Based Search, Inference-Time Planning
2025-09-26VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree SearchICRA 2026ProjectVLA, MCTS, World Model, Value-Guided Search
2025-10-07Verifier-free Test-Time Sampling for Vision-Language-Action ModelsICLR 2026Project / CodeVLA, Verifier-Free Selection, Condition Masking, Best-of-N
2025-10-13RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action ModelarXiv-VLA, Process Reward Model, Candidate Refinement, Test-Time Verification
2025-12-02Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling ApproacharXiv-VLA, Flow Matching, Pseudo-Count Verifier, Best-of-N
2026-02-12Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action AlignmentarXiv-Test-Time Scaling, Action Verification, VLA Alignment, Best-of-N
2026-04-21FASTER: Value-Guided Sampling for Fast RLarXivProject / CodeDiffusion Policy, Value-Guided Sampling, Early Candidate Filtering, Denoising MDP
2026-05-31τ0-WM: A Unified Video-Action World Model for Robotic ManipulationarXivProjectWorld Model, Adaptive Test-Time Compute, Imagined Rollouts, Action Rectification
2026-08-17τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time ComputationarXivProjectHierarchical VLA, World Model, Test-Time Compute, Subtask Verification

6. 🔄 Policy Self-Improvement

These methods use deployment rollouts, failures, corrections, or model-generated experience to improve future behavior across successive data-collection and update rounds.

Typical mechanism: The base policy, residual policy, critic, world model, recovery behavior, or data-generation loop between deployment iterations.

Core advantage: Turns autonomous experience and intervention data into better future behavior, reducing dependence on repeated full demonstrations.

Main limitation: Requires reliable rewards or success labels, safe data collection, practical resets, and controls against policy drift or errors in model-generated data.

Papers (8)

DatePaperVenueResourcesTags
2023-06-20RoboCat: A Self-Improving Generalist Agent for Robotic ManipulationTMLR 2024-Generalist Policy, Self-Generated Data, Iterative Training, Multi-Embodiment
2025-05-28VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language ModelsIROS 2025ProjectEmbodied Visual Tracking, Failure Recovery, Memory-Augmented Self-Reflection, Iterative Recovery Improvement
2025-09-09RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and CorrectionarXivProjectLong-Horizon Manipulation, Human Intervention, Recovery Data, Iterative Imitation Learning
2025-10-30Self-Improving Vision-Language-Action Models with Data Generation via Residual RLICLR 2026ProjectVLA, Residual RL, Deployment-Aligned Data, Policy Distillation
2026-02-12VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelarXivProjectVLA, World Model, Synthetic Rollouts, Iterative Co-Improvement
2026-03-17DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World ModelsarXivProjectVision-Language Planner, Video World Model, Synthetic Rollouts, Reinforcement Fine-Tuning
2026-05-06When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement LearningarXivProjectOffline-to-Online RL, Q-Estimation, Q-Gating, On-Robot Learning
2026-08-21Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-PlanningCoRL 2026Project / CodeFrozen BC Policy, Off-Policy Q-Learning, Failure Rollouts, Value-Guided Planning

🤝 Contributing

Paper additions and corrections are welcome. Please read CONTRIBUTING.md, edit data/papers.json, and run:

python3 scripts/validate_data.py
python3 scripts/build_readme.py --check

You can also use the Add a paper issue template.

🙏 Acknowledgements

The repository structure draws inspiration from community-maintained collections such as Awesome World Models for Robotics, Awesome VLA Post-Training, and Awesome VLA.

📄 License

This repository is released under the MIT License.

Languages

Python

100.0%