News: We add many jev+robot in infrastructure. A curated collection of papers and resources on robot-use agents, tool-based robot control, embodied agent runtimes, and self-evolving robotic systems.
51
25 commits
updated Sep 26, 2026
News: We add many jev+robot demos in Infrastructure and Benchmarks! Please refer these.
A curated list of research on general-purpose AI agents that perceive, program, and operate robots through tools, interfaces, and reusable skills.
Inspired by Phillip Isola's Robot-Use Agents. The emphasis is on how general-purpose intelligence can be connected to different robots and how improvements can spread through models, software interfaces, and reusable capabilities.
General-purpose models operating robots through reusable harnesses, tools, and visual interfaces.
Know Your Body: A Harness for Direct and Self-Improving Robot Control with VLMs — KnowBody; body-grounded robot control and iterative improvement through execution feedback, evaluated on a small real-robot task suite.
Generalizing Manipulation Skills with a Local Coding Agent — A local open-weight VLM writes and executes perception and control code for real-world UR3e manipulation, with documented skills and in-session reuse.
AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation — Agent-ready manipulation through visual grounding, operation tokens, and detect / execute / query interfaces.
Structured World-State Reasoning for Agentic Robotic Search — WORLDS maintains a persistent world-state graph and lets agents request, verify, and revise observations before selecting a target.
Transferring the Intelligence of VLMs to Robotic Control — RoboDawn
RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision
Navi-Agent: Unlocalized Monocular Navigation Agent — Previously reported; coordinate-free spatial memory for closed-loop navigation, progress verification and recovery.
In-Context Robot Learning with VLM Agents — GPT-Policy; in-context adaptation without parameter updates.
WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories — Laboratory robotics.
HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness — Navigation.
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
Guava: An Effective and Universal Harness for Embodied Manipulation
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
Architectures that organize reasoning and robot execution through intermediate plans, coupled reasoning and action modules, or adaptive think/act scheduling. Includes learned-policy building blocks and agent-level systems. For author-described System 1 (e.g., Jev-like system) /System 2 models (VLM-like models), fixed-rate and asynchronous coupling are distinguished from adaptive reasoning. Related task-time memory and recovery methods remain under Planning, Skill Orchestration and Memory.
Systems that turn experience into reusable knowledge, skill programs, improved policies, or validated capability upgrades. Includes human-guided methods and learned-policy precursors where noted.
HarnessPAI: An Evolving Harness for Physical AI — Coding agents refine executable robot harnesses between rollouts using execution feedback, while keeping the underlying action backends frozen.
RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning — Combines reasoning, acting, and code evolution to build reusable robot APIs and harness memory; evaluated in simulation.
AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution — Refines code-based coordination between reasoning agents and frozen VLAs through rollout evidence and persistent revision memory; simulation benchmarks and a real-world quadruped deployment case.
Learning and Transferring Closed-Loop Robot Software — Coding-agent optimization and reuse of closed-loop robot programs in simulation.
Self-Evolving Embodied Agents via Skill-Harness Evolution — SHAPER; frozen-model skill and harness optimization.
Lifelong Robot Library Learning: Bootstrapping Composable and Generalizable Skills for Embodied Control with Language Models
Eureka: Human-Level Reward Design via Coding Large Language Models
Related system-level memory and learning mechanisms: PhyAgentOS, ABot-Claw, and WCM. Execution-time recovery methods are listed under Safe Planning, Verification and Failure Recovery.
Persistent embodied-agent systems that organize robot capabilities, state, resources, execution checks, and feedback across tasks or robots.
RAPID: Robot Agentic Programming from Demonstrations — Converts visual demonstrations into robot programs through task specifications, reusable primitives, and interactive verification.
Coding Agents for Generalized Task and Motion Planning Problems — Coding agents synthesize reusable planning programs for simulated environments; held-out evaluation executes fixed programs without test-time LLM calls.
ManiSkillFormer: Demonstration-Free Compositional Manipulation via Geometric Contracts and Agentic Skill Graph
AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation — VLM selection of executable trajectories through a visual interface and an initialized digital twin.
Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation — Human-in-the-loop interface.
KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation — Learned low-level controller.
AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation
GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use
Code as Policies: Language Model Programs for Embodied Control
ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
LangNav: Language as a Perceptual Representation for Navigation
LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Trust the PRoC3S: Solving Long-Horizon Robotics Problems with LLMs and Constraint Satisfaction
Agents that select and coordinate robot capabilities, track state, verify outcomes, and recover from failures.
Task decomposition, reusable skill orchestration, and task-time memory. Includes learned hierarchical planners and action models where applicable. Cross-cutting reasoning/action coupling and System 1/System 2 designs are listed under Reasoning-Acting and Dual-System Architectures.
NavProbe: Evidence-Grounded Reasoning with Active Memory Retrieval for Zero-Shot Navigation — A hierarchical VLM navigation agent that retrieves visual and geometric evidence to revise subgoals and select parameterized navigation skills.
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models — VLM action plans are refined through action-conditioned world-model imagination, optimization and search; simulation evaluation.
MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration — Natural-language skill orchestration, sensor-event binding, and dialogue-state management on a physical social robot.
Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework
2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation
Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
SayCanPay: Heuristic Planning with Large Language Models using Learnable Domain Knowledge — Offline action-sequence search; simulation evaluation.
Inner Monologue: Embodied Reasoning through Planning with Language Models
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
MOSAIC: Modular Foundation Models for Assistive and Interactive Cooking
Creative Robot Tool Use with Large Language Models — RoboTool; executable plans over parameterized skills.
Methods that assess risks, verify execution, and trigger corrective planning or recovery. Failure-recovery benchmarks are listed under Infrastructure and Benchmarks.
FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation — A planner, VLM monitor, recovery agent, and memory module form a failure-aware supervisory loop for humanoid skills.
CommitFlow: Semantic Commitment Verification and Local Correction for Long-Horizon Robot Manipulation VLA Execution — Frozen-policy execution harness with semantic commitment monitoring, local correction and stage-level verification.
When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence — Evidence-aware choice between autonomous action, additional sensing and human assistance after failure.
Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation — SafeHarness; obstacle-aware route verification, replanning, and contact execution, evaluated in simulation.
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
Safe Task Planning with Long-Term Graph Memory for Embodied Agents
VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies
CoPAL: Corrective Planning of Robot Actions with Large Language Models
REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners — KnowNo; calibrated uncertainty and human clarification.
Task allocation, communication, and organizational structures for teams of robots or embodied agents. ORCH is evaluated in simulation; system-level runtimes are listed under Embodied Agent Operating Systems and Runtimes.
Related learned-policy methods that preserve language interfaces or reduce adaptation to new embodiments; these generally involve robotics training.
Jev+Robot
robo-jev: A 10 Hz Typed-Decision Layer for Physical Robots — A System-One-style decision layer that scores typed robot actions, stop conditions, gripper states, paths, speed, and force for a deterministic executor.
EmbodiedJev: MuJoCo Robot Decision Workbench — Browser-based MuJoCo and Franka Panda workbench supporting Jev, Claude, OpenAI-compatible APIs, and local MiniCPM models with visible observe–decide–execute–feedback loops.
RoboJEV: Two-Stage JEV Control of a Franka Panda in MuJoCo — Two-stage typed decisions for task intent followed by Cartesian motion and gripper commands, evaluated with independent physical success checks.
Jev Robot Control — Reproducible xArm7 MuJoCo comparison of Jev, GPT-6 Astra, and GPT-4.1 mini with archived trajectories, offline verification, and replay.
Robot integration, deployment, latency, runtime reliability, and evaluation of model-plus-interface systems. Includes benchmarks for memory, safety, and recovery, as well as surveys of robot policy verification.
These are essays, research blogs, and evaluations; they are listed separately from papers.
22 commits
3 commits
News: We add many jev+robot in infrastructure. A curated collection of papers and resources on robot-use agents, tool-based robot control, embodied agent runtimes, and self-evolving robotic systems.
51
25 commits
updated Sep 26, 2026
News: We add many jev+robot demos in Infrastructure and Benchmarks! Please refer these.
A curated list of research on general-purpose AI agents that perceive, program, and operate robots through tools, interfaces, and reusable skills.
Inspired by Phillip Isola's Robot-Use Agents. The emphasis is on how general-purpose intelligence can be connected to different robots and how improvements can spread through models, software interfaces, and reusable capabilities.
General-purpose models operating robots through reusable harnesses, tools, and visual interfaces.
Know Your Body: A Harness for Direct and Self-Improving Robot Control with VLMs — KnowBody; body-grounded robot control and iterative improvement through execution feedback, evaluated on a small real-robot task suite.
Generalizing Manipulation Skills with a Local Coding Agent — A local open-weight VLM writes and executes perception and control code for real-world UR3e manipulation, with documented skills and in-session reuse.
AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation — Agent-ready manipulation through visual grounding, operation tokens, and detect / execute / query interfaces.
Structured World-State Reasoning for Agentic Robotic Search — WORLDS maintains a persistent world-state graph and lets agents request, verify, and revise observations before selecting a target.
Transferring the Intelligence of VLMs to Robotic Control — RoboDawn
RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision
Navi-Agent: Unlocalized Monocular Navigation Agent — Previously reported; coordinate-free spatial memory for closed-loop navigation, progress verification and recovery.
In-Context Robot Learning with VLM Agents — GPT-Policy; in-context adaptation without parameter updates.
WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories — Laboratory robotics.
HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness — Navigation.
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
Guava: An Effective and Universal Harness for Embodied Manipulation
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
Architectures that organize reasoning and robot execution through intermediate plans, coupled reasoning and action modules, or adaptive think/act scheduling. Includes learned-policy building blocks and agent-level systems. For author-described System 1 (e.g., Jev-like system) /System 2 models (VLM-like models), fixed-rate and asynchronous coupling are distinguished from adaptive reasoning. Related task-time memory and recovery methods remain under Planning, Skill Orchestration and Memory.
Systems that turn experience into reusable knowledge, skill programs, improved policies, or validated capability upgrades. Includes human-guided methods and learned-policy precursors where noted.
HarnessPAI: An Evolving Harness for Physical AI — Coding agents refine executable robot harnesses between rollouts using execution feedback, while keeping the underlying action backends frozen.
RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning — Combines reasoning, acting, and code evolution to build reusable robot APIs and harness memory; evaluated in simulation.
AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution — Refines code-based coordination between reasoning agents and frozen VLAs through rollout evidence and persistent revision memory; simulation benchmarks and a real-world quadruped deployment case.
Learning and Transferring Closed-Loop Robot Software — Coding-agent optimization and reuse of closed-loop robot programs in simulation.
Self-Evolving Embodied Agents via Skill-Harness Evolution — SHAPER; frozen-model skill and harness optimization.
Lifelong Robot Library Learning: Bootstrapping Composable and Generalizable Skills for Embodied Control with Language Models
Eureka: Human-Level Reward Design via Coding Large Language Models
Related system-level memory and learning mechanisms: PhyAgentOS, ABot-Claw, and WCM. Execution-time recovery methods are listed under Safe Planning, Verification and Failure Recovery.
Persistent embodied-agent systems that organize robot capabilities, state, resources, execution checks, and feedback across tasks or robots.
RAPID: Robot Agentic Programming from Demonstrations — Converts visual demonstrations into robot programs through task specifications, reusable primitives, and interactive verification.
Coding Agents for Generalized Task and Motion Planning Problems — Coding agents synthesize reusable planning programs for simulated environments; held-out evaluation executes fixed programs without test-time LLM calls.
ManiSkillFormer: Demonstration-Free Compositional Manipulation via Geometric Contracts and Agentic Skill Graph
AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation — VLM selection of executable trajectories through a visual interface and an initialized digital twin.
Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation — Human-in-the-loop interface.
KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation — Learned low-level controller.
AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation
GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use
Code as Policies: Language Model Programs for Embodied Control
ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
LangNav: Language as a Perceptual Representation for Navigation
LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Trust the PRoC3S: Solving Long-Horizon Robotics Problems with LLMs and Constraint Satisfaction
Agents that select and coordinate robot capabilities, track state, verify outcomes, and recover from failures.
Task decomposition, reusable skill orchestration, and task-time memory. Includes learned hierarchical planners and action models where applicable. Cross-cutting reasoning/action coupling and System 1/System 2 designs are listed under Reasoning-Acting and Dual-System Architectures.
NavProbe: Evidence-Grounded Reasoning with Active Memory Retrieval for Zero-Shot Navigation — A hierarchical VLM navigation agent that retrieves visual and geometric evidence to revise subgoals and select parameterized navigation skills.
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models — VLM action plans are refined through action-conditioned world-model imagination, optimization and search; simulation evaluation.
MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration — Natural-language skill orchestration, sensor-event binding, and dialogue-state management on a physical social robot.
Bridging Thought and Action: Taming Long-Horizon Instability in Open-Source LLM Agents with a MetaTool-Enhanced ROS Framework
2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation
Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
SayCanPay: Heuristic Planning with Large Language Models using Learnable Domain Knowledge — Offline action-sequence search; simulation evaluation.
Inner Monologue: Embodied Reasoning through Planning with Language Models
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
MOSAIC: Modular Foundation Models for Assistive and Interactive Cooking
Creative Robot Tool Use with Large Language Models — RoboTool; executable plans over parameterized skills.
Methods that assess risks, verify execution, and trigger corrective planning or recovery. Failure-recovery benchmarks are listed under Infrastructure and Benchmarks.
FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation — A planner, VLM monitor, recovery agent, and memory module form a failure-aware supervisory loop for humanoid skills.
CommitFlow: Semantic Commitment Verification and Local Correction for Long-Horizon Robot Manipulation VLA Execution — Frozen-policy execution harness with semantic commitment monitoring, local correction and stage-level verification.
When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence — Evidence-aware choice between autonomous action, additional sensing and human assistance after failure.
Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation — SafeHarness; obstacle-aware route verification, replanning, and contact execution, evaluated in simulation.
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
Safe Task Planning with Long-Term Graph Memory for Embodied Agents
VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies
CoPAL: Corrective Planning of Robot Actions with Large Language Models
REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners — KnowNo; calibrated uncertainty and human clarification.
Task allocation, communication, and organizational structures for teams of robots or embodied agents. ORCH is evaluated in simulation; system-level runtimes are listed under Embodied Agent Operating Systems and Runtimes.
Related learned-policy methods that preserve language interfaces or reduce adaptation to new embodiments; these generally involve robotics training.
Jev+Robot
robo-jev: A 10 Hz Typed-Decision Layer for Physical Robots — A System-One-style decision layer that scores typed robot actions, stop conditions, gripper states, paths, speed, and force for a deterministic executor.
EmbodiedJev: MuJoCo Robot Decision Workbench — Browser-based MuJoCo and Franka Panda workbench supporting Jev, Claude, OpenAI-compatible APIs, and local MiniCPM models with visible observe–decide–execute–feedback loops.
RoboJEV: Two-Stage JEV Control of a Franka Panda in MuJoCo — Two-stage typed decisions for task intent followed by Cartesian motion and gripper commands, evaluated with independent physical success checks.
Jev Robot Control — Reproducible xArm7 MuJoCo comparison of Jev, GPT-6 Astra, and GPT-4.1 mini with archived trajectories, offline verification, and replay.
Robot integration, deployment, latency, runtime reliability, and evaluation of model-plus-interface systems. Includes benchmarks for memory, safety, and recovery, as well as surveys of robot policy verification.
These are essays, research blogs, and evaluations; they are listed separately from papers.
22 commits
3 commits