rhos-ai/IPR-1

Official Repo of IPR-1 Project. https://www.rhos.ai/research/ipr-1

JavaScript

32

6 commits

updated Jun 29, 2026

See the code

README

IPR-1: Interactive Physical Reasoner

This repository contains the official implementation and project page for IPR-1: Interactive Physical Reasoner.

Paper: arXiv:2511.15407
Project Page: https://mybearyZhang.github.io/ipr-1/

Code Layout

The open-source training code is organized into the three IPR components:

  • src/ipr/latent_action/: PhysCode latent action tokenizer. This implements the stage-1 VQ discrete latent action model described in the public IPR paper: adjacent video frames, optional optical flow, and optional action-language embeddings are encoded into discrete physical action codes and trained with next-frame reconstruction plus vector-quantization losses.
  • src/ipr/latent_world_model/: latent-world-model training code. vjepa2_core/ is a minimized V-JEPA2-based vendor copy containing only the modules needed by the game latent-world-model training, data loading, latent evaluation, and frame-decoder utilities.
  • src/ipr/actor_rl/: actor RL code based on GRPO, VLM policy rollout, optional DQN teacher, optional V-JEPA teacher, environment management, and evaluation.

The project page remains in the repository root (index.html, css/, js/, and assets/).

Installation

Create a Python environment with PyTorch appropriate for your CUDA version, then install the repository:

pip install -e .

Some environments need extra packages depending on which component is used:

  • retro / emulator integration for game rollouts.
  • flash-attn if you want flash attention for Qwen-VL.
  • A CUDA-compatible deepspeed build for distributed GRPO.

Required External Artifacts

Large files are intentionally ignored by git. Put them in local paths and pass those paths through config or command-line flags:

  • VLM base model: models/Qwen3-VL-8B-Instruct/ or another transformers-compatible VLM.
  • Gameplay data: data/gameplay/.
  • Prompt data and game list: prompts_nl/ and selected_games.txt.
  • Latent action checkpoints: artifacts/latent_action/latest.pt.
  • V-JEPA/latent-world-model checkpoints and reward heads: artifacts/latent_world_model/.
  • Optional DQN teacher checkpoints: artifacts/dqn_models/.

See docs/OPEN_SOURCE_NOTES.md for the recommended local layout and files that must stay out of git.

Training

Train the latent action tokenizer:

DATA_ROOT=data/gameplay OUTPUT_DIR=artifacts/latent_action bash scripts/train_latent_action.sh

Encode PhysCode ids after training:

python -m ipr.latent_action.encode \
  --checkpoint artifacts/latent_action/latest.pt \
  --data_root data/gameplay \
  --output artifacts/latent_action/physcodes.pt

Train the V-JEPA2-based latent world model:

CONFIG=configs/latent_world_model/vjepa_game.yaml DEVICES=cuda:0 bash scripts/train_latent_world_model.sh

Train the actor with GRPO:

MODEL_PATH=models/Qwen3-VL-8B-Instruct \
PROMPTS_DIR=prompts_nl \
BACK_FILE=selected_games.txt \
bash scripts/train_actor_rl.sh

For multi-GPU GRPO, set NPROC_PER_NODE:

NPROC_PER_NODE=8 bash scripts/train_actor_rl.sh

Configuration

Templates live under configs/:

  • configs/latent_action/physcode.yaml
  • configs/latent_world_model/vjepa_game.yaml
  • configs/actor_rl/grpo.yaml

The actor script currently accepts command-line arguments directly; the YAML file is kept as a documented reference for expected paths and defaults.

Git Hygiene

The .gitignore excludes model weights, datasets, checkpoints, logs, caches, generated evaluations, and prompt corpora. Do not commit ROMs, model caches, replay buffers, or generated .pt/.safetensors/.npy artifacts.

License

Add the final project license before public release. The vendored V-JEPA2-derived code was copied from a Meta-licensed codebase; preserve upstream license notices and include the corresponding license text before publishing.

Significant stargazers

stain lu

99 followers · starred Jun 2026

rhos-ai/IPR-1

Official Repo of IPR-1 Project. https://www.rhos.ai/research/ipr-1

JavaScript

32

6 commits

updated Jun 29, 2026

See the code

README

IPR-1: Interactive Physical Reasoner

This repository contains the official implementation and project page for IPR-1: Interactive Physical Reasoner.

Paper: arXiv:2511.15407
Project Page: https://mybearyZhang.github.io/ipr-1/

Code Layout

The open-source training code is organized into the three IPR components:

  • src/ipr/latent_action/: PhysCode latent action tokenizer. This implements the stage-1 VQ discrete latent action model described in the public IPR paper: adjacent video frames, optional optical flow, and optional action-language embeddings are encoded into discrete physical action codes and trained with next-frame reconstruction plus vector-quantization losses.
  • src/ipr/latent_world_model/: latent-world-model training code. vjepa2_core/ is a minimized V-JEPA2-based vendor copy containing only the modules needed by the game latent-world-model training, data loading, latent evaluation, and frame-decoder utilities.
  • src/ipr/actor_rl/: actor RL code based on GRPO, VLM policy rollout, optional DQN teacher, optional V-JEPA teacher, environment management, and evaluation.

The project page remains in the repository root (index.html, css/, js/, and assets/).

Installation

Create a Python environment with PyTorch appropriate for your CUDA version, then install the repository:

pip install -e .

Some environments need extra packages depending on which component is used:

  • retro / emulator integration for game rollouts.
  • flash-attn if you want flash attention for Qwen-VL.
  • A CUDA-compatible deepspeed build for distributed GRPO.

Required External Artifacts

Large files are intentionally ignored by git. Put them in local paths and pass those paths through config or command-line flags:

  • VLM base model: models/Qwen3-VL-8B-Instruct/ or another transformers-compatible VLM.
  • Gameplay data: data/gameplay/.
  • Prompt data and game list: prompts_nl/ and selected_games.txt.
  • Latent action checkpoints: artifacts/latent_action/latest.pt.
  • V-JEPA/latent-world-model checkpoints and reward heads: artifacts/latent_world_model/.
  • Optional DQN teacher checkpoints: artifacts/dqn_models/.

See docs/OPEN_SOURCE_NOTES.md for the recommended local layout and files that must stay out of git.

Training

Train the latent action tokenizer:

DATA_ROOT=data/gameplay OUTPUT_DIR=artifacts/latent_action bash scripts/train_latent_action.sh

Encode PhysCode ids after training:

python -m ipr.latent_action.encode \
  --checkpoint artifacts/latent_action/latest.pt \
  --data_root data/gameplay \
  --output artifacts/latent_action/physcodes.pt

Train the V-JEPA2-based latent world model:

CONFIG=configs/latent_world_model/vjepa_game.yaml DEVICES=cuda:0 bash scripts/train_latent_world_model.sh

Train the actor with GRPO:

MODEL_PATH=models/Qwen3-VL-8B-Instruct \
PROMPTS_DIR=prompts_nl \
BACK_FILE=selected_games.txt \
bash scripts/train_actor_rl.sh

For multi-GPU GRPO, set NPROC_PER_NODE:

NPROC_PER_NODE=8 bash scripts/train_actor_rl.sh

Configuration

Templates live under configs/:

  • configs/latent_action/physcode.yaml
  • configs/latent_world_model/vjepa_game.yaml
  • configs/actor_rl/grpo.yaml

The actor script currently accepts command-line arguments directly; the YAML file is kept as a documented reference for expected paths and defaults.

Git Hygiene

The .gitignore excludes model weights, datasets, checkpoints, logs, caches, generated evaluations, and prompt corpora. Do not commit ROMs, model caches, replay buffers, or generated .pt/.safetensors/.npy artifacts.

License

Add the final project license before public release. The vendored V-JEPA2-derived code was copied from a Meta-licensed codebase; preserve upstream license notices and include the corresponding license text before publishing.

Significant stargazers

stain lu

99 followers · starred Jun 2026

Languages

JavaScript

50.4%

Python

37.6%

HTML

7.0%

CSS

4.8%