microsoft/Orchard-Agentic

Orchard-Agentic is a collection of open-source work on agentic modeling

Python

535

10 commits

updated Oct 5, 2026

See the code

README

Orchard-Agentic

arXiv Research data on Hugging Face License: MIT

Orchard-Agentic is a collection of open-source work on agentic modeling. It brings together research projects, shared infrastructure, datasets, training recipes, and evaluation artifacts. The collection includes the foundational paper Orchard: An Open-Source Agentic Modeling Framework and subsequent projects such as OpenWebRL and OpenForge RL.

Orchard Env, Orchard-SWE, Orchard-GUI, and Orchard-Claw are established names introduced by the Orchard research and are retained here. Throughout this repository, Orchard-Agentic refers to the broader collection, while Orchard refers to the original paper and framework.

Across these projects, Orchard Env provides a stable service rather than a piece of any one training stack. This lets research reuse the same substrate for trajectory distillation, on-policy RL rollouts, and evaluation, so datasets, training recipes, and evaluation protocols stay portable across harnesses, domains, and projects instead of being rebuilt for each new study.

Orchard framework overview

LayerWhat it is
RecipesThe research. Open SFT + RL recipes explored on top of the foundation — Orchard-SWE, Orchard-GUI, and Orchard-Claw — plus follow-on work such as OpenWebRL and OpenForge RL. Everything built on Orchard is indexed in projects/.
Orchard Env (orchard_env/)The foundation. A Kubernetes-native sandbox service + Python SDK that spins up thousands of isolated containers on demand and drives multi-turn agent ↔ sandbox interaction (exec, file I/O, git patches) over HTTP.
Orchard Eval (orchard_eval/)Evaluation. Runs any agent harness — codex, pi, mini-swe-agent, … — against SWE-bench Verified on Orchard Env sandboxes, graded by the official SWE-bench harness. See Evaluation.
Trainer (trainer/slime/)RL training stack — a slime fork vendored as a git submodule (MSR-Orchard/slime), with Orchard rollout code under examples/orchard_swe/ and examples/orchard_gui/. See Training.

News

  • [2026-07] 🎉 We are excited to release OpenForge RL (code · details), which extends Orchard to train agents inside their real deployment harnesses — ZeroClaw, OpenClaw, Codex — instead of the simplified reimplementations open training stacks usually require, removing the train–deploy mismatch. A lightweight proxy records the harness's own inference calls and reconstructs them into samples for any RL codebase (e.g. veRL), while Orchard Env launches each rollout as a remote container, so any harness pairs with any environment. This is exactly what Orchard Env is built for: every sandbox ships with popular agent harnesses — codex, claude, pi, opencode, hermes — already on PATH, so switching the harness you train or evaluate against is a change of command, not a change of image. OpenForge-GUI (8B) reaches 37.7 on OSWorld-Verified and 72.3 on WebVoyager; OpenForge-Claw (30B-A3B) reaches 33.7 on QwenClawBench and 28.1 on MCPAtlas.

  • [2026-06] 🎉 We are excited to release OpenWebRL (code · details), which extends Orchard-GUI into a full online multi-turn RL study on live websites — covering supervised initialization, multimodal context management, trajectory-level success judging, and multi-turn policy optimization. It builds a fault-tolerant live-browser environment on Orchard Env with navigation retries, timeout handling, and structured failure attribution, so unstable website behavior stays separable from model behavior at training scale. OpenWebRL-4B reaches 67.0% on Online-Mind2Web and 64.0% on DeepShop from only 0.4K initialization trajectories and 2.2K open-ended RL tasks — a new open-source state of the art on live-web benchmarks, competitive with OpenAI CUA and Gemini CUA.

  • [2026-05] 📄 The Orchard paper is on arXiv, introducing Orchard Env and the Orchard-SWE, Orchard-GUI, and Orchard-Claw research recipes.

Projects built on Orchard

projects/ indexes every project that builds on Orchard Env — the three recipes below plus follow-on work with its own papers and repositories.

ProjectDomainPaperCode
OpenForge RL — training inside the real deployment harnessCLI/tool agents · computer use · browserarXiv:2607.21557MSR-Orchard/OpenForge-RL
OpenWebRL — online multi-turn RL on live websitesBrowser usearXiv:2606.02031OpenWebRL/OpenWebRL
Orchard-SWE · Orchard-GUI · Orchard-ClawSWE · browser · CLI/tool agentsarXiv:2605.15040trainer/slime/

Built something on Orchard Env? See Adding your project.

Recipes

Three studies from the Orchard paper — different domains, harnesses, and reward mechanisms, one environment service underneath.

RecipeBackboneTraining dataKey techniquesHeadline result
Orchard-SWEQwen3.5-35B-A3B107K distilled trajectoriesCredit-assignment SFT · Balanced Adaptive Rollout · on-policy distillation · rubric-based process reward · value-model reranking73.0% SWE-bench Verified
Orchard-GUIQwen3-VL-4B-Thinking0.4k SFT + 2.2k RL tasksDistillation, then online RL on live websites68.4% avg — 74.1 / 67.0 / 64.0
Orchard-ClawQwen3-30B-A3B-Thinking0.2k synthetic tasksOpus-synthesized tasks · training across two harnesses59.6% pass@3, 73.9% under ZeroClaw

Orchard-SWE and Orchard-GUI performance vs. total parameters

Orchard agents versus total parameter count. Both approach or match systems 10–30× larger.

The common thread is generalization, not just peak score. Orchard-SWE keeps 51.0 on SWE-bench Multilingual (vs 28.7 for OpenSWE-32B) and still works under a harness never seen in training — 45.0 on SWE-bench Verified and 20.1 on Terminal-Bench 2.0 with Kimi-CLI, where OpenSWE-32B collapses to 3.6 and 0.0. Orchard-Claw gains the most of any model compared when swapped onto a stronger harness after training (+9.3 pass³ / +14.3 pass@3 under ZeroClaw). Orchard-GUI beats its own 235B teacher on two orders of magnitude fewer training tasks. Training against a harness-agnostic environment is what makes these transfers possible.

Orchard Env — the foundation

A thin, Kubernetes-native environment service exposing generic primitives — sandbox lifecycle, command execution, file I/O, network policy, and a REST API — with no assumptions about the harness, trainer, inference backend, or task domain sitting above it.

  • REST API for sandbox lifecycle (create / exec / files / patch / delete)
  • Sync and async Python SDK with auto-cleanup, retries, and context-manager ergonomics
  • In-pod agent for low-latency exec over Pod IP (bypasses the K8s API server hot path)
  • Any base image — the agent is injected by an init container bundling its own self-contained Python interpreter, so user images need no Python
  • Any harness — codex, claude, pi, opencode, and hermes preinstalled on PATH in every sandbox, with no install step or network access needed inside it (details)
  • Multi-replica orchestrator with Redis-backed state and distributed locks
  • Network isolation via Calico NetworkPolicy (deny-egress by default), per-sandbox CPU / memory / timeout limits, TTL cleanup, and API-key auth
MetricOrchard EnvReference
Average command-execution latency0.28 sSkyPilot Code Sandbox 0.284 s · E2B 0.747 s (2.7× slower) · Modal 2.046 s (7.3× slower)
1,000 sandboxes launched in parallel100% success, 26 s end-to-end, ~154 commands/s (11.75 s average create)—
Cost for 128 sandboxes × 240 h (2 vCPU / 8 GiB)$3,362 on-demand (0.47×) · $673 on spot (0.10×, ~10× cheaper)Managed sandbox services $7,078–$10,305

Swapping Docker for Orchard Env on Terminal-Bench 2.0 causes no regression (GPT-4.1 34.1 → 35.1, MiniMax-M2.5 52.6 → 54.4, Qwen3-8B-Thinking 7.0 → 8.8), so the environment is a drop-in substrate for both training and evaluation.

Quick start

pip install -e "orchard_env[dev]"

export SANDBOX_BASE_URL="http://your-orchestrator-host"
export SANDBOX_API_KEY="your-api-key"
from orchard_env import SandboxClient

with SandboxClient() as client:  # reads SANDBOX_BASE_URL / SANDBOX_API_KEY
    with client.create_sandbox("python:3.11-slim") as sandbox:
        result = sandbox.exec("echo 'Hello, Orchard!'")
        print(result.stdout)

No orchestrator yet? Four scripts stand one up on Azure AKS — provision, build and push images, deploy, smoke-test — in about 20 minutes: orchard_env/README.md for the short path, orchard_env/docs/deployment.md for every configuration knob, non-Azure clusters, and cost estimates.

Evaluation

orchard_eval/ runs a benchmark against Orchard Env — one sandbox per task instance, graded by the official SWE-bench harness. The benchmark, the environment, and the agent harness are separate layers, so switching the agent you are measuring is a one-line change rather than a new image or a new runner:

# Both in one command: orchard-eval depends on orchard-env, which is a path in
# this repository rather than a package on PyPI.
pip install -e orchard_env -e "orchard_eval[swebench]"

export SANDBOX_BASE_URL="http://your-orchestrator-host"
export SANDBOX_API_KEY="your-api-key"
export OPENAI_API_KEY="sk-..."

orchard-eval run -c orchard_eval/configs/codex.yaml --limit 5   # smoke test
orchard-eval run -c orchard_eval/configs/codex.yaml             # SWE-bench Verified
orchard-eval run -c orchard_eval/configs/pi.yaml
orchard-eval run -c orchard_eval/configs/mini-swe-agent.yaml

This is the same "any harness" property the environment layer is built for. codex, claude, pi, opencode, hermes and mini are already on PATH in every sandbox, so evaluating any of them costs one exec — nothing is installed at run time and the whole agent loop stays on the pod. Every harness emits a full turn-by-turn trajectory (normalized, plus the raw event stream) for distillation and RL, runs are resumable, results land in SWE-bench's own prediction format for independent re-grading, and infrastructure failures are separated from agent failures so a flaky cluster never quietly depresses a score. See orchard_eval/README.md.

Training

The trainer is a git submodule, so a plain git clone leaves trainer/slime/ empty:

git clone --recursive https://github.com/microsoft/Orchard-Agentic.git

# already cloned without --recursive?
git submodule update --init trainer/slime

Rollout code, reward functions, and launch scripts for each recipe live in the fork:

RecipeDirectory
Orchard-SWEexamples/orchard_swe/
Orchard-GUIexamples/orchard_gui/

Hardware. The two layers have very different requirements, and they scale independently — run them on separate node pools.

  • Orchard Env is CPU-only. Sandbox nodes need no GPUs. The reference AKS deployment uses Standard_D8as_v5, and the paper's 128-sandbox fleet ran on 17 × Standard_D16ads_v5 (16 vCPU / 64 GiB), with the sandbox pool on spot instances.
  • The trainer needs GPUs. examples/orchard_swe/b200_config/ ships 1-node and 2-node Kubernetes job specs for Qwen3.5-35B-A3B, each node requesting 8 × B200, 104 vCPU, 2880 GiB RAM, and 8 RDMA devices. Upstream slime officially supports H100/H200 and B200; the training stack is Megatron-LM with SGLang rollout under Ray.

Beyond hello-world — async client, file I/O, git patches, PTY sessions, and per-sandbox resource limits: SDK reference · REST API · architecture · project overview.

Exploring a new recipe on Orchard Env? The REST API (docs/api.md) is the contract and the SDK is a thin client over it, so a project in any language can depend on the same substrate. Open a PR to add it to projects/.

Paper & dataset

The foundation and the three recipes above are described in Orchard: An Open-Source Agentic Modeling Framework (Peng et al., arXiv:2605.15040). OpenWebRL and OpenForge RL are separate papers that build on the same environment layer — see projects/.

  • 📄 Paper: arXiv:2605.15040
  • 🤗 Research data: Hugging Face dataset — one repository ships two parallel subsets, both produced inside the same Orchard Env sandbox infrastructure:
    • swe config — 107,185 multi-turn SWE rollouts over 19,287 unique task instances across 2,788 repositories, with verified resolve labels (74,649 resolved · 32,536 unresolved) and an average of 47.5 turns per trajectory.
    • gui config — 3,070 judge-verified successful per-step rollouts from a web-browsing GUI agent across 409 WebVoyager-style tasks, each with a rendered screenshot (multimodal).

Roadmap

Stateful sandboxes — pause, resume, and branching

The environment is currently linear: a sandbox is created, driven for N turns, and destroyed. Every rollout re-executes its whole prefix, and a trajectory yields a single outcome reward that has to be spread across all of its turns — in our SWE data, an average of 47.5 turns per trajectory. Attributing that one scalar to the turn that actually mattered is the central credit-assignment problem, and the paper attacks it indirectly, with retrospective value estimation over completed trajectories.

Snapshotting makes that measurable directly instead. We are adding:

  • Pause / resume — checkpoint a sandbox's full state (filesystem, processes, environment) at any turn and restore it later, so a rollout can be suspended and continued instead of restarted.
  • Branching — fork k independent continuations from the same snapshot at turn t. Rolling forward repeatedly from a single state gives a Monte-Carlo estimate of that state's value, which turns per-turn credit assignment into something measured rather than inferred, and makes counterfactuals ("what if the agent had not run that command?") directly observable.
  • Prefix sharing — because a branched prefix is executed once instead of once per rollout, tree-structured search and per-turn advantage estimation get substantially cheaper than the flat-rollout baseline.

This is the environment primitive that the next round of credit-assignment research needs, and it belongs in the environment layer rather than in any one trainer.

Contributing

Contributions are welcome — please open an issue or pull request. Development setup, linting, and the test matrix are documented in orchard_env/README.md and orchard_env/tests/README.md.

This project has adopted the Microsoft Open Source Code of Conduct; see CODE_OF_CONDUCT.md.

Security

Please do not report security vulnerabilities through public GitHub issues. See SECURITY.md for how to report them.

Citation

If you use Orchard or Orchard Env in your research, please cite:

@article{peng2026orchard,
  title={Orchard: An Open-Source Agentic Modeling Framework},
  author={Peng, Baolin and Yao, Wenlin and Wu, Qianhui and Cheng, Hao and
          Yu, Xiao and Yang, Rui and Ge, Tao and Sordoni, Alessandro and
          Yuan, Xingdi and Shen, Yelong and He, Pengcheng and Zhang, Tong and
          Yu, Zhou and Gao, Jianfeng},
  journal={arXiv preprint arXiv:2605.15040},
  year={2026},
  url={https://arxiv.org/abs/2605.15040}
}

License

MIT © Microsoft Corporation.

microsoft/Orchard-Agentic

Orchard-Agentic is a collection of open-source work on agentic modeling

Python

535

10 commits

updated Oct 5, 2026

See the code

README

Orchard-Agentic

arXiv Research data on Hugging Face License: MIT

Orchard-Agentic is a collection of open-source work on agentic modeling. It brings together research projects, shared infrastructure, datasets, training recipes, and evaluation artifacts. The collection includes the foundational paper Orchard: An Open-Source Agentic Modeling Framework and subsequent projects such as OpenWebRL and OpenForge RL.

Orchard Env, Orchard-SWE, Orchard-GUI, and Orchard-Claw are established names introduced by the Orchard research and are retained here. Throughout this repository, Orchard-Agentic refers to the broader collection, while Orchard refers to the original paper and framework.

Across these projects, Orchard Env provides a stable service rather than a piece of any one training stack. This lets research reuse the same substrate for trajectory distillation, on-policy RL rollouts, and evaluation, so datasets, training recipes, and evaluation protocols stay portable across harnesses, domains, and projects instead of being rebuilt for each new study.

Orchard framework overview

LayerWhat it is
RecipesThe research. Open SFT + RL recipes explored on top of the foundation — Orchard-SWE, Orchard-GUI, and Orchard-Claw — plus follow-on work such as OpenWebRL and OpenForge RL. Everything built on Orchard is indexed in projects/.
Orchard Env (orchard_env/)The foundation. A Kubernetes-native sandbox service + Python SDK that spins up thousands of isolated containers on demand and drives multi-turn agent ↔ sandbox interaction (exec, file I/O, git patches) over HTTP.
Orchard Eval (orchard_eval/)Evaluation. Runs any agent harness — codex, pi, mini-swe-agent, … — against SWE-bench Verified on Orchard Env sandboxes, graded by the official SWE-bench harness. See Evaluation.
Trainer (trainer/slime/)RL training stack — a slime fork vendored as a git submodule (MSR-Orchard/slime), with Orchard rollout code under examples/orchard_swe/ and examples/orchard_gui/. See Training.

News

  • [2026-07] 🎉 We are excited to release OpenForge RL (code · details), which extends Orchard to train agents inside their real deployment harnesses — ZeroClaw, OpenClaw, Codex — instead of the simplified reimplementations open training stacks usually require, removing the train–deploy mismatch. A lightweight proxy records the harness's own inference calls and reconstructs them into samples for any RL codebase (e.g. veRL), while Orchard Env launches each rollout as a remote container, so any harness pairs with any environment. This is exactly what Orchard Env is built for: every sandbox ships with popular agent harnesses — codex, claude, pi, opencode, hermes — already on PATH, so switching the harness you train or evaluate against is a change of command, not a change of image. OpenForge-GUI (8B) reaches 37.7 on OSWorld-Verified and 72.3 on WebVoyager; OpenForge-Claw (30B-A3B) reaches 33.7 on QwenClawBench and 28.1 on MCPAtlas.

  • [2026-06] 🎉 We are excited to release OpenWebRL (code · details), which extends Orchard-GUI into a full online multi-turn RL study on live websites — covering supervised initialization, multimodal context management, trajectory-level success judging, and multi-turn policy optimization. It builds a fault-tolerant live-browser environment on Orchard Env with navigation retries, timeout handling, and structured failure attribution, so unstable website behavior stays separable from model behavior at training scale. OpenWebRL-4B reaches 67.0% on Online-Mind2Web and 64.0% on DeepShop from only 0.4K initialization trajectories and 2.2K open-ended RL tasks — a new open-source state of the art on live-web benchmarks, competitive with OpenAI CUA and Gemini CUA.

  • [2026-05] 📄 The Orchard paper is on arXiv, introducing Orchard Env and the Orchard-SWE, Orchard-GUI, and Orchard-Claw research recipes.

Projects built on Orchard

projects/ indexes every project that builds on Orchard Env — the three recipes below plus follow-on work with its own papers and repositories.

ProjectDomainPaperCode
OpenForge RL — training inside the real deployment harnessCLI/tool agents · computer use · browserarXiv:2607.21557MSR-Orchard/OpenForge-RL
OpenWebRL — online multi-turn RL on live websitesBrowser usearXiv:2606.02031OpenWebRL/OpenWebRL
Orchard-SWE · Orchard-GUI · Orchard-ClawSWE · browser · CLI/tool agentsarXiv:2605.15040trainer/slime/

Built something on Orchard Env? See Adding your project.

Recipes

Three studies from the Orchard paper — different domains, harnesses, and reward mechanisms, one environment service underneath.

RecipeBackboneTraining dataKey techniquesHeadline result
Orchard-SWEQwen3.5-35B-A3B107K distilled trajectoriesCredit-assignment SFT · Balanced Adaptive Rollout · on-policy distillation · rubric-based process reward · value-model reranking73.0% SWE-bench Verified
Orchard-GUIQwen3-VL-4B-Thinking0.4k SFT + 2.2k RL tasksDistillation, then online RL on live websites68.4% avg — 74.1 / 67.0 / 64.0
Orchard-ClawQwen3-30B-A3B-Thinking0.2k synthetic tasksOpus-synthesized tasks · training across two harnesses59.6% pass@3, 73.9% under ZeroClaw

Orchard-SWE and Orchard-GUI performance vs. total parameters

Orchard agents versus total parameter count. Both approach or match systems 10–30× larger.

The common thread is generalization, not just peak score. Orchard-SWE keeps 51.0 on SWE-bench Multilingual (vs 28.7 for OpenSWE-32B) and still works under a harness never seen in training — 45.0 on SWE-bench Verified and 20.1 on Terminal-Bench 2.0 with Kimi-CLI, where OpenSWE-32B collapses to 3.6 and 0.0. Orchard-Claw gains the most of any model compared when swapped onto a stronger harness after training (+9.3 pass³ / +14.3 pass@3 under ZeroClaw). Orchard-GUI beats its own 235B teacher on two orders of magnitude fewer training tasks. Training against a harness-agnostic environment is what makes these transfers possible.

Orchard Env — the foundation

A thin, Kubernetes-native environment service exposing generic primitives — sandbox lifecycle, command execution, file I/O, network policy, and a REST API — with no assumptions about the harness, trainer, inference backend, or task domain sitting above it.

  • REST API for sandbox lifecycle (create / exec / files / patch / delete)
  • Sync and async Python SDK with auto-cleanup, retries, and context-manager ergonomics
  • In-pod agent for low-latency exec over Pod IP (bypasses the K8s API server hot path)
  • Any base image — the agent is injected by an init container bundling its own self-contained Python interpreter, so user images need no Python
  • Any harness — codex, claude, pi, opencode, and hermes preinstalled on PATH in every sandbox, with no install step or network access needed inside it (details)
  • Multi-replica orchestrator with Redis-backed state and distributed locks
  • Network isolation via Calico NetworkPolicy (deny-egress by default), per-sandbox CPU / memory / timeout limits, TTL cleanup, and API-key auth
MetricOrchard EnvReference
Average command-execution latency0.28 sSkyPilot Code Sandbox 0.284 s · E2B 0.747 s (2.7× slower) · Modal 2.046 s (7.3× slower)
1,000 sandboxes launched in parallel100% success, 26 s end-to-end, ~154 commands/s (11.75 s average create)—
Cost for 128 sandboxes × 240 h (2 vCPU / 8 GiB)$3,362 on-demand (0.47×) · $673 on spot (0.10×, ~10× cheaper)Managed sandbox services $7,078–$10,305

Swapping Docker for Orchard Env on Terminal-Bench 2.0 causes no regression (GPT-4.1 34.1 → 35.1, MiniMax-M2.5 52.6 → 54.4, Qwen3-8B-Thinking 7.0 → 8.8), so the environment is a drop-in substrate for both training and evaluation.

Quick start

pip install -e "orchard_env[dev]"

export SANDBOX_BASE_URL="http://your-orchestrator-host"
export SANDBOX_API_KEY="your-api-key"
from orchard_env import SandboxClient

with SandboxClient() as client:  # reads SANDBOX_BASE_URL / SANDBOX_API_KEY
    with client.create_sandbox("python:3.11-slim") as sandbox:
        result = sandbox.exec("echo 'Hello, Orchard!'")
        print(result.stdout)

No orchestrator yet? Four scripts stand one up on Azure AKS — provision, build and push images, deploy, smoke-test — in about 20 minutes: orchard_env/README.md for the short path, orchard_env/docs/deployment.md for every configuration knob, non-Azure clusters, and cost estimates.

Evaluation

orchard_eval/ runs a benchmark against Orchard Env — one sandbox per task instance, graded by the official SWE-bench harness. The benchmark, the environment, and the agent harness are separate layers, so switching the agent you are measuring is a one-line change rather than a new image or a new runner:

# Both in one command: orchard-eval depends on orchard-env, which is a path in
# this repository rather than a package on PyPI.
pip install -e orchard_env -e "orchard_eval[swebench]"

export SANDBOX_BASE_URL="http://your-orchestrator-host"
export SANDBOX_API_KEY="your-api-key"
export OPENAI_API_KEY="sk-..."

orchard-eval run -c orchard_eval/configs/codex.yaml --limit 5   # smoke test
orchard-eval run -c orchard_eval/configs/codex.yaml             # SWE-bench Verified
orchard-eval run -c orchard_eval/configs/pi.yaml
orchard-eval run -c orchard_eval/configs/mini-swe-agent.yaml

This is the same "any harness" property the environment layer is built for. codex, claude, pi, opencode, hermes and mini are already on PATH in every sandbox, so evaluating any of them costs one exec — nothing is installed at run time and the whole agent loop stays on the pod. Every harness emits a full turn-by-turn trajectory (normalized, plus the raw event stream) for distillation and RL, runs are resumable, results land in SWE-bench's own prediction format for independent re-grading, and infrastructure failures are separated from agent failures so a flaky cluster never quietly depresses a score. See orchard_eval/README.md.

Training

The trainer is a git submodule, so a plain git clone leaves trainer/slime/ empty:

git clone --recursive https://github.com/microsoft/Orchard-Agentic.git

# already cloned without --recursive?
git submodule update --init trainer/slime

Rollout code, reward functions, and launch scripts for each recipe live in the fork:

RecipeDirectory
Orchard-SWEexamples/orchard_swe/
Orchard-GUIexamples/orchard_gui/

Hardware. The two layers have very different requirements, and they scale independently — run them on separate node pools.

  • Orchard Env is CPU-only. Sandbox nodes need no GPUs. The reference AKS deployment uses Standard_D8as_v5, and the paper's 128-sandbox fleet ran on 17 × Standard_D16ads_v5 (16 vCPU / 64 GiB), with the sandbox pool on spot instances.
  • The trainer needs GPUs. examples/orchard_swe/b200_config/ ships 1-node and 2-node Kubernetes job specs for Qwen3.5-35B-A3B, each node requesting 8 × B200, 104 vCPU, 2880 GiB RAM, and 8 RDMA devices. Upstream slime officially supports H100/H200 and B200; the training stack is Megatron-LM with SGLang rollout under Ray.

Beyond hello-world — async client, file I/O, git patches, PTY sessions, and per-sandbox resource limits: SDK reference · REST API · architecture · project overview.

Exploring a new recipe on Orchard Env? The REST API (docs/api.md) is the contract and the SDK is a thin client over it, so a project in any language can depend on the same substrate. Open a PR to add it to projects/.

Paper & dataset

The foundation and the three recipes above are described in Orchard: An Open-Source Agentic Modeling Framework (Peng et al., arXiv:2605.15040). OpenWebRL and OpenForge RL are separate papers that build on the same environment layer — see projects/.

  • 📄 Paper: arXiv:2605.15040
  • 🤗 Research data: Hugging Face dataset — one repository ships two parallel subsets, both produced inside the same Orchard Env sandbox infrastructure:
    • swe config — 107,185 multi-turn SWE rollouts over 19,287 unique task instances across 2,788 repositories, with verified resolve labels (74,649 resolved · 32,536 unresolved) and an average of 47.5 turns per trajectory.
    • gui config — 3,070 judge-verified successful per-step rollouts from a web-browsing GUI agent across 409 WebVoyager-style tasks, each with a rendered screenshot (multimodal).

Roadmap

Stateful sandboxes — pause, resume, and branching

The environment is currently linear: a sandbox is created, driven for N turns, and destroyed. Every rollout re-executes its whole prefix, and a trajectory yields a single outcome reward that has to be spread across all of its turns — in our SWE data, an average of 47.5 turns per trajectory. Attributing that one scalar to the turn that actually mattered is the central credit-assignment problem, and the paper attacks it indirectly, with retrospective value estimation over completed trajectories.

Snapshotting makes that measurable directly instead. We are adding:

  • Pause / resume — checkpoint a sandbox's full state (filesystem, processes, environment) at any turn and restore it later, so a rollout can be suspended and continued instead of restarted.
  • Branching — fork k independent continuations from the same snapshot at turn t. Rolling forward repeatedly from a single state gives a Monte-Carlo estimate of that state's value, which turns per-turn credit assignment into something measured rather than inferred, and makes counterfactuals ("what if the agent had not run that command?") directly observable.
  • Prefix sharing — because a branched prefix is executed once instead of once per rollout, tree-structured search and per-turn advantage estimation get substantially cheaper than the flat-rollout baseline.

This is the environment primitive that the next round of credit-assignment research needs, and it belongs in the environment layer rather than in any one trainer.

Contributing

Contributions are welcome — please open an issue or pull request. Development setup, linting, and the test matrix are documented in orchard_env/README.md and orchard_env/tests/README.md.

This project has adopted the Microsoft Open Source Code of Conduct; see CODE_OF_CONDUCT.md.

Security

Please do not report security vulnerabilities through public GitHub issues. See SECURITY.md for how to report them.

Citation

If you use Orchard or Orchard Env in your research, please cite:

@article{peng2026orchard,
  title={Orchard: An Open-Source Agentic Modeling Framework},
  author={Peng, Baolin and Yao, Wenlin and Wu, Qianhui and Cheng, Hao and
          Yu, Xiao and Yang, Rui and Ge, Tao and Sordoni, Alessandro and
          Yuan, Xingdi and Shen, Yelong and He, Pengcheng and Zhang, Tong and
          Yu, Zhou and Gao, Jianfeng},
  journal={arXiv preprint arXiv:2605.15040},
  year={2026},
  url={https://arxiv.org/abs/2605.15040}
}

License

MIT © Microsoft Corporation.