Open-source infrastructure for the full inference lifecycle: deploy, observe, scale, optimize, and safely release self-hosted models behind one endpoint.
54
stars
336
commits
Go
primary language
Sep 6, 2026
updated
One endpoint for the complete inference lifecycle.
Deploy or adopt self-hosted models, then route, observe, scale, optimize, release, and recover
them across your infrastructure without changing the application's OpenAI-compatible contract.
See the release loop · Install the public beta · Read the quickstart
Plan before spend · survive disconnects · explain latency · guard every release
Install the v1.0.0-rc.1 public beta CLI with Homebrew or use the matching SDK prerelease:
brew install infercrane/tap/infercrane
python -m pip install 'infercrane==1.0.0rc1'
npm install '@infercrane/sdk@1.0.0-rc.1'
Release archives and Terraform provider binaries are available from the
v1.0.0-rc.1 prerelease.
To run the complete GPU-free product proof without creating cloud resources:
git clone https://github.com/infercrane/infercrane.git
cd infercrane
make demo
The proof connects an OpenAI-compatible worker, sends and inspects a request, creates an isolated candidate, records a deterministic Release Guard rejection, verifies that production traffic did not move, and removes its disposable stack.
Starting a model server can be one command. Operating it while models, runtimes, accelerators, providers, scaling policies, and revisions change is the longer-lived problem.
A candidate does not receive production traffic merely because its health endpoint returns 200.
InferCrane keeps the active revision serving while the candidate is measured and records one of
three explicit Release Guard decisions: ACCEPT, REJECT, or WAIT. An ACCEPT decision permits
an explicit promotion; it does not move traffic by itself.
Stable endpoint ───────────────────────────────▶ active revision
▲
│ explicit promotion after ACCEPT
new serving plan ▶ isolated candidate ▶ evidence ▶ ACCEPT / REJECT / WAIT
└─ reject or wait: active unchanged
The evidence and decision remain inspectable after the operation finishes. If this is an operational gap your team recognizes, star InferCrane and tell us which model, runtime, accelerator, and provider tuple should be qualified next.
See the detailed comparison, including the boundaries InferCrane does not own.
The local proof needs no GPU or cloud account. Real-provider support remains exact-tuple qualified; InferCrane reports missing model/runtime/hardware evidence as unknown instead of turning it into a compatibility claim.
Applications and agents
│
▼
Stable OpenAI-compatible endpoint
│
▼
InferCrane: route · observe · optimize · release · recover
│
├── deploy new inference
├── adopt an existing workload
└── govern a model API or gateway
│
▼
vLLM · SGLang · custom OCI · experimental Dynamo
AWS · GCP · Kubernetes · RunPod · existing infrastructure
| Goal | InferCrane workflow |
|---|---|
| Put a model into production | Initialize a workload, review the serving plan, deploy, then call its stable endpoint. |
| Adopt existing inference | Connect vLLM, SGLang, LiteLLM, or another compatible endpoint without transferring lifecycle ownership. |
| Ship a safer revision | Benchmark and replay an isolated candidate, attach quality evidence, then let Release Guard promote or reject it. |
| Understand production failures | Trace queue wait, attempts, runtime, revision, latency, saturation, and durable operations without storing prompt content. |
| Optimize performance and cost | Propose serving configurations, measure comparable candidates on real hardware, and persist only qualified evidence. |
| Survive infrastructure delays | Submit idempotent durable operations that continue after the CLI or control plane process disconnects. |
Local fixtures prove application and lifecycle behavior. They do not prove GPU performance, cloud capacity, runtime compatibility, or model quality. InferCrane keeps those evidence boundaries explicit.
Start with a curated recipe:
infercrane workload init ./support --recipe qwen3-8b
cd support
infercrane workload plan
infercrane workload deploy --wait
Or bring another compatible immutable model identity:
infercrane workload init ./agent-model --model mistralai/Mistral-7B-Instruct-v0.3
cd agent-model
infercrane workload plan
infercrane workload deploy --wait
Recipes are reproducible configuration starting points, not benchmark claims or an allowlist. Evidence remains bound to the exact model commit, runtime, accelerator, provider, cache state, and workload.
Already operating a workload? Connect it first:
infercrane connect https://vllm.internal/v1 --as support-production
infercrane doctor support-production
infercrane observe support-production
InferCrane can observe an existing endpoint before it manages traffic or infrastructure. Provider credentials and request content do not enter the browser console.
InferCrane separates a modeled proposal from measured and qualified evidence:
infercrane optimize propose llama-3.1-8b-instruct \
--provider aws \
--region eu-central-1 \
--gpu L40S \
--objective interactive \
--write-dir .infercrane/candidates
model + hardware + workload + SLO + cost target
│
▼
candidate serving plans
│
▼
AIPerf + replay + quality evidence
│
▼
performance · errors · cost
│
┌─────┴─────┐
▼ ▼
promote reject
InferCrane composes replaceable execution technology instead of rebuilding it. vLLM, SGLang, and custom OCI are current execution paths. Dynamo is experimental. TensorRT-LLM, LMCache, NIXL, and external optimizers remain capability boundaries until their exact adapters and hardware tuples are qualified. InferCrane owns serving-plan identity, durable operations, comparable evidence, routing policy, promotion, and rollback.
429 and Retry-After, one end-to-end deadline,
and bounded retries prevent unlimited queue growth.Read the architecture, system invariants, and data flows for the complete design.
| Interface | Status and purpose |
|---|---|
| CLI and control API | Primary deployment, operation, evidence, and administration interfaces. |
| OpenAI-compatible gateway | Capability-gated Chat, Completions, Embeddings, Responses, and online batch paths. The pinned vLLM profile currently qualifies Chat plus model-compatible Completions and Embeddings; unsupported capabilities fail before upstream transmission. |
| Python and TypeScript SDKs | Public beta packages: infercrane==1.0.0rc1 and @infercrane/sdk@1.0.0-rc.1. Generated from the checked OpenAPI contract. |
| Terraform provider | Logical deployment lifecycle with guarded updates and import. Release binaries and source are public; Registry publication is pending. |
| Terminal workspace | Fleet attention, evidence inspection, and state-valid guarded actions. |
| Browser console | Separate deny-by-default private-preview application using the same control API. |
| Read-only MCP server | Closed-world operational inspection without deployment, scaling, promotion, deletion, budget, or secret tools. |
InferCrane v1.0.0-rc.1 is the first public beta. The stable v1.0.0 release will promote the exact
qualified product contract after the prerelease cycle; no earlier development tag should be treated
as a supported public release.
See the authoritative compatibility and qualification policy, AWS evidence, and feature qualification matrix before relying on an exact provider, runtime, model, or accelerator combination.
Mintlify generates llms.txt and
llms-full.txt from the public documentation. Every
public documentation page is also available as Markdown by appending .md to its URL.
Contributions are welcome. Start with CONTRIBUTING.md, follow the
Code of Conduct, and sign commits with git commit -s. Changes must include
tests and relevant documentation. Durable architecture, security, storage, and dependency changes
must update their authoritative public documentation.
Never disclose credentials, prompts, model responses, private endpoints, or suspected vulnerabilities in a public issue. Use the private reporting process in SECURITY.md. Questions and reproducible defects follow SUPPORT.md.
InferCrane Community is available under the Apache License 2.0. Hosted and enterprise products are separate distributions and are not licensed by this repository. Release archives also include third-party notices and a release-specific SPDX SBOM. The InferCrane name and crane logo remain subject to the trademark policy.
331 commits
5 commits
Go
87.1%
Shell
8.4%
Python
1.8%
TypeScript
1.2%
Open-source infrastructure for the full inference lifecycle: deploy, observe, scale, optimize, and safely release self-hosted models behind one endpoint.
54
stars
336
commits
Go
primary language
Sep 6, 2026
updated
One endpoint for the complete inference lifecycle.
Deploy or adopt self-hosted models, then route, observe, scale, optimize, release, and recover
them across your infrastructure without changing the application's OpenAI-compatible contract.
See the release loop · Install the public beta · Read the quickstart
Plan before spend · survive disconnects · explain latency · guard every release
Install the v1.0.0-rc.1 public beta CLI with Homebrew or use the matching SDK prerelease:
brew install infercrane/tap/infercrane
python -m pip install 'infercrane==1.0.0rc1'
npm install '@infercrane/sdk@1.0.0-rc.1'
Release archives and Terraform provider binaries are available from the
v1.0.0-rc.1 prerelease.
To run the complete GPU-free product proof without creating cloud resources:
git clone https://github.com/infercrane/infercrane.git
cd infercrane
make demo
The proof connects an OpenAI-compatible worker, sends and inspects a request, creates an isolated candidate, records a deterministic Release Guard rejection, verifies that production traffic did not move, and removes its disposable stack.
Starting a model server can be one command. Operating it while models, runtimes, accelerators, providers, scaling policies, and revisions change is the longer-lived problem.
A candidate does not receive production traffic merely because its health endpoint returns 200.
InferCrane keeps the active revision serving while the candidate is measured and records one of
three explicit Release Guard decisions: ACCEPT, REJECT, or WAIT. An ACCEPT decision permits
an explicit promotion; it does not move traffic by itself.
Stable endpoint ───────────────────────────────▶ active revision
▲
│ explicit promotion after ACCEPT
new serving plan ▶ isolated candidate ▶ evidence ▶ ACCEPT / REJECT / WAIT
└─ reject or wait: active unchanged
The evidence and decision remain inspectable after the operation finishes. If this is an operational gap your team recognizes, star InferCrane and tell us which model, runtime, accelerator, and provider tuple should be qualified next.
See the detailed comparison, including the boundaries InferCrane does not own.
The local proof needs no GPU or cloud account. Real-provider support remains exact-tuple qualified; InferCrane reports missing model/runtime/hardware evidence as unknown instead of turning it into a compatibility claim.
Applications and agents
│
▼
Stable OpenAI-compatible endpoint
│
▼
InferCrane: route · observe · optimize · release · recover
│
├── deploy new inference
├── adopt an existing workload
└── govern a model API or gateway
│
▼
vLLM · SGLang · custom OCI · experimental Dynamo
AWS · GCP · Kubernetes · RunPod · existing infrastructure
| Goal | InferCrane workflow |
|---|---|
| Put a model into production | Initialize a workload, review the serving plan, deploy, then call its stable endpoint. |
| Adopt existing inference | Connect vLLM, SGLang, LiteLLM, or another compatible endpoint without transferring lifecycle ownership. |
| Ship a safer revision | Benchmark and replay an isolated candidate, attach quality evidence, then let Release Guard promote or reject it. |
| Understand production failures | Trace queue wait, attempts, runtime, revision, latency, saturation, and durable operations without storing prompt content. |
| Optimize performance and cost | Propose serving configurations, measure comparable candidates on real hardware, and persist only qualified evidence. |
| Survive infrastructure delays | Submit idempotent durable operations that continue after the CLI or control plane process disconnects. |
Local fixtures prove application and lifecycle behavior. They do not prove GPU performance, cloud capacity, runtime compatibility, or model quality. InferCrane keeps those evidence boundaries explicit.
Start with a curated recipe:
infercrane workload init ./support --recipe qwen3-8b
cd support
infercrane workload plan
infercrane workload deploy --wait
Or bring another compatible immutable model identity:
infercrane workload init ./agent-model --model mistralai/Mistral-7B-Instruct-v0.3
cd agent-model
infercrane workload plan
infercrane workload deploy --wait
Recipes are reproducible configuration starting points, not benchmark claims or an allowlist. Evidence remains bound to the exact model commit, runtime, accelerator, provider, cache state, and workload.
Already operating a workload? Connect it first:
infercrane connect https://vllm.internal/v1 --as support-production
infercrane doctor support-production
infercrane observe support-production
InferCrane can observe an existing endpoint before it manages traffic or infrastructure. Provider credentials and request content do not enter the browser console.
InferCrane separates a modeled proposal from measured and qualified evidence:
infercrane optimize propose llama-3.1-8b-instruct \
--provider aws \
--region eu-central-1 \
--gpu L40S \
--objective interactive \
--write-dir .infercrane/candidates
model + hardware + workload + SLO + cost target
│
▼
candidate serving plans
│
▼
AIPerf + replay + quality evidence
│
▼
performance · errors · cost
│
┌─────┴─────┐
▼ ▼
promote reject
InferCrane composes replaceable execution technology instead of rebuilding it. vLLM, SGLang, and custom OCI are current execution paths. Dynamo is experimental. TensorRT-LLM, LMCache, NIXL, and external optimizers remain capability boundaries until their exact adapters and hardware tuples are qualified. InferCrane owns serving-plan identity, durable operations, comparable evidence, routing policy, promotion, and rollback.
429 and Retry-After, one end-to-end deadline,
and bounded retries prevent unlimited queue growth.Read the architecture, system invariants, and data flows for the complete design.
| Interface | Status and purpose |
|---|---|
| CLI and control API | Primary deployment, operation, evidence, and administration interfaces. |
| OpenAI-compatible gateway | Capability-gated Chat, Completions, Embeddings, Responses, and online batch paths. The pinned vLLM profile currently qualifies Chat plus model-compatible Completions and Embeddings; unsupported capabilities fail before upstream transmission. |
| Python and TypeScript SDKs | Public beta packages: infercrane==1.0.0rc1 and @infercrane/sdk@1.0.0-rc.1. Generated from the checked OpenAPI contract. |
| Terraform provider | Logical deployment lifecycle with guarded updates and import. Release binaries and source are public; Registry publication is pending. |
| Terminal workspace | Fleet attention, evidence inspection, and state-valid guarded actions. |
| Browser console | Separate deny-by-default private-preview application using the same control API. |
| Read-only MCP server | Closed-world operational inspection without deployment, scaling, promotion, deletion, budget, or secret tools. |
InferCrane v1.0.0-rc.1 is the first public beta. The stable v1.0.0 release will promote the exact
qualified product contract after the prerelease cycle; no earlier development tag should be treated
as a supported public release.
See the authoritative compatibility and qualification policy, AWS evidence, and feature qualification matrix before relying on an exact provider, runtime, model, or accelerator combination.
Mintlify generates llms.txt and
llms-full.txt from the public documentation. Every
public documentation page is also available as Markdown by appending .md to its URL.
Contributions are welcome. Start with CONTRIBUTING.md, follow the
Code of Conduct, and sign commits with git commit -s. Changes must include
tests and relevant documentation. Durable architecture, security, storage, and dependency changes
must update their authoritative public documentation.
Never disclose credentials, prompts, model responses, private endpoints, or suspected vulnerabilities in a public issue. Use the private reporting process in SECURITY.md. Questions and reproducible defects follow SUPPORT.md.
InferCrane Community is available under the Apache License 2.0. Hosted and enterprise products are separate distributions and are not licensed by this repository. Release archives also include third-party notices and a release-specific SPDX SBOM. The InferCrane name and crane logo remain subject to the trademark policy.
331 commits
5 commits
Go
87.1%
Shell
8.4%
Python
1.8%
TypeScript
1.2%