Cut AI coding agent bills 90%+ by routing between your local GPU and cloud APIs. OpenAI-compatible.
See the code
Nacho Flow sits between your coding agent and your LLM backends. Routine turns run on your local GPU for $0.00. Complex reasoning escalates to frontier models automatically. Loop detection kills runaway agents in < 3s.
Nacho Flow is a hybrid model dispatcher and agent supervisor written in pure Go. Point Cline, Zoo Code, OpenCode, Aider, Cursor, or Continue at http://localhost:8000/v1 and it handles the rest.
🌐 Website & Documentation: spicebox.dev/nacho-flow 💸 Latest Research: The Frontier Tax: 1:1 Claude Sonnet 5 vs. Nacho Flow (5.7x cheaper, 12 minutes faster) Part of the spicebox.dev developer tool suite by @dixieflatline76.
Tokens, Retries, Keywords, HasTools, HasImages)tool_calls JSON -- no harness crasheslogs/traffic.jsonl to find optimal tier thresholds (nacho-flow tune)-race), and static security audited (gosec).NTS (experimental token compaction) is available but disabled by default to preserve prompt cache stability and prevent diff drift. See NTS docs if you need it for extreme context limits.
flowchart LR
Step1["**Step 1: Local GPU ($0.00)**<br/><code>ollama run gemma4:12b-it-qat</code>"]
Step2["**Step 2: Cloud Gateway**<br/>OpenRouter Key (1 Unified Key)"]
Step3["**Step 3: Point Your Agent**<br/><code>http://127.0.0.1:8000/v1</code>"]
Step1 --> Step2 --> Step3
Local GPU ($0.00) -- Install Ollama and pull a coding model:
ollama run gemma4:12b-it-qat
Set
OLLAMA_CONTEXT_LENGTH=32768to prevent prompt truncation (setx OLLAMA_CONTEXT_LENGTH 32768on Windows). See User Guide.
Cloud Gateway via OpenRouter -- One key for all frontier tiers (Qwen 3 Coder Plus, Gemini 3.8 Flash, Claude Sonnet 5):
export OPENROUTER_API_KEY="sk-or-v1-..."
Start Nacho Flow & Connect Your Agent
Start in the sidebar.nacho-flow in your terminal.http://127.0.0.1:8000/v1, Model ID to nacho-hybrid.VS Code Extension (All-in-One Runtime):
code --install-extension dixieflatline76.nacho-flow
Universal Shell Installer (Linux & macOS):
curl -fsSL https://raw.githubusercontent.com/dixieflatline76/nacho-flow/main/scripts/install.sh | bash
Docker / Podman (Multi-Arch Distroless Container):
docker run -d -p 8000:8000 \
-v $(pwd)/config.yaml:/config/config.yaml \
ghcr.io/dixieflatline76/nacho-flow:latest
Homebrew (macOS & Linux):
brew install dixieflatline76/nacho-flow/nacho-flow
Pre-compiled Binaries: GitHub Releases
Go:
go install github.com/dixieflatline76/nacho-flow/cmd/nacho-flow@latest
config.yaml)Create a config.yaml in your project folder or ~/.config/nacho-flow/config.yaml:
port: 8000
host: "127.0.0.1"
auth_token: "sk-nacho-secret-key"
providers:
ollama:
base_url: "http://127.0.0.1:11434/v1"
type: "local"
openrouter:
base_url: "https://openrouter.ai/api/v1"
api_key: "ENV_OPENROUTER_API_KEY"
headers:
HTTP-Referer: "https://spicebox.dev"
X-Title: "nacho-flow"
langdock:
base_url: "https://api.langdock.com/v1"
api_key: "ENV_LANGDOCK_API_KEY"
anthropic:
base_url: "https://api.anthropic.com"
type: "anthropic"
api_key: "ENV_ANTHROPIC_API_KEY"
# Tiers evaluated top-to-bottom: first match wins
tiers:
- name: "Cloud Reasoning"
model: "deepseek/deepseek-r1"
provider: "openrouter"
when: "any(Keywords, { # in ['deadlock', 'mutex', 'race', 'concurrency', 'atomic'] })"
- name: "Cloud Vision"
model: "google/gemini-2.5-flash-lite"
provider: "openrouter"
when: "HasImages"
- name: "Local GPU"
model: "qwen3.8-coder:14b"
provider: "ollama"
max_context: 16384
when: "Tokens < 16000 && !HasImages && !HasTools && Retries < 2"
strip_images: true
- name: "Cloud Agentic Fast"
model: "qwen/qwen3-coder-30b-a3b-instruct"
provider: "openrouter"
when: "Tokens >= 16000 || HasTools || Retries >= 2"
default_tier:
name: "Cloud Fallback"
model: "deepseek/deepseek-v4-flash-latest"
provider: "openrouter"
when: "true"
cycle_killer:
enabled: true
max_tool_tokens: 8192
repetition_threshold: 3
kickstart_threshold: 5
fairy_dust:
enabled: true
entries:
- name: "Tactical Code Review"
frequency: 15
provider: "openrouter"
model: "anthropic/claude-sonnet-5"
- name: "Strategic Architecture Review"
frequency: 40
provider: "openrouter"
model: "anthropic/claude-sonnet-5"
@nacho:)Control routing and guardrails directly from your editor chat without restarting the gateway:
| Category | Directive | Action |
|---|---|---|
| Session Switches | @nacho:kickstart-off / on | Suspend / resume Kickstart idle stall escalation |
@nacho:cyclekiller-off / on | Suspend / resume Cycle Killer stream loop breaker | |
@nacho:shield-off / on | Suspend / resume synthetic tool-call synthesis | |
@nacho:raw-on / off | Enable / disable raw upstream SSE stream | |
@nacho:fairydust-off / on | Suspend / resume periodic frontier checkpoints | |
| Inspection & Reset | @nacho:toggles | Display live session switches ($0.00 / 0 tokens) |
@nacho:status | Display daemon telemetry, spend & saved dollars | |
@nacho:reset | Hard reset turn counter & restore default switches | |
| Single-Turn Overrides | @nacho:local | Force current turn to Local GPU ($0.00) |
@nacho:cloud | Force current turn to Cloud Fallback tier | |
@nacho:reasoning | Force current turn to DeepSeek-R1 / o1 |
Plan Mode Auto-Detection: When your agent switches into Plan Mode (zero write tools declared), Nacho Flow automatically detects
HasWriteCapability == falseand suspends Kickstart -- no manual toggles required.
nacho-flow tune)Replays your actual logs/traffic.jsonl to find where your local GPU starts failing, prune dead tiers, and stop over-escalating to expensive frontier models:
# Advisory dry-run
nacho-flow tune
# Target specific VRAM ceiling
nacho-flow tune --vram-gb=16
# Apply recommendations with timestamped backup
nacho-flow tune --apply
# Install as native OS service (Windows Service / systemd / launchd)
nacho-flow service install
# Start the background daemon
nacho-flow service start
Nacho Flow exposes a standard OpenAI-compatible proxy at http://localhost:8000/v1. Use nacho-hybrid as your Model ID.
OpenAI Compatiblehttp://localhost:8000/v1sk-nacho-secret-keynacho-hybrid128,000, Max Output 8,192export OPENAI_BASE_URL="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
opencode --model openai/nacho-hybrid
export OPENAI_API_BASE="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
aider --model openai/nacho-hybrid
http://localhost:8000/v1sk-nacho-secret-keynacho-hybrid{
"models": [{
"title": "Nacho Flow (Hybrid Local + Cloud)",
"provider": "openai",
"model": "nacho-hybrid",
"apiBase": "http://127.0.0.1:8000/v1",
"apiKey": "sk-nacho-secret-key"
}]
}
Visit spicebox.dev/nacho-flow for the full documentation portal.
expr routing rulesNacho Flow is 100% free and open source. If it saved your sanity or your API bill:
Dual-Licensing Model:
Free & Open-Source (GNU AGPL-3.0 with API Interoperability Exception): Free for individual developers, open-source projects, and local evaluation. Calling Nacho Flow's OpenAI-compatible APIs from client apps or IDEs does not make your code a derivative work. Modifying and distributing or network-hosting Nacho Flow requires source disclosure under AGPL-3.0.
Spicebox Commercial & Enterprise OEM License: For enterprise fleet deployments, closed-source embedding, commercial SaaS hosting, and IP indemnification. See COMMERCIAL_LICENSE.md or contact karl@spicebox.dev.
Contributions accepted under our Contributor License Agreement (.github/CLA.md).
Copyright © 2026 Karl Kwong / Spicebox · Licensed under GNU AGPL-3.0 with Commercial Dual-Licensing. (VS Code Extension licensed under MIT).
Go
72.7%
TypeScript
12.4%
HTML
6.5%
CSS
5.0%
JavaScript
2.7%
Cut AI coding agent bills 90%+ by routing between your local GPU and cloud APIs. OpenAI-compatible.
See the code
Nacho Flow sits between your coding agent and your LLM backends. Routine turns run on your local GPU for $0.00. Complex reasoning escalates to frontier models automatically. Loop detection kills runaway agents in < 3s.
Nacho Flow is a hybrid model dispatcher and agent supervisor written in pure Go. Point Cline, Zoo Code, OpenCode, Aider, Cursor, or Continue at http://localhost:8000/v1 and it handles the rest.
🌐 Website & Documentation: spicebox.dev/nacho-flow 💸 Latest Research: The Frontier Tax: 1:1 Claude Sonnet 5 vs. Nacho Flow (5.7x cheaper, 12 minutes faster) Part of the spicebox.dev developer tool suite by @dixieflatline76.
Tokens, Retries, Keywords, HasTools, HasImages)tool_calls JSON -- no harness crasheslogs/traffic.jsonl to find optimal tier thresholds (nacho-flow tune)-race), and static security audited (gosec).NTS (experimental token compaction) is available but disabled by default to preserve prompt cache stability and prevent diff drift. See NTS docs if you need it for extreme context limits.
flowchart LR
Step1["**Step 1: Local GPU ($0.00)**<br/><code>ollama run gemma4:12b-it-qat</code>"]
Step2["**Step 2: Cloud Gateway**<br/>OpenRouter Key (1 Unified Key)"]
Step3["**Step 3: Point Your Agent**<br/><code>http://127.0.0.1:8000/v1</code>"]
Step1 --> Step2 --> Step3
Local GPU ($0.00) -- Install Ollama and pull a coding model:
ollama run gemma4:12b-it-qat
Set
OLLAMA_CONTEXT_LENGTH=32768to prevent prompt truncation (setx OLLAMA_CONTEXT_LENGTH 32768on Windows). See User Guide.
Cloud Gateway via OpenRouter -- One key for all frontier tiers (Qwen 3 Coder Plus, Gemini 3.8 Flash, Claude Sonnet 5):
export OPENROUTER_API_KEY="sk-or-v1-..."
Start Nacho Flow & Connect Your Agent
Start in the sidebar.nacho-flow in your terminal.http://127.0.0.1:8000/v1, Model ID to nacho-hybrid.VS Code Extension (All-in-One Runtime):
code --install-extension dixieflatline76.nacho-flow
Universal Shell Installer (Linux & macOS):
curl -fsSL https://raw.githubusercontent.com/dixieflatline76/nacho-flow/main/scripts/install.sh | bash
Docker / Podman (Multi-Arch Distroless Container):
docker run -d -p 8000:8000 \
-v $(pwd)/config.yaml:/config/config.yaml \
ghcr.io/dixieflatline76/nacho-flow:latest
Homebrew (macOS & Linux):
brew install dixieflatline76/nacho-flow/nacho-flow
Pre-compiled Binaries: GitHub Releases
Go:
go install github.com/dixieflatline76/nacho-flow/cmd/nacho-flow@latest
config.yaml)Create a config.yaml in your project folder or ~/.config/nacho-flow/config.yaml:
port: 8000
host: "127.0.0.1"
auth_token: "sk-nacho-secret-key"
providers:
ollama:
base_url: "http://127.0.0.1:11434/v1"
type: "local"
openrouter:
base_url: "https://openrouter.ai/api/v1"
api_key: "ENV_OPENROUTER_API_KEY"
headers:
HTTP-Referer: "https://spicebox.dev"
X-Title: "nacho-flow"
langdock:
base_url: "https://api.langdock.com/v1"
api_key: "ENV_LANGDOCK_API_KEY"
anthropic:
base_url: "https://api.anthropic.com"
type: "anthropic"
api_key: "ENV_ANTHROPIC_API_KEY"
# Tiers evaluated top-to-bottom: first match wins
tiers:
- name: "Cloud Reasoning"
model: "deepseek/deepseek-r1"
provider: "openrouter"
when: "any(Keywords, { # in ['deadlock', 'mutex', 'race', 'concurrency', 'atomic'] })"
- name: "Cloud Vision"
model: "google/gemini-2.5-flash-lite"
provider: "openrouter"
when: "HasImages"
- name: "Local GPU"
model: "qwen3.8-coder:14b"
provider: "ollama"
max_context: 16384
when: "Tokens < 16000 && !HasImages && !HasTools && Retries < 2"
strip_images: true
- name: "Cloud Agentic Fast"
model: "qwen/qwen3-coder-30b-a3b-instruct"
provider: "openrouter"
when: "Tokens >= 16000 || HasTools || Retries >= 2"
default_tier:
name: "Cloud Fallback"
model: "deepseek/deepseek-v4-flash-latest"
provider: "openrouter"
when: "true"
cycle_killer:
enabled: true
max_tool_tokens: 8192
repetition_threshold: 3
kickstart_threshold: 5
fairy_dust:
enabled: true
entries:
- name: "Tactical Code Review"
frequency: 15
provider: "openrouter"
model: "anthropic/claude-sonnet-5"
- name: "Strategic Architecture Review"
frequency: 40
provider: "openrouter"
model: "anthropic/claude-sonnet-5"
@nacho:)Control routing and guardrails directly from your editor chat without restarting the gateway:
| Category | Directive | Action |
|---|---|---|
| Session Switches | @nacho:kickstart-off / on | Suspend / resume Kickstart idle stall escalation |
@nacho:cyclekiller-off / on | Suspend / resume Cycle Killer stream loop breaker | |
@nacho:shield-off / on | Suspend / resume synthetic tool-call synthesis | |
@nacho:raw-on / off | Enable / disable raw upstream SSE stream | |
@nacho:fairydust-off / on | Suspend / resume periodic frontier checkpoints | |
| Inspection & Reset | @nacho:toggles | Display live session switches ($0.00 / 0 tokens) |
@nacho:status | Display daemon telemetry, spend & saved dollars | |
@nacho:reset | Hard reset turn counter & restore default switches | |
| Single-Turn Overrides | @nacho:local | Force current turn to Local GPU ($0.00) |
@nacho:cloud | Force current turn to Cloud Fallback tier | |
@nacho:reasoning | Force current turn to DeepSeek-R1 / o1 |
Plan Mode Auto-Detection: When your agent switches into Plan Mode (zero write tools declared), Nacho Flow automatically detects
HasWriteCapability == falseand suspends Kickstart -- no manual toggles required.
nacho-flow tune)Replays your actual logs/traffic.jsonl to find where your local GPU starts failing, prune dead tiers, and stop over-escalating to expensive frontier models:
# Advisory dry-run
nacho-flow tune
# Target specific VRAM ceiling
nacho-flow tune --vram-gb=16
# Apply recommendations with timestamped backup
nacho-flow tune --apply
# Install as native OS service (Windows Service / systemd / launchd)
nacho-flow service install
# Start the background daemon
nacho-flow service start
Nacho Flow exposes a standard OpenAI-compatible proxy at http://localhost:8000/v1. Use nacho-hybrid as your Model ID.
OpenAI Compatiblehttp://localhost:8000/v1sk-nacho-secret-keynacho-hybrid128,000, Max Output 8,192export OPENAI_BASE_URL="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
opencode --model openai/nacho-hybrid
export OPENAI_API_BASE="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
aider --model openai/nacho-hybrid
http://localhost:8000/v1sk-nacho-secret-keynacho-hybrid{
"models": [{
"title": "Nacho Flow (Hybrid Local + Cloud)",
"provider": "openai",
"model": "nacho-hybrid",
"apiBase": "http://127.0.0.1:8000/v1",
"apiKey": "sk-nacho-secret-key"
}]
}
Visit spicebox.dev/nacho-flow for the full documentation portal.
expr routing rulesNacho Flow is 100% free and open source. If it saved your sanity or your API bill:
Dual-Licensing Model:
Free & Open-Source (GNU AGPL-3.0 with API Interoperability Exception): Free for individual developers, open-source projects, and local evaluation. Calling Nacho Flow's OpenAI-compatible APIs from client apps or IDEs does not make your code a derivative work. Modifying and distributing or network-hosting Nacho Flow requires source disclosure under AGPL-3.0.
Spicebox Commercial & Enterprise OEM License: For enterprise fleet deployments, closed-source embedding, commercial SaaS hosting, and IP indemnification. See COMMERCIAL_LICENSE.md or contact karl@spicebox.dev.
Contributions accepted under our Contributor License Agreement (.github/CLA.md).
Copyright © 2026 Karl Kwong / Spicebox · Licensed under GNU AGPL-3.0 with Commercial Dual-Licensing. (VS Code Extension licensed under MIT).
Go
72.7%
TypeScript
12.4%
HTML
6.5%
CSS
5.0%
JavaScript
2.7%