dixieflatline76/nacho-flow

Cut AI coding agent bills 90%+ by routing between your local GPU and cloud APIs. OpenAI-compatible.

Go

2

202 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a pure Go supervisor proxy to make harnesses like Zoo, Cline and Aider work better (r/golang)

I built a Go supervisor proxy called [**Nacho Flow**](https://github.com/dixieflatline76/nacho-flow) to fix the things that frustrate me about using free coding agent harnesses with cheaper models. Free AI coding agent harnesses like Zoo and Cline work okay with frontier models like Sonnet and…

1

Oct 2, 2026

README

Nacho Flow

🌮 Nacho Flow

CI Status VS Code Extension Go Reference Go Version License: AGPL-3.0 Latest Release Homebrew Docker GHCR

Nacho Flow sits between your coding agent and your LLM backends. Routine turns run on your local GPU for $0.00. Complex reasoning escalates to frontier models automatically. Loop detection kills runaway agents in < 3s.

Nacho Flow is a hybrid model dispatcher and agent supervisor written in pure Go. Point Cline, Zoo Code, OpenCode, Aider, Cursor, or Continue at http://localhost:8000/v1 and it handles the rest.

🌐 Website & Documentation: spicebox.dev/nacho-flow 💸 Latest Research: The Frontier Tax: 1:1 Claude Sonnet 5 vs. Nacho Flow (5.7x cheaper, 12 minutes faster) Part of the spicebox.dev developer tool suite by @dixieflatline76.


What It Does

  • Hybrid routing: Local GPU -> budget cloud -> frontier, evaluated per-turn via deterministic AST bytecode rules (Tokens, Retries, Keywords, HasTools, HasImages)
  • Cycle Killer: Monitors live SSE streams across thinking/prose/tool lanes. Kills repetition loops in < 3s, injects a local $0.00 override to force tool action, never interrupts file writes
  • Kickstart: Detects consecutive non-write turns (analysis paralysis) and injects resuscitation prompts or escalates to a smarter model
  • Tool Normalizer: Converts 8 raw tool-call format families (Hermes, Mistral, Llama 3, ReAct, Markdown fences, bare JSON) into standard OpenAI tool_calls JSON -- no harness crashes
  • Fairy Dust: Programmable frontier checkpoints every N file writes for quality audits (without running expensive models all day)
  • Cost engine: Prompt-cache-aware billing tracking input, output, and cache tokens at exact provider rates
  • Auto-Tune: Replays your logs/traffic.jsonl to find optimal tier thresholds (nacho-flow tune)
  • VS Code extension: Bundles the Go binary, live financial telemetry, route inspector, one-click preset switching. Zero CLI setup. Install from Marketplace
  • Wire-speed core: < 0.19ms routing overhead, 30,000+ req/s (peak 30,284 req/s), lock-free atomic RCU state, zero heap churn during proxying
  • Zero runtime dependencies: Single static binary, CGO_ENABLED=0, no Node or Python required
  • 🧪 Engineered for Reliability: Strictly $\ge 95.0%\text{--}100%$ statement test coverage across all packages (95.4% global coverage), 100% race-detector clean (-race), and static security audited (gosec).

NTS (experimental token compaction) is available but disabled by default to preserve prompt cache stability and prevent diff drift. See NTS docs if you need it for extreme context limits.


Quickstart (3 Steps)

flowchart LR
    Step1["**Step 1: Local GPU ($0.00)**<br/><code>ollama run gemma4:12b-it-qat</code>"]
    Step2["**Step 2: Cloud Gateway**<br/>OpenRouter Key (1 Unified Key)"]
    Step3["**Step 3: Point Your Agent**<br/><code>http://127.0.0.1:8000/v1</code>"]

    Step1 --> Step2 --> Step3
  1. Local GPU ($0.00) -- Install Ollama and pull a coding model:

    ollama run gemma4:12b-it-qat
    

    Set OLLAMA_CONTEXT_LENGTH=32768 to prevent prompt truncation (setx OLLAMA_CONTEXT_LENGTH 32768 on Windows). See User Guide.

  2. Cloud Gateway via OpenRouter -- One key for all frontier tiers (Qwen 3 Coder Plus, Gemini 3.8 Flash, Claude Sonnet 5):

    export OPENROUTER_API_KEY="sk-or-v1-..."
    
  3. Start Nacho Flow & Connect Your Agent

    • VS Code Extension: Install from Marketplace, click Start in the sidebar.
    • Standalone CLI: Run nacho-flow in your terminal.
    • Set your agent's Base URL to http://127.0.0.1:8000/v1, Model ID to nacho-hybrid.

Installation

VS Code Extension (All-in-One Runtime):

code --install-extension dixieflatline76.nacho-flow

Universal Shell Installer (Linux & macOS):

curl -fsSL https://raw.githubusercontent.com/dixieflatline76/nacho-flow/main/scripts/install.sh | bash

Docker / Podman (Multi-Arch Distroless Container):

docker run -d -p 8000:8000 \
  -v $(pwd)/config.yaml:/config/config.yaml \
  ghcr.io/dixieflatline76/nacho-flow:latest

Homebrew (macOS & Linux):

brew install dixieflatline76/nacho-flow/nacho-flow

Pre-compiled Binaries: GitHub Releases

Go:

go install github.com/dixieflatline76/nacho-flow/cmd/nacho-flow@latest

Configuration (config.yaml)

Create a config.yaml in your project folder or ~/.config/nacho-flow/config.yaml:

port: 8000
host: "127.0.0.1"
auth_token: "sk-nacho-secret-key"

providers:
  ollama:
    base_url: "http://127.0.0.1:11434/v1"
    type: "local"

  openrouter:
    base_url: "https://openrouter.ai/api/v1"
    api_key: "ENV_OPENROUTER_API_KEY"
    headers:
      HTTP-Referer: "https://spicebox.dev"
      X-Title: "nacho-flow"

  langdock:
    base_url: "https://api.langdock.com/v1"
    api_key: "ENV_LANGDOCK_API_KEY"

  anthropic:
    base_url: "https://api.anthropic.com"
    type: "anthropic"
    api_key: "ENV_ANTHROPIC_API_KEY"

# Tiers evaluated top-to-bottom: first match wins
tiers:
  - name: "Cloud Reasoning"
    model: "deepseek/deepseek-r1"
    provider: "openrouter"
    when: "any(Keywords, { # in ['deadlock', 'mutex', 'race', 'concurrency', 'atomic'] })"

  - name: "Cloud Vision"
    model: "google/gemini-2.5-flash-lite"
    provider: "openrouter"
    when: "HasImages"

  - name: "Local GPU"
    model: "qwen3.8-coder:14b"
    provider: "ollama"
    max_context: 16384
    when: "Tokens < 16000 && !HasImages && !HasTools && Retries < 2"
    strip_images: true

  - name: "Cloud Agentic Fast"
    model: "qwen/qwen3-coder-30b-a3b-instruct"
    provider: "openrouter"
    when: "Tokens >= 16000 || HasTools || Retries >= 2"

default_tier:
  name: "Cloud Fallback"
  model: "deepseek/deepseek-v4-flash-latest"
  provider: "openrouter"
  when: "true"

cycle_killer:
  enabled: true
  max_tool_tokens: 8192
  repetition_threshold: 3
  kickstart_threshold: 5

fairy_dust:
  enabled: true
  entries:
    - name: "Tactical Code Review"
      frequency: 15
      provider: "openrouter"
      model: "anthropic/claude-sonnet-5"
    - name: "Strategic Architecture Review"
      frequency: 40
      provider: "openrouter"
      model: "anthropic/claude-sonnet-5"

In-Chat Control Directives (@nacho:)

Control routing and guardrails directly from your editor chat without restarting the gateway:

CategoryDirectiveAction
Session Switches@nacho:kickstart-off / onSuspend / resume Kickstart idle stall escalation
@nacho:cyclekiller-off / onSuspend / resume Cycle Killer stream loop breaker
@nacho:shield-off / onSuspend / resume synthetic tool-call synthesis
@nacho:raw-on / offEnable / disable raw upstream SSE stream
@nacho:fairydust-off / onSuspend / resume periodic frontier checkpoints
Inspection & Reset@nacho:togglesDisplay live session switches ($0.00 / 0 tokens)
@nacho:statusDisplay daemon telemetry, spend & saved dollars
@nacho:resetHard reset turn counter & restore default switches
Single-Turn Overrides@nacho:localForce current turn to Local GPU ($0.00)
@nacho:cloudForce current turn to Cloud Fallback tier
@nacho:reasoningForce current turn to DeepSeek-R1 / o1

Plan Mode Auto-Detection: When your agent switches into Plan Mode (zero write tools declared), Nacho Flow automatically detects HasWriteCapability == false and suspends Kickstart -- no manual toggles required.


Auto-Tune (nacho-flow tune)

Replays your actual logs/traffic.jsonl to find where your local GPU starts failing, prune dead tiers, and stop over-escalating to expensive frontier models:

# Advisory dry-run
nacho-flow tune

# Target specific VRAM ceiling
nacho-flow tune --vram-gb=16

# Apply recommendations with timestamped backup
nacho-flow tune --apply

Running as a Background Daemon

# Install as native OS service (Windows Service / systemd / launchd)
nacho-flow service install

# Start the background daemon
nacho-flow service start

Connect Your IDE

Nacho Flow exposes a standard OpenAI-compatible proxy at http://localhost:8000/v1. Use nacho-hybrid as your Model ID.

Zoo Code & Cline

  • API Provider: OpenAI Compatible
  • Base URL: http://localhost:8000/v1
  • API Key: sk-nacho-secret-key
  • Model ID: nacho-hybrid
  • Enable Supports Images, Supports Tools, Context Window 128,000, Max Output 8,192

OpenCode

export OPENAI_BASE_URL="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
opencode --model openai/nacho-hybrid

Aider

export OPENAI_API_BASE="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
aider --model openai/nacho-hybrid

Cursor

  • OpenAI API Base URL: http://localhost:8000/v1
  • OpenAI API Key: sk-nacho-secret-key
  • Add custom model: nacho-hybrid

Continue.dev

{
  "models": [{
    "title": "Nacho Flow (Hybrid Local + Cloud)",
    "provider": "openai",
    "model": "nacho-hybrid",
    "apiBase": "http://127.0.0.1:8000/v1",
    "apiKey": "sk-nacho-secret-key"
  }]
}

Documentation

Visit spicebox.dev/nacho-flow for the full documentation portal.


Support Nacho Flow

Nacho Flow is 100% free and open source. If it saved your sanity or your API bill:


Licensing

Dual-Licensing Model:

  1. Free & Open-Source (GNU AGPL-3.0 with API Interoperability Exception): Free for individual developers, open-source projects, and local evaluation. Calling Nacho Flow's OpenAI-compatible APIs from client apps or IDEs does not make your code a derivative work. Modifying and distributing or network-hosting Nacho Flow requires source disclosure under AGPL-3.0.

  2. Spicebox Commercial & Enterprise OEM License: For enterprise fleet deployments, closed-source embedding, commercial SaaS hosting, and IP indemnification. See COMMERCIAL_LICENSE.md or contact karl@spicebox.dev.

Contributions accepted under our Contributor License Agreement (.github/CLA.md).


Copyright © 2026 Karl Kwong / Spicebox · Licensed under GNU AGPL-3.0 with Commercial Dual-Licensing. (VS Code Extension licensed under MIT).

agent-supervisor
ai-coding
ai-coding-assistant
ai-coding-tools
cost-optimization
go
golang
llm
llms
ollama-client
openai-compatible
proxy
vscode-extension

dixieflatline76/nacho-flow

Cut AI coding agent bills 90%+ by routing between your local GPU and cloud APIs. OpenAI-compatible.

Go

2

202 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a pure Go supervisor proxy to make harnesses like Zoo, Cline and Aider work better (r/golang)

I built a Go supervisor proxy called [**Nacho Flow**](https://github.com/dixieflatline76/nacho-flow) to fix the things that frustrate me about using free coding agent harnesses with cheaper models. Free AI coding agent harnesses like Zoo and Cline work okay with frontier models like Sonnet and…

1

Oct 2, 2026

README

Nacho Flow

🌮 Nacho Flow

CI Status VS Code Extension Go Reference Go Version License: AGPL-3.0 Latest Release Homebrew Docker GHCR

Nacho Flow sits between your coding agent and your LLM backends. Routine turns run on your local GPU for $0.00. Complex reasoning escalates to frontier models automatically. Loop detection kills runaway agents in < 3s.

Nacho Flow is a hybrid model dispatcher and agent supervisor written in pure Go. Point Cline, Zoo Code, OpenCode, Aider, Cursor, or Continue at http://localhost:8000/v1 and it handles the rest.

🌐 Website & Documentation: spicebox.dev/nacho-flow 💸 Latest Research: The Frontier Tax: 1:1 Claude Sonnet 5 vs. Nacho Flow (5.7x cheaper, 12 minutes faster) Part of the spicebox.dev developer tool suite by @dixieflatline76.


What It Does

  • Hybrid routing: Local GPU -> budget cloud -> frontier, evaluated per-turn via deterministic AST bytecode rules (Tokens, Retries, Keywords, HasTools, HasImages)
  • Cycle Killer: Monitors live SSE streams across thinking/prose/tool lanes. Kills repetition loops in < 3s, injects a local $0.00 override to force tool action, never interrupts file writes
  • Kickstart: Detects consecutive non-write turns (analysis paralysis) and injects resuscitation prompts or escalates to a smarter model
  • Tool Normalizer: Converts 8 raw tool-call format families (Hermes, Mistral, Llama 3, ReAct, Markdown fences, bare JSON) into standard OpenAI tool_calls JSON -- no harness crashes
  • Fairy Dust: Programmable frontier checkpoints every N file writes for quality audits (without running expensive models all day)
  • Cost engine: Prompt-cache-aware billing tracking input, output, and cache tokens at exact provider rates
  • Auto-Tune: Replays your logs/traffic.jsonl to find optimal tier thresholds (nacho-flow tune)
  • VS Code extension: Bundles the Go binary, live financial telemetry, route inspector, one-click preset switching. Zero CLI setup. Install from Marketplace
  • Wire-speed core: < 0.19ms routing overhead, 30,000+ req/s (peak 30,284 req/s), lock-free atomic RCU state, zero heap churn during proxying
  • Zero runtime dependencies: Single static binary, CGO_ENABLED=0, no Node or Python required
  • 🧪 Engineered for Reliability: Strictly $\ge 95.0%\text{--}100%$ statement test coverage across all packages (95.4% global coverage), 100% race-detector clean (-race), and static security audited (gosec).

NTS (experimental token compaction) is available but disabled by default to preserve prompt cache stability and prevent diff drift. See NTS docs if you need it for extreme context limits.


Quickstart (3 Steps)

flowchart LR
    Step1["**Step 1: Local GPU ($0.00)**<br/><code>ollama run gemma4:12b-it-qat</code>"]
    Step2["**Step 2: Cloud Gateway**<br/>OpenRouter Key (1 Unified Key)"]
    Step3["**Step 3: Point Your Agent**<br/><code>http://127.0.0.1:8000/v1</code>"]

    Step1 --> Step2 --> Step3
  1. Local GPU ($0.00) -- Install Ollama and pull a coding model:

    ollama run gemma4:12b-it-qat
    

    Set OLLAMA_CONTEXT_LENGTH=32768 to prevent prompt truncation (setx OLLAMA_CONTEXT_LENGTH 32768 on Windows). See User Guide.

  2. Cloud Gateway via OpenRouter -- One key for all frontier tiers (Qwen 3 Coder Plus, Gemini 3.8 Flash, Claude Sonnet 5):

    export OPENROUTER_API_KEY="sk-or-v1-..."
    
  3. Start Nacho Flow & Connect Your Agent

    • VS Code Extension: Install from Marketplace, click Start in the sidebar.
    • Standalone CLI: Run nacho-flow in your terminal.
    • Set your agent's Base URL to http://127.0.0.1:8000/v1, Model ID to nacho-hybrid.

Installation

VS Code Extension (All-in-One Runtime):

code --install-extension dixieflatline76.nacho-flow

Universal Shell Installer (Linux & macOS):

curl -fsSL https://raw.githubusercontent.com/dixieflatline76/nacho-flow/main/scripts/install.sh | bash

Docker / Podman (Multi-Arch Distroless Container):

docker run -d -p 8000:8000 \
  -v $(pwd)/config.yaml:/config/config.yaml \
  ghcr.io/dixieflatline76/nacho-flow:latest

Homebrew (macOS & Linux):

brew install dixieflatline76/nacho-flow/nacho-flow

Pre-compiled Binaries: GitHub Releases

Go:

go install github.com/dixieflatline76/nacho-flow/cmd/nacho-flow@latest

Configuration (config.yaml)

Create a config.yaml in your project folder or ~/.config/nacho-flow/config.yaml:

port: 8000
host: "127.0.0.1"
auth_token: "sk-nacho-secret-key"

providers:
  ollama:
    base_url: "http://127.0.0.1:11434/v1"
    type: "local"

  openrouter:
    base_url: "https://openrouter.ai/api/v1"
    api_key: "ENV_OPENROUTER_API_KEY"
    headers:
      HTTP-Referer: "https://spicebox.dev"
      X-Title: "nacho-flow"

  langdock:
    base_url: "https://api.langdock.com/v1"
    api_key: "ENV_LANGDOCK_API_KEY"

  anthropic:
    base_url: "https://api.anthropic.com"
    type: "anthropic"
    api_key: "ENV_ANTHROPIC_API_KEY"

# Tiers evaluated top-to-bottom: first match wins
tiers:
  - name: "Cloud Reasoning"
    model: "deepseek/deepseek-r1"
    provider: "openrouter"
    when: "any(Keywords, { # in ['deadlock', 'mutex', 'race', 'concurrency', 'atomic'] })"

  - name: "Cloud Vision"
    model: "google/gemini-2.5-flash-lite"
    provider: "openrouter"
    when: "HasImages"

  - name: "Local GPU"
    model: "qwen3.8-coder:14b"
    provider: "ollama"
    max_context: 16384
    when: "Tokens < 16000 && !HasImages && !HasTools && Retries < 2"
    strip_images: true

  - name: "Cloud Agentic Fast"
    model: "qwen/qwen3-coder-30b-a3b-instruct"
    provider: "openrouter"
    when: "Tokens >= 16000 || HasTools || Retries >= 2"

default_tier:
  name: "Cloud Fallback"
  model: "deepseek/deepseek-v4-flash-latest"
  provider: "openrouter"
  when: "true"

cycle_killer:
  enabled: true
  max_tool_tokens: 8192
  repetition_threshold: 3
  kickstart_threshold: 5

fairy_dust:
  enabled: true
  entries:
    - name: "Tactical Code Review"
      frequency: 15
      provider: "openrouter"
      model: "anthropic/claude-sonnet-5"
    - name: "Strategic Architecture Review"
      frequency: 40
      provider: "openrouter"
      model: "anthropic/claude-sonnet-5"

In-Chat Control Directives (@nacho:)

Control routing and guardrails directly from your editor chat without restarting the gateway:

CategoryDirectiveAction
Session Switches@nacho:kickstart-off / onSuspend / resume Kickstart idle stall escalation
@nacho:cyclekiller-off / onSuspend / resume Cycle Killer stream loop breaker
@nacho:shield-off / onSuspend / resume synthetic tool-call synthesis
@nacho:raw-on / offEnable / disable raw upstream SSE stream
@nacho:fairydust-off / onSuspend / resume periodic frontier checkpoints
Inspection & Reset@nacho:togglesDisplay live session switches ($0.00 / 0 tokens)
@nacho:statusDisplay daemon telemetry, spend & saved dollars
@nacho:resetHard reset turn counter & restore default switches
Single-Turn Overrides@nacho:localForce current turn to Local GPU ($0.00)
@nacho:cloudForce current turn to Cloud Fallback tier
@nacho:reasoningForce current turn to DeepSeek-R1 / o1

Plan Mode Auto-Detection: When your agent switches into Plan Mode (zero write tools declared), Nacho Flow automatically detects HasWriteCapability == false and suspends Kickstart -- no manual toggles required.


Auto-Tune (nacho-flow tune)

Replays your actual logs/traffic.jsonl to find where your local GPU starts failing, prune dead tiers, and stop over-escalating to expensive frontier models:

# Advisory dry-run
nacho-flow tune

# Target specific VRAM ceiling
nacho-flow tune --vram-gb=16

# Apply recommendations with timestamped backup
nacho-flow tune --apply

Running as a Background Daemon

# Install as native OS service (Windows Service / systemd / launchd)
nacho-flow service install

# Start the background daemon
nacho-flow service start

Connect Your IDE

Nacho Flow exposes a standard OpenAI-compatible proxy at http://localhost:8000/v1. Use nacho-hybrid as your Model ID.

Zoo Code & Cline

  • API Provider: OpenAI Compatible
  • Base URL: http://localhost:8000/v1
  • API Key: sk-nacho-secret-key
  • Model ID: nacho-hybrid
  • Enable Supports Images, Supports Tools, Context Window 128,000, Max Output 8,192

OpenCode

export OPENAI_BASE_URL="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
opencode --model openai/nacho-hybrid

Aider

export OPENAI_API_BASE="http://127.0.0.1:8000/v1"
export OPENAI_API_KEY="sk-nacho-secret-key"
aider --model openai/nacho-hybrid

Cursor

  • OpenAI API Base URL: http://localhost:8000/v1
  • OpenAI API Key: sk-nacho-secret-key
  • Add custom model: nacho-hybrid

Continue.dev

{
  "models": [{
    "title": "Nacho Flow (Hybrid Local + Cloud)",
    "provider": "openai",
    "model": "nacho-hybrid",
    "apiBase": "http://127.0.0.1:8000/v1",
    "apiKey": "sk-nacho-secret-key"
  }]
}

Documentation

Visit spicebox.dev/nacho-flow for the full documentation portal.


Support Nacho Flow

Nacho Flow is 100% free and open source. If it saved your sanity or your API bill:


Licensing

Dual-Licensing Model:

  1. Free & Open-Source (GNU AGPL-3.0 with API Interoperability Exception): Free for individual developers, open-source projects, and local evaluation. Calling Nacho Flow's OpenAI-compatible APIs from client apps or IDEs does not make your code a derivative work. Modifying and distributing or network-hosting Nacho Flow requires source disclosure under AGPL-3.0.

  2. Spicebox Commercial & Enterprise OEM License: For enterprise fleet deployments, closed-source embedding, commercial SaaS hosting, and IP indemnification. See COMMERCIAL_LICENSE.md or contact karl@spicebox.dev.

Contributions accepted under our Contributor License Agreement (.github/CLA.md).


Copyright © 2026 Karl Kwong / Spicebox · Licensed under GNU AGPL-3.0 with Commercial Dual-Licensing. (VS Code Extension licensed under MIT).

agent-supervisor
ai-coding
ai-coding-assistant
ai-coding-tools
cost-optimization
go
golang
llm
llms
ollama-client
openai-compatible
proxy
vscode-extension

Languages

Go

72.7%

TypeScript

12.4%

HTML

6.5%

CSS

5.0%

JavaScript

2.7%