The agent-aware model router built for hybrid inference
Python
0
7 commits
updated Oct 5, 2026
Cut your AI bill by a 30%+, keep personal data private, and seamlessly run agents across private and public inference.
Recursant sits between your AI agents and the AI models they use. Every time an agent asks for its next step, Recursant picks who answers: a top model for the hard steps, a cheaper model for the routine ones, and your own private model whenever personal data is involved. Your agent doesn't change. You point it at Recursant instead of at the model provider, and Recursant does the rest.
You need to run Recursant. macOS support is underway, Windows is tricky.
On Linux or macOS, run:
curl -fsSL https://raw.githubusercontent.com/ajensenwaud/recursant/main/install.sh | bash
This downloads the source code, installs anything needed to build it (through your package manager, so it may ask for your password), builds Recursant, and writes a starter setup. It doesn't start anything and doesn't send anything anywhere. The script is short if you want to read it first.
You end up with:
~/.local/bin/recursant~/.config/recursant/config.json~/.config/recursant/recursant.env, readable only by you. This includes a key for your agents to use, created for you.If your shell says recursant: command not found, add ~/.local/bin to your path: export PATH="$HOME/.local/bin:$PATH".
The starter setup uses:
| Name | Model | Used for |
|---|---|---|
baseline | Claude Sonnet 5.5 (via OpenRouter) | the main model for anything not routine |
economy | GPT-6 luna (via OpenRouter) | routine steps and simple questions |
strong | Claude Opus 5.5 (via OpenRouter) | when the agent keeps getting stuck |
local | qwen3:8b on Ollama, on this machine | anything containing personal data |
1. Add your OpenRouter key. Recursant asks for it and doesn't show it as you type:
recursant configure --set-key OPENROUTER_API_KEY
2. Point it at your private model. The starter setup expects Ollama on this machine. If your private model runs somewhere else, add it and make it the place personal data goes. For example, for a GPU server called gpu-box:
recursant configure --add-provider gpu --url http://gpu-box:8000/v1 --trust private --private-default gpu:my-model-name
3. Check everything.
recursant check
It tells you in plain words what's missing or wrong.
Prefer menus? Run recursant configure on its own for step-by-step setup of model services, models, privacy patterns, network and keys. Every change is checked before it's saved, and your previous settings are kept as a dated backup next to the file.
Other things you can change:
recursant configure --alias economy=openrouter:openai/gpt-6-luna # use a different cheap model
recursant configure --add-pattern 'CUST-[0-9]{6}' # treat your own IDs as private
recursant configure --listen tailnet # reachable from your Tailscale network
recursant configure --list-providers # 28 known model services
So far Recursant has been tested live with OpenRouter and with models on your own servers (vLLM, Ollama and similar). The other services in the list should work the same way, but haven't been tested against their real systems yet.
Run it as a background service in systemd that starts with your computer:
recursant install --user
recursant start
recursant status
status shows whether it's running, its address, how many requests it has handled, and where they went.
sudo recursant install --system and sudo recursant start instead.recursant serve, and press Ctrl-C to stop it.Your agent needs three settings:
| Setting | Value |
|---|---|
| Address (base URL) | http://127.0.0.1:8080/v1 |
| Key | RECURSANT_API_KEY from ~/.config/recursant/recursant.env |
| Model | auto (let Recursant choose) |
Hermes: run hermes model, choose Custom endpoint (enter URL manually), and enter the three values above.
pi: add Recursant to ~/.pi/agent/models.json:
{
"providers": {
"recursant": {
"baseUrl": "http://127.0.0.1:8080/v1",
"api": "openai-completions",
"apiKey": "${RECURSANT_API_KEY}",
"models": [{ "id": "auto", "contextWindow": 131072, "maxTokens": 8192 }]
}
}
}
Then export your key (export RECURSANT_API_KEY=...) and run pi --provider recursant --model auto.
Anything else: wherever the tool asks for an OpenAI address, key and model, enter the values above. To see it working from the command line:
source ~/.config/recursant/recursant.env
curl -s http://127.0.0.1:8080/v1/chat/completions \
-H "Authorization: Bearer $RECURSANT_API_KEY" -H 'Content-Type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"Say hello"}]}' \
-D - -o /dev/null | grep -i '^x-recursant'
Every answer says which model handled it and why (in the X-Recursant-Model and X-Recursant-Decision headers), so you can always see what happened. To choose a model yourself, ask for baseline, economy, strong or local instead of auto.
| Command | What it does |
|---|---|
recursant status | Is it running, where, and what has it done |
recursant check | Check your settings before restarting |
recursant restart | Load new settings. If they're broken, the running copy carries on untouched |
recursant start / stop | Start or stop the service |
recursant configure | Change settings, with menus or the switches above |
recursant install / uninstall | Add or remove the background service. uninstall --purge also deletes your settings and keys |
recursant serve | Run in this window instead of in the background |
Logs go to the system journal: journalctl --user -u recursant, or sudo journalctl -u recursant for the whole-machine service. Logs record decisions only, never what your agents wrote.
For every request, Recursant asks four questions in this order:
Recursant decides all this itself, in a few milliseconds, from the request the agent already sent. There's no second AI model making the call, and no extra service to run.
recursant.env, readable only by you.--listen). If other people will reach it, put it behind a secure proxy.Answering 980 general-knowledge exam questions (MMLU-Pro), one request each:
| Model | Correct | Cost |
|---|---|---|
| GLM-5.3-Flash on your own GPU | 80.1% | US$0 |
| GPT-6 luna (cheaper, via OpenRouter) | 84.6% | US$0.15 |
| GPT-6.1 sol (top, via OpenRouter) | 88.3% | US$1.53 |
The top model does better on hard questions, but nothing we tried could tell in advance which questions those are. So for single questions Recursant uses the cheapest model you allow, and the trade-off is your choice (details). We also tested small "decision models" (Jev, Strands Decider) as advisers. They didn't save money reliably, so they're switched off (details).
docker build -f deploy/Dockerfile -t recursant:local .
docker run --rm --user "$(id -u):$(id -g)" --read-only --cap-drop ALL --security-opt no-new-privileges \
-p 127.0.0.1:8080:8080 -v "$HOME/.config/recursant/config.json:/etc/recursant/config.json:ro" \
--env-file "$HOME/.config/recursant/recursant.env" recursant:local
| Part | State |
|---|---|
| One address for your own models and paid services | Done, tested live |
| Privacy rules: personal data stays on your machines | Done, including Australian ID numbers with check digits |
| Step-by-step model choice for agents | Done: 29–39% cheaper at the same quality in our tests |
| Web console, audit history, tracing | Planned |
The agent-aware model router built for hybrid inference
Python
0
7 commits
updated Oct 5, 2026
Cut your AI bill by a 30%+, keep personal data private, and seamlessly run agents across private and public inference.
Recursant sits between your AI agents and the AI models they use. Every time an agent asks for its next step, Recursant picks who answers: a top model for the hard steps, a cheaper model for the routine ones, and your own private model whenever personal data is involved. Your agent doesn't change. You point it at Recursant instead of at the model provider, and Recursant does the rest.
You need to run Recursant. macOS support is underway, Windows is tricky.
On Linux or macOS, run:
curl -fsSL https://raw.githubusercontent.com/ajensenwaud/recursant/main/install.sh | bash
This downloads the source code, installs anything needed to build it (through your package manager, so it may ask for your password), builds Recursant, and writes a starter setup. It doesn't start anything and doesn't send anything anywhere. The script is short if you want to read it first.
You end up with:
~/.local/bin/recursant~/.config/recursant/config.json~/.config/recursant/recursant.env, readable only by you. This includes a key for your agents to use, created for you.If your shell says recursant: command not found, add ~/.local/bin to your path: export PATH="$HOME/.local/bin:$PATH".
The starter setup uses:
| Name | Model | Used for |
|---|---|---|
baseline | Claude Sonnet 5.5 (via OpenRouter) | the main model for anything not routine |
economy | GPT-6 luna (via OpenRouter) | routine steps and simple questions |
strong | Claude Opus 5.5 (via OpenRouter) | when the agent keeps getting stuck |
local | qwen3:8b on Ollama, on this machine | anything containing personal data |
1. Add your OpenRouter key. Recursant asks for it and doesn't show it as you type:
recursant configure --set-key OPENROUTER_API_KEY
2. Point it at your private model. The starter setup expects Ollama on this machine. If your private model runs somewhere else, add it and make it the place personal data goes. For example, for a GPU server called gpu-box:
recursant configure --add-provider gpu --url http://gpu-box:8000/v1 --trust private --private-default gpu:my-model-name
3. Check everything.
recursant check
It tells you in plain words what's missing or wrong.
Prefer menus? Run recursant configure on its own for step-by-step setup of model services, models, privacy patterns, network and keys. Every change is checked before it's saved, and your previous settings are kept as a dated backup next to the file.
Other things you can change:
recursant configure --alias economy=openrouter:openai/gpt-6-luna # use a different cheap model
recursant configure --add-pattern 'CUST-[0-9]{6}' # treat your own IDs as private
recursant configure --listen tailnet # reachable from your Tailscale network
recursant configure --list-providers # 28 known model services
So far Recursant has been tested live with OpenRouter and with models on your own servers (vLLM, Ollama and similar). The other services in the list should work the same way, but haven't been tested against their real systems yet.
Run it as a background service in systemd that starts with your computer:
recursant install --user
recursant start
recursant status
status shows whether it's running, its address, how many requests it has handled, and where they went.
sudo recursant install --system and sudo recursant start instead.recursant serve, and press Ctrl-C to stop it.Your agent needs three settings:
| Setting | Value |
|---|---|
| Address (base URL) | http://127.0.0.1:8080/v1 |
| Key | RECURSANT_API_KEY from ~/.config/recursant/recursant.env |
| Model | auto (let Recursant choose) |
Hermes: run hermes model, choose Custom endpoint (enter URL manually), and enter the three values above.
pi: add Recursant to ~/.pi/agent/models.json:
{
"providers": {
"recursant": {
"baseUrl": "http://127.0.0.1:8080/v1",
"api": "openai-completions",
"apiKey": "${RECURSANT_API_KEY}",
"models": [{ "id": "auto", "contextWindow": 131072, "maxTokens": 8192 }]
}
}
}
Then export your key (export RECURSANT_API_KEY=...) and run pi --provider recursant --model auto.
Anything else: wherever the tool asks for an OpenAI address, key and model, enter the values above. To see it working from the command line:
source ~/.config/recursant/recursant.env
curl -s http://127.0.0.1:8080/v1/chat/completions \
-H "Authorization: Bearer $RECURSANT_API_KEY" -H 'Content-Type: application/json' \
-d '{"model":"auto","messages":[{"role":"user","content":"Say hello"}]}' \
-D - -o /dev/null | grep -i '^x-recursant'
Every answer says which model handled it and why (in the X-Recursant-Model and X-Recursant-Decision headers), so you can always see what happened. To choose a model yourself, ask for baseline, economy, strong or local instead of auto.
| Command | What it does |
|---|---|
recursant status | Is it running, where, and what has it done |
recursant check | Check your settings before restarting |
recursant restart | Load new settings. If they're broken, the running copy carries on untouched |
recursant start / stop | Start or stop the service |
recursant configure | Change settings, with menus or the switches above |
recursant install / uninstall | Add or remove the background service. uninstall --purge also deletes your settings and keys |
recursant serve | Run in this window instead of in the background |
Logs go to the system journal: journalctl --user -u recursant, or sudo journalctl -u recursant for the whole-machine service. Logs record decisions only, never what your agents wrote.
For every request, Recursant asks four questions in this order:
Recursant decides all this itself, in a few milliseconds, from the request the agent already sent. There's no second AI model making the call, and no extra service to run.
recursant.env, readable only by you.--listen). If other people will reach it, put it behind a secure proxy.Answering 980 general-knowledge exam questions (MMLU-Pro), one request each:
| Model | Correct | Cost |
|---|---|---|
| GLM-5.3-Flash on your own GPU | 80.1% | US$0 |
| GPT-6 luna (cheaper, via OpenRouter) | 84.6% | US$0.15 |
| GPT-6.1 sol (top, via OpenRouter) | 88.3% | US$1.53 |
The top model does better on hard questions, but nothing we tried could tell in advance which questions those are. So for single questions Recursant uses the cheapest model you allow, and the trade-off is your choice (details). We also tested small "decision models" (Jev, Strands Decider) as advisers. They didn't save money reliably, so they're switched off (details).
docker build -f deploy/Dockerfile -t recursant:local .
docker run --rm --user "$(id -u):$(id -g)" --read-only --cap-drop ALL --security-opt no-new-privileges \
-p 127.0.0.1:8080:8080 -v "$HOME/.config/recursant/config.json:/etc/recursant/config.json:ro" \
--env-file "$HOME/.config/recursant/recursant.env" recursant:local
| Part | State |
|---|---|
| One address for your own models and paid services | Done, tested live |
| Privacy rules: personal data stays on your machines | Done, including Australian ID numbers with check digits |
| Step-by-step model choice for agents | Done: 29–39% cheaper at the same quality in our tests |
| Web console, audit history, tracing | Planned |