DustinVerzal/magic-router

Routes each Claude Code session to Sonnet or Opus, and each prompt to an effort level, with a classifier that runs locally.

Python

0

60 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

Let Claude Code pick its own model/effort level (locally, no API key) (r/ClaudeAI)

I kept making the same mistake in both directions: burning Opus at xhigh on "fix this typo", then forgetting to bump effort when I asked for a real design. Switching by hand with `/model` and `/effort` got old, and switching models mid-conversation throws away the prompt cache anyway. All the Jev…

0

Oct 4, 2026

README

magic-router

Sonnet or Opus for the session. An effort level for every prompt.
Picked by a classifier that runs on your machine.

version Claude Code 2.1.287+ macOS · Linux MIT license


One session with the router band under each prompt. The first prompt, about designing a sharded job queue, scores 1.86 and picks Opus 5.5 at xhigh effort in 326 ms. A short 'go ahead' keeps xhigh. A prompt to add retry with backoff scores 0.49 and drops to medium, and fixing a typo scores 0.21 and drops to low, while the model stays on Opus 5.5.

One session. Scores and latencies are real classifier output for these prompts.


Install · How it works · Privacy · Tuning · Troubleshooting


A typo fix doesn't need Opus at xhigh, and a distributed-systems design shouldn't get Sonnet at low. magic-router is a Claude Code plugin that makes that call for you: it routes each session to Sonnet 5.5 or Opus 5.5 and each prompt to an effort level, using GLiNER2.5-Decide running locally.

  • The model is picked once, from the session's first prompt, and then stays put: switching models mid-conversation throws away the prompt cache.
  • Effort is picked on every prompt. Short follow-ups ("yes", "go ahead") keep the last effort.
  • It runs on your machine. About 0.3 s of CPU per prompt, and no prompt goes to a third party to be routed.
  • You can see every decision. A band above the prompt shows the model, the effort, how long classifying took, and last turn's cache-read %.
  • It routes subagents too. Each subagent gets its own model and effort from the task it was given. A model the caller or the agent's definition named is kept.
  • It gets out of your way for the rest of the session once you pick a model or effort by hand (details).
  • It learns how you work. Optionally, train it on your own prompt history: on held-out prompts, a tuned adapter matched the labelled effort 66% of the time, against 40% for the default score.

Install

Requirements

  • Claude Code 2.1.287 or later (claude --version)
  • macOS or Linux, with curl
  • About 2 GB of free RAM while the classifier runs, and about 2 GB of disk for torch and the weights

Python and uv are optional: if uv is missing, the first session installs it to ~/.local/bin, and uv fetches Python 3.10–3.13 itself.

Add the plugin

In Claude Code:

/plugin marketplace add DustinVerzal/magic-router
/plugin install magic-router@magic-router

Restart Claude Code. There's nothing to configure.

First run

The first session starts the classifier daemon in the background. It keeps running after the session ends, and every later session reuses it.

WhenWhat happens
First run everuv installs (if missing), torch and gliner2 install, and the weights download (~1.7 GB). Until then the band shows classifier unreachable and the session keeps its own model.
First session after a rebootThe model loads in about 10 s, and the first prompt waits for it.
Every later promptAbout 0.3 s of CPU.
Warm it up first: skip the first-run download wait
git clone https://github.com/DustinVerzal/magic-router
magic-router/scripts/gliner.sh setup

setup installs uv if it's missing, installs torch and gliner2, downloads the weights, and leaves the daemon running. The downloads land in uv's and Hugging Face's shared caches, so the installed plugin reuses them; the clone is only needed to run the script.

From source: load a checkout to hack on it
git clone https://github.com/DustinVerzal/magic-router
claude --plugin-dir ./magic-router

This loads the checkout for one session instead of installing it. Edits to hooks/ reload while the session runs.

How it works

All routing policy lives in hooks/route.ts. The daemon only answers the questions the plugin sends it.

  1. Task type: one probability distribution over seven labels, each mapped to an eval family of the Artificial Analysis Intelligence Index v4.1.
  2. Complexity: seven independent yes/no signals, such as multi_file, planning, deep_reasoning, large_scope, and quick.
  3. Score: Σ weight × P(signal) + Σ bias × P(task). The task bias leans toward Opus where its lead on the matching evals is widest.
  4. Effort and model: the score maps to low < 0.4 ≤ medium < 1.0 ≤ high < 1.5 ≤ xhigh < 2.0 ≤ max. On the first prompt, a score ≥ 1.1 picks Opus; anything lower picks Sonnet.
The score scale. Effort bands run low below 0.4, medium to 1.0, high to 1.5, xhigh to 2.0, and max above. A dashed line at 1.1 splits Sonnet from Opus. The session's prompts sit at 0.21 (fix the typo, low), 0.49 (add retry and tests, medium), and 1.86 (design the job queue, xhigh).
Task labels and the evals behind them
LabelEval family
agentic_codingTerminal-Bench
scientific_codingSciCode
tool_useτ³-Bench
knowledge_workGDPval-AA
long_contextAA-LCR
knowledge_qaAA-Omniscience, GPQA
reasoningHLE, CritPt

[!NOTE] The default weights and thresholds were tuned by eye on 15 prompts. To fit them to how you work, see Tune it to your prompts.

When it stands aside

  • Subagents: each one is classified once, on its first request, and keeps that route. Only what it would inherit from the session is rewritten: a model or effort set by the Agent call or the agent's definition is kept. Forks are left alone, because they share their parent's context and prompt cache.
  • /model, /effort, or a fallback: if any of these changes what the engine asks for, the router stops rewriting for the rest of the session.
  • Classifier unreachable on the first prompt: the session keeps its own model, and effort routing starts once the daemon answers.
  • /clear: the next prompt picks a model again.

Privacy

  • Routing never leaves your machine. Each prompt's first 2,000 characters go to the daemon on 127.0.0.1 and nowhere else (or to your own ssh host, if you run the daemon there).
  • One background daemon, on port 8765, shared by all your sessions. It outlives them; stop it any time.
  • Files stay in ~/.cache/magic-router, plus the weights in Hugging Face's cache, downloaded once.
  • uv comes from Astral's installer (curl … | sh, to ~/.local/bin, no shell profile edits) if it's missing.

Only the opt-in tools reach further: just tune reads your local transcripts and sends prompts to Claude through claude -p (and to your ssh host, if you set one), and just bench sends prompts to the hosted models whose keys you set.

Configuration

Nothing needs configuring. These environment variables change the defaults:

VariableDefaultEffect
ROUTER_PORT8765The daemon's port. Set it for Claude Code and for scripts/gliner.sh.
ROUTER_MODELfastino/GLiNER2.5-DecideThe Hugging Face model the daemon loads.
ROUTER_WAIT900Seconds gliner.sh start waits for the model on a first run.
TUNE_HOSTnoneAn ssh host with an NVIDIA GPU to tune on.
AA_API_KEYnoneArtificial Analysis key for just benchmarks.
OPENROUTER_KEY, FASTINO_API_KEYnoneHosted-model keys for just bench.

Weights and thresholds are constants in hooks/route.ts. Fable 5.1 is wired in but off: set FABLE_AT there to a score above OPUS_AT to send the hardest first prompts to it.

Managing the classifier

These need a checkout of this repo:

scripts/gliner.sh setup         # install uv if missing, install torch + gliner2, download the weights, start
scripts/gliner.sh start         # start the daemon (if it isn't up) and wait until the model is loaded
scripts/gliner.sh stop          # stop the daemon
scripts/gliner.sh status        # print /health
scripts/gliner.sh check         # run the classifier's offline self-check
scripts/gliner.sh logs          # follow the daemon log
scripts/gliner.sh remote HOST   # run the daemon on an ssh host instead (macOS), such as a box with a GPU
scripts/gliner.sh local         # run it here again

Without a checkout, stop the daemon with pkill -f server/classifier.py. The log is at ~/.cache/magic-router/classifier.log.

Running the daemon on another machine

remote HOST copies this checkout and your tuned adapter to ~/.cache/magic-router on the host, starts the daemon there (on its NVIDIA GPU if it has one), and adds a launch agent that forwards port 8765 to it over ssh. The plugin keeps calling 127.0.0.1:8765 and needs no change, and your Mac no longer runs the model. After that, setup, start, stop, check and logs act on the host.

The host needs key-based ssh and rsync. If it's down when a session starts, the plugin starts a local daemon, which then holds the port until you stop it.

Tune it to your prompts

just tune trains the classifier on how you actually work. It's opt-in and needs a checkout. Labelling runs through claude -p, so it counts against your Claude subscription's usage, or bills your API key if that's how you're signed in. Run just benchmarks first: the labeller uses its scores.

  1. Collect. It reads every prompt you've sent Claude Code or Codex, from ~/.claude/projects, ~/.claude/history.jsonl and ~/.codex. Slash commands, $skills and short follow-ups are skipped, because the router never classifies them, and so are subagents, codex exec runs and orchestrators. A session's first prompt is kept however short, because it picks the model. just tune 500 uses only your newest 500.
  2. Label. Opus 5.5 at xhigh decides which model and effort each prompt needed, given both models' Artificial Analysis scores and response times at every effort, and is asked for the best answer without overthinking, not the cheapest.
  3. Train. It trains a LoRA adapter on 70% of your sessions, on CPU, or on TUNE_HOST if set. Sessions are split whole, so related prompts never land on both sides. Effort is learned from every prompt. The model is learned only from each session's first prompt, labelled with what the whole session needed: a session that opens with a typo fix and turns into design work counts as Opus, because the model picked on that first prompt has to last the session.
  4. Gate. On the other 30%, the adapter must beat both the score and any adapter already installed on effort, and pick the model no worse than the score on those sessions' first prompts. Only then does it install to ~/.cache/magic-router/tuned and restart the daemon; otherwise nothing changes.

Results from the author's history (3,344 prompts over six months), scored on 972 held-out prompts:

Matches labelMean levels off
Adapter trained on 3,344 prompts66%0.36
Adapter trained on 404 Claude Code prompts54%0.49
Always medium44%0.59
The default score40%0.69

Labelling cost $10.33 at API prices. Training took about 15 minutes on an RTX 4080; on CPU it runs about 15 minutes per 400 prompts. Your numbers will differ.

With an adapter installed, the daemon answers effort and model from it, which adds a second pass of about 0.1 s. The model is still picked only on a session's first prompt. An adapter trained before model labelling existed answers effort only; retune to get both. To undo, delete ~/.cache/magic-router/tuned and run scripts/gliner.sh stop.

What tuning sends and stores
  • Labelling sends 25 prompts per claude -p call, with no tools, no settings and no saved session.
  • Each label and a one-line reason go to ~/.cache/magic-router/tune/labels.jsonl. A rerun only labels new prompts, and any labelled under an older labeller prompt. Skim them there.
  • With TUNE_HOST, training and scoring run on that host (it needs an NVIDIA GPU, uv and rsync), and the prompts sent there are deleted afterwards.
Refreshing the benchmarks: just benchmarks

just benchmarks pulls the Artificial Analysis evals for Sonnet 5.5, Opus 5.5 and Fable 5.1 at every effort level, prints them as a table, and caches the raw JSON in ~/.cache/magic-router. With no arguments it also rewrites TASK_BIAS in hooks/route.ts from Opus's per-family lead over Sonnet; model names as arguments only filter the printout. Put AA_API_KEY=... in a git-ignored .env, or export it.

Against hosted decision models: just bench

just bench scores the trained adapter against hosted decision models on the same 972 held-out prompts, without retraining, sending each prompt's first 2,000 characters to the provider. OPENROUTER_KEY adds Jev (TypeSafe) and FASTINO_API_KEY adds GLiDE. Jev got the labeller's definition of each level, and was asked for the effort directly, both as a choice among the five levels and as a score on the ordered scale. A third run put Jev's answers to the router's own task and signal questions through the score's weights.

Matches labelMean levels off
Tuned adapter66%0.36
Jev, effort as a score63%0.42
Jev, effort as a choice56%0.53
Always medium44%0.59
The default score40%0.69
Jev through the score's weights27%1.06

Asked directly, Jev comes within 3 points of the adapter without seeing any of your prompts. The adapter stays ahead, though it was trained on labels from the same labeller, so the margin favours it. It also runs locally, while Jev is a hosted call (median 234 ms, $0.049 for all 972 prompts) that sends your prompts to a third party. Through the score's weights Jev does worse than always guessing medium, because those weights were set for GLiNER, so it can't drop into the router as is.

Troubleshooting

SymptomFix
Band says classifier unreachableRun scripts/gliner.sh start and read the log it names. On a first run, it's usually still downloading.
Daemon never starts from Claude Code, but gliner.sh start worksRead ~/.cache/magic-router/classifier.log: the uv install or a torch download likely failed (offline, proxy).
Band says stood asideYou picked a model or effort by hand (/model, /effort). /clear to let the router pick again.
Port 8765 is taken by something elseSet ROUTER_PORT to a free port for Claude Code and the shell you run gliner.sh from, then restart Claude Code.

Updating and uninstalling

A new version ships with every merge to main. Third-party marketplaces don't auto-update by default, so turn it on once: run /plugin, open Marketplaces, select magic-router and choose Enable auto-update. Claude Code then pulls new versions at startup and asks you to restart or run /reload-plugins. To update by hand:

/plugin marketplace update magic-router
/plugin update magic-router@magic-router

To uninstall:

/plugin uninstall magic-router@magic-router
/plugin marketplace remove magic-router

Then stop the daemon (pkill -f server/classifier.py) and, to reclaim the disk, delete ~/.cache/magic-router and ~/.cache/huggingface/hub/models--fastino--GLiNER2.5-Decide.

Contributing

Issues and pull requests are welcome. AGENTS.md covers the layout, the gotchas, and the conventions for both humans and coding agents. Before opening a PR, run what CI runs:

claude plugin validate .
claude plugin test .                           # routing, stickiness, stand-aside, band
uv run --script server/classifier.py --check   # classifier self-check
shellcheck scripts/*.sh

Acknowledgements

GLiNER2.5-Decide by Fastino does the classifying. Artificial Analysis supplies the evals that set the task bias and inform the labeller.

License

MIT

anthropic
claude-code
claude-code-plugin
gliner
llm-routing
model-routing

DustinVerzal/magic-router

Routes each Claude Code session to Sonnet or Opus, and each prompt to an effort level, with a classifier that runs locally.

Python

0

60 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

Let Claude Code pick its own model/effort level (locally, no API key) (r/ClaudeAI)

I kept making the same mistake in both directions: burning Opus at xhigh on "fix this typo", then forgetting to bump effort when I asked for a real design. Switching by hand with `/model` and `/effort` got old, and switching models mid-conversation throws away the prompt cache anyway. All the Jev…

0

Oct 4, 2026

README

magic-router

Sonnet or Opus for the session. An effort level for every prompt.
Picked by a classifier that runs on your machine.

version Claude Code 2.1.287+ macOS · Linux MIT license


One session with the router band under each prompt. The first prompt, about designing a sharded job queue, scores 1.86 and picks Opus 5.5 at xhigh effort in 326 ms. A short 'go ahead' keeps xhigh. A prompt to add retry with backoff scores 0.49 and drops to medium, and fixing a typo scores 0.21 and drops to low, while the model stays on Opus 5.5.

One session. Scores and latencies are real classifier output for these prompts.


Install · How it works · Privacy · Tuning · Troubleshooting


A typo fix doesn't need Opus at xhigh, and a distributed-systems design shouldn't get Sonnet at low. magic-router is a Claude Code plugin that makes that call for you: it routes each session to Sonnet 5.5 or Opus 5.5 and each prompt to an effort level, using GLiNER2.5-Decide running locally.

  • The model is picked once, from the session's first prompt, and then stays put: switching models mid-conversation throws away the prompt cache.
  • Effort is picked on every prompt. Short follow-ups ("yes", "go ahead") keep the last effort.
  • It runs on your machine. About 0.3 s of CPU per prompt, and no prompt goes to a third party to be routed.
  • You can see every decision. A band above the prompt shows the model, the effort, how long classifying took, and last turn's cache-read %.
  • It routes subagents too. Each subagent gets its own model and effort from the task it was given. A model the caller or the agent's definition named is kept.
  • It gets out of your way for the rest of the session once you pick a model or effort by hand (details).
  • It learns how you work. Optionally, train it on your own prompt history: on held-out prompts, a tuned adapter matched the labelled effort 66% of the time, against 40% for the default score.

Install

Requirements

  • Claude Code 2.1.287 or later (claude --version)
  • macOS or Linux, with curl
  • About 2 GB of free RAM while the classifier runs, and about 2 GB of disk for torch and the weights

Python and uv are optional: if uv is missing, the first session installs it to ~/.local/bin, and uv fetches Python 3.10–3.13 itself.

Add the plugin

In Claude Code:

/plugin marketplace add DustinVerzal/magic-router
/plugin install magic-router@magic-router

Restart Claude Code. There's nothing to configure.

First run

The first session starts the classifier daemon in the background. It keeps running after the session ends, and every later session reuses it.

WhenWhat happens
First run everuv installs (if missing), torch and gliner2 install, and the weights download (~1.7 GB). Until then the band shows classifier unreachable and the session keeps its own model.
First session after a rebootThe model loads in about 10 s, and the first prompt waits for it.
Every later promptAbout 0.3 s of CPU.
Warm it up first: skip the first-run download wait
git clone https://github.com/DustinVerzal/magic-router
magic-router/scripts/gliner.sh setup

setup installs uv if it's missing, installs torch and gliner2, downloads the weights, and leaves the daemon running. The downloads land in uv's and Hugging Face's shared caches, so the installed plugin reuses them; the clone is only needed to run the script.

From source: load a checkout to hack on it
git clone https://github.com/DustinVerzal/magic-router
claude --plugin-dir ./magic-router

This loads the checkout for one session instead of installing it. Edits to hooks/ reload while the session runs.

How it works

All routing policy lives in hooks/route.ts. The daemon only answers the questions the plugin sends it.

  1. Task type: one probability distribution over seven labels, each mapped to an eval family of the Artificial Analysis Intelligence Index v4.1.
  2. Complexity: seven independent yes/no signals, such as multi_file, planning, deep_reasoning, large_scope, and quick.
  3. Score: Σ weight × P(signal) + Σ bias × P(task). The task bias leans toward Opus where its lead on the matching evals is widest.
  4. Effort and model: the score maps to low < 0.4 ≤ medium < 1.0 ≤ high < 1.5 ≤ xhigh < 2.0 ≤ max. On the first prompt, a score ≥ 1.1 picks Opus; anything lower picks Sonnet.
The score scale. Effort bands run low below 0.4, medium to 1.0, high to 1.5, xhigh to 2.0, and max above. A dashed line at 1.1 splits Sonnet from Opus. The session's prompts sit at 0.21 (fix the typo, low), 0.49 (add retry and tests, medium), and 1.86 (design the job queue, xhigh).
Task labels and the evals behind them
LabelEval family
agentic_codingTerminal-Bench
scientific_codingSciCode
tool_useτ³-Bench
knowledge_workGDPval-AA
long_contextAA-LCR
knowledge_qaAA-Omniscience, GPQA
reasoningHLE, CritPt

[!NOTE] The default weights and thresholds were tuned by eye on 15 prompts. To fit them to how you work, see Tune it to your prompts.

When it stands aside

  • Subagents: each one is classified once, on its first request, and keeps that route. Only what it would inherit from the session is rewritten: a model or effort set by the Agent call or the agent's definition is kept. Forks are left alone, because they share their parent's context and prompt cache.
  • /model, /effort, or a fallback: if any of these changes what the engine asks for, the router stops rewriting for the rest of the session.
  • Classifier unreachable on the first prompt: the session keeps its own model, and effort routing starts once the daemon answers.
  • /clear: the next prompt picks a model again.

Privacy

  • Routing never leaves your machine. Each prompt's first 2,000 characters go to the daemon on 127.0.0.1 and nowhere else (or to your own ssh host, if you run the daemon there).
  • One background daemon, on port 8765, shared by all your sessions. It outlives them; stop it any time.
  • Files stay in ~/.cache/magic-router, plus the weights in Hugging Face's cache, downloaded once.
  • uv comes from Astral's installer (curl … | sh, to ~/.local/bin, no shell profile edits) if it's missing.

Only the opt-in tools reach further: just tune reads your local transcripts and sends prompts to Claude through claude -p (and to your ssh host, if you set one), and just bench sends prompts to the hosted models whose keys you set.

Configuration

Nothing needs configuring. These environment variables change the defaults:

VariableDefaultEffect
ROUTER_PORT8765The daemon's port. Set it for Claude Code and for scripts/gliner.sh.
ROUTER_MODELfastino/GLiNER2.5-DecideThe Hugging Face model the daemon loads.
ROUTER_WAIT900Seconds gliner.sh start waits for the model on a first run.
TUNE_HOSTnoneAn ssh host with an NVIDIA GPU to tune on.
AA_API_KEYnoneArtificial Analysis key for just benchmarks.
OPENROUTER_KEY, FASTINO_API_KEYnoneHosted-model keys for just bench.

Weights and thresholds are constants in hooks/route.ts. Fable 5.1 is wired in but off: set FABLE_AT there to a score above OPUS_AT to send the hardest first prompts to it.

Managing the classifier

These need a checkout of this repo:

scripts/gliner.sh setup         # install uv if missing, install torch + gliner2, download the weights, start
scripts/gliner.sh start         # start the daemon (if it isn't up) and wait until the model is loaded
scripts/gliner.sh stop          # stop the daemon
scripts/gliner.sh status        # print /health
scripts/gliner.sh check         # run the classifier's offline self-check
scripts/gliner.sh logs          # follow the daemon log
scripts/gliner.sh remote HOST   # run the daemon on an ssh host instead (macOS), such as a box with a GPU
scripts/gliner.sh local         # run it here again

Without a checkout, stop the daemon with pkill -f server/classifier.py. The log is at ~/.cache/magic-router/classifier.log.

Running the daemon on another machine

remote HOST copies this checkout and your tuned adapter to ~/.cache/magic-router on the host, starts the daemon there (on its NVIDIA GPU if it has one), and adds a launch agent that forwards port 8765 to it over ssh. The plugin keeps calling 127.0.0.1:8765 and needs no change, and your Mac no longer runs the model. After that, setup, start, stop, check and logs act on the host.

The host needs key-based ssh and rsync. If it's down when a session starts, the plugin starts a local daemon, which then holds the port until you stop it.

Tune it to your prompts

just tune trains the classifier on how you actually work. It's opt-in and needs a checkout. Labelling runs through claude -p, so it counts against your Claude subscription's usage, or bills your API key if that's how you're signed in. Run just benchmarks first: the labeller uses its scores.

  1. Collect. It reads every prompt you've sent Claude Code or Codex, from ~/.claude/projects, ~/.claude/history.jsonl and ~/.codex. Slash commands, $skills and short follow-ups are skipped, because the router never classifies them, and so are subagents, codex exec runs and orchestrators. A session's first prompt is kept however short, because it picks the model. just tune 500 uses only your newest 500.
  2. Label. Opus 5.5 at xhigh decides which model and effort each prompt needed, given both models' Artificial Analysis scores and response times at every effort, and is asked for the best answer without overthinking, not the cheapest.
  3. Train. It trains a LoRA adapter on 70% of your sessions, on CPU, or on TUNE_HOST if set. Sessions are split whole, so related prompts never land on both sides. Effort is learned from every prompt. The model is learned only from each session's first prompt, labelled with what the whole session needed: a session that opens with a typo fix and turns into design work counts as Opus, because the model picked on that first prompt has to last the session.
  4. Gate. On the other 30%, the adapter must beat both the score and any adapter already installed on effort, and pick the model no worse than the score on those sessions' first prompts. Only then does it install to ~/.cache/magic-router/tuned and restart the daemon; otherwise nothing changes.

Results from the author's history (3,344 prompts over six months), scored on 972 held-out prompts:

Matches labelMean levels off
Adapter trained on 3,344 prompts66%0.36
Adapter trained on 404 Claude Code prompts54%0.49
Always medium44%0.59
The default score40%0.69

Labelling cost $10.33 at API prices. Training took about 15 minutes on an RTX 4080; on CPU it runs about 15 minutes per 400 prompts. Your numbers will differ.

With an adapter installed, the daemon answers effort and model from it, which adds a second pass of about 0.1 s. The model is still picked only on a session's first prompt. An adapter trained before model labelling existed answers effort only; retune to get both. To undo, delete ~/.cache/magic-router/tuned and run scripts/gliner.sh stop.

What tuning sends and stores
  • Labelling sends 25 prompts per claude -p call, with no tools, no settings and no saved session.
  • Each label and a one-line reason go to ~/.cache/magic-router/tune/labels.jsonl. A rerun only labels new prompts, and any labelled under an older labeller prompt. Skim them there.
  • With TUNE_HOST, training and scoring run on that host (it needs an NVIDIA GPU, uv and rsync), and the prompts sent there are deleted afterwards.
Refreshing the benchmarks: just benchmarks

just benchmarks pulls the Artificial Analysis evals for Sonnet 5.5, Opus 5.5 and Fable 5.1 at every effort level, prints them as a table, and caches the raw JSON in ~/.cache/magic-router. With no arguments it also rewrites TASK_BIAS in hooks/route.ts from Opus's per-family lead over Sonnet; model names as arguments only filter the printout. Put AA_API_KEY=... in a git-ignored .env, or export it.

Against hosted decision models: just bench

just bench scores the trained adapter against hosted decision models on the same 972 held-out prompts, without retraining, sending each prompt's first 2,000 characters to the provider. OPENROUTER_KEY adds Jev (TypeSafe) and FASTINO_API_KEY adds GLiDE. Jev got the labeller's definition of each level, and was asked for the effort directly, both as a choice among the five levels and as a score on the ordered scale. A third run put Jev's answers to the router's own task and signal questions through the score's weights.

Matches labelMean levels off
Tuned adapter66%0.36
Jev, effort as a score63%0.42
Jev, effort as a choice56%0.53
Always medium44%0.59
The default score40%0.69
Jev through the score's weights27%1.06

Asked directly, Jev comes within 3 points of the adapter without seeing any of your prompts. The adapter stays ahead, though it was trained on labels from the same labeller, so the margin favours it. It also runs locally, while Jev is a hosted call (median 234 ms, $0.049 for all 972 prompts) that sends your prompts to a third party. Through the score's weights Jev does worse than always guessing medium, because those weights were set for GLiNER, so it can't drop into the router as is.

Troubleshooting

SymptomFix
Band says classifier unreachableRun scripts/gliner.sh start and read the log it names. On a first run, it's usually still downloading.
Daemon never starts from Claude Code, but gliner.sh start worksRead ~/.cache/magic-router/classifier.log: the uv install or a torch download likely failed (offline, proxy).
Band says stood asideYou picked a model or effort by hand (/model, /effort). /clear to let the router pick again.
Port 8765 is taken by something elseSet ROUTER_PORT to a free port for Claude Code and the shell you run gliner.sh from, then restart Claude Code.

Updating and uninstalling

A new version ships with every merge to main. Third-party marketplaces don't auto-update by default, so turn it on once: run /plugin, open Marketplaces, select magic-router and choose Enable auto-update. Claude Code then pulls new versions at startup and asks you to restart or run /reload-plugins. To update by hand:

/plugin marketplace update magic-router
/plugin update magic-router@magic-router

To uninstall:

/plugin uninstall magic-router@magic-router
/plugin marketplace remove magic-router

Then stop the daemon (pkill -f server/classifier.py) and, to reclaim the disk, delete ~/.cache/magic-router and ~/.cache/huggingface/hub/models--fastino--GLiNER2.5-Decide.

Contributing

Issues and pull requests are welcome. AGENTS.md covers the layout, the gotchas, and the conventions for both humans and coding agents. Before opening a PR, run what CI runs:

claude plugin validate .
claude plugin test .                           # routing, stickiness, stand-aside, band
uv run --script server/classifier.py --check   # classifier self-check
shellcheck scripts/*.sh

Acknowledgements

GLiNER2.5-Decide by Fastino does the classifying. Artificial Analysis supplies the evals that set the task bias and inform the labeller.

License

MIT

anthropic
claude-code
claude-code-plugin
gliner
llm-routing
model-routing

Languages

Python

55.3%

TypeScript

35.0%

Shell

8.5%

Just

1.3%