Coding agent where a 0.8B decision model works alongside Qwen 3.8-27B: Jeff takes the routine steps itself and decides when Qwen should think. 47% faster (32% less time) per task at the same pass rate. A fork of Pi.
See the code
A coding agent where a 0.8B decision model works alongside Qwen 3.8-27B.
Coding tasks 47% faster (32% less time) on average, at the same pass rate.
jeffhub.ai · Jeff · code adapter · code-router adapter · Pi's README
Jeff-Code is a fork of Pi, the coding agent by Mario Zechner and the Pi contributors. It puts Jeff, a 0.8B decision model, inside Pi's agent loop. Around every Qwen turn, Jeff makes two quick decisions (about 0.2 s each):
Each decision is a small LoRA adapter on the same Jeff v1.3 base.
Run side by side in paired blocks: each task ran under every setting at the same time, on the same Qwen server, and every comparison is paired by task. The baseline is Qwen 3.8-27B alone in the same build with every Jeff feature switched off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it.
| Jeff-Code (threshold 0.6) vs Qwen alone | |
|---|---|
| Pass rate | 62.4% vs 62.8%; paired difference −0.2 points (95% interval −2.6 to +2.1), 1,242 paired tasks |
| Time per task, on average | 0.68× (0.64-0.72; geometric mean of per-task time ratios), median 0.70× |
| Total time, all tasks combined | 0.86× (0.80-0.93) |
| Per benchmark | SWE-bench Verified 0.63×, SWE-rebench 0.66×, Terminal-Bench Pro 0.64×, Harbor Index 0.71×; no clear speed-up on Terminal-Bench 2.0 (0.96×) or SkillsBench (0.91×) |
| Thinking off on every turn instead | faster still, but −7.6 points (−10.6 to −4.5); −13.5 on Terminal-Bench 2.0 |
| A less cautious Jeff (threshold 0.7) | −4.5 points |
| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
|---|---|---|---|---|---|
| SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (−3.5 to +4.3) | 0.63× (0.57-0.69) |
| SWE-rebench (2 rounds) | 370 | 58.9% | 58.1% | −0.8 (−5.7 to +3.5) | 0.66× (0.61-0.73) |
| Terminal-Bench Pro (2 rounds) | 195 | 61.2% | 62.8% | +1.5 (−4.6 to +7.7) | 0.64× (0.55-0.76) |
| Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | −4.6 (−12.1 to +3.7) | 0.96× (0.78-1.16) |
| SkillsBench | 42 | 28.6% | 31.0% | +2.4 (−11.9 to +16.7) | 0.91× (0.68-1.20) |
| Harbor Index | 41 | 12.2% | 9.8% | −2.4 (−12.2 to +7.3) | 0.71× (0.51-0.99) |
| All six, pooled | 1,242 | 62.8% | 62.4% | −0.2 (−2.6 to +2.1) | 0.68× (0.64-0.72) |
No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's), below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates. Terminal-Bench (original) and Terminal-Bench Science also ran, but Qwen alone and Jeff-Code both solve 0% of their tasks, so they are left out.
Full report: results/imitation/eval-tonight.md.
Jeff predicts what Qwen would do next.
jeff-adapter-code): each label is the step Qwen actually took next in recorded Qwen sessions, built by
code, with no other model involved.jeff-adapter-code-router): each recorded Qwen turn at full thinking was asked again with thinking
off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full
thinking if none was. "As good" is decided by code wherever possible (the same kind of step on the same target);
otherwise Qwen3.8-Max, with thinking off, judges whether the cheaper step would serve the task just as well. These
strict labels make the router cautious, and that is what keeps the pass rate.The code is in packages/coding-agent/src/core/jeff-first/; the evaluation and training tools are in tools/jeff-first/.
You need Qwen 3.8-27B behind an OpenAI-compatible server (for example vLLM) and a Jeff server with the two adapters.
1. Jeff server. Install Jeff, then download the base and both adapters (the folder names are the adapter names Jeff-Code asks for):
hf download mstrasser/jeff-base --revision v1.3 --local-dir jeff-base-v1.3
hf download mstrasser/jeff-adapter-code --revision v1.3 --local-dir adapters/jeff-step
hf download mstrasser/jeff-adapter-code-router --revision v1.3 --local-dir adapters/jeff-router
JEFF_CHECKPOINT=jeff-base-v1.3 JEFF_ADAPTERS=adapters JEFF_DEVICE=cuda JEFF_HOST=0.0.0.0 PORT=8920 \
JEFF_CPU_THREADS=8 python tools/jeff-first/jeff_serve.py # with the Python of the Jeff environment
jeff_serve.py starts Jeff's own server and adds the endpoint that fits long prompts into Jeff's 8,192 tokens. For
many parallel sessions, run one server per GPU behind tools/jeff-first/jeff_pool.py.
GGUF versions for llama.cpp are on Hugging Face (-gguf); run the router on the Q8_0 base, since about 6% of its
decisions change at Q4_K_M.
2. Jeff-Code with Qwen. Build the repository (npm install && npm run build) and add Qwen to ~/.jeff/agent/models.json
as an OpenAI-compatible model with "reasoning": true, "compat": {"thinkingFormat": "qwen-chat-template"} and
"maxTokens": 32768 (see the model docs).
3. Switch Jeff on with the settings the evaluation used:
export JEFF_FIRST_MODE=jeff
export JEFF_FIRST_JEFF_URL=http://localhost:8920
export JEFF_FIRST_JEFF_STEP_ADAPTER=jeff-step
export JEFF_FIRST_JEFF_STEP_THRESHOLD=0.40 # Jeff's top option needs 0.40, otherwise Qwen takes over
export JEFF_FIRST_THINKING_ROUTER=jeff-off-unless:jeff-router:0.6
export JEFF_FIRST_THINKING_LIMIT=8000
export JEFF_FIRST_OUTPUT_TRIM=off # required; leave off
export JEFF_FIRST_RUN_APPROVAL=all # all, seen or never: may Jeff run scripts Qwen wrote and install packages
export JEFF_FIRST_DRIVER_BUILD=qwen3.8-27b # the exact Qwen build, written to the trace
export JEFF_FIRST_TRACE_FILE=$HOME/jeff-code/trace.jsonl # its folder must exist
export JEFF_FIRST_TASK_ID=my-project
./jeff-test.sh
Every setting is required in jeff mode, and a missing or invalid one stops Jeff-Code with an error starting with
JeffFirst:. Unset JEFF_FIRST_MODE for plain Pi.
Status: research code. It was built and measured inside benchmark containers (Harbor); interactive use works, but has had far less testing.
Jeff-Code is a fork of Pi and keeps Pi's MIT licence (LICENSE). All credit for the agent itself goes to Pi's authors; Pi's own README is in README-pi.md. We'd happily upstream whatever Pi wants to take. The Jeff adapters are Apache 2.0.
Coding agent where a 0.8B decision model works alongside Qwen 3.8-27B: Jeff takes the routine steps itself and decides when Qwen should think. 47% faster (32% less time) per task at the same pass rate. A fork of Pi.
See the code
A coding agent where a 0.8B decision model works alongside Qwen 3.8-27B.
Coding tasks 47% faster (32% less time) on average, at the same pass rate.
jeffhub.ai · Jeff · code adapter · code-router adapter · Pi's README
Jeff-Code is a fork of Pi, the coding agent by Mario Zechner and the Pi contributors. It puts Jeff, a 0.8B decision model, inside Pi's agent loop. Around every Qwen turn, Jeff makes two quick decisions (about 0.2 s each):
Each decision is a small LoRA adapter on the same Jeff v1.3 base.
Run side by side in paired blocks: each task ran under every setting at the same time, on the same Qwen server, and every comparison is paired by task. The baseline is Qwen 3.8-27B alone in the same build with every Jeff feature switched off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it.
| Jeff-Code (threshold 0.6) vs Qwen alone | |
|---|---|
| Pass rate | 62.4% vs 62.8%; paired difference −0.2 points (95% interval −2.6 to +2.1), 1,242 paired tasks |
| Time per task, on average | 0.68× (0.64-0.72; geometric mean of per-task time ratios), median 0.70× |
| Total time, all tasks combined | 0.86× (0.80-0.93) |
| Per benchmark | SWE-bench Verified 0.63×, SWE-rebench 0.66×, Terminal-Bench Pro 0.64×, Harbor Index 0.71×; no clear speed-up on Terminal-Bench 2.0 (0.96×) or SkillsBench (0.91×) |
| Thinking off on every turn instead | faster still, but −7.6 points (−10.6 to −4.5); −13.5 on Terminal-Bench 2.0 |
| A less cautious Jeff (threshold 0.7) | −4.5 points |
| Benchmark | Paired tasks | Qwen alone | Jeff-Code | Difference, points (95% interval) | Time per task |
|---|---|---|---|---|---|
| SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (−3.5 to +4.3) | 0.63× (0.57-0.69) |
| SWE-rebench (2 rounds) | 370 | 58.9% | 58.1% | −0.8 (−5.7 to +3.5) | 0.66× (0.61-0.73) |
| Terminal-Bench Pro (2 rounds) | 195 | 61.2% | 62.8% | +1.5 (−4.6 to +7.7) | 0.64× (0.55-0.76) |
| Terminal-Bench 2.0 (40 tasks, 3 attempts each) | 108 | 75.9% | 70.0% | −4.6 (−12.1 to +3.7) | 0.96× (0.78-1.16) |
| SkillsBench | 42 | 28.6% | 31.0% | +2.4 (−11.9 to +16.7) | 0.91× (0.68-1.20) |
| Harbor Index | 41 | 12.2% | 9.8% | −2.4 (−12.2 to +7.3) | 0.71× (0.51-0.99) |
| All six, pooled | 1,242 | 62.8% | 62.4% | −0.2 (−2.6 to +2.1) | 0.68× (0.64-0.72) |
No benchmark shows a clear pass-rate difference: every interval includes zero. SWE-rebench and Terminal-Bench Pro ran twice; their two rounds are combined, with intervals computed over both. Time per task is the geometric mean of the per-task time ratios (Jeff-Code's time divided by Qwen alone's), below 1 is faster. The pass rates count every finished session; the difference counts only tasks finished in both settings, so it is not exactly the gap between the two pass rates. Terminal-Bench (original) and Terminal-Bench Science also ran, but Qwen alone and Jeff-Code both solve 0% of their tasks, so they are left out.
Full report: results/imitation/eval-tonight.md.
Jeff predicts what Qwen would do next.
jeff-adapter-code): each label is the step Qwen actually took next in recorded Qwen sessions, built by
code, with no other model involved.jeff-adapter-code-router): each recorded Qwen turn at full thinking was asked again with thinking
off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full
thinking if none was. "As good" is decided by code wherever possible (the same kind of step on the same target);
otherwise Qwen3.8-Max, with thinking off, judges whether the cheaper step would serve the task just as well. These
strict labels make the router cautious, and that is what keeps the pass rate.The code is in packages/coding-agent/src/core/jeff-first/; the evaluation and training tools are in tools/jeff-first/.
You need Qwen 3.8-27B behind an OpenAI-compatible server (for example vLLM) and a Jeff server with the two adapters.
1. Jeff server. Install Jeff, then download the base and both adapters (the folder names are the adapter names Jeff-Code asks for):
hf download mstrasser/jeff-base --revision v1.3 --local-dir jeff-base-v1.3
hf download mstrasser/jeff-adapter-code --revision v1.3 --local-dir adapters/jeff-step
hf download mstrasser/jeff-adapter-code-router --revision v1.3 --local-dir adapters/jeff-router
JEFF_CHECKPOINT=jeff-base-v1.3 JEFF_ADAPTERS=adapters JEFF_DEVICE=cuda JEFF_HOST=0.0.0.0 PORT=8920 \
JEFF_CPU_THREADS=8 python tools/jeff-first/jeff_serve.py # with the Python of the Jeff environment
jeff_serve.py starts Jeff's own server and adds the endpoint that fits long prompts into Jeff's 8,192 tokens. For
many parallel sessions, run one server per GPU behind tools/jeff-first/jeff_pool.py.
GGUF versions for llama.cpp are on Hugging Face (-gguf); run the router on the Q8_0 base, since about 6% of its
decisions change at Q4_K_M.
2. Jeff-Code with Qwen. Build the repository (npm install && npm run build) and add Qwen to ~/.jeff/agent/models.json
as an OpenAI-compatible model with "reasoning": true, "compat": {"thinkingFormat": "qwen-chat-template"} and
"maxTokens": 32768 (see the model docs).
3. Switch Jeff on with the settings the evaluation used:
export JEFF_FIRST_MODE=jeff
export JEFF_FIRST_JEFF_URL=http://localhost:8920
export JEFF_FIRST_JEFF_STEP_ADAPTER=jeff-step
export JEFF_FIRST_JEFF_STEP_THRESHOLD=0.40 # Jeff's top option needs 0.40, otherwise Qwen takes over
export JEFF_FIRST_THINKING_ROUTER=jeff-off-unless:jeff-router:0.6
export JEFF_FIRST_THINKING_LIMIT=8000
export JEFF_FIRST_OUTPUT_TRIM=off # required; leave off
export JEFF_FIRST_RUN_APPROVAL=all # all, seen or never: may Jeff run scripts Qwen wrote and install packages
export JEFF_FIRST_DRIVER_BUILD=qwen3.8-27b # the exact Qwen build, written to the trace
export JEFF_FIRST_TRACE_FILE=$HOME/jeff-code/trace.jsonl # its folder must exist
export JEFF_FIRST_TASK_ID=my-project
./jeff-test.sh
Every setting is required in jeff mode, and a missing or invalid one stops Jeff-Code with an error starting with
JeffFirst:. Unset JEFF_FIRST_MODE for plain Pi.
Status: research code. It was built and measured inside benchmark containers (Harbor); interactive use works, but has had far less testing.
Jeff-Code is a fork of Pi and keeps Pi's MIT licence (LICENSE). All credit for the agent itself goes to Pi's authors; Pi's own README is in README-pi.md. We'd happily upstream whatever Pi wants to take. The Jeff adapters are Apache 2.0.