voxlo-dev/qwen-agent-8gb

Run a coding agent on an 8 GB consumer GPU. 3 qwen models optimized for best performance, wired into the pi harness.

Shell

1

100 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

I turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card (r/LocalLLaMA)

Everyone keeps saying you need expensive dedicated hardware for local agents. I have an RTX 4060 Ti with 8 GB and 64 GB of system RAM, and I wanted to see how far a normal gaming PC gets if you stop running defaults. So I let Claude (Opus 5.5) go through the whole setup, change one thing at a time…

0

Oct 6, 2026

README

Qwen Agent for 8 GB VRAM

A local coding agent on an ordinary 8 GB graphics card, two to nine times faster than download-and-go. The card was never the limit: the RAM next to it decides how large a model it runs.

One command builds llama.cpp, fetches a model and sets up the pi coding agent against it, tuned to the last few hundred megabytes of the card. Then qwen-pi in your project: an agent that reads your files, runs your tests and edits your code. No API key, no rate limit, nothing leaves the machine.

Which model

Mixture-of-experts models only use a few billion parameters per token, so their experts can live in system RAM while the card holds the rest. More RAM, a bigger model:

RAM next to the cardModelSpeedWindowGood for
anyTernary-Bonsai-2-27B, all on the card36 tok/s64ksmall, well-scoped tasks
32 GBQwen3.6-35B-A3B52-65 tok/s131kmost work: fast, and the default
64 GBQwen3.8-Flash-Next, 125B~19 tok/s131kthe strongest, for when it may take longer

./install.sh reads your RAM, shows which of the three run, and asks. How they did as agents on the same task: the agent sessions.

Install

You need an 8 GB NVIDIA or AMD card, Linux or WSL2, and the RAM from the table above.

git clone https://github.com/voxlo-dev/qwen-agent-8gb.git
cd qwen-agent-8gb
./install.sh                    # NVIDIA
BACKEND=vulkan ./install.sh     # AMD

It checks your machine first and stops with a list of anything missing, before it downloads or compiles. Then a 10-30 minute build and the download; re-running is safe. On Windows, set up WSL2 first: Windows. Disk and the rest: Requirements.

Use

qwen-pi            # in your project directory; arguments go to pi

qwen-pi starts the server in the background, waits for the model and stops it after the last session. To keep it running, start it yourself:

qwen-server        # terminal 1, ready at "listening on http://127.0.0.1:8080"
qwen-pi            # terminal 2

The server is also a plain OpenAI-compatible endpoint at http://127.0.0.1:8080/v1 for any other tool. qwen-studio opens the same model with the same settings in Unsloth Studio's chat UI, at the same speed and VRAM: details.

Speed

The same card and the same models, three levels of effort. Generation in tok/s, RTX 4060 Ti 8 GB, Ryzen 7 7800X3D, 64 GB RAM:

BonsaiQwen3.6Qwen3.8-Flash
Download and go: the GGUF with default settings~4~25~4
This repo's configuration: the right llama.cpp tree, expert placement, KV cache, multi-token prediction; under WSL23639-459-10
Plus the system: native Linux, no display on the card36.652-65~19

The biggest single step after the configuration: do not let this GPU drive your monitor. A desktop takes 0.5-1.2 GB of VRAM and GPU time, both out of the model; plug the monitor into the mainboard's iGPU. If you cannot, PROFILE=display makes room. Where each number comes from: docs/performance.md and the pages per model.

Configure

Each model has two profiles: dedicated (default) for a card that drives no display, display for one that does. Settings live in config.env, each with a comment; an environment variable always wins, and arguments go straight to llama-server:

PROFILE=display ./install.sh pi && PROFILE=display qwen-pi
MODEL=qwen38-flash qwen-pi
qwen-server --port 9000

After changing the profile, the window or the port, run ./install.sh pi again so pi's copy matches the server. The window and pi's budget values constrain each other: Context budget.

Requirements

Measured, each with a logged run:

CardSystemQwen3.6Qwen3.8-FlashBonsai
RTX 4060 Ti 8 GBLinux Mint 22.3, native52-65 tok/s~19 tok/s36.6 tok/s
RTX 4060 Ti 8 GBWindows 11 + WSL2, Ubuntu 26.0439-45 tok/s9-10 tok/s36 tok/s
AMD RX 570 8 GBDebian 13, Vulkan (Mesa 26.1)~25 tok/s-7 tok/s with this repo's Vulkan patch, 0.94 without

Expected to work, unmeasured: RTX 20xx to 40xx and RDNA2/3 cards with 8 GB or more; RTX 50xx needs CUDA >= 12.8. If you run one, a hardware report is the most useful thing you can send. Not supported: less than 8 GB of VRAM, ROCm, Metal, CPU-only.

Disk: 22 GB for Qwen3.6, 88 GB for Qwen3.8-Flash, 6 GB for Bonsai, plus ~10 GB for the build and the CUDA toolkit. Drivers, toolkit versions, Node for pi and what deps installs per distro: docs/setup.md.

Install in detail

./install.sh runs deps build model pi link in order; name steps to run only those.

StepDoes
depsthe apt toolchain: CUDA or Vulkan, cmake, gcc (asks for sudo)
buildthe model's llama.cpp tree at its pinned version (FORCE=1 rebuilds)
modelthe GGUF, from the Hugging Face cache or downloaded, checksummed
piits own pinned pi with a config for the model; your own pi and ~/.pi stay untouched
linkqwen-server, qwen-pi and qwen-studio into ~/.local/bin

The model you pick is recorded in $QWEN_HOME/model.env and used by every qwen-* command. A second one installs next to it with MODEL=bonsai ./install.sh and runs with MODEL=bonsai qwen-pi. Everything lands in ~/.local/share/qwen-local (QWEN_HOME); an install from before the rename, in ~/.local/share/bonsai-local, stays where it is and keeps running Bonsai.

How it was measured

Every non-default setting in this repo answers a failure seen on real hardware, and the measurement behind it is in docs/. Speed, VRAM and windows are reproducible. Agent quality is a handful of single sessions on one prompt, enough to put Qwen3.6 ahead of Bonsai and not more; a reproducible test is next. The starting point was a study of eight local models as coding agents on an 8 GB card: docs/model-comparison.md.

Uninstall

rm -rf ~/.local/share/qwen-local ~/.local/bin/qwen-server ~/.local/bin/qwen-pi ~/.local/bin/qwen-studio

That includes the models, pi and its sessions (bonsai-local for an install from before the rename). The apt packages from deps stay.

Contributing

Issues, pull requests and hardware reports welcome. CONTRIBUTING.md has the one rule: a non-default choice arrives with the measurement that justifies it. Working on the code: AGENTS.md.

Your hardware

This fills an 8 GB card to within a few hundred megabytes and keeps the GPU, and for the Qwen models tens of gigabytes of RAM, busy for as long as a session runs. Nothing here overclocks, raises a power limit or touches a fan curve. Your cooling, power supply and driver are yours. Provided as is, no warranty: see LICENSE.

Credits

  • The Qwen team for Qwen3.6-35B-A3B and Qwen3.8-Flash-Next
  • Unsloth for the GGUFs, the llama.cpp tree both Qwen models run on, and Unsloth Studio
  • PrismML for Ternary Bonsai 2 27B and the llama.cpp fork that loads it
  • ggml-org/llama.cpp, including whoever wrote the IQ-grid Vulkan shaders the patch follows, and the reporter of PrismML-Eng/llama.cpp#185, who made that work findable
  • Mesa and RADV, and pi by earendil-works

The weights are their authors', under the licenses on their model cards, downloaded at install time and never redistributed here. The model comparison and the localagent workflow come from a project thesis at Technische Hochschule Mittelhessen by this repo's author.

License: MIT, including the PTQ1_0 Vulkan decode in patches/vulkan, which is offered upstream under the same terms.

8gb-vram
bonsai
coding-agent
llama-cpp
local-llm
quantization
qwen
vulkan

voxlo-dev/qwen-agent-8gb

Run a coding agent on an 8 GB consumer GPU. 3 qwen models optimized for best performance, wired into the pi harness.

Shell

1

100 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

I turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card (r/LocalLLaMA)

Everyone keeps saying you need expensive dedicated hardware for local agents. I have an RTX 4060 Ti with 8 GB and 64 GB of system RAM, and I wanted to see how far a normal gaming PC gets if you stop running defaults. So I let Claude (Opus 5.5) go through the whole setup, change one thing at a time…

0

Oct 6, 2026

README

Qwen Agent for 8 GB VRAM

A local coding agent on an ordinary 8 GB graphics card, two to nine times faster than download-and-go. The card was never the limit: the RAM next to it decides how large a model it runs.

One command builds llama.cpp, fetches a model and sets up the pi coding agent against it, tuned to the last few hundred megabytes of the card. Then qwen-pi in your project: an agent that reads your files, runs your tests and edits your code. No API key, no rate limit, nothing leaves the machine.

Which model

Mixture-of-experts models only use a few billion parameters per token, so their experts can live in system RAM while the card holds the rest. More RAM, a bigger model:

RAM next to the cardModelSpeedWindowGood for
anyTernary-Bonsai-2-27B, all on the card36 tok/s64ksmall, well-scoped tasks
32 GBQwen3.6-35B-A3B52-65 tok/s131kmost work: fast, and the default
64 GBQwen3.8-Flash-Next, 125B~19 tok/s131kthe strongest, for when it may take longer

./install.sh reads your RAM, shows which of the three run, and asks. How they did as agents on the same task: the agent sessions.

Install

You need an 8 GB NVIDIA or AMD card, Linux or WSL2, and the RAM from the table above.

git clone https://github.com/voxlo-dev/qwen-agent-8gb.git
cd qwen-agent-8gb
./install.sh                    # NVIDIA
BACKEND=vulkan ./install.sh     # AMD

It checks your machine first and stops with a list of anything missing, before it downloads or compiles. Then a 10-30 minute build and the download; re-running is safe. On Windows, set up WSL2 first: Windows. Disk and the rest: Requirements.

Use

qwen-pi            # in your project directory; arguments go to pi

qwen-pi starts the server in the background, waits for the model and stops it after the last session. To keep it running, start it yourself:

qwen-server        # terminal 1, ready at "listening on http://127.0.0.1:8080"
qwen-pi            # terminal 2

The server is also a plain OpenAI-compatible endpoint at http://127.0.0.1:8080/v1 for any other tool. qwen-studio opens the same model with the same settings in Unsloth Studio's chat UI, at the same speed and VRAM: details.

Speed

The same card and the same models, three levels of effort. Generation in tok/s, RTX 4060 Ti 8 GB, Ryzen 7 7800X3D, 64 GB RAM:

BonsaiQwen3.6Qwen3.8-Flash
Download and go: the GGUF with default settings~4~25~4
This repo's configuration: the right llama.cpp tree, expert placement, KV cache, multi-token prediction; under WSL23639-459-10
Plus the system: native Linux, no display on the card36.652-65~19

The biggest single step after the configuration: do not let this GPU drive your monitor. A desktop takes 0.5-1.2 GB of VRAM and GPU time, both out of the model; plug the monitor into the mainboard's iGPU. If you cannot, PROFILE=display makes room. Where each number comes from: docs/performance.md and the pages per model.

Configure

Each model has two profiles: dedicated (default) for a card that drives no display, display for one that does. Settings live in config.env, each with a comment; an environment variable always wins, and arguments go straight to llama-server:

PROFILE=display ./install.sh pi && PROFILE=display qwen-pi
MODEL=qwen38-flash qwen-pi
qwen-server --port 9000

After changing the profile, the window or the port, run ./install.sh pi again so pi's copy matches the server. The window and pi's budget values constrain each other: Context budget.

Requirements

Measured, each with a logged run:

CardSystemQwen3.6Qwen3.8-FlashBonsai
RTX 4060 Ti 8 GBLinux Mint 22.3, native52-65 tok/s~19 tok/s36.6 tok/s
RTX 4060 Ti 8 GBWindows 11 + WSL2, Ubuntu 26.0439-45 tok/s9-10 tok/s36 tok/s
AMD RX 570 8 GBDebian 13, Vulkan (Mesa 26.1)~25 tok/s-7 tok/s with this repo's Vulkan patch, 0.94 without

Expected to work, unmeasured: RTX 20xx to 40xx and RDNA2/3 cards with 8 GB or more; RTX 50xx needs CUDA >= 12.8. If you run one, a hardware report is the most useful thing you can send. Not supported: less than 8 GB of VRAM, ROCm, Metal, CPU-only.

Disk: 22 GB for Qwen3.6, 88 GB for Qwen3.8-Flash, 6 GB for Bonsai, plus ~10 GB for the build and the CUDA toolkit. Drivers, toolkit versions, Node for pi and what deps installs per distro: docs/setup.md.

Install in detail

./install.sh runs deps build model pi link in order; name steps to run only those.

StepDoes
depsthe apt toolchain: CUDA or Vulkan, cmake, gcc (asks for sudo)
buildthe model's llama.cpp tree at its pinned version (FORCE=1 rebuilds)
modelthe GGUF, from the Hugging Face cache or downloaded, checksummed
piits own pinned pi with a config for the model; your own pi and ~/.pi stay untouched
linkqwen-server, qwen-pi and qwen-studio into ~/.local/bin

The model you pick is recorded in $QWEN_HOME/model.env and used by every qwen-* command. A second one installs next to it with MODEL=bonsai ./install.sh and runs with MODEL=bonsai qwen-pi. Everything lands in ~/.local/share/qwen-local (QWEN_HOME); an install from before the rename, in ~/.local/share/bonsai-local, stays where it is and keeps running Bonsai.

How it was measured

Every non-default setting in this repo answers a failure seen on real hardware, and the measurement behind it is in docs/. Speed, VRAM and windows are reproducible. Agent quality is a handful of single sessions on one prompt, enough to put Qwen3.6 ahead of Bonsai and not more; a reproducible test is next. The starting point was a study of eight local models as coding agents on an 8 GB card: docs/model-comparison.md.

Uninstall

rm -rf ~/.local/share/qwen-local ~/.local/bin/qwen-server ~/.local/bin/qwen-pi ~/.local/bin/qwen-studio

That includes the models, pi and its sessions (bonsai-local for an install from before the rename). The apt packages from deps stay.

Contributing

Issues, pull requests and hardware reports welcome. CONTRIBUTING.md has the one rule: a non-default choice arrives with the measurement that justifies it. Working on the code: AGENTS.md.

Your hardware

This fills an 8 GB card to within a few hundred megabytes and keeps the GPU, and for the Qwen models tens of gigabytes of RAM, busy for as long as a session runs. Nothing here overclocks, raises a power limit or touches a fan curve. Your cooling, power supply and driver are yours. Provided as is, no warranty: see LICENSE.

Credits

  • The Qwen team for Qwen3.6-35B-A3B and Qwen3.8-Flash-Next
  • Unsloth for the GGUFs, the llama.cpp tree both Qwen models run on, and Unsloth Studio
  • PrismML for Ternary Bonsai 2 27B and the llama.cpp fork that loads it
  • ggml-org/llama.cpp, including whoever wrote the IQ-grid Vulkan shaders the patch follows, and the reporter of PrismML-Eng/llama.cpp#185, who made that work findable
  • Mesa and RADV, and pi by earendil-works

The weights are their authors', under the licenses on their model cards, downloaded at install time and never redistributed here. The model comparison and the localagent workflow come from a project thesis at Technische Hochschule Mittelhessen by this repo's author.

License: MIT, including the PTQ1_0 Vulkan decode in patches/vulkan, which is offered upstream under the same terms.

8gb-vram
bonsai
coding-agent
llama-cpp
local-llm
quantization
qwen
vulkan