The fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090: up to 500 tokens/s for one agent and up to 2,000 tokens/s for many, contexts up to 1M tokens, and a launcher that shows what every setting costs. Windows and Linux.
See the code
The fastest inference engine for Qwen3.8-27B on the NVIDIA GeForce RTX 5090.
Up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, on Windows 11 and Linux.
A launcher sets it up and shows, before you load, what every setting costs in accuracy, speed and memory.
Download for Windows · Download for Linux · Models on Hugging Face · Usage guide
Twelve coding agents working at once on one RTX 5090 with the Tiny model, at about 1,500 tokens/s.
State-of-the-art inference for Qwen3.8-27B on the RTX 5090, for one agent and for many at once:
| Tokens per second | Tiny | Small | Medium | Large | XXL |
|---|---|---|---|---|---|
| One agent coding | 540 | 495 | 440 | 377 | 313 |
| 8 agents coding, in total | 1,890 | 1,754 | 1,638 | 1,427 | 1,254 |
Greedy decoding with thinking off, on an RTX 5090 with its memory overclocked by 20%. A stock card is somewhat slower.
Agents never work in step: one reads a 100K-token codebase while another writes a two-line fix. So instead of a fixed slice each, MegaCapybara gives them one shared pool of context:
Fixed slots are one switch away if you prefer them.
Before every GPU step, the scheduler decides what runs:
When an agent sends its next request, its conversation is usually still cached: in VRAM, in RAM (back in about 3 ms) or on disk (back in seconds). The reply starts right away instead of reading a long prompt again from the start.
Every model file carries its own accuracy scores, measured against Qwen's original weights, and every setting has a measured cost. The launcher adds them up for exactly the settings you pick:
Switch to a smaller model or a 4-bit KV cache and you see right away what you trade for the extra speed or context.

Top-1 agreement is how often the model picks the same next token as Qwen's original weights: higher is closer. Unsloth's GGUF quants are widely considered the state of the art for running models locally, so here is how our model files compare with them, size for size.
Unsloth's stay a little closer at the same size. Ours use the 4- and 6-bit formats that the RTX 5090 computes natively, which is where MegaCapybara's speed comes from.
Hover any ? in the launcher for a plain explanation of its setting: an animated picture, the options compared on
your PC, and their pros and cons. The animations on this page are taken from those panels.

Press the arrow next to the model to list every model file on Hugging Face, then download one with a click:

MegaCapybara speaks both OpenAI's and Anthropic's APIs, including tool calls, thinking and images. Point your client at it:
| Client | Base URL |
|---|---|
| OpenCode, Cline, Continue, Open WebUI and other OpenAI-compatible clients | http://127.0.0.1:8080/v1 |
| Claude Code and Anthropic's SDKs | http://127.0.0.1:8080 |
Setup for each client: docs/USAGE.md.
The launcher is optional. The presets folder has ready-made start scripts, .bat on Windows and .sh on Linux:
the launcher's three presets and six more, one for each model size and number of agents. Run one and the server
starts with a live dashboard showing the speed, prompt reading, each agent's context, and what every conversation is
doing.


Every option is in docs/USAGE.md and in megacapybara --help.
MegaCapybaraLauncher. Windows may warn you once because the programs are not code-signed: click
More info, then Run anyway.http://127.0.0.1:8080/v1 and set its context size to the number under
Context size for your front end.Nothing else to install: the programs are self-contained and the folder can live anywhere you can write to. On Linux, the launcher uses the desktop's own X11, Cairo, Pango and libcurl, and offers to install any that are missing.
Converted from Qwen/Qwen3.8-27B, in five sizes: perkel/Qwen3.8-27B-MC.
| Size | VRAM | KL divergence | Top-1 agreement | Good for |
|---|---|---|---|---|
| Tiny | 12.70 GiB | 0.0582 | 92.7% | the most speed and context |
| Small | 13.67 GiB | 0.0478 | 93.8% | speed, with a little more accuracy |
| Medium | 15.39 GiB | 0.0298 | 95.2% | the balance, and the presets' choice |
| Large | 18.66 GiB | 0.0081 | 97.7% | answers close to the original |
| XXL | 23.58 GiB | 0.0051 | 98.3% | the closest to the original, less room for context |
Both measured against Qwen's original BF16 weights on 81,880 held-out tokens. KL divergence: 0 is identical, lower is closer. Top-1 agreement: how often the most likely next token is the same.
The same five sizes also come uncensored, converted from an abliterated release (one with the refusals removed): perkel/Qwen3.8-27B-Uncensored-MC.
MegaCapybara is released as programs for Windows and Linux for now; the source code will be published later.
MegaCapybara is under the MIT License, (c) 2026 Perkel's Software Corner. The model weights and the
libraries it includes keep their own licenses: see THIRD-PARTY-NOTICES.md. Each release
lists its changes in its notes and in the package's CHANGELOG.txt.
If MegaCapybara is useful to you, you can support it on Patreon. The money goes to the GPUs needed to support more cards and multi-GPU.
The fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090: up to 500 tokens/s for one agent and up to 2,000 tokens/s for many, contexts up to 1M tokens, and a launcher that shows what every setting costs. Windows and Linux.
See the code
The fastest inference engine for Qwen3.8-27B on the NVIDIA GeForce RTX 5090.
Up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, on Windows 11 and Linux.
A launcher sets it up and shows, before you load, what every setting costs in accuracy, speed and memory.
Download for Windows · Download for Linux · Models on Hugging Face · Usage guide
Twelve coding agents working at once on one RTX 5090 with the Tiny model, at about 1,500 tokens/s.
State-of-the-art inference for Qwen3.8-27B on the RTX 5090, for one agent and for many at once:
| Tokens per second | Tiny | Small | Medium | Large | XXL |
|---|---|---|---|---|---|
| One agent coding | 540 | 495 | 440 | 377 | 313 |
| 8 agents coding, in total | 1,890 | 1,754 | 1,638 | 1,427 | 1,254 |
Greedy decoding with thinking off, on an RTX 5090 with its memory overclocked by 20%. A stock card is somewhat slower.
Agents never work in step: one reads a 100K-token codebase while another writes a two-line fix. So instead of a fixed slice each, MegaCapybara gives them one shared pool of context:
Fixed slots are one switch away if you prefer them.
Before every GPU step, the scheduler decides what runs:
When an agent sends its next request, its conversation is usually still cached: in VRAM, in RAM (back in about 3 ms) or on disk (back in seconds). The reply starts right away instead of reading a long prompt again from the start.
Every model file carries its own accuracy scores, measured against Qwen's original weights, and every setting has a measured cost. The launcher adds them up for exactly the settings you pick:
Switch to a smaller model or a 4-bit KV cache and you see right away what you trade for the extra speed or context.

Top-1 agreement is how often the model picks the same next token as Qwen's original weights: higher is closer. Unsloth's GGUF quants are widely considered the state of the art for running models locally, so here is how our model files compare with them, size for size.
Unsloth's stay a little closer at the same size. Ours use the 4- and 6-bit formats that the RTX 5090 computes natively, which is where MegaCapybara's speed comes from.
Hover any ? in the launcher for a plain explanation of its setting: an animated picture, the options compared on
your PC, and their pros and cons. The animations on this page are taken from those panels.

Press the arrow next to the model to list every model file on Hugging Face, then download one with a click:

MegaCapybara speaks both OpenAI's and Anthropic's APIs, including tool calls, thinking and images. Point your client at it:
| Client | Base URL |
|---|---|
| OpenCode, Cline, Continue, Open WebUI and other OpenAI-compatible clients | http://127.0.0.1:8080/v1 |
| Claude Code and Anthropic's SDKs | http://127.0.0.1:8080 |
Setup for each client: docs/USAGE.md.
The launcher is optional. The presets folder has ready-made start scripts, .bat on Windows and .sh on Linux:
the launcher's three presets and six more, one for each model size and number of agents. Run one and the server
starts with a live dashboard showing the speed, prompt reading, each agent's context, and what every conversation is
doing.


Every option is in docs/USAGE.md and in megacapybara --help.
MegaCapybaraLauncher. Windows may warn you once because the programs are not code-signed: click
More info, then Run anyway.http://127.0.0.1:8080/v1 and set its context size to the number under
Context size for your front end.Nothing else to install: the programs are self-contained and the folder can live anywhere you can write to. On Linux, the launcher uses the desktop's own X11, Cairo, Pango and libcurl, and offers to install any that are missing.
Converted from Qwen/Qwen3.8-27B, in five sizes: perkel/Qwen3.8-27B-MC.
| Size | VRAM | KL divergence | Top-1 agreement | Good for |
|---|---|---|---|---|
| Tiny | 12.70 GiB | 0.0582 | 92.7% | the most speed and context |
| Small | 13.67 GiB | 0.0478 | 93.8% | speed, with a little more accuracy |
| Medium | 15.39 GiB | 0.0298 | 95.2% | the balance, and the presets' choice |
| Large | 18.66 GiB | 0.0081 | 97.7% | answers close to the original |
| XXL | 23.58 GiB | 0.0051 | 98.3% | the closest to the original, less room for context |
Both measured against Qwen's original BF16 weights on 81,880 held-out tokens. KL divergence: 0 is identical, lower is closer. Top-1 agreement: how often the most likely next token is the same.
The same five sizes also come uncensored, converted from an abliterated release (one with the refusals removed): perkel/Qwen3.8-27B-Uncensored-MC.
MegaCapybara is released as programs for Windows and Linux for now; the source code will be published later.
MegaCapybara is under the MIT License, (c) 2026 Perkel's Software Corner. The model weights and the
libraries it includes keep their own licenses: see THIRD-PARTY-NOTICES.md. Each release
lists its changes in its notes and in the package's CHANGELOG.txt.
If MegaCapybara is useful to you, you can support it on Patreon. The money goes to the GPUs needed to support more cards and multi-GPU.