Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
C++
0
711 commits
updated Oct 4, 2026
English · 简体中文 · 日本語 · Deutsch · Français · Español · Português
Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) ·
full video (49 s)
Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.
We measured it on two ordinary gaming PCs. A token is about ¾ of a word.
| NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM | AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model, community results.

Strata is free. If it runs well on your PC, a coffee keeps the work on it going.
| Graphics card | NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more. |
| RAM | 32 GB or more. Your RAM decides which model fits. 64 GB runs every size. |
| Disk | About 80 GB free. Use an SSD if you can: the first start is much faster. |
| System | Windows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD. |
The installer sets up everything else. Two or three cards can share the model (multi-GPU).
Experimental, written and tested by community members on their own machines:
The full list: docs/INSTALL.md.
Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server.
Download Strata and unzip it (or git clone it).
Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder.
The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions:
Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the
download stops, run it again: it continues where it left off. Your browser opens the Strata app at
http://127.0.0.1:8080.
While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The window shows what Strata is doing.
Next time, run START-HERE.bat (or ./setup.sh) again. It starts right away and downloads nothing twice. Close
its window to stop the model. UPDATE.bat (./update.sh) updates Strata without starting it. Updating, Docker,
several cards, where the files go and every option: docs/INSTALL.md.
The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter.
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too) |
| 48 GB | IQ2_XS (or Q2_0, the fastest) | the larger sizes do not fit |
| 64 GB | IQ2_XS (recommended), or IQ3_XXS / IQ3_S | every size fits; IQ3_S is the best and the slowest |
| 96 GB or more | IQ3_S, or Unsloth's UD-IQ4_XS (~4-bit) | room for the largest sizes with everything else open |
Sizes, downloads and what fits where: docs/MODELS.md. To add another model later, run
SETUP.bat (Linux: ./setup.sh --setup).

The Strata app's Monitor (left) while a coding agent writes the pagoda garden from the video (right)
http://127.0.0.1:8080. It has Chat, a live Monitor of the model and your
GPU/CPU/RAM, and About with the settings and addresses.http://127.0.0.1:8080/v1. Any API key and any model name work.
http://127.0.0.1:8080/v1/messages (Claude Code:
ANTHROPIC_BASE_URL=http://127.0.0.1:8080)./v1/responses
(setup).START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>. Always set a key."parallel": 2 (BATCHING.md). On a 12 GB card this makes each answer slower.More: where your chats are stored, the API.
START-HERE.bat (or ./setup.sh) again. It continues where
it stopped.More problems and their fixes: docs/TROUBLESHOOTING.md. Still stuck? Open an
issue and attach strata-<model>.log from the Strata folder. Found a
security problem? Report it privately: SECURITY.md.
Models like this one usually run on servers with hundreds of gigabytes of graphics memory. Your graphics card has 12-24 GB. Strata makes the model fit by sharing the work across your whole PC. Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.
The longer explanation: docs/HOW_IT_WORKS.md. Every part and its numbers: the details and the paper.
The model is Qwen3.8-Flash-Next by the Qwen team. It was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. Strata uses parts of llama.cpp / ggml. All credits: docs/HOW_IT_WORKS.md. Strata is open source under the MIT License. A few parts and every model have their own licenses (which ones).
Strata is free and open source. If it is useful to you, you can support its development:
C++
73.3%
Python
14.1%
Cuda
10.3%
Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
C++
0
711 commits
updated Oct 4, 2026
English · 简体中文 · 日本語 · Deutsch · Français · Español · Português
Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) ·
full video (49 s)
Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.
We measured it on two ordinary gaming PCs. A token is about ¾ of a word.
| NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM | AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model, community results.

Strata is free. If it runs well on your PC, a coffee keeps the work on it going.
| Graphics card | NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more. |
| RAM | 32 GB or more. Your RAM decides which model fits. 64 GB runs every size. |
| Disk | About 80 GB free. Use an SSD if you can: the first start is much faster. |
| System | Windows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD. |
The installer sets up everything else. Two or three cards can share the model (multi-GPU).
Experimental, written and tested by community members on their own machines:
The full list: docs/INSTALL.md.
Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server.
Download Strata and unzip it (or git clone it).
Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder.
The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions:
Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the
download stops, run it again: it continues where it left off. Your browser opens the Strata app at
http://127.0.0.1:8080.
While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The window shows what Strata is doing.
Next time, run START-HERE.bat (or ./setup.sh) again. It starts right away and downloads nothing twice. Close
its window to stop the model. UPDATE.bat (./update.sh) updates Strata without starting it. Updating, Docker,
several cards, where the files go and every option: docs/INSTALL.md.
The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter.
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too) |
| 48 GB | IQ2_XS (or Q2_0, the fastest) | the larger sizes do not fit |
| 64 GB | IQ2_XS (recommended), or IQ3_XXS / IQ3_S | every size fits; IQ3_S is the best and the slowest |
| 96 GB or more | IQ3_S, or Unsloth's UD-IQ4_XS (~4-bit) | room for the largest sizes with everything else open |
Sizes, downloads and what fits where: docs/MODELS.md. To add another model later, run
SETUP.bat (Linux: ./setup.sh --setup).

The Strata app's Monitor (left) while a coding agent writes the pagoda garden from the video (right)
http://127.0.0.1:8080. It has Chat, a live Monitor of the model and your
GPU/CPU/RAM, and About with the settings and addresses.http://127.0.0.1:8080/v1. Any API key and any model name work.
http://127.0.0.1:8080/v1/messages (Claude Code:
ANTHROPIC_BASE_URL=http://127.0.0.1:8080)./v1/responses
(setup).START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>. Always set a key."parallel": 2 (BATCHING.md). On a 12 GB card this makes each answer slower.More: where your chats are stored, the API.
START-HERE.bat (or ./setup.sh) again. It continues where
it stopped.More problems and their fixes: docs/TROUBLESHOOTING.md. Still stuck? Open an
issue and attach strata-<model>.log from the Strata folder. Found a
security problem? Report it privately: SECURITY.md.
Models like this one usually run on servers with hundreds of gigabytes of graphics memory. Your graphics card has 12-24 GB. Strata makes the model fit by sharing the work across your whole PC. Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.
The longer explanation: docs/HOW_IT_WORKS.md. Every part and its numbers: the details and the paper.
The model is Qwen3.8-Flash-Next by the Qwen team. It was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. Strata uses parts of llama.cpp / ggml. All credits: docs/HOW_IT_WORKS.md. Strata is open source under the MIT License. A few parts and every model have their own licenses (which ones).
Strata is free and open source. If it is useful to you, you can support its development:
C++
73.3%
Python
14.1%
Cuda
10.3%