eelgaev/Strata-AC922

Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

C++

0

764 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode (r/LocalLLaMA)

I forked Strata and worked with Claude code with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V100 SXM2 GPUs, the CPU-GPU BW advertised as 150GB/s…

5

Oct 4, 2026

README

[!NOTE] This is the ac922 fork of Niko1221/Strata for the IBM Power System AC922 (2x POWER9, 4x V100-SXM2 16 GB, NVLink 2.0, ppc64le, unified memory). It is experimental and not supported upstream. Everything below the line is upstream's README, unchanged.

Qwen3.8-Flash-Next UD-Q4_K_XL (llama-benchy pp2048/pp8192 at depth 0-64K, MTP --spec 4):

Reads your promptWrites answers
4x V1001,967-6,353 tok/s (2K-74K); 7,299 on a 249K prompt83.4 tok/s (85-87 greedy)
2x V100 (one socket)1,845-3,785 tok/s; 4,027 at 123K73.4 tok/s (72-75 greedy)

Peaks on 4x V100 with a 256K context (--max-context 262144; natural text, one request at a time):

Peak
Reads your prompt7,357 tok/s (135K tokens); 7,188 at 202K; 7,089 at 252K (35.5 s)
Writes answers (512 tokens, greedy)113 tok/s JSON, 103 code, 107 counting, 84 prose
Follow-up at 252K depththe whole prompt reused: first token after 0.26 s, then 60 tok/s

What the fork adds: a per-socket page-locked arena and NUMA-aware expert placement, the idle peer GPU fetching over its own NVLink, Volta tensor-core kernels (FP16 weights, prompt attention, fused W4A16 prompt experts, QSA selection), a pipelined layer split, and POWER9 VSX / SMT / per-socket CPU expert pools. Build, run, speed, quality: docs/IBM_AC922.md.

The fork also runs on an NVIDIA DGX Spark (GB10, ARM; experimental, tested with IQ2_XS and UD-Q4_K_XL): ./setup.sh compiles the engine there, and every expert fits on its GPU (decode 55-62 tok/s, prefill 928-1,515 tok/s). Details: docs/DGX_SPARK.md.


Strata

English · 简体中文 · 日本語 · Deutsch · Français · Español · Português

Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden that Strata's model wrote, running in the browser
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s)

Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.

How fast is it?

We measured it on two ordinary gaming PCs. A token is about ¾ of a word.

  • Writes answers: how fast the reply appears in a short chat. 60 tokens per second is faster than you can read.
  • Reads your prompt: how fast it takes in what you send (here a 32K-token document, code or chat history).
NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAMAMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
SizeWrites answersReads your prompt
Q2_094 tokens/s2,650 tokens/s
IQ2_XS79 tokens/s2,090 tokens/s
IQ3_XXS62 tokens/s1,750 tokens/s
IQ3_S53 tokens/s1,620 tokens/s
Coder55 tokens/s2,180 tokens/s
SizeWrites answersReads your prompt
Q2_060 tokens/s1,160 tokens/s
IQ2_XS52 tokens/s1,110 tokens/s
Coder44 tokens/s1,420 tokens/s

NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model, community results.

Buy Me A Coffee
Strata is free. If it runs well on your PC, a coffee keeps the work on it going.

What you need

Graphics cardNVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more.
RAM32 GB or more. Your RAM decides which model fits. 64 GB runs every size.
DiskAbout 80 GB free. Use an SSD if you can: the first start is much faster.
SystemWindows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD.

The installer sets up everything else. Two or three cards can share the model (multi-GPU).

Experimental, written and tested by community members on their own machines:

  • Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
  • Intel Arc, built from source on Linux: Intel Arc.
  • Older processors without AVX2: they work, but slowly. Older CPUs.

The full list: docs/INSTALL.md.

Install

Let your AI set it up

Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:

Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server.

Or do it yourself

Download Strata and unzip it (or git clone it). Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder.

The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions:

  • which model and which size,
  • how much context (how much text the model keeps in mind),
  • whether it should read pictures.

Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the download stops, run it again: it continues where it left off. Your browser opens the Strata app at http://127.0.0.1:8080.

While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The window shows what Strata is doing.

Next time, run START-HERE.bat (or ./setup.sh) again. It starts right away and downloads nothing twice. Close its window to stop the model. UPDATE.bat (./update.sh) updates Strata without starting it. Updating, Docker, several cards, where the files go and every option: docs/INSTALL.md.

Which model should I pick?

The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter.

Your RAMTakeWhy
32 GBCoderit fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too)
48 GBIQ2_XS (or Q2_0, the fastest)the larger sizes do not fit
64 GBIQ2_XS (recommended), or IQ3_XXS / IQ3_Severy size fits; IQ3_S is the best and the slowest
96 GB or moreIQ3_S, or Unsloth's UD-IQ4_XS (~4-bit)room for the largest sizes with everything else open
  • Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. It is weaker outside code, including Chinese and other CJK text (#438). For those, take Q2_0, IQ2_XS or IQ3_S, which keep every expert.
  • Swift 1.5: a fine-tune that thinks for a much shorter time before it answers. You get the answer sooner, at about the same quality.
  • Unsloth UD-IQ4_XS: Unsloth's ~4-bit version, between IQ3_S and UD-Q4_K_XL in quality. A 94 GB download. With less than ~80 GB of RAM, Strata reads part of it from the SSD while it answers, so it is slower there (an NVMe SSD helps).
  • Unsloth UD-Q4_K_XL (experimental): the closest to the full model. But Strata reads most of it from the SSD while it answers, so it writes only 7-8.5 tokens/s on a 64 GB PC.
  • OrcaRouter's Uncensored IQ3_XXS: you set it up by hand. It is not in the installer's menu.

Sizes, downloads and what fits where: docs/MODELS.md. To add another model later, run SETUP.bat (Linux: ./setup.sh --setup).

Using it

The Strata app's Monitor tab next to a coding agent
The Strata app's Monitor (left) while a coding agent writes the pagoda garden from the video (right)

  • In the browser: open http://127.0.0.1:8080. It has Chat, a live Monitor of the model and your GPU/CPU/RAM, and About with the settings and addresses.
  • Your apps and coding agents: add an "OpenAI-compatible" provider with the base URL http://127.0.0.1:8080/v1. Any API key and any model name work.
    • Apps that use Anthropic's API: http://127.0.0.1:8080/v1/messages (Claude Code: ANTHROPIC_BASE_URL=http://127.0.0.1:8080).
    • Codex CLI and other apps that use the OpenAI Responses API: /v1/responses (setup).
  • Thinking: choose off, low, medium or high in the chat menu or in your app's "reasoning effort". Off is the fastest. High is best for hard questions.
  • Pictures: say yes to "Images?" in setup. Then click Picture in the chat, or attach pictures in your app. AMD cards read pictures on Linux through the processor; on Windows they can't yet.
  • From your phone or another PC: START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>. Always set a key.
  • One request at a time: by default Strata answers one request, and the others wait. To answer several at once, set "parallel": 2 (BATCHING.md). On a 12 GB card this makes each answer slower.
  • Long prompts: Strata reads the first message of a chat in full, about 1 minute per 30,000 tokens. Follow-up messages start in seconds.

More: where your chats are stored, the API.

Something went wrong?

  • My PC froze the first time Strata started. This is normal while it loads the model. Wait, and don't close the window. Still frozen after 10 minutes? Restart the PC, close other programs and try again, or pick a smaller size.
  • It stopped while downloading or installing. Run START-HERE.bat (or ./setup.sh) again. It continues where it stopped.
  • It's very slow and the disk light keeps blinking, or it says "the engine stopped unexpectedly". Your PC does not have enough free RAM. Close other programs (browsers use a lot), or pick a smaller size (Q2_0 or IQ2_XS).
  • It says port 8080 is already in use. Strata is already running. Look for its window.

More problems and their fixes: docs/TROUBLESHOOTING.md. Still stuck? Open an issue and attach strata-<model>.log from the Strata folder. Found a security problem? Report it privately: SECURITY.md.

How does it work?

Models like this one usually run on servers with hundreds of gigabytes of graphics memory. Your graphics card has 12-24 GB. Strata makes the model fit by sharing the work across your whole PC. Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.

The model's 24,576 experts: the busiest on the graphics card, all of them in RAM, a lookup table on the SSD

  • The model is a team of 24,576 small specialists ("experts"). Each word needs only 10 of them.
  • Your graphics card keeps the few thousand experts that are used most often. Your RAM holds all of them, and your processor works on the rest at the same time. Your SSD holds a big lookup table.

A small helper guesses the next words; the big model checks them all at once and keeps the right ones

  • Guess, then check: a small helper guesses the next few words. The big model checks them all at once. You get the same answer, 1.6-1.8x sooner.
  • Long texts are read in big pieces (up to 8,192 tokens at a time), at over 1,000 tokens per second.

The longer explanation: docs/HOW_IT_WORKS.md. Every part and its numbers: the details and the paper.

Credits and license

The model is Qwen3.8-Flash-Next by the Qwen team. It was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. Strata uses parts of llama.cpp / ggml. All credits: docs/HOW_IT_WORKS.md. Strata is open source under the MIT License. A few parts and every model have their own licenses (which ones).

Support Strata

Strata is free and open source. If it is useful to you, you can support its development:

Buy Me A Coffee

eelgaev/Strata-AC922

Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

C++

0

764 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode (r/LocalLLaMA)

I forked Strata and worked with Claude code with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V100 SXM2 GPUs, the CPU-GPU BW advertised as 150GB/s…

5

Oct 4, 2026

README

[!NOTE] This is the ac922 fork of Niko1221/Strata for the IBM Power System AC922 (2x POWER9, 4x V100-SXM2 16 GB, NVLink 2.0, ppc64le, unified memory). It is experimental and not supported upstream. Everything below the line is upstream's README, unchanged.

Qwen3.8-Flash-Next UD-Q4_K_XL (llama-benchy pp2048/pp8192 at depth 0-64K, MTP --spec 4):

Reads your promptWrites answers
4x V1001,967-6,353 tok/s (2K-74K); 7,299 on a 249K prompt83.4 tok/s (85-87 greedy)
2x V100 (one socket)1,845-3,785 tok/s; 4,027 at 123K73.4 tok/s (72-75 greedy)

Peaks on 4x V100 with a 256K context (--max-context 262144; natural text, one request at a time):

Peak
Reads your prompt7,357 tok/s (135K tokens); 7,188 at 202K; 7,089 at 252K (35.5 s)
Writes answers (512 tokens, greedy)113 tok/s JSON, 103 code, 107 counting, 84 prose
Follow-up at 252K depththe whole prompt reused: first token after 0.26 s, then 60 tok/s

What the fork adds: a per-socket page-locked arena and NUMA-aware expert placement, the idle peer GPU fetching over its own NVLink, Volta tensor-core kernels (FP16 weights, prompt attention, fused W4A16 prompt experts, QSA selection), a pipelined layer split, and POWER9 VSX / SMT / per-socket CPU expert pools. Build, run, speed, quality: docs/IBM_AC922.md.

The fork also runs on an NVIDIA DGX Spark (GB10, ARM; experimental, tested with IQ2_XS and UD-Q4_K_XL): ./setup.sh compiles the engine there, and every expert fits on its GPU (decode 55-62 tok/s, prefill 928-1,515 tok/s). Details: docs/DGX_SPARK.md.


Strata

English · 简体中文 · 日本語 · Deutsch · Français · Español · Português

Run a 125-billion-parameter AI model on your own gaming PC
NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden that Strata's model wrote, running in the browser
A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s)

Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.

How fast is it?

We measured it on two ordinary gaming PCs. A token is about ¾ of a word.

  • Writes answers: how fast the reply appears in a short chat. 60 tokens per second is faster than you can read.
  • Reads your prompt: how fast it takes in what you send (here a 32K-token document, code or chat history).
NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAMAMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
SizeWrites answersReads your prompt
Q2_094 tokens/s2,650 tokens/s
IQ2_XS79 tokens/s2,090 tokens/s
IQ3_XXS62 tokens/s1,750 tokens/s
IQ3_S53 tokens/s1,620 tokens/s
Coder55 tokens/s2,180 tokens/s
SizeWrites answersReads your prompt
Q2_060 tokens/s1,160 tokens/s
IQ2_XS52 tokens/s1,110 tokens/s
Coder44 tokens/s1,420 tokens/s

NVIDIA: Q2_0 with engine 0.1.36, the other rows with 0.1.26 (4K answers, 32K prompts). The full tables are in DETAILS.md. A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and other cards: speed of each model, community results.

Buy Me A Coffee
Strata is free. If it runs well on your PC, a coffee keeps the work on it going.

What you need

Graphics cardNVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more.
RAM32 GB or more. Your RAM decides which model fits. 64 GB runs every size.
DiskAbout 80 GB free. Use an SSD if you can: the first start is much faster.
SystemWindows 10 / 11 or Linux, and a current graphics driver from NVIDIA or AMD.

The installer sets up everything else. Two or three cards can share the model (multi-GPU).

Experimental, written and tested by community members on their own machines:

  • Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
  • Intel Arc, built from source on Linux: Intel Arc.
  • Older processors without AVX2: they work, but slowly. Older CPUs.

The full list: docs/INSTALL.md.

Install

Let your AI set it up

Do you use an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:

Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

It checks your graphics card, RAM and disk and picks the model that fits. Then it installs and starts it and tells you how to connect your apps. AI tools can also install, start and stop Strata through its MCP server.

Or do it yourself

Download Strata and unzip it (or git clone it). Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder.

The steps are the same for NVIDIA and AMD. The installer finds your card and sets up the right engine for it. It asks you a few questions:

  • which model and which size,
  • how much context (how much text the model keeps in mind),
  • whether it should read pictures.

Press Enter each time for the recommended answer. Then it downloads the model (about 70 GB) and starts it. If the download stops, run it again: it continues where it left off. Your browser opens the Strata app at http://127.0.0.1:8080.

While the model starts, your PC can be slow or stop responding for 1-3 minutes (longest the first time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The window shows what Strata is doing.

Next time, run START-HERE.bat (or ./setup.sh) again. It starts right away and downloads nothing twice. Close its window to stop the model. UPDATE.bat (./update.sh) updates Strata without starting it. Updating, Docker, several cards, where the files go and every option: docs/INSTALL.md.

Which model should I pick?

The installer recommends one for your RAM. The same model comes in several sizes, compressed more or less. Smaller sizes are faster. Larger sizes are a bit smarter.

Your RAMTakeWhy
32 GBCoderit fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too)
48 GBIQ2_XS (or Q2_0, the fastest)the larger sizes do not fit
64 GBIQ2_XS (recommended), or IQ3_XXS / IQ3_Severy size fits; IQ3_S is the best and the slowest
96 GB or moreIQ3_S, or Unsloth's UD-IQ4_XS (~4-bit)room for the largest sizes with everything else open
  • Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM. It is weaker outside code, including Chinese and other CJK text (#438). For those, take Q2_0, IQ2_XS or IQ3_S, which keep every expert.
  • Swift 1.5: a fine-tune that thinks for a much shorter time before it answers. You get the answer sooner, at about the same quality.
  • Unsloth UD-IQ4_XS: Unsloth's ~4-bit version, between IQ3_S and UD-Q4_K_XL in quality. A 94 GB download. With less than ~80 GB of RAM, Strata reads part of it from the SSD while it answers, so it is slower there (an NVMe SSD helps).
  • Unsloth UD-Q4_K_XL (experimental): the closest to the full model. But Strata reads most of it from the SSD while it answers, so it writes only 7-8.5 tokens/s on a 64 GB PC.
  • OrcaRouter's Uncensored IQ3_XXS: you set it up by hand. It is not in the installer's menu.

Sizes, downloads and what fits where: docs/MODELS.md. To add another model later, run SETUP.bat (Linux: ./setup.sh --setup).

Using it

The Strata app's Monitor tab next to a coding agent
The Strata app's Monitor (left) while a coding agent writes the pagoda garden from the video (right)

  • In the browser: open http://127.0.0.1:8080. It has Chat, a live Monitor of the model and your GPU/CPU/RAM, and About with the settings and addresses.
  • Your apps and coding agents: add an "OpenAI-compatible" provider with the base URL http://127.0.0.1:8080/v1. Any API key and any model name work.
    • Apps that use Anthropic's API: http://127.0.0.1:8080/v1/messages (Claude Code: ANTHROPIC_BASE_URL=http://127.0.0.1:8080).
    • Codex CLI and other apps that use the OpenAI Responses API: /v1/responses (setup).
  • Thinking: choose off, low, medium or high in the chat menu or in your app's "reasoning effort". Off is the fastest. High is best for hard questions.
  • Pictures: say yes to "Images?" in setup. Then click Picture in the chat, or attach pictures in your app. AMD cards read pictures on Linux through the processor; on Windows they can't yet.
  • From your phone or another PC: START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>. Always set a key.
  • One request at a time: by default Strata answers one request, and the others wait. To answer several at once, set "parallel": 2 (BATCHING.md). On a 12 GB card this makes each answer slower.
  • Long prompts: Strata reads the first message of a chat in full, about 1 minute per 30,000 tokens. Follow-up messages start in seconds.

More: where your chats are stored, the API.

Something went wrong?

  • My PC froze the first time Strata started. This is normal while it loads the model. Wait, and don't close the window. Still frozen after 10 minutes? Restart the PC, close other programs and try again, or pick a smaller size.
  • It stopped while downloading or installing. Run START-HERE.bat (or ./setup.sh) again. It continues where it stopped.
  • It's very slow and the disk light keeps blinking, or it says "the engine stopped unexpectedly". Your PC does not have enough free RAM. Close other programs (browsers use a lot), or pick a smaller size (Q2_0 or IQ2_XS).
  • It says port 8080 is already in use. Strata is already running. Look for its window.

More problems and their fixes: docs/TROUBLESHOOTING.md. Still stuck? Open an issue and attach strata-<model>.log from the Strata folder. Found a security problem? Report it privately: SECURITY.md.

How does it work?

Models like this one usually run on servers with hundreds of gigabytes of graphics memory. Your graphics card has 12-24 GB. Strata makes the model fit by sharing the work across your whole PC. Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.

The model's 24,576 experts: the busiest on the graphics card, all of them in RAM, a lookup table on the SSD

  • The model is a team of 24,576 small specialists ("experts"). Each word needs only 10 of them.
  • Your graphics card keeps the few thousand experts that are used most often. Your RAM holds all of them, and your processor works on the rest at the same time. Your SSD holds a big lookup table.

A small helper guesses the next words; the big model checks them all at once and keeps the right ones

  • Guess, then check: a small helper guesses the next few words. The big model checks them all at once. You get the same answer, 1.6-1.8x sooner.
  • Long texts are read in big pieces (up to 8,192 tokens at a time), at over 1,000 tokens per second.

The longer explanation: docs/HOW_IT_WORKS.md. Every part and its numbers: the details and the paper.

Credits and license

The model is Qwen3.8-Flash-Next by the Qwen team. It was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. Strata uses parts of llama.cpp / ggml. All credits: docs/HOW_IT_WORKS.md. Strata is open source under the MIT License. A few parts and every model have their own licenses (which ones).

Support Strata

Strata is free and open source. If it is useful to you, you can support its development:

Buy Me A Coffee

Languages

C++

73.0%

Python

14.0%

Cuda

10.8%