A 27-billion-parameter text-generation model designed for a smaller GPU-memory footprint. The model file is 8.62 GB. The recommended starting point is an NVIDIA RTX 30, 40, or 50 series GPU with at least 12 GB of VRAM, using the L0xRE BeeLLama runtime.
“27b” is the parameter count, not the download size. “Low” describes the memory-focused model build. Context length and available GPU memory depend on the card, driver, runtime profile, and other applications.
Start here: runtime downloads, installation, and GPU test results. There is one Linux/WSL runtime archive and one Windows runtime ZIP for the supported GPU families.
| File | What it does | Download size |
|---|---|---|
| L0xRE-27b-Low.gguf | The main model. This is the recommended model file. | 8.62 GB |
| Qwen3.8-27B-DFlash2-Q4_K_M.gguf | Optional speedup helper, recommended with the main model. The runtime uses it to propose tokens that the main model verifies. | 1.14 GB |
For the recommended setup, download both files: 9.76 GB in total, plus the runtime package. The helper is called a drafter; it is an additional file, not another name or edition of L0xRE-27b-Low. The main model can also run without a drafter, with speculative decoding disabled.
The smaller Qwen3.8-27B-DFlash2-Q2_K.gguf file is an optional 705 MB drafter for memory-constrained experiments. Start with the recommended Q4_K_M drafter to match the published default configuration.
models folder.# Linux / WSL
./l0xre serve --profile 12gb -m models/L0xRE-27b-Low.gguf \
-md models/Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080
# Windows
.\l0xre.cmd serve --profile 12gb -m models\L0xRE-27b-Low.gguf -md models\Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080
When ready, the server's health endpoint is http://127.0.0.1:8080/health and its API base is http://127.0.0.1:8080/v1. For a compatible chat client, use that API base and select L0xRE-27b-Low from the server's /v1/models endpoint.
This file requires the L0xRE runtime. Direct use with stock llama.cpp, Ollama, LM Studio, Transformers, or other inference engines is outside this release's supported setup. The GGUF extension alone does not establish compatibility.
The runtime targets NVIDIA compute capabilities 8.6, 8.9, and 12.0. A 12 GB VRAM target does not guarantee that the longest context fits on every card. Close other GPU-heavy applications or reduce context when memory is tight.
| GPU / OS | Default 12gb profile | Evidence status |
|---|---|---|
| RTX 3060 12 GB / Linux | 98,304-token context with the recommended drafter | Measured model/runtime components; only 215 MiB minimum free VRAM in the long-context test |
| RTX 4090 / Linux | 81,920-token context with the recommended drafter | Published configuration measurements; additional 16gb and full32k profiles documented in the runtime guide |
| RTX 50 series / Linux or WSL | 8,192-token context with drafter; larger ordinary-decode profile available | GPU runtime payload included; the current non-MTP model path still awaits hardware qualification |
| RTX 30–50 series / Windows | GPU-specific presets selected automatically | Builds, dependencies and launchers verified; GPU inference, throughput and memory headroom still await qualification |
The universal runtime is a release candidate. The runtime page lists exactly what was tested. Windows and unmeasured GPU/model combinations should not inherit Linux benchmark claims.
These are workload-specific measurements, not a controlled comparison between GPUs or a guaranteed speed for every prompt.
| Hardware / configuration | Code generation | Narrative generation | Measurement scope |
|---|---|---|---|
| RTX 3060 12 GB, Linux, 96K context, recommended drafter | 39.046 tokens/s | 29.422 tokens/s | Temperature 0, seed 0, exactly 800 output tokens; two measured runs per prompt |
RTX 4090, Linux, 12gb profile, recommended drafter | 115.4 ± 2.4 tokens/s | Not reported for this configuration | Five measured code-generation runs; 81,920-token allocated context |
The RTX 3060 long-context test processed 92,879 input tokens plus 128 generated tokens and matched the recorded reference output. Prefill was 272.60 tokens/s and decode 21.06 tokens/s, with just 215 MiB minimum free VRAM. This checks a particular runtime workload; it is not a broad model-quality evaluation.
Quality scores from other checkpoints in the model family are not presented as an evaluation of this exact downloadable file. No external leaderboard result is claimed here.
Download sizes use decimal GB. Model weights have not changed during the runtime documentation update.
| File | Bytes | SHA-256 |
|---|---|---|
L0xRE-27b-Low.gguf | 8,619,127,680 | b0849250c633aa93853bf119a877dbdafdd7b1a4ebb672bd6b6a1439906f3543 |
Qwen3.8-27B-DFlash2-Q4_K_M.gguf | 1,143,006,816 | 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd |
Qwen3.8-27B-DFlash2-Q2_K.gguf | 705,430,880 | e3eb7705404817cdbcdabe56049a1952b3b37bcc8df6e4d4efaec5d41563fb7e |
L0xRE-27b-Low-MTP is a separate, larger checkpoint with an additional prediction block. Its file hash, runtime setup and qualification are distinct. Use the main model file on this page for the recommended drafter-based setup.
L0xRE-27b-Low is a quantized derivative of Qwen3.8-27B, with low-bit projections and higher-precision embeddings/output tensors. The model's configured native context capacity is 262,144 tokens; supported and tested serving contexts are determined by the runtime profiles above.
Model weights remain under the base model's research/community terms: license. The L0xRE runtime tooling is Apache-2.0; applicable runtime and third-party notices ship with its packages. Model licensing and runtime licensing are separate.
A 27-billion-parameter text-generation model designed for a smaller GPU-memory footprint. The model file is 8.62 GB. The recommended starting point is an NVIDIA RTX 30, 40, or 50 series GPU with at least 12 GB of VRAM, using the L0xRE BeeLLama runtime.
“27b” is the parameter count, not the download size. “Low” describes the memory-focused model build. Context length and available GPU memory depend on the card, driver, runtime profile, and other applications.
Start here: runtime downloads, installation, and GPU test results. There is one Linux/WSL runtime archive and one Windows runtime ZIP for the supported GPU families.
| File | What it does | Download size |
|---|---|---|
| L0xRE-27b-Low.gguf | The main model. This is the recommended model file. | 8.62 GB |
| Qwen3.8-27B-DFlash2-Q4_K_M.gguf | Optional speedup helper, recommended with the main model. The runtime uses it to propose tokens that the main model verifies. | 1.14 GB |
For the recommended setup, download both files: 9.76 GB in total, plus the runtime package. The helper is called a drafter; it is an additional file, not another name or edition of L0xRE-27b-Low. The main model can also run without a drafter, with speculative decoding disabled.
The smaller Qwen3.8-27B-DFlash2-Q2_K.gguf file is an optional 705 MB drafter for memory-constrained experiments. Start with the recommended Q4_K_M drafter to match the published default configuration.
models folder.# Linux / WSL
./l0xre serve --profile 12gb -m models/L0xRE-27b-Low.gguf \
-md models/Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080
# Windows
.\l0xre.cmd serve --profile 12gb -m models\L0xRE-27b-Low.gguf -md models\Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080
When ready, the server's health endpoint is http://127.0.0.1:8080/health and its API base is http://127.0.0.1:8080/v1. For a compatible chat client, use that API base and select L0xRE-27b-Low from the server's /v1/models endpoint.
This file requires the L0xRE runtime. Direct use with stock llama.cpp, Ollama, LM Studio, Transformers, or other inference engines is outside this release's supported setup. The GGUF extension alone does not establish compatibility.
The runtime targets NVIDIA compute capabilities 8.6, 8.9, and 12.0. A 12 GB VRAM target does not guarantee that the longest context fits on every card. Close other GPU-heavy applications or reduce context when memory is tight.
| GPU / OS | Default 12gb profile | Evidence status |
|---|---|---|
| RTX 3060 12 GB / Linux | 98,304-token context with the recommended drafter | Measured model/runtime components; only 215 MiB minimum free VRAM in the long-context test |
| RTX 4090 / Linux | 81,920-token context with the recommended drafter | Published configuration measurements; additional 16gb and full32k profiles documented in the runtime guide |
| RTX 50 series / Linux or WSL | 8,192-token context with drafter; larger ordinary-decode profile available | GPU runtime payload included; the current non-MTP model path still awaits hardware qualification |
| RTX 30–50 series / Windows | GPU-specific presets selected automatically | Builds, dependencies and launchers verified; GPU inference, throughput and memory headroom still await qualification |
The universal runtime is a release candidate. The runtime page lists exactly what was tested. Windows and unmeasured GPU/model combinations should not inherit Linux benchmark claims.
These are workload-specific measurements, not a controlled comparison between GPUs or a guaranteed speed for every prompt.
| Hardware / configuration | Code generation | Narrative generation | Measurement scope |
|---|---|---|---|
| RTX 3060 12 GB, Linux, 96K context, recommended drafter | 39.046 tokens/s | 29.422 tokens/s | Temperature 0, seed 0, exactly 800 output tokens; two measured runs per prompt |
RTX 4090, Linux, 12gb profile, recommended drafter | 115.4 ± 2.4 tokens/s | Not reported for this configuration | Five measured code-generation runs; 81,920-token allocated context |
The RTX 3060 long-context test processed 92,879 input tokens plus 128 generated tokens and matched the recorded reference output. Prefill was 272.60 tokens/s and decode 21.06 tokens/s, with just 215 MiB minimum free VRAM. This checks a particular runtime workload; it is not a broad model-quality evaluation.
Quality scores from other checkpoints in the model family are not presented as an evaluation of this exact downloadable file. No external leaderboard result is claimed here.
Download sizes use decimal GB. Model weights have not changed during the runtime documentation update.
| File | Bytes | SHA-256 |
|---|---|---|
L0xRE-27b-Low.gguf | 8,619,127,680 | b0849250c633aa93853bf119a877dbdafdd7b1a4ebb672bd6b6a1439906f3543 |
Qwen3.8-27B-DFlash2-Q4_K_M.gguf | 1,143,006,816 | 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd |
Qwen3.8-27B-DFlash2-Q2_K.gguf | 705,430,880 | e3eb7705404817cdbcdabe56049a1952b3b37bcc8df6e4d4efaec5d41563fb7e |
L0xRE-27b-Low-MTP is a separate, larger checkpoint with an additional prediction block. Its file hash, runtime setup and qualification are distinct. Use the main model file on this page for the recommended drafter-based setup.
L0xRE-27b-Low is a quantized derivative of Qwen3.8-27B, with low-bit projections and higher-precision embeddings/output tensors. The model's configured native context capacity is 262,144 tokens; supported and tested serving contexts are determined by the runtime profiles above.
Model weights remain under the base model's research/community terms: license. The L0xRE runtime tooling is Apache-2.0; applicable runtime and third-party notices ship with its packages. Model licensing and runtime licensing are separate.