YourHighnessLA/L0xRE-27b-Low

Model

L0xRE-27b-Low

0

6 commits

2 linked in READMEs

updated Sep 30, 2026

See the code

README

L0xRE-27b-Low

A 27-billion-parameter text-generation model designed for a smaller GPU-memory footprint. The model file is 8.62 GB. The recommended starting point is an NVIDIA RTX 30, 40, or 50 series GPU with at least 12 GB of VRAM, using the L0xRE BeeLLama runtime.

“27b” is the parameter count, not the download size. “Low” describes the memory-focused model build. Context length and available GPU memory depend on the card, driver, runtime profile, and other applications.

Start here: runtime downloads, installation, and GPU test results. There is one Linux/WSL runtime archive and one Windows runtime ZIP for the supported GPU families.

Download these two files

FileWhat it doesDownload size
L0xRE-27b-Low.ggufThe main model. This is the recommended model file.8.62 GB
Qwen3.8-27B-DFlash2-Q4_K_M.ggufOptional speedup helper, recommended with the main model. The runtime uses it to propose tokens that the main model verifies.1.14 GB

For the recommended setup, download both files: 9.76 GB in total, plus the runtime package. The helper is called a drafter; it is an additional file, not another name or edition of L0xRE-27b-Low. The main model can also run without a drafter, with speculative decoding disabled.

The smaller Qwen3.8-27B-DFlash2-Q2_K.gguf file is an optional 705 MB drafter for memory-constrained experiments. Start with the recommended Q4_K_M drafter to match the published default configuration.

Install and run

  1. Download and extract the runtime for your OS from the universal release.
  2. Download the two recommended files above into the extracted runtime's models folder.
  3. Verify the runtime and model checksums using the installation guide.
  4. Start the server from the extracted runtime folder:
# Linux / WSL
./l0xre serve --profile 12gb -m models/L0xRE-27b-Low.gguf \
  -md models/Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080
# Windows
.\l0xre.cmd serve --profile 12gb -m models\L0xRE-27b-Low.gguf -md models\Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080

When ready, the server's health endpoint is http://127.0.0.1:8080/health and its API base is http://127.0.0.1:8080/v1. For a compatible chat client, use that API base and select L0xRE-27b-Low from the server's /v1/models endpoint.

This file requires the L0xRE runtime. Direct use with stock llama.cpp, Ollama, LM Studio, Transformers, or other inference engines is outside this release's supported setup. The GGUF extension alone does not establish compatibility.

GPU support and context

The runtime targets NVIDIA compute capabilities 8.6, 8.9, and 12.0. A 12 GB VRAM target does not guarantee that the longest context fits on every card. Close other GPU-heavy applications or reduce context when memory is tight.

GPU / OSDefault 12gb profileEvidence status
RTX 3060 12 GB / Linux98,304-token context with the recommended drafterMeasured model/runtime components; only 215 MiB minimum free VRAM in the long-context test
RTX 4090 / Linux81,920-token context with the recommended drafterPublished configuration measurements; additional 16gb and full32k profiles documented in the runtime guide
RTX 50 series / Linux or WSL8,192-token context with drafter; larger ordinary-decode profile availableGPU runtime payload included; the current non-MTP model path still awaits hardware qualification
RTX 30–50 series / WindowsGPU-specific presets selected automaticallyBuilds, dependencies and launchers verified; GPU inference, throughput and memory headroom still await qualification

The universal runtime is a release candidate. The runtime page lists exactly what was tested. Windows and unmeasured GPU/model combinations should not inherit Linux benchmark claims.

Measured speed

These are workload-specific measurements, not a controlled comparison between GPUs or a guaranteed speed for every prompt.

Hardware / configurationCode generationNarrative generationMeasurement scope
RTX 3060 12 GB, Linux, 96K context, recommended drafter39.046 tokens/s29.422 tokens/sTemperature 0, seed 0, exactly 800 output tokens; two measured runs per prompt
RTX 4090, Linux, 12gb profile, recommended drafter115.4 ± 2.4 tokens/sNot reported for this configurationFive measured code-generation runs; 81,920-token allocated context

The RTX 3060 long-context test processed 92,879 input tokens plus 128 generated tokens and matched the recorded reference output. Prefill was 272.60 tokens/s and decode 21.06 tokens/s, with just 215 MiB minimum free VRAM. This checks a particular runtime workload; it is not a broad model-quality evaluation.

Quality scores from other checkpoints in the model family are not presented as an evaluation of this exact downloadable file. No external leaderboard result is claimed here.

File integrity

Download sizes use decimal GB. Model weights have not changed during the runtime documentation update.

FileBytesSHA-256
L0xRE-27b-Low.gguf8,619,127,680b0849250c633aa93853bf119a877dbdafdd7b1a4ebb672bd6b6a1439906f3543
Qwen3.8-27B-DFlash2-Q4_K_M.gguf1,143,006,8161a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd
Qwen3.8-27B-DFlash2-Q2_K.gguf705,430,880e3eb7705404817cdbcdabe56049a1952b3b37bcc8df6e4d4efaec5d41563fb7e

L0xRE-27b-Low-MTP is a separate, larger checkpoint with an additional prediction block. Its file hash, runtime setup and qualification are distinct. Use the main model file on this page for the recommended drafter-based setup.

Model origin and license

L0xRE-27b-Low is a quantized derivative of Qwen3.8-27B, with low-bit projections and higher-precision embeddings/output tensors. The model's configured native context capacity is 262,144 tokens; supported and tested serving contexts are determined by the runtime profiles above.

Model weights remain under the base model's research/community terms: license. The L0xRE runtime tooling is Apache-2.0; applicable runtime and third-party notices ship with its packages. Model licensing and runtime licensing are separate.

12gb-vram
conversational
endpoints_compatible
gguf
l0xre
speculative-decoding
text-generation

YourHighnessLA/L0xRE-27b-Low

Model

L0xRE-27b-Low

0

6 commits

2 linked in READMEs

updated Sep 30, 2026

See the code

README

L0xRE-27b-Low

A 27-billion-parameter text-generation model designed for a smaller GPU-memory footprint. The model file is 8.62 GB. The recommended starting point is an NVIDIA RTX 30, 40, or 50 series GPU with at least 12 GB of VRAM, using the L0xRE BeeLLama runtime.

“27b” is the parameter count, not the download size. “Low” describes the memory-focused model build. Context length and available GPU memory depend on the card, driver, runtime profile, and other applications.

Start here: runtime downloads, installation, and GPU test results. There is one Linux/WSL runtime archive and one Windows runtime ZIP for the supported GPU families.

Download these two files

FileWhat it doesDownload size
L0xRE-27b-Low.ggufThe main model. This is the recommended model file.8.62 GB
Qwen3.8-27B-DFlash2-Q4_K_M.ggufOptional speedup helper, recommended with the main model. The runtime uses it to propose tokens that the main model verifies.1.14 GB

For the recommended setup, download both files: 9.76 GB in total, plus the runtime package. The helper is called a drafter; it is an additional file, not another name or edition of L0xRE-27b-Low. The main model can also run without a drafter, with speculative decoding disabled.

The smaller Qwen3.8-27B-DFlash2-Q2_K.gguf file is an optional 705 MB drafter for memory-constrained experiments. Start with the recommended Q4_K_M drafter to match the published default configuration.

Install and run

  1. Download and extract the runtime for your OS from the universal release.
  2. Download the two recommended files above into the extracted runtime's models folder.
  3. Verify the runtime and model checksums using the installation guide.
  4. Start the server from the extracted runtime folder:
# Linux / WSL
./l0xre serve --profile 12gb -m models/L0xRE-27b-Low.gguf \
  -md models/Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080
# Windows
.\l0xre.cmd serve --profile 12gb -m models\L0xRE-27b-Low.gguf -md models\Qwen3.8-27B-DFlash2-Q4_K_M.gguf --host 127.0.0.1 --port 8080

When ready, the server's health endpoint is http://127.0.0.1:8080/health and its API base is http://127.0.0.1:8080/v1. For a compatible chat client, use that API base and select L0xRE-27b-Low from the server's /v1/models endpoint.

This file requires the L0xRE runtime. Direct use with stock llama.cpp, Ollama, LM Studio, Transformers, or other inference engines is outside this release's supported setup. The GGUF extension alone does not establish compatibility.

GPU support and context

The runtime targets NVIDIA compute capabilities 8.6, 8.9, and 12.0. A 12 GB VRAM target does not guarantee that the longest context fits on every card. Close other GPU-heavy applications or reduce context when memory is tight.

GPU / OSDefault 12gb profileEvidence status
RTX 3060 12 GB / Linux98,304-token context with the recommended drafterMeasured model/runtime components; only 215 MiB minimum free VRAM in the long-context test
RTX 4090 / Linux81,920-token context with the recommended drafterPublished configuration measurements; additional 16gb and full32k profiles documented in the runtime guide
RTX 50 series / Linux or WSL8,192-token context with drafter; larger ordinary-decode profile availableGPU runtime payload included; the current non-MTP model path still awaits hardware qualification
RTX 30–50 series / WindowsGPU-specific presets selected automaticallyBuilds, dependencies and launchers verified; GPU inference, throughput and memory headroom still await qualification

The universal runtime is a release candidate. The runtime page lists exactly what was tested. Windows and unmeasured GPU/model combinations should not inherit Linux benchmark claims.

Measured speed

These are workload-specific measurements, not a controlled comparison between GPUs or a guaranteed speed for every prompt.

Hardware / configurationCode generationNarrative generationMeasurement scope
RTX 3060 12 GB, Linux, 96K context, recommended drafter39.046 tokens/s29.422 tokens/sTemperature 0, seed 0, exactly 800 output tokens; two measured runs per prompt
RTX 4090, Linux, 12gb profile, recommended drafter115.4 ± 2.4 tokens/sNot reported for this configurationFive measured code-generation runs; 81,920-token allocated context

The RTX 3060 long-context test processed 92,879 input tokens plus 128 generated tokens and matched the recorded reference output. Prefill was 272.60 tokens/s and decode 21.06 tokens/s, with just 215 MiB minimum free VRAM. This checks a particular runtime workload; it is not a broad model-quality evaluation.

Quality scores from other checkpoints in the model family are not presented as an evaluation of this exact downloadable file. No external leaderboard result is claimed here.

File integrity

Download sizes use decimal GB. Model weights have not changed during the runtime documentation update.

FileBytesSHA-256
L0xRE-27b-Low.gguf8,619,127,680b0849250c633aa93853bf119a877dbdafdd7b1a4ebb672bd6b6a1439906f3543
Qwen3.8-27B-DFlash2-Q4_K_M.gguf1,143,006,8161a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd
Qwen3.8-27B-DFlash2-Q2_K.gguf705,430,880e3eb7705404817cdbcdabe56049a1952b3b37bcc8df6e4d4efaec5d41563fb7e

L0xRE-27b-Low-MTP is a separate, larger checkpoint with an additional prediction block. Its file hash, runtime setup and qualification are distinct. Use the main model file on this page for the recommended drafter-based setup.

Model origin and license

L0xRE-27b-Low is a quantized derivative of Qwen3.8-27B, with low-bit projections and higher-precision embeddings/output tensors. The model's configured native context capacity is 262,144 tokens; supported and tested serving contexts are determined by the runtime profiles above.

Model weights remain under the base model's research/community terms: license. The L0xRE runtime tooling is Apache-2.0; applicable runtime and third-party notices ship with its packages. Model licensing and runtime licensing are separate.

12gb-vram
conversational
endpoints_compatible
gguf
l0xre
speculative-decoding
text-generation