Lucebox/Laguna-XS-2.1-DFlash-GGUF

Model

4

stars

4

commits

2

linked in READMEs

Jul 9, 2026

updated

dflash
endpoints_compatible
gguf
laguna
lucebox
speculative-decoding

README

Laguna-XS-2.1-DFlash-GGUF (quantized drafter)

Quantized GGUF builds of the official poolside/Laguna-XS-2.1-DFlash speculator, packaged for the Lucebox inference engine's speculative decoding of poolside/Laguna-XS-2.1.

fileschemesizenotes
laguna-xs21-dflash-q4.ggufQ4_0 projections, Q8_0 feature projection (fc), F32 norms271 MBrecommended
laguna-xs21-dflash-q8.ggufQ8_0 projections, F32 norms492 MBconservative

Because speculative decoding verifies every draft against the target model, drafter quantization cannot change the target's greedy output quality; the only quantity at stake is acceptance length. Measured on an RTX 3090 with the Laguna XS 2.1 Q4_K_M target (HumanEval/GSM8K/Math screens): acceptance unchanged vs the BF16 drafter, end-to-end decode ~+3% (q4), gold-scored task accuracy identical.

Produced with server/scripts/requant_dflash_draft.py from the lucebox engine tree (GGUF-to-GGUF requant; all metadata, gates and aux norms preserved).

License: OpenMDW-1.1, inherited from the base speculator. All credit for the DFlash speculator weights to poolside.

Contributors

davide221

4 commits

Lucebox/Laguna-XS-2.1-DFlash-GGUF

Model

4

stars

4

commits

2

linked in READMEs

Jul 9, 2026

updated

dflash
endpoints_compatible
gguf
laguna
lucebox
speculative-decoding

README

Laguna-XS-2.1-DFlash-GGUF (quantized drafter)

Quantized GGUF builds of the official poolside/Laguna-XS-2.1-DFlash speculator, packaged for the Lucebox inference engine's speculative decoding of poolside/Laguna-XS-2.1.

fileschemesizenotes
laguna-xs21-dflash-q4.ggufQ4_0 projections, Q8_0 feature projection (fc), F32 norms271 MBrecommended
laguna-xs21-dflash-q8.ggufQ8_0 projections, F32 norms492 MBconservative

Because speculative decoding verifies every draft against the target model, drafter quantization cannot change the target's greedy output quality; the only quantity at stake is acceptance length. Measured on an RTX 3090 with the Laguna XS 2.1 Q4_K_M target (HumanEval/GSM8K/Math screens): acceptance unchanged vs the BF16 drafter, end-to-end decode ~+3% (q4), gold-scored task accuracy identical.

Produced with server/scripts/requant_dflash_draft.py from the lucebox engine tree (GGUF-to-GGUF requant; all metadata, gates and aux norms preserved).

License: OpenMDW-1.1, inherited from the base speculator. All credit for the DFlash speculator weights to poolside.

Contributors

davide221

4 commits