shisa-ai/FastDMS

Production-speed compact Dynamic Memory Sparsification (DMS) for KV cache compression

Python

27

23 commits

updated May 6, 2026

See the code

README

FastDMS

A production-speed implementation of compact Dynamic Memory Sparsification (DMS) for KV cache compression (arXiv:2506.05345, NeurIPS 2025 poster).

FastDMS runs faster than vLLM (v0.19.2rc1, nightly) with BF16, FP8, and TurboQuant KV cache settings while using significantly less KV memory. On Llama-3.2-1B at c=8, the zero-BF16 default decodes 1.53x faster than vLLM BF16 with 4.8x less KV memory. On Qwen3-8B at c=8, it decodes 2.06x faster with 6.35x less KV memory. All vLLM baselines use exact-sized token pools (KV cache sized to the workload, not over-provisioned). All tests were run on NVIDIA RTX PRO 6000 (Blackwell, sm120).

Note, while the speed is faster than vLLM, this should be considered a fast reference implementation as only two DMS-trained checkpoints have been tested/validated:

To train your own DMS checkpoints with the eviction-head retrofit recipe, see the training/ folder. The in-repo trainer is a fast for single-GPU training. For larger models or multi-GPU training, start from NVIDIA's Model-Optimizer DMS trainer.

About DMS

DMS trains learned per-head token eviction via logit distillation. This takes only a very short time to train, with our Llama 3.2 1B DMS model taking only about 20 minutes on a single RTX PRO 6000.

FastDMS implements a compact KV cache layout that reclaims evicted slots. The result is FP8 compact KV with allocator-visible compression ratios of 5-8x smaller than vLLM BF16 KV at 8K context, while also decoding 1.5-2x faster.

Initial Work

This project started as a research pack that reviewed a broad range of academic and "folk" KV-cache compression techniques. Among all the combinations tested, the strongest near-lossless research-stack result was DMS + AQUA-KV + HIGGS 4-bit: 25.6x theoretical KV compression at +0.09% PPL on Llama-3.2-1B, trained in ~20 minutes on a single GPU:

ConfigurationPPLDeltaKLD (nats/tok)Compression
Vanilla Llama-3.2-1B9.226--1x
DMS (trained, eviction active)9.200-0.28%0.0266.4x
DMS + AQUA9.205-0.23%0.0266.4x
DMS + HIGGS 4-bit9.621+4.28%0.05825.6x
DMS + AQUA + HIGGS 4-bit9.234+0.09%0.03225.6x

A fair amount of effort was spent optimizing performance for this stack, but ultimately HIGGS was tabled due to our best efforts only hitting about 50% of BF16/FP8 prefill/decode speed. AQUA-KV was not required for best FP8+DMS quality. HIGGS+AQUA composes naturally with DMS, and is an obvious target for future kernel work.

Initial bring-up of DMS was done on a HuggingFace-based correctness/PPL harness (training/dms_eval.py) that ran full-forward evaluation at ~18 tok/s. That harness validated quality but didn't reclaim any memory (evicted tokens were masked in attention but their KV slots stayed allocated).

Optimized Performance

FastDMS is able to bring DMS up to production speeds, beating vLLM's BF16 and FP8 KV cache speeds at both prefill and decode, and being nearly 40x faster than our initial HF implementation.

FastDMS is benchmarked on WikiText-2 with ctx_len=8192, gen_len=128, and post-warmup timing. vLLM baselines use exact-sized token pools — the KV cache is sized to fit exactly the sequences being tested (e.g. 10240 tokens for c=1, 67584 tokens for c=8), so reported KV memory reflects what the workload actually needs, not pre-allocated headroom. Ratio columns are against vLLM BF16 KV because many external KV-cache papers report against FP16/BF16 cache baselines; FP8 rows, however, are probably best for direct serving-engine comparison.

shisa-ai/Llama-3.2-1B-DMS-8x

The zero-BF16 FastDMS default is 1.52x / 1.53x faster than vLLM BF16 decode at c=1 / c=8, while using 5.6x / 4.8x less KV memory. Against vLLM FP8, it is 1.43x / 1.25x faster decode and 2.8x / 2.4x smaller in KV. vLLM's TurboQuant 4-bit is actually slower than BF16 (0.73x / 0.72x decode) for only 2.2x KV savings — and with worse output quality. The default-off B46 c=1 speed profile reaches 2.30x BF16 decode at the same compact-KV footprint, with 0.719 GiB of int4 shadow storage.

PathcPrefill tok/sPrefill vs BF16Decode tok/sDecode vs BF16KV / stage memoryStatus
vLLM BF161123098.01.00x459.41.00x0.312 GiB BF16 KVdense BF16-KV baseline
vLLM FP81119991.30.97x489.41.07x0.156 GiB FP8 KVdense FP8-KV baseline
vLLM TurboQuant 4bit_nc1126429.01.03x333.40.73x0.142 GiB TQ4 KV4-bit KV baseline
FastDMS FP8 compact-DMS default1123194.61.00x698.91.52x0.056 GiBpromoted zero-BF16 row
FastDMS B46 int4 speed profile1121489.90.99x1060.02.31x0.056 GiB + 0.719 GiB int4 shadowdefault-off storage-for-speed
vLLM BF168103668.51.00x2357.51.00x2.062 GiB BF16 KVdense BF16-KV baseline
vLLM FP88102959.50.99x2888.71.23x1.031 GiB FP8 KVdense FP8-KV baseline
vLLM TurboQuant 4bit_nc8104409.91.01x1696.00.72x0.939 GiB TQ4 KV4-bit KV baseline
FastDMS FP8 compact-DMS default8105531.71.02x3606.91.53x0.431 GiBpromoted zero-BF16 row
FastDMS B25 narrow int4 speed profile8104753.71.01x3640.71.54x0.431 GiB + 0.078 GiB int4 shadowdefault-off storage-for-speed
FastDMS BF16-attention speed control8108070.51.04x3745.31.59x0.429 GiB + 0.312 GiB BF16 backingexplicit speed control

nvidia/Qwen3-8B-DMS-8x

FastDMS performs similarly well with Nvidia's Qwen3-8B example model. It is 1.54x / 2.06x faster than vLLM BF16 decode and 7.64x / 6.35x smaller than vLLM BF16 KV at c=1 / c=8. Against vLLM FP8, it is 1.48x / 1.57x faster decode and 3.82x / 3.17x smaller in KV. Versus same-engine FastDMS dense FP8, it is 3.26x / 3.46x faster decode and 11.47x / 9.52x smaller in allocator-visible KV/stage memory.

PathcPrefill tok/sPrefill vs BF16Decode tok/sDecode vs BF16KV / stage memoryStatus
vLLM BF16117143.61.00x89.441.00x1.406 GiB BF16 KVdense BF16-KV baseline
vLLM FP8116700.10.97x93.121.04x0.703 GiB FP8 KVdense FP8-KV baseline
FastDMS dense FP8122125.71.29x42.240.47x0.703 GiB FP8 KV + 1.406 GiB BF16 stagingsame-engine dense baseline
FastDMS compact DMS121610.11.26x137.761.54x0.184 GiB compact+metadatazero retained LM-head BF16 backing
vLLM BF16815800.21.00x444.591.00x9.281 GiB BF16 KVdense BF16-KV baseline
vLLM FP8815659.10.99x583.391.31x4.641 GiB FP8 KVdense FP8-KV baseline
FastDMS dense FP8819502.41.23x265.000.60x4.641 GiB FP8 KV + 9.281 GiB BF16 stagingsame-engine dense baseline
FastDMS compact DMS819366.51.23x917.902.06x1.462 GiB compact+metadatazero retained LM-head BF16 backing

Memory Savings vs vLLM

Compact DMS saves real allocator/device memory, not just theoretical KV bytes. The table below uses the same WikiText-2 ctx_len=8192, gen_len=128 rows as the speed tables above. All vLLM baselines use exact-sized token pools matching the workload. KV/stage memory is the cache or cache-plus-staging footprint. vLLM BF16 means dtype=bfloat16 with kv_cache_dtype=auto; vLLM FP8 means kv_cache_dtype=fp8.

Model / compact-DMS rowcvLLM BF16 KV → FastDMS KVBF16 KV savedvLLM FP8 KV → FastDMS KVFP8 KV savedvLLM TQ4 KV → FastDMS KVTQ4 KV saved
Llama-3.2-1B FastDMS default10.312 → 0.056 GiB5.6x0.156 → 0.056 GiB2.8x0.142 → 0.056 GiB2.5x
Llama-3.2-1B FastDMS default82.062 → 0.431 GiB4.8x1.031 → 0.431 GiB2.4x0.939 → 0.431 GiB2.2x
Qwen3-8B FastDMS compact DMS11.406 → 0.184 GiB7.6x0.703 → 0.184 GiB3.8x——
Qwen3-8B FastDMS compact DMS89.281 → 1.462 GiB6.3x4.641 → 1.462 GiB3.2x——

Per token, Qwen3 KV is larger: about 72 KiB/token for FP8 KV (36 layers × 8 KV heads × 128 head dim × K/V) versus Llama-3.2-1B’s 16 KiB/token.

Max Context

Max-context rows use each model's config-supported final window while leaving 128 generated tokens: Llama-3.2-1B at 130944+128 of 131072, and Qwen3-8B at 40832+128 of 40960. vLLM rows use vllm-nightly with post-warmup timing, wrapped WikiText-2 prompts, and a small KV-pool margin above the exact context budget to avoid exact-full-pool scheduler waits. vLLM FP8 max-context CUDA peaks are upper bounds from paired BF16/FP8 runs; KV memory is dtype-specific.

shisa-ai/Llama-3.2-1B-DMS-8x

Pathcctx+genPrefill tok/sDecode tok/sKV / compact+meta GiBCUDA peak GiBDecode vs BF16KV vs BF16
vLLM BF16 KV1130944+12822906.4189.14.12510.921.00x1.00x
vLLM FP8 KV1130944+12821639.8265.92.063<=10.921.41x2.00x smaller
FastDMS compact DMS1130944+12829059.6286.00.7666915.201.51x5.38x smaller
vLLM BF16 KV8130944+12820384.0178.833.00038.901.00x1.00x
vLLM FP8 KV8130944+12820448.6492.216.500<=38.902.75x2.00x smaller
FastDMS compact DMS8130944+12827655.0273.61.3099512.171.53x25.2x smaller

FastDMS is faster than vLLM BF16 at both Llama max-context concurrencies and is dramatically smaller in KV. Against vLLM FP8, c=1 still wins decode (1.08x) while c=8 trades lower decode (0.56x) for 12.6x lower KV memory and 3.2x lower paired-run CUDA peak. The c=1 FastDMS CUDA peak is higher than vLLM BF16 because the max-context FastDMS row carries graph/runtime overhead even though compact KV is much smaller.

nvidia/Qwen3-8B-DMS-8x

Pathcctx+genPrefill tok/sDecode tok/sKV / compact+meta GiBCUDA peak GiBDecode vs BF16KV vs BF16
vLLM BF16 KV140832+1289614.869.246.18826.041.00x1.00x
vLLM FP8 KV140832+1289321.579.463.094<=26.041.15x2.00x smaller
FastDMS compact DMS140832+12812544.1111.90.5831314.641.62x10.6x smaller
vLLM BF16 KV840832+1289116.6161.649.50069.221.00x1.00x
vLLM FP8 KV840832+1288996.9286.324.750<=69.221.77x2.00x smaller
FastDMS compact DMS840832+12810517.679.31.0069820.690.49x49.2x smaller

FastDMS transfers cleanly at Qwen c=1 max context: 1.62x BF16 decode, 10.6x lower KV, and 1.78x lower CUDA peak. Qwen c=8 max context is the one long-context throughput caveat: it is a large memory/capacity win (49.2x lower BF16 KV, 3.35x lower CUDA peak), but decode is 0.49x of vLLM BF16 and 0.28x of vLLM FP8 at this context.

Compression Quality

Of course, none of this matters if the compression tanks output quality. In theory, DMS eviction is applied before FP8 quantization, deciding which tokens to keep or evict, so the quality comparison for FastDMS compact-DMS should be the same versus FP8 quantization alone, but it's still worth double-checking quality.

We measure this by generating tokens with a compressed KV cache and comparing against an uncompressed reference, token by token. Lower KLD (KL divergence) is better - it means the compressed model's next-token probabilities are closer to the reference. Higher token match is better - it means greedy decoding produces the same output.

How to read the columns:

  • KLD vs ref - KL divergence in nats/token between the compressed and reference logits. Measures how much the probability distribution over next tokens shifts due to compression. Lower is better; 0.000 means identical.
  • Token match - percentage of greedy-decoded tokens that are identical to the reference. 96.9% means ~2 out of 64 tokens differed.
  • Tokens scored - how many decode steps could be compared. Once the candidate produces a different token than the reference, the sequences diverge and later steps aren't comparable. 33/60 means quality metrics only cover the first 33 tokens before divergence - the reported KLD and PPL are over that prefix, not the full generation. A higher ratio means the comparison is more complete.

Test setup: ctx_len=1024, decode_len=16, four prompts (60-64 total decode steps). vLLM rows compare against vLLM BF16 full-KV logits. FastDMS rows compare against FastDMS with eviction disabled (reference window of 1M tokens, effectively keeping the full KV cache).

shisa-ai/Llama-3.2-1B-DMS-8x

PathReferenceKLD vs refToken matchPPLTokens scored
vLLM BF16 full KVself0.000000100.0%2.374860/60
vLLM FP8 KVvLLM BF160.00511092.2%2.089333/60
vLLM TurboQuant 4bit_ncvLLM BF160.01273076.6%1.960622/60
FastDMS FP8 compact-DMSFastDMS no-evict0.00300996.9%2.281064/64

nvidia/Qwen3-8B-DMS-8x

PathReferenceKLD vs refToken matchPPLTokens scored
vLLM BF16 full KVself0.000000100.0%1.673860/60
vLLM FP8 KVvLLM BF160.00104270.3%1.197132/60
vLLM TurboQuant 4bit_ncvLLM BF160.00603984.4%1.491045/60
FastDMS FP8 compact-DMSFastDMS no-evict0.00528495.3%1.830164/64

FastDMS compact-DMS scores 64/64 tokens on both models - every decode step was comparable to the reference, and the KLD is lower than or comparable to vLLM's own FP8 and TurboQuant compression. Note that PPL values across rows are not directly comparable when Tokens scored differs, because each row's PPL is computed over a different-length prefix.

Why a Standalone Engine Instead of a vLLM Plugin?

We investigated porting compact DMS directly into vLLM and concluded it is major surgery, not a plugin. DMS compact KV touches nearly every serving-engine subsystem:

SubsystemWhat changes for DMS
PagedAttention / KV memory poolDMS needs per-layer, per-head variable token counts with partial block deallocation - not standard fixed-page blocks
Prefill kernelMust stream surviving K/V into compact per-layer storage after DMS extraction, rather than writing dense KV pages
Decode kernelEach decode step evaluates per-head keep/evict, manages a sliding retention window, and appends to compact storage
Attention scoringReplaced entirely: split-K grouped compact decode attention over variable-length per-head live spans
Scheduler / admissionMust admit requests based on compact KV capacity, not dense full-sequence page count - this is the hardest boundary
Prefix cachingDMS eviction is per-sequence and per-head; shared prefix blocks need per-sequence eviction overlays or must be disabled
Continuous batchingMemory accounting must reflect actual surviving token count, not logical sequence length

vLLM's TurboQuant backend provides a useful template (custom cache dtype, custom KVCacheSpec, backend-owned metadata builder, custom store/decode ops). But TurboQuant still uses vLLM's paged block table with one compressed slot per logical token. DMS needs per-layer, per-sequence, per-KV-head live spans and eviction state - the decode metadata is fundamentally different.

The critical gap is scheduler/cache accounting: until vLLM stops reserving dense full-sequence pages, a DMS backend is only a functional/speed experiment and not the compact-memory serving path that delivers the real value. That scheduler change touches scheduler.py, kv_cache_manager.py, kv_cache_coordinator.py, single_type_kv_cache_manager.py, and block_pool.py - the core of vLLM's memory management.

A proper vLLM port would need to:

  1. Add DMS as a first-class cache dtype with config gates and hard-disable unsupported features (prefix cache, chunked prefill, spec decode, distributed)
  2. Port DMS metadata loading and per-model borrowed-channel extraction (Llama-only first)
  3. Build a custom DMSCompactAttentionBackend with sidecar compact arena
  4. Port compact append-store, DMS expiry, fused DMS/RoPE/store, and split-K compact decode kernels
  5. Replace sidecar-only accounting with native DMS reservation-cap admission in vLLM's cache manager
  6. Only then re-enable chunked prefill, prefix cache overlays, and distributed execution

This is a viable path but a large, bounded engineering project. FastDMS exists mainly to show why this effort might be worthwhile and it proves that DMS can be served efficiently.

Install

pip install fastdms

Or from source:

git clone https://github.com/shisa-ai/FastDMS
cd FastDMS
pip install -e .

Requires recent CUDA, torch, triton, and flash-attn.

Quick Start

from fastdms import LLM, SamplingParams

llm = LLM("shisa-ai/Llama-3.2-1B-DMS-8x", enforce_eager=True)
sampling_params = SamplingParams(temperature=0.6, max_tokens=256)
outputs = llm.generate(["Hello, FastDMS."], sampling_params)
print(outputs[0]["text"])

Canonical Papers

Acknowledgements

FastDMS is built on top of nano-vLLM by Xingkai Yu. The clean, readable nano-vLLM codebase was used as the starting harness for the compact-DMS implementation. We're grateful for the upstream work that made rapid iteration on the KV-cache layout possible.

Citation

@misc{fastdms2026,
  title        = {FastDMS: Production-Speed Compact DMS for KV Cache Compression},
  author       = {{Leonard Lin}},
  year         = {2026},
  url          = {https://github.com/shisa-ai/FastDMS},
  note         = {Fast reference implementation of compact Dynamic Memory Sparsification with FP8 KV cache}
}

License

MIT

shisa-ai/FastDMS

Production-speed compact Dynamic Memory Sparsification (DMS) for KV cache compression

Python

27

23 commits

updated May 6, 2026

See the code

README

FastDMS

A production-speed implementation of compact Dynamic Memory Sparsification (DMS) for KV cache compression (arXiv:2506.05345, NeurIPS 2025 poster).

FastDMS runs faster than vLLM (v0.19.2rc1, nightly) with BF16, FP8, and TurboQuant KV cache settings while using significantly less KV memory. On Llama-3.2-1B at c=8, the zero-BF16 default decodes 1.53x faster than vLLM BF16 with 4.8x less KV memory. On Qwen3-8B at c=8, it decodes 2.06x faster with 6.35x less KV memory. All vLLM baselines use exact-sized token pools (KV cache sized to the workload, not over-provisioned). All tests were run on NVIDIA RTX PRO 6000 (Blackwell, sm120).

Note, while the speed is faster than vLLM, this should be considered a fast reference implementation as only two DMS-trained checkpoints have been tested/validated:

To train your own DMS checkpoints with the eviction-head retrofit recipe, see the training/ folder. The in-repo trainer is a fast for single-GPU training. For larger models or multi-GPU training, start from NVIDIA's Model-Optimizer DMS trainer.

About DMS

DMS trains learned per-head token eviction via logit distillation. This takes only a very short time to train, with our Llama 3.2 1B DMS model taking only about 20 minutes on a single RTX PRO 6000.

FastDMS implements a compact KV cache layout that reclaims evicted slots. The result is FP8 compact KV with allocator-visible compression ratios of 5-8x smaller than vLLM BF16 KV at 8K context, while also decoding 1.5-2x faster.

Initial Work

This project started as a research pack that reviewed a broad range of academic and "folk" KV-cache compression techniques. Among all the combinations tested, the strongest near-lossless research-stack result was DMS + AQUA-KV + HIGGS 4-bit: 25.6x theoretical KV compression at +0.09% PPL on Llama-3.2-1B, trained in ~20 minutes on a single GPU:

ConfigurationPPLDeltaKLD (nats/tok)Compression
Vanilla Llama-3.2-1B9.226--1x
DMS (trained, eviction active)9.200-0.28%0.0266.4x
DMS + AQUA9.205-0.23%0.0266.4x
DMS + HIGGS 4-bit9.621+4.28%0.05825.6x
DMS + AQUA + HIGGS 4-bit9.234+0.09%0.03225.6x

A fair amount of effort was spent optimizing performance for this stack, but ultimately HIGGS was tabled due to our best efforts only hitting about 50% of BF16/FP8 prefill/decode speed. AQUA-KV was not required for best FP8+DMS quality. HIGGS+AQUA composes naturally with DMS, and is an obvious target for future kernel work.

Initial bring-up of DMS was done on a HuggingFace-based correctness/PPL harness (training/dms_eval.py) that ran full-forward evaluation at ~18 tok/s. That harness validated quality but didn't reclaim any memory (evicted tokens were masked in attention but their KV slots stayed allocated).

Optimized Performance

FastDMS is able to bring DMS up to production speeds, beating vLLM's BF16 and FP8 KV cache speeds at both prefill and decode, and being nearly 40x faster than our initial HF implementation.

FastDMS is benchmarked on WikiText-2 with ctx_len=8192, gen_len=128, and post-warmup timing. vLLM baselines use exact-sized token pools — the KV cache is sized to fit exactly the sequences being tested (e.g. 10240 tokens for c=1, 67584 tokens for c=8), so reported KV memory reflects what the workload actually needs, not pre-allocated headroom. Ratio columns are against vLLM BF16 KV because many external KV-cache papers report against FP16/BF16 cache baselines; FP8 rows, however, are probably best for direct serving-engine comparison.

shisa-ai/Llama-3.2-1B-DMS-8x

The zero-BF16 FastDMS default is 1.52x / 1.53x faster than vLLM BF16 decode at c=1 / c=8, while using 5.6x / 4.8x less KV memory. Against vLLM FP8, it is 1.43x / 1.25x faster decode and 2.8x / 2.4x smaller in KV. vLLM's TurboQuant 4-bit is actually slower than BF16 (0.73x / 0.72x decode) for only 2.2x KV savings — and with worse output quality. The default-off B46 c=1 speed profile reaches 2.30x BF16 decode at the same compact-KV footprint, with 0.719 GiB of int4 shadow storage.

PathcPrefill tok/sPrefill vs BF16Decode tok/sDecode vs BF16KV / stage memoryStatus
vLLM BF161123098.01.00x459.41.00x0.312 GiB BF16 KVdense BF16-KV baseline
vLLM FP81119991.30.97x489.41.07x0.156 GiB FP8 KVdense FP8-KV baseline
vLLM TurboQuant 4bit_nc1126429.01.03x333.40.73x0.142 GiB TQ4 KV4-bit KV baseline
FastDMS FP8 compact-DMS default1123194.61.00x698.91.52x0.056 GiBpromoted zero-BF16 row
FastDMS B46 int4 speed profile1121489.90.99x1060.02.31x0.056 GiB + 0.719 GiB int4 shadowdefault-off storage-for-speed
vLLM BF168103668.51.00x2357.51.00x2.062 GiB BF16 KVdense BF16-KV baseline
vLLM FP88102959.50.99x2888.71.23x1.031 GiB FP8 KVdense FP8-KV baseline
vLLM TurboQuant 4bit_nc8104409.91.01x1696.00.72x0.939 GiB TQ4 KV4-bit KV baseline
FastDMS FP8 compact-DMS default8105531.71.02x3606.91.53x0.431 GiBpromoted zero-BF16 row
FastDMS B25 narrow int4 speed profile8104753.71.01x3640.71.54x0.431 GiB + 0.078 GiB int4 shadowdefault-off storage-for-speed
FastDMS BF16-attention speed control8108070.51.04x3745.31.59x0.429 GiB + 0.312 GiB BF16 backingexplicit speed control

nvidia/Qwen3-8B-DMS-8x

FastDMS performs similarly well with Nvidia's Qwen3-8B example model. It is 1.54x / 2.06x faster than vLLM BF16 decode and 7.64x / 6.35x smaller than vLLM BF16 KV at c=1 / c=8. Against vLLM FP8, it is 1.48x / 1.57x faster decode and 3.82x / 3.17x smaller in KV. Versus same-engine FastDMS dense FP8, it is 3.26x / 3.46x faster decode and 11.47x / 9.52x smaller in allocator-visible KV/stage memory.

PathcPrefill tok/sPrefill vs BF16Decode tok/sDecode vs BF16KV / stage memoryStatus
vLLM BF16117143.61.00x89.441.00x1.406 GiB BF16 KVdense BF16-KV baseline
vLLM FP8116700.10.97x93.121.04x0.703 GiB FP8 KVdense FP8-KV baseline
FastDMS dense FP8122125.71.29x42.240.47x0.703 GiB FP8 KV + 1.406 GiB BF16 stagingsame-engine dense baseline
FastDMS compact DMS121610.11.26x137.761.54x0.184 GiB compact+metadatazero retained LM-head BF16 backing
vLLM BF16815800.21.00x444.591.00x9.281 GiB BF16 KVdense BF16-KV baseline
vLLM FP8815659.10.99x583.391.31x4.641 GiB FP8 KVdense FP8-KV baseline
FastDMS dense FP8819502.41.23x265.000.60x4.641 GiB FP8 KV + 9.281 GiB BF16 stagingsame-engine dense baseline
FastDMS compact DMS819366.51.23x917.902.06x1.462 GiB compact+metadatazero retained LM-head BF16 backing

Memory Savings vs vLLM

Compact DMS saves real allocator/device memory, not just theoretical KV bytes. The table below uses the same WikiText-2 ctx_len=8192, gen_len=128 rows as the speed tables above. All vLLM baselines use exact-sized token pools matching the workload. KV/stage memory is the cache or cache-plus-staging footprint. vLLM BF16 means dtype=bfloat16 with kv_cache_dtype=auto; vLLM FP8 means kv_cache_dtype=fp8.

Model / compact-DMS rowcvLLM BF16 KV → FastDMS KVBF16 KV savedvLLM FP8 KV → FastDMS KVFP8 KV savedvLLM TQ4 KV → FastDMS KVTQ4 KV saved
Llama-3.2-1B FastDMS default10.312 → 0.056 GiB5.6x0.156 → 0.056 GiB2.8x0.142 → 0.056 GiB2.5x
Llama-3.2-1B FastDMS default82.062 → 0.431 GiB4.8x1.031 → 0.431 GiB2.4x0.939 → 0.431 GiB2.2x
Qwen3-8B FastDMS compact DMS11.406 → 0.184 GiB7.6x0.703 → 0.184 GiB3.8x——
Qwen3-8B FastDMS compact DMS89.281 → 1.462 GiB6.3x4.641 → 1.462 GiB3.2x——

Per token, Qwen3 KV is larger: about 72 KiB/token for FP8 KV (36 layers × 8 KV heads × 128 head dim × K/V) versus Llama-3.2-1B’s 16 KiB/token.

Max Context

Max-context rows use each model's config-supported final window while leaving 128 generated tokens: Llama-3.2-1B at 130944+128 of 131072, and Qwen3-8B at 40832+128 of 40960. vLLM rows use vllm-nightly with post-warmup timing, wrapped WikiText-2 prompts, and a small KV-pool margin above the exact context budget to avoid exact-full-pool scheduler waits. vLLM FP8 max-context CUDA peaks are upper bounds from paired BF16/FP8 runs; KV memory is dtype-specific.

shisa-ai/Llama-3.2-1B-DMS-8x

Pathcctx+genPrefill tok/sDecode tok/sKV / compact+meta GiBCUDA peak GiBDecode vs BF16KV vs BF16
vLLM BF16 KV1130944+12822906.4189.14.12510.921.00x1.00x
vLLM FP8 KV1130944+12821639.8265.92.063<=10.921.41x2.00x smaller
FastDMS compact DMS1130944+12829059.6286.00.7666915.201.51x5.38x smaller
vLLM BF16 KV8130944+12820384.0178.833.00038.901.00x1.00x
vLLM FP8 KV8130944+12820448.6492.216.500<=38.902.75x2.00x smaller
FastDMS compact DMS8130944+12827655.0273.61.3099512.171.53x25.2x smaller

FastDMS is faster than vLLM BF16 at both Llama max-context concurrencies and is dramatically smaller in KV. Against vLLM FP8, c=1 still wins decode (1.08x) while c=8 trades lower decode (0.56x) for 12.6x lower KV memory and 3.2x lower paired-run CUDA peak. The c=1 FastDMS CUDA peak is higher than vLLM BF16 because the max-context FastDMS row carries graph/runtime overhead even though compact KV is much smaller.

nvidia/Qwen3-8B-DMS-8x

Pathcctx+genPrefill tok/sDecode tok/sKV / compact+meta GiBCUDA peak GiBDecode vs BF16KV vs BF16
vLLM BF16 KV140832+1289614.869.246.18826.041.00x1.00x
vLLM FP8 KV140832+1289321.579.463.094<=26.041.15x2.00x smaller
FastDMS compact DMS140832+12812544.1111.90.5831314.641.62x10.6x smaller
vLLM BF16 KV840832+1289116.6161.649.50069.221.00x1.00x
vLLM FP8 KV840832+1288996.9286.324.750<=69.221.77x2.00x smaller
FastDMS compact DMS840832+12810517.679.31.0069820.690.49x49.2x smaller

FastDMS transfers cleanly at Qwen c=1 max context: 1.62x BF16 decode, 10.6x lower KV, and 1.78x lower CUDA peak. Qwen c=8 max context is the one long-context throughput caveat: it is a large memory/capacity win (49.2x lower BF16 KV, 3.35x lower CUDA peak), but decode is 0.49x of vLLM BF16 and 0.28x of vLLM FP8 at this context.

Compression Quality

Of course, none of this matters if the compression tanks output quality. In theory, DMS eviction is applied before FP8 quantization, deciding which tokens to keep or evict, so the quality comparison for FastDMS compact-DMS should be the same versus FP8 quantization alone, but it's still worth double-checking quality.

We measure this by generating tokens with a compressed KV cache and comparing against an uncompressed reference, token by token. Lower KLD (KL divergence) is better - it means the compressed model's next-token probabilities are closer to the reference. Higher token match is better - it means greedy decoding produces the same output.

How to read the columns:

  • KLD vs ref - KL divergence in nats/token between the compressed and reference logits. Measures how much the probability distribution over next tokens shifts due to compression. Lower is better; 0.000 means identical.
  • Token match - percentage of greedy-decoded tokens that are identical to the reference. 96.9% means ~2 out of 64 tokens differed.
  • Tokens scored - how many decode steps could be compared. Once the candidate produces a different token than the reference, the sequences diverge and later steps aren't comparable. 33/60 means quality metrics only cover the first 33 tokens before divergence - the reported KLD and PPL are over that prefix, not the full generation. A higher ratio means the comparison is more complete.

Test setup: ctx_len=1024, decode_len=16, four prompts (60-64 total decode steps). vLLM rows compare against vLLM BF16 full-KV logits. FastDMS rows compare against FastDMS with eviction disabled (reference window of 1M tokens, effectively keeping the full KV cache).

shisa-ai/Llama-3.2-1B-DMS-8x

PathReferenceKLD vs refToken matchPPLTokens scored
vLLM BF16 full KVself0.000000100.0%2.374860/60
vLLM FP8 KVvLLM BF160.00511092.2%2.089333/60
vLLM TurboQuant 4bit_ncvLLM BF160.01273076.6%1.960622/60
FastDMS FP8 compact-DMSFastDMS no-evict0.00300996.9%2.281064/64

nvidia/Qwen3-8B-DMS-8x

PathReferenceKLD vs refToken matchPPLTokens scored
vLLM BF16 full KVself0.000000100.0%1.673860/60
vLLM FP8 KVvLLM BF160.00104270.3%1.197132/60
vLLM TurboQuant 4bit_ncvLLM BF160.00603984.4%1.491045/60
FastDMS FP8 compact-DMSFastDMS no-evict0.00528495.3%1.830164/64

FastDMS compact-DMS scores 64/64 tokens on both models - every decode step was comparable to the reference, and the KLD is lower than or comparable to vLLM's own FP8 and TurboQuant compression. Note that PPL values across rows are not directly comparable when Tokens scored differs, because each row's PPL is computed over a different-length prefix.

Why a Standalone Engine Instead of a vLLM Plugin?

We investigated porting compact DMS directly into vLLM and concluded it is major surgery, not a plugin. DMS compact KV touches nearly every serving-engine subsystem:

SubsystemWhat changes for DMS
PagedAttention / KV memory poolDMS needs per-layer, per-head variable token counts with partial block deallocation - not standard fixed-page blocks
Prefill kernelMust stream surviving K/V into compact per-layer storage after DMS extraction, rather than writing dense KV pages
Decode kernelEach decode step evaluates per-head keep/evict, manages a sliding retention window, and appends to compact storage
Attention scoringReplaced entirely: split-K grouped compact decode attention over variable-length per-head live spans
Scheduler / admissionMust admit requests based on compact KV capacity, not dense full-sequence page count - this is the hardest boundary
Prefix cachingDMS eviction is per-sequence and per-head; shared prefix blocks need per-sequence eviction overlays or must be disabled
Continuous batchingMemory accounting must reflect actual surviving token count, not logical sequence length

vLLM's TurboQuant backend provides a useful template (custom cache dtype, custom KVCacheSpec, backend-owned metadata builder, custom store/decode ops). But TurboQuant still uses vLLM's paged block table with one compressed slot per logical token. DMS needs per-layer, per-sequence, per-KV-head live spans and eviction state - the decode metadata is fundamentally different.

The critical gap is scheduler/cache accounting: until vLLM stops reserving dense full-sequence pages, a DMS backend is only a functional/speed experiment and not the compact-memory serving path that delivers the real value. That scheduler change touches scheduler.py, kv_cache_manager.py, kv_cache_coordinator.py, single_type_kv_cache_manager.py, and block_pool.py - the core of vLLM's memory management.

A proper vLLM port would need to:

  1. Add DMS as a first-class cache dtype with config gates and hard-disable unsupported features (prefix cache, chunked prefill, spec decode, distributed)
  2. Port DMS metadata loading and per-model borrowed-channel extraction (Llama-only first)
  3. Build a custom DMSCompactAttentionBackend with sidecar compact arena
  4. Port compact append-store, DMS expiry, fused DMS/RoPE/store, and split-K compact decode kernels
  5. Replace sidecar-only accounting with native DMS reservation-cap admission in vLLM's cache manager
  6. Only then re-enable chunked prefill, prefix cache overlays, and distributed execution

This is a viable path but a large, bounded engineering project. FastDMS exists mainly to show why this effort might be worthwhile and it proves that DMS can be served efficiently.

Install

pip install fastdms

Or from source:

git clone https://github.com/shisa-ai/FastDMS
cd FastDMS
pip install -e .

Requires recent CUDA, torch, triton, and flash-attn.

Quick Start

from fastdms import LLM, SamplingParams

llm = LLM("shisa-ai/Llama-3.2-1B-DMS-8x", enforce_eager=True)
sampling_params = SamplingParams(temperature=0.6, max_tokens=256)
outputs = llm.generate(["Hello, FastDMS."], sampling_params)
print(outputs[0]["text"])

Canonical Papers

Acknowledgements

FastDMS is built on top of nano-vLLM by Xingkai Yu. The clean, readable nano-vLLM codebase was used as the starting harness for the compact-DMS implementation. We're grateful for the upstream work that made rapid iteration on the KV-cache layout possible.

Citation

@misc{fastdms2026,
  title        = {FastDMS: Production-Speed Compact DMS for KV Cache Compression},
  author       = {{Leonard Lin}},
  year         = {2026},
  url          = {https://github.com/shisa-ai/FastDMS},
  note         = {Fast reference implementation of compact Dynamic Memory Sparsification with FP8 KV cache}
}

License

MIT