5 repos
huawei-csl/KVarN
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput…
496
225 commits
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the…
25
642 commits
orkait/orka.py
Vector-quantization compiler for LLM weights. Fits per-tensor RVQ codebooks, stores indices as…
1
262 commits
quantumaikr/quant.cpp
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression +…
403
788 commits
liventruth/UL-SMF-Cache-Compression
The Unified Latent-State Memory Fabric (UL-SMF) is a hardware-software co-designed memory…
9
35 commits