11 repos
huawei-csl/KVarN
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput…
488
225 commits
ahb-sjsu/turboquant-pro
Consumer-aware compression for embedding indexes and LLM KV caches — compress by the metric the…
25
642 commits
intel/auto-round
A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包
1,610
1,400 commits
gauravapiscean/agentic-kv-cache
Reproducing agentic KV-cache policy claims on real traces. 68k requests from 393 Claude Code…
0
1 commits
orkait/orka.py
Vector-quantization compiler for LLM weights. Fits per-tensor RVQ codebooks, stores indices as…
1
262 commits
quantumaikr/quant.cpp
LLM inference with 7x longer context. Pure C, zero dependencies. Lossless KV cache compression +…
401
788 commits
gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090
No description
154
43 commits
jaeseok614/llm-gpu-checker-ko
AI hardware fit calculator for LLM, embedding, reranker, OCR and VLM workloads — VRAM, throughput,…
41
304 commits
Zhongzhu/OSCAR-RotationZoo
8
23 commits
liventruth/UL-SMF-Cache-Compression
The Unified Latent-State Memory Fabric (UL-SMF) is a hardware-software co-designed memory…
9
35 commits
LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
11,745
2,250 commits