5 repos
caiovicentino/polarengine-vllm
PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM
36
12 commits
caiovicentino/eoq-quantization
EOQ: Entropy-Optimal Quantization for LLMs. 11-41% smaller than GGUF Q4_K_M with near-FP16…
46
7 commits
varjoranta/turboquant-vllm
TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused…
80
330 commits
caiovicentino1/Qwen3.5-9B-PolarEngine-v4
> [!IMPORTANT]
0
16 commits
TheTom/turboquant_plus
No description
7,030
342 commits