4 repos
fvliang/DART
Official Implementation of DART (DART: Low-Latency Parallel Drafting with Continuity-Aware Tree…
71
1 commits
FasterDecoding/REST
REST: Retrieval-Based Speculative Decoding, NAACL 2024
220
12 commits
Infini-AI-Lab/UMbreLLa
LLM Inference on consumer devices
132
149 commits
Scottcjn/rpi-inference
RPI (Resonant Permutation Inference) — Zero-multiply text generation. 18K tok/s. 868 KB models.…
41
19 commits