4 repos
microsoft/RetrievalAttention
[VLDB 26, NeurIPS 25] Scalable long-context LLM decoding that leverages sparsity—by treating the KV…
152
17 commits
santosardr/riskernel
RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
73
57 commits
rednote-machine-learning/RedKnot
Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
2,482
10 commits
libertywing/FlashMemory-Deepseek-V4
FlashMemory DS-V4 Retriever: a lightweight retriever that sparsifies DeepSeek-V4 CSA KV-cache.…
110
11 commits