4 repos
ims-kdks/TIDE
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
7
22 commits
Maknee/Kimi-K3.cpp
Run Kimi-K3 (2.8T MoE) from SSD with a tiny readable C++17 engine—native MXFP4, io_uring, and…
5
4 commits
flashserve/Mosaic
MOSAIC: Unlocking Over 30× Context Length for Diffusion LLMs Inference via Global Memory Planning…
1 commits
argonautlabsai/deltafin
ARGODRIVE Deltafin: Kimi K3 (2.8T MoE) streamed from SSDs on Apple Silicon — fork of…
84
41 commits