13 repos
lvyufeng/PocketLLM
Inference engine for large language models on consumer GPUs, with deep optimization for older…
53
334 commits
giannisanni/pulsar
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3…
221
416 commits
Helldez/BigMoeOnEdge
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on…
576
373 commits
FlashML-org/FreeToken
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast…
13,068
43 commits
Entrpi/ds4-on-spark
Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x…
400
63 commits
MatN23/AdaptiveTrainingSystem
BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES
21
602 commits
siris9476/pulsarforge
No description
2 commits
pjordanandrsn/experts4bit-qlora
Train and serve MoE models that do not fit in VRAM: fused 4-bit experts, QLoRA, CPU/NVMe offload,…
3
734 commits
pytorch/ao
PyTorch native quantization for training and inference
2,979
2,570 commits
ims-kdks/TIDE
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
7
22 commits
Maknee/Kimi-K3.cpp
Run Kimi-K3 (2.8T MoE) from SSD with a tiny readable C++17 engine—native MXFP4, io_uring, and…
6
4 commits
jd-opensource/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI…
1,574
1,371 commits
SamsungSAILMontreal/ream
REAM: Merging Improves Pruning of Experts in LLMs
26
9 commits