5 repos
EfficientMoE/MoE-Infinity
PyTorch library for cost-effective, fast and easy serving of MoE models.
360
129 commits
batchgen-project/batchgen
High-Throughput Batch Inference
14
357 commits
cmavro/PackLLM
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
15
1 commits
QuixiAI/laserRMT
This is our own implementation of 'Layer Selective Rank Reduction'
240
57 commits
PSNbst/PAseer-TDD-Accelerator
LoRA 的结构:
0
6 commits