11
stars
0
commits
9
repos using this model
4
linked in READMEs
Sep 3, 2025
updated
meta-llama/Llama-4-Maverick-17B-128E
98
bingyang-lei/qwen3-4b-eagle3-thinking-draftopd
nvidia/Qwen3.5-397B-A17B-NVFP4
107
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8
62
nvidia/Qwen3.8-Flash-Next-NVFP4
197
nvidia/Qwen3-235B-A22B-Eagle3
13
nvidia/gpt-oss-120b-Eagle3
75
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4
185
FutureMLS-Lab/OSCAR
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
557
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
35,766
rednote-machine-learning/RedKnot
Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
2,482
OpenMOSS/MOSS-VL
An open-weight 11B model series for long-form and real-time video understanding
593
OpenMOSS/MOSS-Music
An open-source model for music captioning, lyrics transcription, structural analysis, and musical…
158
Introspective-Diffusion/I-DLM
154
guqiong96/Lsglang
Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with…
124
libertywing/FlashMemory-Deepseek-V4
FlashMemory DS-V4 Retriever: a lightweight retriever that sparsifies DeepSeek-V4 CSA KV-cache.…
110
jpezzulli/sglang-rtxpro6000
Personal Optimized SGLang runtime for Qwen3.8-27B/DFlash2 and Qwen3.8 Flash-Next on one RTX PRO…
60
SafeAILab/EAGLE
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
2,527
Okapi-dog/quantized-eagle
BradMcDanel/EAGLE
EEErinzzz/Enhanced-Eagle