23 repos
Frameworks and tools for efficiently serving and running large language models in production, with emphasis on inference optimization, speculative decoding, and specialized architectures like MoE and Llama variants. The cluster centers on SGLang and related systems that enable fast, flexible LLM inference through techniques like structured generation, dynamic batching, and hardware-aware optimization. Most repositories are Python-based implementations spanning model serving, inference engines, and compiler-level optimizations.
huoji120/QWEN-EXO-booster
QWEN-BASE model-native memory and continuous learning system,base on SGLang