Qwen3 Model Optimization & Inference

13 repos

Optimized implementations and variants of Qwen3 language models across different scales (4B, 8B, 14B parameters), focusing on efficient inference techniques like speculative decoding (Eagle3), dynamic flash attention (dflash), and block-level optimizations. These repositories provide safetensors-compatible model weights and custom inference code for deploying Qwen3 and related architectures like Llama, addressing the engineering challenge of running capable LLMs with reduced latency and memory overhead.

qwen3 ·98
safetensors ·98
custom_code ·0
qwen3_vl ·0