Qwen LLM Inference Optimization

9 repos

Optimized inference implementations and deployments of Alibaba's Qwen language models, with focus on techniques like speculative decoding, quantization (FP8), and hardware-specific acceleration for consumer GPUs (RTX 3090) and AMD ROCm platforms. The cluster centers on vLLM-based serving frameworks and model-specific optimizations that reduce latency and memory overhead for running large language models in resource-constrained environments.

Python · 4
C++ · 2
HTML · 1
Rust · 1
speculative-decoding ·7,478
qwen ·7,189
cuda ·7,188
luce ·5,734
local-ai ·5,734
heterogeneous-computing ·5,734
kernel ·5,734
cuda-kernels ·5,734
dflash ·5,734
llama-cpp ·5,734