9 repos
Optimized inference implementations and deployments of Alibaba's Qwen language models, with focus on techniques like speculative decoding, quantization (FP8), and hardware-specific acceleration for consumer GPUs (RTX 3090) and AMD ROCm platforms. The cluster centers on vLLM-based serving frameworks and model-specific optimizations that reduce latency and memory overhead for running large language models in resource-constrained environments.