18 repos
Python-based frameworks and tools for optimizing large language model inference performance, memory efficiency, and deployment at scale. This cluster covers serving systems (vLLM, SGLang), inference engines with techniques like quantization and batching, and specialized optimizations for running LLMs on resource-constrained hardware. Repositories span reference implementations, production serving platforms, and experimental inference acceleration methods.
gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090
RTX5090 Blackwell GPU Optimized Qwen-3.8-27B-NVFP4