11 repos
Python-based tools and frameworks for running, optimizing, and serving large language models in production environments. This cluster focuses on practical LLM inference infrastructure, including model hosting platforms, chat interfaces, and deployment solutions. Repositories range from end-to-end LLM application frameworks to specialized inference engines, with particular emphasis on making LLMs accessible and efficient across different hardware configurations.