Large Language Model Inference & Deployment

11 repos

Python-based tools and frameworks for running, optimizing, and serving large language models in production environments. This cluster focuses on practical LLM inference infrastructure, including model hosting platforms, chat interfaces, and deployment solutions. Repositories range from end-to-end LLM application frameworks to specialized inference engines, with particular emphasis on making LLMs accessible and efficient across different hardware configurations.

Python · 10
C++ · 1
ai-chat ·77,385
llm-inference ·77,385