Language Model Inference and Deployment

10 repos

Libraries, models, and tools for running and serving large language models with a focus on efficient inference, text generation, and conversational AI. The cluster centers on transformer-based language models (particularly IBM's Granite models) optimized for deployment, along with supporting infrastructure for model serialization, loading, and execution. Resources here cover both model artifacts and the frameworks needed to integrate them into applications.

conversational ·642
language ·642
safetensors ·642
text-generation ·642
transformers ·642
granite ·526
endpoints_compatible ·458
granite-4.1 ·458
eval-results ·388
granite-3.3 ·161