Large Language Model Inference & Deployment

13 repos

PyTorch-based systems for serving and optimizing large transformer language models, with a focus on text generation inference at scale. The cluster centers on fine-tuned and quantized versions of foundational models (GPT-J, GPT-Neo, OPT, fairseq) optimized for inference efficiency, alongside inference frameworks and deployment tooling. This is a practical engineering area for practitioners looking to run large models in production environments with constrained resources.

pytorch ·553
transformers ·553
text-generation ·553
en ·538
text-generation-inference ·410
opt ·386
endpoints_compatible ·191
xglm ·76
gptj ·54
safetensors ·31