13 repos
PyTorch-based systems for serving and optimizing large transformer language models, with a focus on text generation inference at scale. The cluster centers on fine-tuned and quantized versions of foundational models (GPT-J, GPT-Neo, OPT, fairseq) optimized for inference efficiency, alongside inference frameworks and deployment tooling. This is a practical engineering area for practitioners looking to run large models in production environments with constrained resources.