Large Language Model Implementations & Inference

18 repos

Large language model implementations, inference optimization, and evaluation frameworks. This cluster centers on modern LLM architectures (DeepSeek, MiniMax) with emphasis on efficient quantization (FP8), model serialization via safetensors, and transformer-based implementations. Repositories focus on deployment-ready models with standardized endpoints, evaluation results tracking, and inference optimization techniques for production use.

safetensors ·47,957
endpoints_compatible ·47,954
transformers ·47,954
fp8 ·47,954
eval-results ·46,618
conversational ·45,759
text-generation ·45,723
custom_code ·29,799
deepseek_v3 ·23,430
text-generation-inference ·23,430