20 repos
Tools and frameworks for optimizing, quantizing, and deploying machine learning models using ONNX format and ONNX Runtime. The cluster focuses on making models smaller, faster, and more portable—particularly for edge inference and resource-constrained environments. Key themes include model compression, format conversion (especially GGUF and llama.cpp integration), and runtime optimization across different platforms.