GGML inference and quantization

28 repos

C++ libraries and tools for efficient machine learning inference, centered on the GGML tensor library and GGUF quantization format. The cluster includes llama.cpp as its most prominent project—a portable inference engine for large language models—alongside related tooling for model quantization, speech-to-text processing, and cross-platform deployment. Most repos are low-level C++ implementations optimized for running ML models on consumer hardware with minimal dependencies.

C++ · 19
Python · 8
C · 1
ggml ·267,191
gguf ·63,531
cross-platform ·51,878
local-inference ·51,878
open-source-ai ·51,878
single-file-executable ·51,878
llama-cpp ·51,878
local-llm ·51,878
local-ai ·51,878
speech-to-text ·51,878