Local LLM Inference Optimization

19 repos

C++ implementations and optimizations for running large language models efficiently on consumer hardware without cloud dependencies. The cluster focuses on inference acceleration techniques including quantization, key-value cache optimization, and hybrid execution strategies applied to llama.cpp and related inference engines. Developers exploring this area will find optimized model serving code, performance tuning frameworks, and techniques for reducing memory footprint and latency in edge and local deployment scenarios.

C++ · 17
Python · 2
llm-inference ·12,353
large-language-models ·9,788
llama ·9,788
local-inference ·9,788
llm ·9,788
speculative-decoding ·2,659
intel-optimized-llamacpp ·2,171
autoround ·2,171
chatpdf ·2,171
gaudi3 ·2,171

dorogit/inteLearn_ML

No description

Python

0

1,933 commits