LLM Inference Optimization & Quantization

20 repos

C++ implementations and optimizations for running large language models locally, with a focus on inference acceleration, memory efficiency, and quantization techniques. The cluster centers on llama.cpp and its variants, exploring adaptive key-value caching, turboquantization, and hybrid inference strategies to reduce computational and memory overhead while maintaining model quality.

C++ · 18
Python · 2
llm-inference ·22,190
large-language-models ·19,610
llama ·19,610
local-inference ·19,610
llm ·19,610
speculative-decoding ·2,676
intel-optimized-llamacpp ·2,172
autoround ·2,172
chatpdf ·2,172
gaudi3 ·2,172

dorogit/inteLearn_ML

No description

Python

0

1,933 commits