20 repos
C++ implementations and optimizations for running large language models locally, with a focus on inference acceleration, memory efficiency, and quantization techniques. The cluster centers on llama.cpp and its variants, exploring adaptive key-value caching, turboquantization, and hybrid inference strategies to reduce computational and memory overhead while maintaining model quality.