Persist and reuse KV Cache to speedup your LLM.
See the code
Documentation · Quickstart · Website · 中文
Unified Cache Manager (UCM) is a unified KV Cache layer between inference engines and storage. It manages cache reuse, placement, and movement so that compatible requests and inference instances can reuse existing computation.
Agent workflows, multi-turn conversations, and multimodal applications often revisit the same context. UCM makes reusable KV Cache available beyond an individual inference process and extends cache capacity through external storage. This helps reduce repeated computation and the pressure on accelerator memory.
Engine adapters, cache mechanisms, and storage backends can evolve independently through pluggable interfaces.
Feature availability and storage guarantees depend on the selected engine, model, platform, and backend; see supported configurations.
Installation commands, image tags, engine versions, and configuration examples are maintained in the documentation. For deployment problems, see troubleshooting.
Contributions to engine adapters, storage backends, cache mechanisms, tests, and documentation are welcome.
UCM is licensed under the MIT license with additional conditions. Read LICENSE for the full terms.
C++
56.7%
Python
38.7%
CMake
1.9%
Persist and reuse KV Cache to speedup your LLM.
See the code
Documentation · Quickstart · Website · 中文
Unified Cache Manager (UCM) is a unified KV Cache layer between inference engines and storage. It manages cache reuse, placement, and movement so that compatible requests and inference instances can reuse existing computation.
Agent workflows, multi-turn conversations, and multimodal applications often revisit the same context. UCM makes reusable KV Cache available beyond an individual inference process and extends cache capacity through external storage. This helps reduce repeated computation and the pressure on accelerator memory.
Engine adapters, cache mechanisms, and storage backends can evolve independently through pluggable interfaces.
Feature availability and storage guarantees depend on the selected engine, model, platform, and backend; see supported configurations.
Installation commands, image tags, engine versions, and configuration examples are maintained in the documentation. For deployment problems, see troubleshooting.
Contributions to engine adapters, storage backends, cache mechanisms, tests, and documentation are welcome.
UCM is licensed under the MIT license with additional conditions. Read LICENSE for the full terms.
C++
56.7%
Python
38.7%
CMake
1.9%