ModelEngine-Group/unified-cache-management

Persist and reuse KV Cache to speedup your LLM.

C++

339

740 commits

updated Sep 28, 2026

See the code

README

UCM

Unified KV Cache management for LLM inference

Documentation · Quickstart · Website · 中文

DeepWiki

What is UCM?

Unified Cache Manager (UCM) is a unified KV Cache layer between inference engines and storage. It manages cache reuse, placement, and movement so that compatible requests and inference instances can reuse existing computation.

Agent workflows, multi-turn conversations, and multimodal applications often revisit the same context. UCM makes reusable KV Cache available beyond an individual inference process and extends cache capacity through external storage. This helps reduce repeated computation and the pressure on accelerator memory.

Engine adapters, cache mechanisms, and storage backends can evolve independently through pluggable interfaces.

Architecture

UCM logical architecture: inference workloads and engines, unified cache lifecycle management, and a global KV store with pluggable backends.

  • Applications and inference engines produce and consume KV Cache. Reuse patterns include multimodal context, shared prefixes, and cached chunks used by mechanisms such as KV Bridge.
  • Unified Cache manages the cache's usage lifecycle: what to reuse, where to retain it, and when to move it. Prefill/decode (PD) transfer is one form of cache movement. Pluggable retrieval and loading strategies also support sparse attention.
  • Global KV Store provides a common storage abstraction over file systems, memory pools, and other backends. Store implementations handle retention, reclamation, and reliability according to their capabilities.

Feature availability and storage guarantees depend on the selected engine, model, platform, and backend; see supported configurations.

Getting Started

  1. Check the support matrix for a compatible configuration.
  2. Follow the Quickstart to select installation artifacts, configure storage, and run an engine with UCM.
  3. Use the operations guide to verify cache reuse and inspect runtime behavior.

Installation commands, image tags, engine versions, and configuration examples are maintained in the documentation. For deployment problems, see troubleshooting.

Contributing & Community

Contributions to engine adapters, storage backends, cache mechanisms, tests, and documentation are welcome.

License

UCM is licensed under the MIT license with additional conditions. Read LICENSE for the full terms.

ascend
cuda
deepseek
dram
gpu
hbm
kvcache
llm
nfs
npu
ssd
torch
ucm
vllm

Significant stargazers

samzong

204 followers · starred Nov 2025

Gang Chen

45 followers · starred Nov 2025

Fisher Xu

271 followers · starred Aug 2026

ModelEngine-Group/unified-cache-management

Persist and reuse KV Cache to speedup your LLM.

C++

339

740 commits

updated Sep 28, 2026

See the code

README

UCM

Unified KV Cache management for LLM inference

Documentation · Quickstart · Website · 中文

DeepWiki

What is UCM?

Unified Cache Manager (UCM) is a unified KV Cache layer between inference engines and storage. It manages cache reuse, placement, and movement so that compatible requests and inference instances can reuse existing computation.

Agent workflows, multi-turn conversations, and multimodal applications often revisit the same context. UCM makes reusable KV Cache available beyond an individual inference process and extends cache capacity through external storage. This helps reduce repeated computation and the pressure on accelerator memory.

Engine adapters, cache mechanisms, and storage backends can evolve independently through pluggable interfaces.

Architecture

UCM logical architecture: inference workloads and engines, unified cache lifecycle management, and a global KV store with pluggable backends.

  • Applications and inference engines produce and consume KV Cache. Reuse patterns include multimodal context, shared prefixes, and cached chunks used by mechanisms such as KV Bridge.
  • Unified Cache manages the cache's usage lifecycle: what to reuse, where to retain it, and when to move it. Prefill/decode (PD) transfer is one form of cache movement. Pluggable retrieval and loading strategies also support sparse attention.
  • Global KV Store provides a common storage abstraction over file systems, memory pools, and other backends. Store implementations handle retention, reclamation, and reliability according to their capabilities.

Feature availability and storage guarantees depend on the selected engine, model, platform, and backend; see supported configurations.

Getting Started

  1. Check the support matrix for a compatible configuration.
  2. Follow the Quickstart to select installation artifacts, configure storage, and run an engine with UCM.
  3. Use the operations guide to verify cache reuse and inspect runtime behavior.

Installation commands, image tags, engine versions, and configuration examples are maintained in the documentation. For deployment problems, see troubleshooting.

Contributing & Community

Contributions to engine adapters, storage backends, cache mechanisms, tests, and documentation are welcome.

License

UCM is licensed under the MIT license with additional conditions. Read LICENSE for the full terms.

ascend
cuda
deepseek
dram
gpu
hbm
kvcache
llm
nfs
npu
ssd
torch
ucm
vllm

Significant stargazers

samzong

204 followers · starred Nov 2025

Gang Chen

45 followers · starred Nov 2025

Fisher Xu

271 followers · starred Aug 2026

Languages

C++

56.7%

Python

38.7%

CMake

1.9%