HarryWu99/llm_kvcache_sparsity

Implement some method of LLM KV Cache Sparsity

Python

41

2 commits

updated Jun 6, 2024

See the code

README

LLM KV Cache Sparsity

Implement some method of LLM KV Cache Sparsity, including:

  1. Efficient Streaming Language Models with Attention Sinks, also called "SinkCache"
  2. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
  3. SnapKV: LLM Knows What You are Looking for Before Generation

To Run

pip install -r requirements.txt
# edit longbench loading method `load_from_disk` in example/test.py
python example/test.py --sparsity_method snapkv

The result file will write to results folder.

Then you can use longbench_eval/eval.py to get the scores.

The core code for KV Cache eviction is in models/kv_clusters.py

Results

Todo

HarryWu99/llm_kvcache_sparsity

Implement some method of LLM KV Cache Sparsity

Python

41

2 commits

updated Jun 6, 2024

See the code

README

LLM KV Cache Sparsity

Implement some method of LLM KV Cache Sparsity, including:

  1. Efficient Streaming Language Models with Attention Sinks, also called "SinkCache"
  2. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
  3. SnapKV: LLM Knows What You are Looking for Before Generation

To Run

pip install -r requirements.txt
# edit longbench loading method `load_from_disk` in example/test.py
python example/test.py --sparsity_method snapkv

The result file will write to results folder.

Then you can use longbench_eval/eval.py to get the scores.

The core code for KV Cache eviction is in models/kv_clusters.py

Results

Todo

Languages

Python

100.0%