Mergeable, privacy-safe sketch primitives for LLM/GenAI telemetry - HLL++, weighted frequent-items, Bloom, MinHash. Domain-separated keyed hashing, byte-deterministic wire format, conformant pure-Go and pure-Python implementations.
2
stars
16
commits
Go
primary language
Sep 5, 2026
updated
llm-sketchkit provides continuous, bounded answers about high-cardinality agent
traffic without exporting or indexing every underlying value. It is a small Go and
Python library with matching semantics for text canonicalization, privacy-preserving
keyed hashes, mergeable sketches, and a deterministic protobuf wire format.
It provides bounded measurement primitives for AI agent observability pipelines, especially when exporting and indexing every identity or event would be expensive, slow to query, or inappropriate to retain.
Raw prompts and identifiers do not need to enter sketch state. Producers can summarize locally and merge compatible sketches across processes or languages.
Use llm-sketchkit inside telemetry producers and processing components when
exporting and indexing every key would create uncontrolled cardinality, make
operational queries slow or unpredictable, or retain values that should not enter
aggregate state.
Agent fleets and multi-agent systems can produce more identities and events than an observability backend should continuously index. The library provides bounded, mergeable summaries across workers and windows.
Questions it can help answer include:
For investigations described as "tokenmaxxing" (also written "token-maxing"), reported token counts can be used as weights in the frequent-items sketch to identify which keyed values account for the most token volume. The library measures concentration; it does not infer task value, enforce budgets, or stop agent loops. See Token-Volume Heavy Hitters for a runnable example.
Inputs can be canonicalized and keyed before entering sketch state. This keeps raw values out of the sketch, but the resulting hashes remain pseudonymous and linkable while the same secret is in use.
This is a sketch library, not a trace collector, sampling processor, storage backend, dashboard, or differential-privacy system. The FAQ answers common questions about fit, operational bounds, accuracy, and interoperability.
| Your pipeline | Use | Why |
|---|---|---|
| GenAI spans already flow through an OpenTelemetry Collector | OpenTelemetry Collector connector | It applies keyed hashing, bounded windows, cardinality controls, and trace-to-metrics conversion at the collector boundary. |
| A custom Go or Python streaming service processes events | llm-sketchkit directly | Update sketches inside each bounded window, then serialize or merge compatible summaries. |
| A batch or warehouse job reads stored events | llm-sketchkit directly | Build bounded summaries per partition or window and merge them before publishing results. |
| You only need to store or visualize finished metrics | Your existing backend integration | The library is not an exporter; ClickHouse, Datadog, Prometheus, and similar systems normally receive the aggregated results. |
Use the collector path when the source spans already flow through OpenTelemetry. Use the library when you own the event-processing code or need matching Go and Python summaries outside an OpenTelemetry pipeline. In either case, compatible producers must agree on profile, hash domain, hash algorithm, secret, and window boundaries.
| Component | Use it for | Important property |
|---|---|---|
| HLL++ | Approximate distinct counts | Bounded, mergeable state |
| Weighted frequent-items | Heavy hitters and top items | Deterministic lower and upper bounds |
| Bloom filter | Set membership | No false negatives; configurable false-positive rate |
| MinHash | Approximate Jaccard similarity | Bounded, mergeable signatures |
The Go and Python implementations share the same profiles, hash domains, test vectors, and serialized representation.
Measurements use deterministic workloads and report the least favorable of five Linux benchmark runs where applicable.
small profile's maximum observed relative error was 2.3301%
across the characterization grid, within its 2.4375% enforced bound.k=128 to 0.02009 at
k=256, closely following the expected inverse-square-root relationship.See the visual scorecard, raw measurement records, and general-purpose library comparison for methods, limitations, and reproduction commands.
python -m pip install llm-sketchkit
Generate a process secret:
export LLM_SKETCHKIT_SECRET="$(python -c 'import secrets; print(secrets.token_hex(32))')"
Run a distinct-count example:
python - <<'PY'
from llm_sketchkit import PROMPT_V1, canonicalize_text_v1, hash64, hllpp
from llm_sketchkit import secret_from_env
secret = secret_from_env("LLM_SKETCHKIT_SECRET")
sketch = hllpp.new("small", PROMPT_V1)
canonical = canonicalize_text_v1(" cafe\u0301\r\n")
sketch.add_hash(hash64(secret, PROMPT_V1, canonical))
print(f"estimated distinct prompts: {sketch.estimate():.0f}")
PY
Expected output:
estimated distinct prompts: 1
From an existing Go module, add the package used by the equivalent example:
go get github.com/llm-measurement/llm-sketchkit/go/sketchkit/hllpp@latest
Use the same LLM_SKETCHKIT_SECRET and run:
package main
import (
"fmt"
"log"
"github.com/llm-measurement/llm-sketchkit/go/sketchkit/canon"
sketchhash "github.com/llm-measurement/llm-sketchkit/go/sketchkit/hash"
"github.com/llm-measurement/llm-sketchkit/go/sketchkit/hllpp"
)
func main() {
secret, err := sketchhash.SecretFromEnv("LLM_SKETCHKIT_SECRET")
if err != nil {
log.Fatal(err)
}
sketch, err := hllpp.New(
hllpp.ProfileSmall,
sketchhash.PromptV1,
sketchhash.HMACSHA25664,
)
if err != nil {
log.Fatal(err)
}
canonical, err := canon.CanonicalizeString(canon.TextV1, " cafe\u0301\r\n")
if err != nil {
log.Fatal(err)
}
digest, err := sketchhash.Hash64(secret, sketchhash.PromptV1, canonical)
if err != nil {
log.Fatal(err)
}
sketch.AddHash(digest)
fmt.Printf("estimated distinct prompts: %.0f\n", sketch.Estimate())
}
The runnable notebook produces per-service HLL++ and weighted frequent-items summaries in Go, then loads, validates, merges, and plots them in Python. It checks that serialization round trips return the same bytes, explicit merge rejection for incompatible profiles, distinct-count estimates, and deterministic bounds around token-heavy pseudonymous keys.
The tutorial uses an exact synthetic side channel only to check its estimates. Raw synthetic identifiers do not enter the emitted files. See the example guide for setup and security details.
If by "tokenmaxxing" (also written "token-maxing") you mean unexpected or runaway
token consumption, use a bounded frequent-items sketch to find where reported volume
is concentrated. Hash the value being investigated, such as a prompt template, tool,
or user, with the registered domain for that entity class, and use the reported token
count as its weight. This example measures prompt templates with prompt:v1:
from llm_sketchkit import PROMPT_V1, canonicalize_text_v1, frequentitems
from llm_sketchkit import hash64, secret_from_env
secret = secret_from_env("LLM_SKETCHKIT_SECRET")
sketch = frequentitems.Sketch("small", PROMPT_V1)
events = [
("support/refund", 1_240),
("research/synthesis", 8_900),
("support/refund", 980),
("coding/review", 3_600),
]
for prompt_template, reported_tokens in events:
canonical = canonicalize_text_v1(prompt_template)
digest = hash64(secret, PROMPT_V1, canonical)
sketch.add_hash(digest, reported_tokens)
for item in sketch.frequent_items(frequentitems.NO_FALSE_NEGATIVES)[:10]:
print(
f"{item.hash:016x} estimate={item.estimate} "
f"bounds=[{item.lower_bound}, {item.upper_bound}]"
)
The sketch retains at most the selected profile's bounded map size. Returned hashes are pseudonymous and remain linkable while the same secret and domain are in use. The estimate is not a billing total; use the lower and upper bounds when deciding whether an item is meaningfully heavy.
Do not substitute guessed weights when token usage is missing. Count missing usage
separately so operators know how complete the reported totals are. For an OTLP
pipeline with ready-made metrics, bounded slices, missing-usage accounting, and
token-weighted top-k snapshots, use
otelcol-genai-sketches.
Sketches merge only when their kind, profile, hash domain, hash algorithm, and shape metadata match. A mismatch is an error rather than an implicit conversion.
left = hllpp.new("small", PROMPT_V1)
right = hllpp.new("small", PROMPT_V1)
left.add_hash(hash64(secret, PROMPT_V1, canonicalize_text_v1("alpha")))
right.add_hash(hash64(secret, PROMPT_V1, canonicalize_text_v1("beta")))
left.merge(right)
print(f"merged distinct prompts: {left.estimate():.0f}")
See SECURITY.md for private vulnerability reporting.
git clone https://github.com/llm-measurement/llm-sketchkit.git
cd llm-sketchkit
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip==26.2 setuptools==83.0.0
python -m pip install -e '.[dev]'
Deterministic protobuf encoding is part of the compatibility surface. Go and
Python are checked against the same canonicalization, hashing, sketch, and
cross-language merge fixtures in
vectors/.
Run all local checks:
go test ./... -race
python -m pytest -q
ruff check .
mypy --strict
The optional Apache DataSketches comparison checks weighted frequent-items query behavior against an independent implementation:
python -m pip install -e '.[oracle]'
python scripts/datasketches_oracle.py --check
spec/ defines canonicalization, hashing, profiles, and wire encoding.vectors/ contains executable conformance fixtures.reports/ contains benchmark, accuracy, and oracle results.bench/ contains the Go benchmark harnesses.docs/FAQ.md answers common adoption questions.CHANGELOG.md records release-level changes.llm-sketchkit is an alpha library. The compatibility surface consists of the
specifications, the Go and Python APIs exercised by the vectors, and the checked-in
conformance fixtures.
Apache-2.0.
16 commits
Go
55.3%
Python
44.1%
Mergeable, privacy-safe sketch primitives for LLM/GenAI telemetry - HLL++, weighted frequent-items, Bloom, MinHash. Domain-separated keyed hashing, byte-deterministic wire format, conformant pure-Go and pure-Python implementations.
2
stars
16
commits
Go
primary language
Sep 5, 2026
updated
llm-sketchkit provides continuous, bounded answers about high-cardinality agent
traffic without exporting or indexing every underlying value. It is a small Go and
Python library with matching semantics for text canonicalization, privacy-preserving
keyed hashes, mergeable sketches, and a deterministic protobuf wire format.
It provides bounded measurement primitives for AI agent observability pipelines, especially when exporting and indexing every identity or event would be expensive, slow to query, or inappropriate to retain.
Raw prompts and identifiers do not need to enter sketch state. Producers can summarize locally and merge compatible sketches across processes or languages.
Use llm-sketchkit inside telemetry producers and processing components when
exporting and indexing every key would create uncontrolled cardinality, make
operational queries slow or unpredictable, or retain values that should not enter
aggregate state.
Agent fleets and multi-agent systems can produce more identities and events than an observability backend should continuously index. The library provides bounded, mergeable summaries across workers and windows.
Questions it can help answer include:
For investigations described as "tokenmaxxing" (also written "token-maxing"), reported token counts can be used as weights in the frequent-items sketch to identify which keyed values account for the most token volume. The library measures concentration; it does not infer task value, enforce budgets, or stop agent loops. See Token-Volume Heavy Hitters for a runnable example.
Inputs can be canonicalized and keyed before entering sketch state. This keeps raw values out of the sketch, but the resulting hashes remain pseudonymous and linkable while the same secret is in use.
This is a sketch library, not a trace collector, sampling processor, storage backend, dashboard, or differential-privacy system. The FAQ answers common questions about fit, operational bounds, accuracy, and interoperability.
| Your pipeline | Use | Why |
|---|---|---|
| GenAI spans already flow through an OpenTelemetry Collector | OpenTelemetry Collector connector | It applies keyed hashing, bounded windows, cardinality controls, and trace-to-metrics conversion at the collector boundary. |
| A custom Go or Python streaming service processes events | llm-sketchkit directly | Update sketches inside each bounded window, then serialize or merge compatible summaries. |
| A batch or warehouse job reads stored events | llm-sketchkit directly | Build bounded summaries per partition or window and merge them before publishing results. |
| You only need to store or visualize finished metrics | Your existing backend integration | The library is not an exporter; ClickHouse, Datadog, Prometheus, and similar systems normally receive the aggregated results. |
Use the collector path when the source spans already flow through OpenTelemetry. Use the library when you own the event-processing code or need matching Go and Python summaries outside an OpenTelemetry pipeline. In either case, compatible producers must agree on profile, hash domain, hash algorithm, secret, and window boundaries.
| Component | Use it for | Important property |
|---|---|---|
| HLL++ | Approximate distinct counts | Bounded, mergeable state |
| Weighted frequent-items | Heavy hitters and top items | Deterministic lower and upper bounds |
| Bloom filter | Set membership | No false negatives; configurable false-positive rate |
| MinHash | Approximate Jaccard similarity | Bounded, mergeable signatures |
The Go and Python implementations share the same profiles, hash domains, test vectors, and serialized representation.
Measurements use deterministic workloads and report the least favorable of five Linux benchmark runs where applicable.
small profile's maximum observed relative error was 2.3301%
across the characterization grid, within its 2.4375% enforced bound.k=128 to 0.02009 at
k=256, closely following the expected inverse-square-root relationship.See the visual scorecard, raw measurement records, and general-purpose library comparison for methods, limitations, and reproduction commands.
python -m pip install llm-sketchkit
Generate a process secret:
export LLM_SKETCHKIT_SECRET="$(python -c 'import secrets; print(secrets.token_hex(32))')"
Run a distinct-count example:
python - <<'PY'
from llm_sketchkit import PROMPT_V1, canonicalize_text_v1, hash64, hllpp
from llm_sketchkit import secret_from_env
secret = secret_from_env("LLM_SKETCHKIT_SECRET")
sketch = hllpp.new("small", PROMPT_V1)
canonical = canonicalize_text_v1(" cafe\u0301\r\n")
sketch.add_hash(hash64(secret, PROMPT_V1, canonical))
print(f"estimated distinct prompts: {sketch.estimate():.0f}")
PY
Expected output:
estimated distinct prompts: 1
From an existing Go module, add the package used by the equivalent example:
go get github.com/llm-measurement/llm-sketchkit/go/sketchkit/hllpp@latest
Use the same LLM_SKETCHKIT_SECRET and run:
package main
import (
"fmt"
"log"
"github.com/llm-measurement/llm-sketchkit/go/sketchkit/canon"
sketchhash "github.com/llm-measurement/llm-sketchkit/go/sketchkit/hash"
"github.com/llm-measurement/llm-sketchkit/go/sketchkit/hllpp"
)
func main() {
secret, err := sketchhash.SecretFromEnv("LLM_SKETCHKIT_SECRET")
if err != nil {
log.Fatal(err)
}
sketch, err := hllpp.New(
hllpp.ProfileSmall,
sketchhash.PromptV1,
sketchhash.HMACSHA25664,
)
if err != nil {
log.Fatal(err)
}
canonical, err := canon.CanonicalizeString(canon.TextV1, " cafe\u0301\r\n")
if err != nil {
log.Fatal(err)
}
digest, err := sketchhash.Hash64(secret, sketchhash.PromptV1, canonical)
if err != nil {
log.Fatal(err)
}
sketch.AddHash(digest)
fmt.Printf("estimated distinct prompts: %.0f\n", sketch.Estimate())
}
The runnable notebook produces per-service HLL++ and weighted frequent-items summaries in Go, then loads, validates, merges, and plots them in Python. It checks that serialization round trips return the same bytes, explicit merge rejection for incompatible profiles, distinct-count estimates, and deterministic bounds around token-heavy pseudonymous keys.
The tutorial uses an exact synthetic side channel only to check its estimates. Raw synthetic identifiers do not enter the emitted files. See the example guide for setup and security details.
If by "tokenmaxxing" (also written "token-maxing") you mean unexpected or runaway
token consumption, use a bounded frequent-items sketch to find where reported volume
is concentrated. Hash the value being investigated, such as a prompt template, tool,
or user, with the registered domain for that entity class, and use the reported token
count as its weight. This example measures prompt templates with prompt:v1:
from llm_sketchkit import PROMPT_V1, canonicalize_text_v1, frequentitems
from llm_sketchkit import hash64, secret_from_env
secret = secret_from_env("LLM_SKETCHKIT_SECRET")
sketch = frequentitems.Sketch("small", PROMPT_V1)
events = [
("support/refund", 1_240),
("research/synthesis", 8_900),
("support/refund", 980),
("coding/review", 3_600),
]
for prompt_template, reported_tokens in events:
canonical = canonicalize_text_v1(prompt_template)
digest = hash64(secret, PROMPT_V1, canonical)
sketch.add_hash(digest, reported_tokens)
for item in sketch.frequent_items(frequentitems.NO_FALSE_NEGATIVES)[:10]:
print(
f"{item.hash:016x} estimate={item.estimate} "
f"bounds=[{item.lower_bound}, {item.upper_bound}]"
)
The sketch retains at most the selected profile's bounded map size. Returned hashes are pseudonymous and remain linkable while the same secret and domain are in use. The estimate is not a billing total; use the lower and upper bounds when deciding whether an item is meaningfully heavy.
Do not substitute guessed weights when token usage is missing. Count missing usage
separately so operators know how complete the reported totals are. For an OTLP
pipeline with ready-made metrics, bounded slices, missing-usage accounting, and
token-weighted top-k snapshots, use
otelcol-genai-sketches.
Sketches merge only when their kind, profile, hash domain, hash algorithm, and shape metadata match. A mismatch is an error rather than an implicit conversion.
left = hllpp.new("small", PROMPT_V1)
right = hllpp.new("small", PROMPT_V1)
left.add_hash(hash64(secret, PROMPT_V1, canonicalize_text_v1("alpha")))
right.add_hash(hash64(secret, PROMPT_V1, canonicalize_text_v1("beta")))
left.merge(right)
print(f"merged distinct prompts: {left.estimate():.0f}")
See SECURITY.md for private vulnerability reporting.
git clone https://github.com/llm-measurement/llm-sketchkit.git
cd llm-sketchkit
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip==26.2 setuptools==83.0.0
python -m pip install -e '.[dev]'
Deterministic protobuf encoding is part of the compatibility surface. Go and
Python are checked against the same canonicalization, hashing, sketch, and
cross-language merge fixtures in
vectors/.
Run all local checks:
go test ./... -race
python -m pytest -q
ruff check .
mypy --strict
The optional Apache DataSketches comparison checks weighted frequent-items query behavior against an independent implementation:
python -m pip install -e '.[oracle]'
python scripts/datasketches_oracle.py --check
spec/ defines canonicalization, hashing, profiles, and wire encoding.vectors/ contains executable conformance fixtures.reports/ contains benchmark, accuracy, and oracle results.bench/ contains the Go benchmark harnesses.docs/FAQ.md answers common adoption questions.CHANGELOG.md records release-level changes.llm-sketchkit is an alpha library. The compatibility surface consists of the
specifications, the Go and Python APIs exercised by the vectors, and the checked-in
conformance fixtures.
Apache-2.0.
16 commits
Go
55.3%
Python
44.1%