Research alpha. MLX-RFSN is a KV-cache compression research and benchmarking system for Apple Silicon. It is not production-ready, not a general AI server, and not a full vector database.
Official promoted candidate: NONE
Best practical baseline: rfsn_direct_packed_k8v8
Promotion allowed: false (pending runtime-instrumented cache trace and token-sequence-hash provenance)
Release ID: alpha-8.4 (from release.toml)
Package Version: 10.2.0a84 (from release.toml)
Artifact Schema: 3.0 (from release.toml)
To verify the current state locally:
python -m compileall -q rfsn_v10 rfsn_v11 tests benchmarks scripts
bash scripts/release_gate.sh
| Path | Status |
|---|---|
| Stable runtime (MLX 8-bit KV compression) | Alpha — validated on Apple Silicon, quality gates passing |
| Package installation (subpackages) | Fixed — rfsn_v10.kernels, rfsn_v10.runtime install correctly |
CLI health check (python -m rfsn_v10 healthcheck) | Working |
| Sparse decode | Disabled by default — not end-to-end proven |
| QJL score correction | Experimental — disabled by default, requires explicit opt-in |
| Polar / hybrid quantization | Experimental — disabled by default, requires explicit opt-in |
| Adaptive sparse controller | Experimental — disabled by default, requires explicit opt-in |
| CUDA backend | Not implemented |
| Full portable runtime | Not implemented — MLX required for core runtime |
| End-to-end speedup | Not proven — decode TPS comparable, compression overhead makes total slower at short contexts |
| Research server | FastAPI server — /v1/chat/completions with SSE streaming, authentication, and proper wheel packaging (mlx/numpy). Not production-hardened. |
| Docker | Healthcheck validation + ClickHouse telemetry (CPU-only, no inference) |
| >8-bit compression | Uses raw uint32 fallback — bit-packing is real for 2-8 bit only |
| Metal kernel | Scaffold/stub with CPU fallback — actual Metal GPU computation NOT yet implemented |
| Experimental throughput | No experimental throughput speedup is proven |
| Platform | Status |
|---|---|
| Apple Silicon + MLX | Supported (primary runtime) |
| NumPy CPU | Supported — kernel validation, config, security tests pass; MLX-dependent runtime tests skip |
| Linux / CI | CPU tests pass; MLX suites skip cleanly |
| CUDA | Not implemented |
| macOS x86 (Intel) | MLX not supported on Intel Macs — NumPy-only |
# Three install modes — do not install everything at once
# Basic: core only, no MLX, no memory (any platform)
pip install -e ".[basic]"
# Fusion: MLX + KV compression benchmarks (macOS Apple Silicon)
pip install -e ".[fusion]"
# Memory: external vector memory (any platform)
pip install -e ".[memory]"
# Dev: everything for development
pip install -e ".[fusion,memory,dev]"
# Verify install
python -m rfsn_v10 version
# Health check
python -m rfsn_v10 healthcheck
# Validate config
python -m rfsn_v10 validate-config --config configs/default_runtime.yaml
# CPU-safe tests (no MLX required)
pytest tests/test_config.py tests/test_config_strict.py \
tests/test_kernels_validation.py \
tests/test_quantization_lazy_imports.py \
tests/test_experimental_flags.py \
tests/test_clickhouse_security.py \
tests/test_no_runtime_raw_sdpa.py -q
# MLX-dependent tests (Apple Silicon required)
pytest tests/test_attention.py tests/test_bitpack.py \
tests/test_bitpack_fuzz.py tests/test_kv_manager.py \
tests/test_drift.py tests/test_attention_causal_mask.py \
tests/test_short_prompt_decode_drift.py \
tests/test_prefill_decode_split.py -q
# Central benchmark command
python benchmarks/kv_shootout.py --promotion-report
Run the OpenAI-compatible FastAPI server locally:
export RFSN_MODEL_ID=mlx-community/Llama-3-8B-Instruct-4bit
python -m rfsn_v10.server
# Or with uvicorn directly:
# uvicorn rfsn_v10.server.app:app --host 0.0.0.0 --port 8000
Test the endpoint:
curl http://localhost:8000/health
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Hello"}],"stream":true}'
Or with Docker Compose (container validation + telemetry):
export CLICKHOUSE_PASSWORD=your-password
docker compose up -d
# Runs healthcheck and exits — inference server requires mlx/torch backend natively
rfsn_v10/bitpack.py — Bit-packed quantizer (2–8 bit widths) with exact roundtrip guaranteesrfsn_v10/kv_manager.py — TurboQuant KV manager with grouped symmetric quantization and WHT preconditioningrfsn_v10/attention.py — Adaptive block-sparse attention; all dense fallbacks route through attention_reference.pyrfsn_v10/attention_reference.py — Canonical causal attention reference (always applies causal mask for T_q > 1)rfsn_v10/runtime/engine.py — Orchestrator integrating KV cache, sparse attention, audit mode, and telemetryrfsn_v10/runtime/__init__.py — Re-exports RFSNRuntime from engine.pyrfsn_v10/runtime/generation.py — RFSNGenerator with prefill, decode, sampling, and telemetryrfsn_v10/config.py — Strict Pydantic config (extra='forbid' on all models)rfsn_v10/health.py — Health check system; returns UNHEALTHY until checks have been runrfsn_v10/clickhouse_client.py — HMAC-SHA256 prompt sanitization, recursive sanitizer, retry queuerfsn_v10/model_loader.py — Unified model/tokenizer loading (mlx-lm / transformers)rfsn_v10/server/app.py — FastAPI OpenAI-compatible server with SSE streamingAll experimental paths require explicit opt-in. The runtime will not activate them
silently — attempting to use an experimental feature without enabling it raises a
RuntimeError.
# config.yaml
experimental:
enable_qjl: false # QJL score correction
enable_polar: false # Polar / hybrid quantization
enable_adaptive: false # Adaptive sparse controller
Or via environment:
RFSN_EXPERIMENTAL_QJL=true
RFSN_EXPERIMENTAL_POLAR=true
RFSN_EXPERIMENTAL_ADAPTIVE=true
Warning: Experimental features are not validated for production or quality-critical generation.
rfsn_v10/quantization/polar_quant.py — Iterative hierarchical polar quantizationrfsn_v10/quantization/hybrid_polar_cartesian.py — Hybrid polar-cartesian quantizerrfsn_v10/quantization/qjl_score_correction.py — QJL sketch-based score correctionrfsn_v10/quantization/isoquant_precondition.py — IsoQuant quaternion preconditionerValidated at the beta level on Apple Silicon. Quality thresholds (cosine ≥ 0.998 vs.
FP32 reference) measured with tests/test_short_prompt_decode_drift.py.
| Config | K bits | V bits | Group size | Status |
|---|---|---|---|---|
k8_v5_gs32 | 8 | 5 | 32 | Default — slightly better cosine |
k8_v5_gs64 | 8 | 5 | 64 | Validated |
k8_v4_gs64 | 8 | 4 | 64 | Validated |
RFSN_TELEMETRY_HMAC_KEY is required when events with sensitive fields are written.
An absent key raises RuntimeError rather than silently hashing with an empty key.(table, event) tuples — routing metadata is never mixed into event payloads.# CPU-safe tests (no MLX required, runs on Linux CI)
RFSN_BACKEND=numpy pytest \
tests/test_config.py \
tests/test_config_strict.py \
tests/test_kernels_validation.py \
tests/test_quantization_lazy_imports.py \
tests/test_experimental_flags.py \
tests/test_clickhouse_security.py \
tests/test_no_runtime_raw_sdpa.py -q
# MLX-dependent tests (Apple Silicon required)
pytest \
tests/test_attention.py \
tests/test_bitpack.py \
tests/test_bitpack_fuzz.py \
tests/test_kv_manager.py \
tests/test_drift.py \
tests/test_attention_causal_mask.py \
tests/test_short_prompt_decode_drift.py \
tests/test_prefill_decode_split.py -q
# Security tests
pytest tests/test_clickhouse_security.py \
tests/test_clickhouse_routing.py \
tests/test_tool_runner_security.py -q
# Full CI (mirrors .github/workflows/ci.yml linux-cpu job)
RFSN_BACKEND=numpy pytest --collect-only -q
# Fast run (attention + bitpack)
python benchmarks/run_all.py --fast
# Full benchmark suite
python benchmarks/run_all.py
# Regression check against production_baseline.json
python benchmarks/run_all.py --check
Results are written to benchmarks/results/run_all_<timestamp>.json and benchmarks/results/latest.json.
docker build -t rfsn-qjl .
# Health check (Dockerfile default CMD now runs healthcheck)
docker run --rm rfsn-qjl
# With ClickHouse (production compose)
cp .env.example .env # fill CLICKHOUSE_PASSWORD
docker-compose up
# With ClickHouse ports exposed (local dev only)
docker-compose -f docker-compose.yml -f docker-compose.dev.yml up
Requires Python 3.11. Core runtime does not function in Docker without Apple Silicon + MLX.
| Variable | Default | Description |
|---|---|---|
RFSN_BACKEND | (auto) | Force backend: mlx or numpy |
RFSN_LOG_LEVEL | INFO | Logging level |
RFSN_CACHE_DIR | ~/.cache/rfsn | KV cache directory |
RFSN_TELEMETRY_HMAC_KEY | (required if telemetry enabled) | HMAC key for prompt sanitisation |
RFSN_CLICKHOUSE_HOST | localhost | ClickHouse hostname |
RFSN_CLICKHOUSE_PORT | 8123 | ClickHouse HTTP port |
RFSN_CLICKHOUSE_SECURE | true | Use HTTPS |
RFSN_CLICKHOUSE_TOKEN | (empty) | Bearer token for RFSN-Auth header |
RFSN_EXPERIMENTAL_QJL | false | Enable QJL score correction |
RFSN_EXPERIMENTAL_POLAR | false | Enable polar/hybrid quantization |
RFSN_EXPERIMENTAL_ADAPTIVE | false | Enable adaptive sparse controller |
10 commits
Python
98.7%
Metal
1.0%
Research alpha. MLX-RFSN is a KV-cache compression research and benchmarking system for Apple Silicon. It is not production-ready, not a general AI server, and not a full vector database.
Official promoted candidate: NONE
Best practical baseline: rfsn_direct_packed_k8v8
Promotion allowed: false (pending runtime-instrumented cache trace and token-sequence-hash provenance)
Release ID: alpha-8.4 (from release.toml)
Package Version: 10.2.0a84 (from release.toml)
Artifact Schema: 3.0 (from release.toml)
To verify the current state locally:
python -m compileall -q rfsn_v10 rfsn_v11 tests benchmarks scripts
bash scripts/release_gate.sh
| Path | Status |
|---|---|
| Stable runtime (MLX 8-bit KV compression) | Alpha — validated on Apple Silicon, quality gates passing |
| Package installation (subpackages) | Fixed — rfsn_v10.kernels, rfsn_v10.runtime install correctly |
CLI health check (python -m rfsn_v10 healthcheck) | Working |
| Sparse decode | Disabled by default — not end-to-end proven |
| QJL score correction | Experimental — disabled by default, requires explicit opt-in |
| Polar / hybrid quantization | Experimental — disabled by default, requires explicit opt-in |
| Adaptive sparse controller | Experimental — disabled by default, requires explicit opt-in |
| CUDA backend | Not implemented |
| Full portable runtime | Not implemented — MLX required for core runtime |
| End-to-end speedup | Not proven — decode TPS comparable, compression overhead makes total slower at short contexts |
| Research server | FastAPI server — /v1/chat/completions with SSE streaming, authentication, and proper wheel packaging (mlx/numpy). Not production-hardened. |
| Docker | Healthcheck validation + ClickHouse telemetry (CPU-only, no inference) |
| >8-bit compression | Uses raw uint32 fallback — bit-packing is real for 2-8 bit only |
| Metal kernel | Scaffold/stub with CPU fallback — actual Metal GPU computation NOT yet implemented |
| Experimental throughput | No experimental throughput speedup is proven |
| Platform | Status |
|---|---|
| Apple Silicon + MLX | Supported (primary runtime) |
| NumPy CPU | Supported — kernel validation, config, security tests pass; MLX-dependent runtime tests skip |
| Linux / CI | CPU tests pass; MLX suites skip cleanly |
| CUDA | Not implemented |
| macOS x86 (Intel) | MLX not supported on Intel Macs — NumPy-only |
# Three install modes — do not install everything at once
# Basic: core only, no MLX, no memory (any platform)
pip install -e ".[basic]"
# Fusion: MLX + KV compression benchmarks (macOS Apple Silicon)
pip install -e ".[fusion]"
# Memory: external vector memory (any platform)
pip install -e ".[memory]"
# Dev: everything for development
pip install -e ".[fusion,memory,dev]"
# Verify install
python -m rfsn_v10 version
# Health check
python -m rfsn_v10 healthcheck
# Validate config
python -m rfsn_v10 validate-config --config configs/default_runtime.yaml
# CPU-safe tests (no MLX required)
pytest tests/test_config.py tests/test_config_strict.py \
tests/test_kernels_validation.py \
tests/test_quantization_lazy_imports.py \
tests/test_experimental_flags.py \
tests/test_clickhouse_security.py \
tests/test_no_runtime_raw_sdpa.py -q
# MLX-dependent tests (Apple Silicon required)
pytest tests/test_attention.py tests/test_bitpack.py \
tests/test_bitpack_fuzz.py tests/test_kv_manager.py \
tests/test_drift.py tests/test_attention_causal_mask.py \
tests/test_short_prompt_decode_drift.py \
tests/test_prefill_decode_split.py -q
# Central benchmark command
python benchmarks/kv_shootout.py --promotion-report
Run the OpenAI-compatible FastAPI server locally:
export RFSN_MODEL_ID=mlx-community/Llama-3-8B-Instruct-4bit
python -m rfsn_v10.server
# Or with uvicorn directly:
# uvicorn rfsn_v10.server.app:app --host 0.0.0.0 --port 8000
Test the endpoint:
curl http://localhost:8000/health
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Hello"}],"stream":true}'
Or with Docker Compose (container validation + telemetry):
export CLICKHOUSE_PASSWORD=your-password
docker compose up -d
# Runs healthcheck and exits — inference server requires mlx/torch backend natively
rfsn_v10/bitpack.py — Bit-packed quantizer (2–8 bit widths) with exact roundtrip guaranteesrfsn_v10/kv_manager.py — TurboQuant KV manager with grouped symmetric quantization and WHT preconditioningrfsn_v10/attention.py — Adaptive block-sparse attention; all dense fallbacks route through attention_reference.pyrfsn_v10/attention_reference.py — Canonical causal attention reference (always applies causal mask for T_q > 1)rfsn_v10/runtime/engine.py — Orchestrator integrating KV cache, sparse attention, audit mode, and telemetryrfsn_v10/runtime/__init__.py — Re-exports RFSNRuntime from engine.pyrfsn_v10/runtime/generation.py — RFSNGenerator with prefill, decode, sampling, and telemetryrfsn_v10/config.py — Strict Pydantic config (extra='forbid' on all models)rfsn_v10/health.py — Health check system; returns UNHEALTHY until checks have been runrfsn_v10/clickhouse_client.py — HMAC-SHA256 prompt sanitization, recursive sanitizer, retry queuerfsn_v10/model_loader.py — Unified model/tokenizer loading (mlx-lm / transformers)rfsn_v10/server/app.py — FastAPI OpenAI-compatible server with SSE streamingAll experimental paths require explicit opt-in. The runtime will not activate them
silently — attempting to use an experimental feature without enabling it raises a
RuntimeError.
# config.yaml
experimental:
enable_qjl: false # QJL score correction
enable_polar: false # Polar / hybrid quantization
enable_adaptive: false # Adaptive sparse controller
Or via environment:
RFSN_EXPERIMENTAL_QJL=true
RFSN_EXPERIMENTAL_POLAR=true
RFSN_EXPERIMENTAL_ADAPTIVE=true
Warning: Experimental features are not validated for production or quality-critical generation.
rfsn_v10/quantization/polar_quant.py — Iterative hierarchical polar quantizationrfsn_v10/quantization/hybrid_polar_cartesian.py — Hybrid polar-cartesian quantizerrfsn_v10/quantization/qjl_score_correction.py — QJL sketch-based score correctionrfsn_v10/quantization/isoquant_precondition.py — IsoQuant quaternion preconditionerValidated at the beta level on Apple Silicon. Quality thresholds (cosine ≥ 0.998 vs.
FP32 reference) measured with tests/test_short_prompt_decode_drift.py.
| Config | K bits | V bits | Group size | Status |
|---|---|---|---|---|
k8_v5_gs32 | 8 | 5 | 32 | Default — slightly better cosine |
k8_v5_gs64 | 8 | 5 | 64 | Validated |
k8_v4_gs64 | 8 | 4 | 64 | Validated |
RFSN_TELEMETRY_HMAC_KEY is required when events with sensitive fields are written.
An absent key raises RuntimeError rather than silently hashing with an empty key.(table, event) tuples — routing metadata is never mixed into event payloads.# CPU-safe tests (no MLX required, runs on Linux CI)
RFSN_BACKEND=numpy pytest \
tests/test_config.py \
tests/test_config_strict.py \
tests/test_kernels_validation.py \
tests/test_quantization_lazy_imports.py \
tests/test_experimental_flags.py \
tests/test_clickhouse_security.py \
tests/test_no_runtime_raw_sdpa.py -q
# MLX-dependent tests (Apple Silicon required)
pytest \
tests/test_attention.py \
tests/test_bitpack.py \
tests/test_bitpack_fuzz.py \
tests/test_kv_manager.py \
tests/test_drift.py \
tests/test_attention_causal_mask.py \
tests/test_short_prompt_decode_drift.py \
tests/test_prefill_decode_split.py -q
# Security tests
pytest tests/test_clickhouse_security.py \
tests/test_clickhouse_routing.py \
tests/test_tool_runner_security.py -q
# Full CI (mirrors .github/workflows/ci.yml linux-cpu job)
RFSN_BACKEND=numpy pytest --collect-only -q
# Fast run (attention + bitpack)
python benchmarks/run_all.py --fast
# Full benchmark suite
python benchmarks/run_all.py
# Regression check against production_baseline.json
python benchmarks/run_all.py --check
Results are written to benchmarks/results/run_all_<timestamp>.json and benchmarks/results/latest.json.
docker build -t rfsn-qjl .
# Health check (Dockerfile default CMD now runs healthcheck)
docker run --rm rfsn-qjl
# With ClickHouse (production compose)
cp .env.example .env # fill CLICKHOUSE_PASSWORD
docker-compose up
# With ClickHouse ports exposed (local dev only)
docker-compose -f docker-compose.yml -f docker-compose.dev.yml up
Requires Python 3.11. Core runtime does not function in Docker without Apple Silicon + MLX.
| Variable | Default | Description |
|---|---|---|
RFSN_BACKEND | (auto) | Force backend: mlx or numpy |
RFSN_LOG_LEVEL | INFO | Logging level |
RFSN_CACHE_DIR | ~/.cache/rfsn | KV cache directory |
RFSN_TELEMETRY_HMAC_KEY | (required if telemetry enabled) | HMAC key for prompt sanitisation |
RFSN_CLICKHOUSE_HOST | localhost | ClickHouse hostname |
RFSN_CLICKHOUSE_PORT | 8123 | ClickHouse HTTP port |
RFSN_CLICKHOUSE_SECURE | true | Use HTTPS |
RFSN_CLICKHOUSE_TOKEN | (empty) | Bearer token for RFSN-Auth header |
RFSN_EXPERIMENTAL_QJL | false | Enable QJL score correction |
RFSN_EXPERIMENTAL_POLAR | false | Enable polar/hybrid quantization |
RFSN_EXPERIMENTAL_ADAPTIVE | false | Enable adaptive sparse controller |
10 commits
Python
98.7%
Metal
1.0%