T-Crypt/speculum

Mirror-glass monitor for a local LLM runtime (llama-swap, NInfer, Strata) on one RTX 4090

Python

1

27 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

Inference Dashboard (r/LocalLLM)

Nothing fancy but more of what works for me with how fast projects are moving, flexible enough that if another project like strata spins up then I'm not missing inference metrics. I'm aware there are 100's of these on github, throw a webfetch and you'll find 10. Project is live:…

1

Oct 5, 2026

README

Speculum

A lightweight observability dashboard for a local LLM runtime.

Speculum runs on the box doing the work. It watches a local LLM stack — one GPU serving local models through llama-swap, plus engine backends such as NInfer or Strata (or Ollama, vLLM, LM Studio, Unsloth Studio, any OpenAI-compatible or custom endpoint) — and renders everything on one page, refreshed at 1 Hz:

Speculum — KPI strip, throughput, context residency, GPU and host

  • No build step, no dependencies. The front end is plain HTML, CSS and ES modules; the collector is one Python 3 process using only the standard library. Nothing to pip install, nothing to compile.
  • No infrastructure. The only storage is an optional local SQLite file.
  • Local by design. Binds 127.0.0.1 only. Off-box viewing goes through tailscale serve (tailnet-only) — never to the LAN or the public internet.
  • Honest degradation. Every source is optional. A down engine or backend shows as an honest "no data" — never a fabricated number, and never enough to take the page down.

Quick start

All you need is Python 3.9 or newer (3.11+ only if you use a TOML config file) — the collector is stdlib-only, so there is nothing to install.

Live — the collector serves the static files and the feed:

python3 collector/speculum.py            # → http://127.0.0.1:8792/

For 24/7, run it as a systemd --user service:

cp collector/speculum.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now speculum
journalctl --user -u speculum -f

The unit binds 127.0.0.1:8792 only, runs with CUDA_VISIBLE_DEVICES= (the collector never touches the CUDA context — it reads nvidia-smi as one long-lived subprocess) and a MemoryMax=64M ceiling. To view it off-box, publish tailnet-only with tailscale serve.

The unit assumes the checkout is at ~/speculum — if yours lives elsewhere, point ExecStart at it.

Configuration is optional: copy speculum.example.toml to speculum.toml (repo root or ~/.config/speculum/) and edit. With no config file, Speculum probes the default ports on localhost and shows what it finds. The config sets the bind address/port, enables or disables discovery, sets the history retention (30/60/90 days, or 0 to turn history off), and adds engine adapters (ollama, vllm, lmstudio, unsloth, openai, custom) on any host.

Demo — ?demo runs a seeded simulation with no backend at all; serve the directory any way you like (e.g. python3 -m http.server 8792) and open http://localhost:8792/?demo. The simulator fills exactly the same state shape as live data, so the page is identical in behaviour.

How it works

nvidia-smi · /proc · llama-swap :9090 · NInfer backend · Strata :8080 · other engines
        │  (each optional; every poll degrades to "no data")
        ▼
collector/speculum.py   one Python 3 process, daemon poll threads,
                        one shared state under one lock
        │
        ├─ static files from the repo root
        ├─ GET /api/snapshot   one-shot full JSON
        └─ GET /api/stream     SSE: 1 Hz tick (hot values + new events/requests)
        ▼
app.js   EventSource with snapshot-poll fallback; paint loop; all panels

The collector keeps bounded in-memory buffers (60 min of GPU/host/KPI samples at 1 Hz, the last 500 requests, the last 200 events), and collector/history.py persists the long view into SQLite: per-request rows, minute/hour rollups, model load spans and events, with daily pruning at the configured retention. Long-range views (throughput, timeline, leaderboard beyond 1 h) are answered from that database, so the page works across collector restarts.

Screenshots

Captured live on the monitor box (dark theme; the settings menu also has a light theme).

Token ledger and the storage panel, with GPU and host readouts

Model timeline and request history

The whole deck zoomed out: leaderboard, engines, events, KV pool

The advanced view with every panel in play

Request history: each inference with its cached/fresh split, TTFT and decode rate

Engine cards, the event stream and the KV pool

What it watches

SourceFeeds
nvidia-smi (one long-lived -lms 1000 subprocess, read line by line)GPU name/driver, temperature, utilisation, power, VRAM
/proc/stat /proc/meminfo /proc/loadavg /proc/uptimehost CPU, RAM, load, uptime, top processes
llama-swap :9090 (/running, /v1/models, /api/metrics/activity, /api/events)models in use, per-request records (cached vs fresh prompt, outputs, TTFT), event stream
NInfer backend (port from llama-swap /running, e.g. :5803) — /metrics, /slots, log streamdecode/prefill tok/s, KV slots, per-request TTFT/queue/prefill/decode/MTP, engine latch alerts
Strata :8080 (/v1/models, /metrics, /slots)second engine + context slots
any engine from speculum.tomlsame, via the pluggable adapters

Panels

The top bar shows the overall state (Live / Degraded / Offline), GPU, driver, host uptime, feed type, the view toggle and the settings menu; an alert strip appears only while an alert is active (e.g. an engine latched red after a run of "service unavailable"). The deck then holds twelve panels:

PanelShows
KPIten live KPIs in four groups — Throughput (decode rate, requests/min), Latency (first token, time/token, p95), Cache (cache hit, re-ingest tax, MTP acceptance), Capacity (VRAM, queue depth) — each with a 60-min sparkline and a delta vs 15 min ago
Throughputrolling decode tok/s, one trace per engine, ranges 15 min → 30 d (beyond 1 h from the history DB)
GPU & hosthorizontal meters for temp, utilisation, power draw and VRAM (warn/crit ticks from the real power limit and VRAM total), plus device readouts (SM/mem clock, fan, PCIe) and host (CPU, RAM, load, top processes)
Contextone bar per live session: prompt tokens vs the model window, with 75 % / 90 % ticks
Model timelinewhich model was loaded on which engine over time (6 h → 30 d; spans survive collector restarts)
Token ledgergenerated / fresh / cached token volume over the last hour, last 24 h and since engine start, with per-column mix bars
Storagehistory database size, retention, what longer retention would cost; CSV export and prune-now
Requestsrecent inferences: model, prompt vs window (cached vs fresh), window, TTFT, decode rate
Leaderboardper-model requests, output tokens, median decode, cache hit, MTP and tokens per Wh (6 h → 30 d)
Enginesone card per engine: state, origin, window, backend, decode rate, queue, MTP, history sparkline; dimmed with a reason when stopped
Eventsruntime event stream (request completions, engine lifecycle, warnings, alerts) with severity filters and pause-autoscroll
KV poolone cell per live KV slot, fill = prompt tokens vs window

Views. The top bar toggles Basic (KPI, throughput, GPU, timeline, engines) and Advanced (everything); the choice is bookmarkable via ?view=basic|advanced. Settings (top-bar menu) persist in the browser: theme (dark default / light), reduce motion (follows the OS preference, or forced on/off), and the optional ripple effect (off by default).

Keys: P pause · R reseed (demo) / resync (live).

HTTP API

EndpointPurpose
GET /api/snapshotone-shot full JSON (state + history)
GET /api/streamSSE: 1 Hz tick plus push events
GET /api/history?range=…rollup rows (minute/hour)
GET /api/requests?range=…raw request rows
GET /api/spans?range=…model load spans
GET /api/leaderboard?range=…per-model statistics
GET /api/storagedatabase size, retention, projections
GET /api/export?range=…&format=csvCSV export
POST /api/pruneapply retention now

Files

FileRole
index.htmlpage shell: top bar, twelve panel shells, footer, one module script
tokens.css · components.css · layout.cssdesign tokens (dark + light themes, WCAG AA) · UI primitives · layout — no framework, no build
app.jsfeed (SSE + poll fallback), ?demo simulator, paint loop, all panels
ui.js · viewmodel.js · ripple.jsDOM primitives and formatters · pure view-model adapters · optional ripple effect
collector/speculum.pystdlib-only collector and HTTP server (Python 3, one process)
collector/history.pySQLite history: writer thread, rollups, spans, storage/export API
collector/engines.py · collector/config.pyengine adapters + backoff scheduler · speculum.toml loader
collector/speculum.servicesystemd --user unit
speculum.example.tomlexample configuration
smoke/Node DOM-stub smoke tests (no browser needed)
LICENSEMIT license

Development

python3 -m py_compile collector/speculum.py    # collector parses
node smoke/ui-primitives.mjs                   # primitives + ripple state machine
node smoke/shell.mjs                           # boots the real app.js shell over real payload shapes
python3 -m http.server 8792                    # → http://localhost:8792/?demo (simulator)

The UI files are ES modules, so node --check does not apply to them; the smoke scripts import and execute them, which is the parse-and-run check. Demo values are deliberately invented — that is the simulator's job.

Origin

Speculum started on 2026-10-03 as a prototype generated in a single agent turn on a local model — one RTX 4090 box serving local models through llama-swap and NInfer — and has been iterated on that same box since.

License

MIT — see LICENSE.

hacktoberfest
lightweight
llama-cpp
llama-swap
llm
llm-tools
lmstudio
monitoring
ninfer
runtime
strata
unsloth

T-Crypt/speculum

Mirror-glass monitor for a local LLM runtime (llama-swap, NInfer, Strata) on one RTX 4090

Python

1

27 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

Inference Dashboard (r/LocalLLM)

Nothing fancy but more of what works for me with how fast projects are moving, flexible enough that if another project like strata spins up then I'm not missing inference metrics. I'm aware there are 100's of these on github, throw a webfetch and you'll find 10. Project is live:…

1

Oct 5, 2026

README

Speculum

A lightweight observability dashboard for a local LLM runtime.

Speculum runs on the box doing the work. It watches a local LLM stack — one GPU serving local models through llama-swap, plus engine backends such as NInfer or Strata (or Ollama, vLLM, LM Studio, Unsloth Studio, any OpenAI-compatible or custom endpoint) — and renders everything on one page, refreshed at 1 Hz:

Speculum — KPI strip, throughput, context residency, GPU and host

  • No build step, no dependencies. The front end is plain HTML, CSS and ES modules; the collector is one Python 3 process using only the standard library. Nothing to pip install, nothing to compile.
  • No infrastructure. The only storage is an optional local SQLite file.
  • Local by design. Binds 127.0.0.1 only. Off-box viewing goes through tailscale serve (tailnet-only) — never to the LAN or the public internet.
  • Honest degradation. Every source is optional. A down engine or backend shows as an honest "no data" — never a fabricated number, and never enough to take the page down.

Quick start

All you need is Python 3.9 or newer (3.11+ only if you use a TOML config file) — the collector is stdlib-only, so there is nothing to install.

Live — the collector serves the static files and the feed:

python3 collector/speculum.py            # → http://127.0.0.1:8792/

For 24/7, run it as a systemd --user service:

cp collector/speculum.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now speculum
journalctl --user -u speculum -f

The unit binds 127.0.0.1:8792 only, runs with CUDA_VISIBLE_DEVICES= (the collector never touches the CUDA context — it reads nvidia-smi as one long-lived subprocess) and a MemoryMax=64M ceiling. To view it off-box, publish tailnet-only with tailscale serve.

The unit assumes the checkout is at ~/speculum — if yours lives elsewhere, point ExecStart at it.

Configuration is optional: copy speculum.example.toml to speculum.toml (repo root or ~/.config/speculum/) and edit. With no config file, Speculum probes the default ports on localhost and shows what it finds. The config sets the bind address/port, enables or disables discovery, sets the history retention (30/60/90 days, or 0 to turn history off), and adds engine adapters (ollama, vllm, lmstudio, unsloth, openai, custom) on any host.

Demo — ?demo runs a seeded simulation with no backend at all; serve the directory any way you like (e.g. python3 -m http.server 8792) and open http://localhost:8792/?demo. The simulator fills exactly the same state shape as live data, so the page is identical in behaviour.

How it works

nvidia-smi · /proc · llama-swap :9090 · NInfer backend · Strata :8080 · other engines
        │  (each optional; every poll degrades to "no data")
        ▼
collector/speculum.py   one Python 3 process, daemon poll threads,
                        one shared state under one lock
        │
        ├─ static files from the repo root
        ├─ GET /api/snapshot   one-shot full JSON
        └─ GET /api/stream     SSE: 1 Hz tick (hot values + new events/requests)
        ▼
app.js   EventSource with snapshot-poll fallback; paint loop; all panels

The collector keeps bounded in-memory buffers (60 min of GPU/host/KPI samples at 1 Hz, the last 500 requests, the last 200 events), and collector/history.py persists the long view into SQLite: per-request rows, minute/hour rollups, model load spans and events, with daily pruning at the configured retention. Long-range views (throughput, timeline, leaderboard beyond 1 h) are answered from that database, so the page works across collector restarts.

Screenshots

Captured live on the monitor box (dark theme; the settings menu also has a light theme).

Token ledger and the storage panel, with GPU and host readouts

Model timeline and request history

The whole deck zoomed out: leaderboard, engines, events, KV pool

The advanced view with every panel in play

Request history: each inference with its cached/fresh split, TTFT and decode rate

Engine cards, the event stream and the KV pool

What it watches

SourceFeeds
nvidia-smi (one long-lived -lms 1000 subprocess, read line by line)GPU name/driver, temperature, utilisation, power, VRAM
/proc/stat /proc/meminfo /proc/loadavg /proc/uptimehost CPU, RAM, load, uptime, top processes
llama-swap :9090 (/running, /v1/models, /api/metrics/activity, /api/events)models in use, per-request records (cached vs fresh prompt, outputs, TTFT), event stream
NInfer backend (port from llama-swap /running, e.g. :5803) — /metrics, /slots, log streamdecode/prefill tok/s, KV slots, per-request TTFT/queue/prefill/decode/MTP, engine latch alerts
Strata :8080 (/v1/models, /metrics, /slots)second engine + context slots
any engine from speculum.tomlsame, via the pluggable adapters

Panels

The top bar shows the overall state (Live / Degraded / Offline), GPU, driver, host uptime, feed type, the view toggle and the settings menu; an alert strip appears only while an alert is active (e.g. an engine latched red after a run of "service unavailable"). The deck then holds twelve panels:

PanelShows
KPIten live KPIs in four groups — Throughput (decode rate, requests/min), Latency (first token, time/token, p95), Cache (cache hit, re-ingest tax, MTP acceptance), Capacity (VRAM, queue depth) — each with a 60-min sparkline and a delta vs 15 min ago
Throughputrolling decode tok/s, one trace per engine, ranges 15 min → 30 d (beyond 1 h from the history DB)
GPU & hosthorizontal meters for temp, utilisation, power draw and VRAM (warn/crit ticks from the real power limit and VRAM total), plus device readouts (SM/mem clock, fan, PCIe) and host (CPU, RAM, load, top processes)
Contextone bar per live session: prompt tokens vs the model window, with 75 % / 90 % ticks
Model timelinewhich model was loaded on which engine over time (6 h → 30 d; spans survive collector restarts)
Token ledgergenerated / fresh / cached token volume over the last hour, last 24 h and since engine start, with per-column mix bars
Storagehistory database size, retention, what longer retention would cost; CSV export and prune-now
Requestsrecent inferences: model, prompt vs window (cached vs fresh), window, TTFT, decode rate
Leaderboardper-model requests, output tokens, median decode, cache hit, MTP and tokens per Wh (6 h → 30 d)
Enginesone card per engine: state, origin, window, backend, decode rate, queue, MTP, history sparkline; dimmed with a reason when stopped
Eventsruntime event stream (request completions, engine lifecycle, warnings, alerts) with severity filters and pause-autoscroll
KV poolone cell per live KV slot, fill = prompt tokens vs window

Views. The top bar toggles Basic (KPI, throughput, GPU, timeline, engines) and Advanced (everything); the choice is bookmarkable via ?view=basic|advanced. Settings (top-bar menu) persist in the browser: theme (dark default / light), reduce motion (follows the OS preference, or forced on/off), and the optional ripple effect (off by default).

Keys: P pause · R reseed (demo) / resync (live).

HTTP API

EndpointPurpose
GET /api/snapshotone-shot full JSON (state + history)
GET /api/streamSSE: 1 Hz tick plus push events
GET /api/history?range=…rollup rows (minute/hour)
GET /api/requests?range=…raw request rows
GET /api/spans?range=…model load spans
GET /api/leaderboard?range=…per-model statistics
GET /api/storagedatabase size, retention, projections
GET /api/export?range=…&format=csvCSV export
POST /api/pruneapply retention now

Files

FileRole
index.htmlpage shell: top bar, twelve panel shells, footer, one module script
tokens.css · components.css · layout.cssdesign tokens (dark + light themes, WCAG AA) · UI primitives · layout — no framework, no build
app.jsfeed (SSE + poll fallback), ?demo simulator, paint loop, all panels
ui.js · viewmodel.js · ripple.jsDOM primitives and formatters · pure view-model adapters · optional ripple effect
collector/speculum.pystdlib-only collector and HTTP server (Python 3, one process)
collector/history.pySQLite history: writer thread, rollups, spans, storage/export API
collector/engines.py · collector/config.pyengine adapters + backoff scheduler · speculum.toml loader
collector/speculum.servicesystemd --user unit
speculum.example.tomlexample configuration
smoke/Node DOM-stub smoke tests (no browser needed)
LICENSEMIT license

Development

python3 -m py_compile collector/speculum.py    # collector parses
node smoke/ui-primitives.mjs                   # primitives + ripple state machine
node smoke/shell.mjs                           # boots the real app.js shell over real payload shapes
python3 -m http.server 8792                    # → http://localhost:8792/?demo (simulator)

The UI files are ES modules, so node --check does not apply to them; the smoke scripts import and execute them, which is the parse-and-run check. Demo values are deliberately invented — that is the simulator's job.

Origin

Speculum started on 2026-10-03 as a prototype generated in a single agent turn on a local model — one RTX 4090 box serving local models through llama-swap and NInfer — and has been iterated on that same box since.

License

MIT — see LICENSE.

hacktoberfest
lightweight
llama-cpp
llama-swap
llm
llm-tools
lmstudio
monitoring
ninfer
runtime
strata
unsloth