Unified TUI and CLI to index and search your local coding agent session history across 11+ providers (Codex, Claude, Gemini, Cursor, Aider, etc.)
Rust
1,158
5,529 commits
updated Oct 2, 2026
Unified, high-performance TUI to index and search your local coding agent history. Aggregates sessions from Codex, Claude Code, Gemini CLI, Cline, OpenCode, Amp, Cursor, ChatGPT, Aider, Pi-Agent, Prime Agent, Oh My Pi, GitHub Copilot Chat, Copilot CLI, OpenClaw, Clawdbot, Vibe, Crush, Goose, Hermes, Kimi Code, Muse Code, Qwen Code, Factory (Droid), Antigravity, OpenHands, Grok Build, Grok Bot, Codebuff/Freebuff, Devin CLI, Shelley, and Kiro CLI into a single, searchable timeline.
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/coding_agent_session_search/main/install.sh?$(date +%s)" \
| bash -s -- --easy-mode --verify
# Windows (PowerShell)
& ([scriptblock]::Create((irm "https://raw.githubusercontent.com/Dicklesworthstone/coding_agent_session_search/main/install.ps1"))) -EasyMode -Verify
Installs the latest release by default. Pass --version <tag> / -Version <tag> to pin a specific version.
Or via package managers:
# Homebrew (Apple Silicon macOS + Linux)
brew install dicklesworthstone/tap/cass
# Windows (Scoop)
scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cass
The Homebrew tap installs prebuilt release tarballs (not bottles) for Linux and Apple Silicon macOS. On Intel macOS, use the install script with --from-source.
⚠️ Never run bare cass in an agent context — it launches the interactive TUI. Always use --robot or --json.
# 1) Check the installed interface once per version (recipe verified on 0.8.0).
cass --version
cass search --help
# Verify a newly installed executable without opening the configured archive.
cass selftest --json
# `health --binary-only` still reports (and therefore probes) archive readiness.
# 2) For a quick history question, start with scoped read-only lexical retrieval.
# Hybrid remains the product default; lexical is explicit for this workflow.
cass search "performance regression" --workspace /path/to/project --days 7 \
--mode lexical --no-maintenance --robot --robot-meta --fields minimal \
--limit 5 --max-tokens 2000 --timeout 2000
# 3) Find the current or recent session for this workspace
cass sessions --current --json
cass sessions --workspace "$(pwd)" --json --limit 5
# 4) View + expand a hit (use source_path/line_number from search output)
cass view /path/to/session.jsonl -n 42 -C 3 --json --timeout 2000
# 5) Discover the full machine API
cass capabilities --json
cass robot-docs guide
cass robot-docs schemas
# 6) Exclude a noisy agent harness from future indexing
cass sources agents list --json
cass sources agents exclude openclaw
cass sources agents include openclaw
The retrieval flags above are available in 0.8.0. On older builds, check help;
if --no-maintenance is absent, report the mismatch instead of dropping the
read-only constraint. --timeout is in milliseconds, while --max-tokens limits
approximate output size. Also set a caller-side deadline (for example, GNU
timeout 10s); an externally interrupted command may leave incomplete JSON.
Inspect budget.timed_out even after exit 0: timed-out empty hits are not proof
that no history exists. A maintenance-required response ends the retrieval
attempt; indexing or repair is a separate mutating task. Use triage/health/status
for readiness diagnosis, not as repeated prerequisites to a short summary.
Broaden scope deliberately, expand useful hits, and preserve source/line citations.
view -C bounds context lines, not bytes; check excerpt size before including
a long JSONL record in an agent prompt.
Output conventions
Search asset contract
--robot --robot-meta) reports the requested mode, realized mode, semantic refinement status, and any lexical fallback reason when semantic assets are not ready.cass models install downloads the default all-minilm-l6-v2 (alias minilm, ~90 MB) only on explicit request; --model multilingual-minilm selects the larger multilingual MiniLM L12 model (~480 MB) for CJK/mixed-language archives. Cass never auto-downloads or auto-selects the multilingual space. Air-gapped installs use --from-file <dir>. While the selected model is absent, hybrid search uses lexical-only and reports fallback_mode="lexical" in health/status.cass triage --json combines readiness, next_command, recommended_commands[], docs/schema pointers, starter workflows, and accepted recoveries for diagnosis. Review recommended mutations before executing them. cass health --json and cass status --json remain the narrower truth surfaces for readiness, active rebuilds, and recovery.Lexical publish durability (atomic-swap)
src/indexer/mod.rs::publish_staged_lexical_index.<data_dir>/index/.lexical-publish-backups/<dated>/ for a bounded retention window. Default cap is 1 (keep just the most-recent prior generation for one-step rollback); override via the CASS_LEXICAL_PUBLISH_BACKUP_RETENTION env var (0 disables retention entirely, higher N keeps deeper history). Pruning runs after every successful publish and emits structured tracing::info! events with freed_bytes + retention_limit for observability.recover_or_finalize_interrupted_lexical_publish_backup at the start of the next lexical publish or rebuild (not at process startup), which moves any orphaned canonical sidecar (.<name>.publish-in-progress.bak) into .lexical-publish-backups/ before the next publish lands.Quarantine, GC, and the doctor/diag surface
cass diag --json --quarantine enumerates every quarantined artifact (failed seed bundles, retained publish backups, quarantined lexical generations) with size_bytes, age_seconds, safe_to_gc, and a human-readable gc_reason. The safe_to_gc flag is advisory — it reflects retention policy + cleanup dry-run eligibility and is not wired to any automatic deletion path.cass doctor --json surfaces the same quarantine summary plus checks[] status for every diagnostic the tool runs. Without --fix, doctor is read-only (auto_fix_applied=false, auto_fix_actions=[], issues_fixed=0); with --fix it applies only the repairs whose dry-run plans are proven safe (currently: Track A analytics rebuild, Track B rollup rebuild via rebuild_token_daily_stats when the token_usage ledger is intact).cass doctor --fix never have a generation reclaimed silently — every quarantine stays on disk until an explicit derived-asset rebuild (cass models backfill or an index refresh recommended by cass health --json) supersedes it.cass index runs escalates from a warning to a non-zero exit (#434): the counter persists in <data_dir>/index/.fts-repair-failure-streak.json, watch daemons log the escalation instead of exiting, and any run whose repair succeeds — or fails differently — resets it. Canonical rows and the Tantivy index are unaffected; run cass doctor --rebuild-canonical-fts --yes --json for the explicit repair.Schema stability guarantees
tests/golden/robot/: capabilities, selftest, health, status, diag, models status/verify/check-update, introspect, doctor, api-version, stats, search, export-html, onboarding, the quarantine and dedup commands and analytics incidents, plus sessions and pack on their missing-database and error paths only. swarm status scenarios are pinned under tests/golden/swarm_status/. triage, swarm work-packet and swarm lint have no golden files; their shape is covered only by assertion tests. A change to any field name, type, or nullability fails the golden test suite and requires a deliberate regeneration pass (UPDATE_GOLDENS=1 rch exec -- env CARGO_TARGET_DIR=/data/tmp/cass-golden-target cargo test --test golden_robot_json --test golden_robot_docs).cass introspect --json's response_schemas block enumerates every schema in a stable alphabetical order (BTreeMap-backed — see bead coding_agent_session_search-8sl73).{error: {code, kind, message, hint, retryable}}) have a fixed shape. kind values are kebab-case; branch on err.kind, not on the numeric code, for codes ≥ 10 (see the Error Handling section below).If your runtime does not expose built-in mcp-agent-mail tools (for example, list_mcp_resources is empty), you can still coordinate via direct MCP HTTP calls.
~/.local/pipx/venvs/mcp-agent-mail/bin/python -m mcp_agent_mail.cli serve-http --host 127.0.0.1 --port 8765
/mcp)curl -sS -X POST http://127.0.0.1:8765/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":"health","method":"tools/call","params":{"name":"health_check","arguments":{}}}'
# Ensure project
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"ensure","method":"tools/call","params":{"name":"ensure_project","arguments":{"human_key":"/data/projects/coding_agent_session_search"}}}'
# Register agent
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"register","method":"tools/call","params":{"name":"register_agent","arguments":{"project_key":"/data/projects/coding_agent_session_search","program":"codex","model":"gpt-5","name":"YourAgentName"}}}'
# Send message
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"send","method":"tools/call","params":{"name":"send_message","arguments":{"project_key":"/data/projects/coding_agent_session_search","sender_name":"YourAgentName","to":["PeerAgent"],"subject":"[coord] hello","thread_id":"coord-2026-02-13","ack_required":true,"body_md":"Online and starting work."}}}'
# Fetch inbox
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"inbox","method":"tools/call","params":{"name":"fetch_inbox","arguments":{"project_key":"/data/projects/coding_agent_session_search","agent_name":"YourAgentName","limit":50,"include_bodies":true}}}'
# Acknowledge message id 42
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"ack","method":"tools/call","params":{"name":"call_extended_tool","arguments":{"tool_name":"acknowledge_message","arguments":{"project_key":"/data/projects/coding_agent_session_search","agent_name":"YourAgentName","message_id":42}}}}'
mcp_agent_mail defaults to sqlite+aiosqlite:///./storage.sqlite3. That means the server working directory determines which mailbox database you are using. To avoid "project not found" confusion, start the server from the same directory your team expects for mailbox state.
Three-pane layout with semantic styling: filter bar with pills, results list with color-coded agents and score tiers, and syntax-highlighted detail preview with tab navigation
Full conversation rendering with markdown formatting, code blocks, headers, and structured content
Built-in help screen (press F1 or ?) with all shortcuts, filters, modes, and navigation tips
AI coding agents are transforming how we write software. Claude Code, Codex, Cursor, Copilot, Aider, Pi-Agent; each creates a trail of conversations, debugging sessions, and problem-solving attempts. But this wealth of knowledge is scattered and unsearchable:
cass treats your coding agent history as a unified knowledge base. It:
snake_case ("my_var" matches "my" and "var"), hyphenated terms, and code symbols (c++, foo.bar) correctly.reader.reload() ensures new messages appear in the search bar immediately without restarting.cass search --robot currently spends roughly a second in archive open and integrity preflight on a ~10 GB archive; --robot-meta reports that separately as _meta.timing.other_ms, while search_ms stays in the tens of milliseconds.Local inference: Uses frankensearch's pure-Rust native MiniLM implementation with local safetensors weights. Once MiniLM is installed, no network traffic is required to answer queries.
Warm-daemon reuse: Semantic and hybrid CLI searches automatically use an
already-running local embedding daemon (including a socket selected with
CASS_DAEMON_SOCKET) and only initialize the installed in-process model if
daemon inference fails. Pass --daemon to permit auto-spawning a missing
daemon in human-mode searches (robot/JSON searches never spawn one, even
with --daemon, because their bounded budget cannot wait for a daemon to
start; start cass daemon yourself first; _meta.effective.daemon shows
the request and what applied), or --no-daemon to force
direct inference. --fast-only stays in
the deterministic hash-vector space. Each data directory gets a distinct
default socket and owner-private pinned key; fresh handshake, health,
embedding, batch, and rerank challenges authenticate the exact response and
immutable Frankensearch embedding identity before any daemon output is used.
--two-tier progressive refinement (fast results refined in place by the
quality tier) is experimental and currently inactive: the one-shot CLI
collapses it to a single-tier quality search and the TUI's progressive lanes
are disabled at HEAD, so hybrid search today is lexical plus one MiniLM
refinement pass when the model is installed.
Opt-in acquisition: cass models install downloads all-minilm-l6-v2 from Hugging Face on explicit request and verifies SHA256 checksums. cass models install --model multilingual-minilm explicitly selects paraphrase-multilingual-MiniLM-L12-v2 for CJK and mixed-language retrieval. Nothing is fetched until an install command runs, and merely installing the multilingual model never changes the active space.
Air-gapped install: cass models install --model <minilm|multilingual-minilm> --from-file <dir> accepts a pre-downloaded model directory so you can bring the assets in yourself.
Switching spaces: both models output 384 values, but their identities and vectors are incompatible. Set CASS_SEMANTIC_EMBEDDER=multilingual-minilm, then run cass models backfill --tier quality --embedder multilingual-minilm; cass keeps lexical fail-open active until the complete new generation is atomically published.
Required files (all must be present after install; cass models verify --model <minilm|multilingual-minilm> checks the selected model):
model.safetensorstokenizer.jsonconfig.jsonspecial_tokens_map.jsontokenizer_config.jsonVector index: Stored as vector_index/index-<embedder>.fsvi in the data directory.
Lexical fail-open: While the model is absent, cass returns lexical-only results and reports fallback_mode="lexical" in health/status; search never blocks on semantic assets.
The deterministic hash embedder is available only when explicitly selected, such as with --fast-only, --embedder hash, or CASS_SEMANTIC_EMBEDDER=hash. It is a separate lexical-feature vector space, not a silent substitute for missing MiniLM vectors:
| Feature | ML Model (MiniLM) | Hash Embedder (FNV-1a) |
|---|---|---|
| Meaning Understanding | ✅ "car" ≈ "automobile" | ❌ Exact tokens only |
| Initialization Time | ~500ms (model loading) | <1ms (instant) |
| Network Dependency | None (after install) | None |
| Disk Footprint | ~90MB model files | 0 bytes |
| Deterministic | ✅ Same input = same output | ✅ Same input = same output |
Algorithm:
When to Use:
Override: Set CASS_SEMANTIC_EMBEDDER=hash to force hash mode even when ML model is available.
cass uses the frankensearch FSVI vector index format (.fsvi) for storing semantic embeddings.
Features:
f32 and f16 storage for smaller on-disk size--approximate is passed and the HNSW sidecar file exists. hnsw_ready in status --json means only that the sidecar file is present, not that ANN is in useIndex Location: ~/.local/share/coding-agent-search/vector_index/index-<embedder>.fsvi
cass supports three search modes, selectable via --mode flag or Alt+S in the TUI:
| Mode | Algorithm | Best For |
|---|---|---|
| Lexical | BM25 full-text | Exact term matching, code searches |
| Semantic | Vector similarity | Conceptual queries, "find similar" |
| Hybrid (default) | Lexical + single-tier semantic refinement fused with RRF; lexical fail-open | Balanced precision and recall |
Lexical Search: Uses Quill's BM25 implementation with prefix matching. Best when you know the exact terms you're looking for. The lexical index is derived from SQLite; if it is missing, stale, or incompatible, cass reports the state and rebuilds through the normal indexing path from the canonical database.
Semantic Search: Computes vector similarity between query and indexed MiniLM embeddings. Finds conceptually related content even without exact term overlap. Explicit semantic mode requires the MiniLM model and a compatible MiniLM vector index; it never substitutes same-dimensional hash vectors.
Hybrid Search: The default. It combines lexical and semantic results using Reciprocal Rank Fusion (RRF) when semantic assets are ready, and it fails open to lexical when semantic enrichment is still catching up or disabled:
RRF_score = Σ 1 / (K + rank_i)
Where K=60 (tuning constant) and rank_i is the position in each result list. This balances the precision of lexical search with the recall of semantic search. Semantic refinement is a single pass over the installed MiniLM index; progressive two-tier refinement (--two-tier) is experimental and currently inactive.
# CLI examples
cass search "authentication" --mode lexical --robot
cass search "how to handle user login" --mode semantic --robot
cass search "auth error handling" --mode hybrid --robot
foo* - Prefix match (finds "foobar", "foo123")*foo - Suffix match (finds "barfoo", "configfoo")*foo* - Substring match (finds "afoob", "configuration")*term* wildcards to broaden matches. A visual indicator shows when the fallback is active.Up/Down arrows.F12) that prioritizes exact matches over wildcard/fuzzy results.**bold**, in human-readable and robot/JSON output alike; --highlight also marks the query terms' other occurrences, never marking a term twice.Powered by FrankenTUI (ftui) — a high-performance Elm-architecture TUI framework with adaptive frame budgets, Bayesian diff selection, and spring-based animations.
Indexing 150/2000 (7%)—plus active filters.Ctrl+Enter, then open all in your editor with Ctrl+O. Confirmation prompt for large batches (≥12 items)./ to search within the detail pane; matches highlighted with n/N navigation.recent/balanced/relevance/quality with F12; quality mode penalizes fuzzy matches.Alt+A; Esc returns to search.cass tui --inline to keep terminal scrollback intact. The UI anchors to a region of the terminal while logs scroll normally. Configure with --ui-height <rows> and --anchor top|bottom.cass tui --record-macro session.macro for reproducible bug reports and workflow automation. Events are saved as human-readable JSONL with full timing data.cass tui --asciicast demo.cast.
Export conversations as styled, portable HTML files with optional encryption:
cdn.jsdelivr.net, pinned with SRI hashes.onerror="...no-prism" — code blocks remain readable offline in plain monospace, and the page layout never depends on a network resource.TUI Usage: Press Ctrl+E in the detail view to open the export modal, or Ctrl+Shift+E to export Markdown immediately with defaults. On the detail pane's Export tab, e/h open the HTML export modal and m runs the Markdown export.
CLI Usage:
# Basic export
cass export-html /path/to/session.jsonl
# With encryption
printf '%s\n' "secret" | cass export-html /path/to/session.jsonl --encrypt --password-stdin
# Custom output location
cass export-html session.jsonl --output-dir ~/exports --filename "my-session"
# Open in browser after export
cass export-html session.jsonl --open
# Robot mode (JSON output)
cass export-html session.jsonl --json
Ingests history from 32 local agent connectors, normalizing them into a unified Conversation -> Message -> Snippet model. cass capabilities --json | jq .connectors is the canonical machine-readable inventory (kept in lockstep with the runtime registry):
~/.codex/sessions (Rollout JSONL)~/.gemini/tmp (Chat JSON)~/.claude/projects (Session JSONL), plus macOS Desktop metadata sidecars under
~/Library/Application Support/Claude/claude-code-sessions and
~/Library/Application Support/Claude/local-agent-mode-sessions~/.clawdbot/sessions (Session JSONL)~/.vibe/logs/session/*/messages.jsonl (Session JSONL).opencode directories (SQLite)~/.local/share/amp & VS Code storage~/Library/Application Support/Cursor/User/ global + workspace storage (SQLite state.vscdb)~/Library/Application Support/com.openai.chat (v1 unencrypted JSON; v2/v3 encrypted—see Environment)~/.aider.chat.history.md and per-project .aider.chat.history.md files (Markdown)~/.pi/agent/sessions (Session JSONL with thinking content)prime_agent): ~/.prime/agent/sessions/<session-id>.jsonl (versions 1–3). Indexes the active branch with omission counts for abandoned siblings; preserves thinking, tool results and context summaries. Overrides, in precedence order: PRIME_AGENT_SESSION_DIR, legacy PRIME_AGENT_CODING_AGENT_SESSION_DIR, then PRIME_AGENT_CODING_AGENT_DIR (with /sessions appended). Prime retains its own agent identity.omp): OMP v18's default ~/.omp/agent/sessions, named profiles under ~/.omp/profiles/<name>/agent/sessions, XDG stores under $XDG_DATA_HOME/omp, and explicit OMP-only archive roots via CASS_OMP_DATA_ROOT (pi-family JSONL, including per-session sub-agent transcripts)github.copilot-chat (JSON)~/.copilot/session-state, legacy ~/.copilot/history-session-state, and gh copilot config paths (JSONL/JSON)~/.openclaw/agents/*/sessions (Session JSONL)~/.local/share/goose/sessions/sessions.db (SQLite, v1.20+), plus the earlier per-session *.jsonl layout under ~/.goose/sessions~/.crush/crush.db and per-project .crush/crush.db (SQLite)~/.hermes/state.db and project-local .hermes/state.db (SQLite)~/.local/share/devin/cli/sessions.db (SQLite; override with CASS_DEVIN_DATA_ROOT). Indexes visible local sessions along their active parent chain, preserving tool messages and excluding abandoned branches and inline image payloads. Cloud-only sessions are outside this connector's scope.CASS_SHELLEY_DB=/absolute/path/to/shelley.db, or add that file to the paths of a type = "local" source in sources.toml. Any filename is accepted after schema validation. Defaults include ~/.config/shelley/shelley.db and shelley.db in the current directory. Live indexing watches the database and its WAL/SHM sidecars; metadata changes refresh existing sessions. CASS_SKIP_SUBAGENTS=1 excludes conversations with a Shelley parent ID. The database can also contain credentials and application settings, so raw mirroring and remote database ingestion are disabled; keep the database on its original machine.CASS_GROK_BOT_DATA_ROOT to its persistence directory. Native message IDs preserve already indexed history as older messages leave the application's window; repeat scans do not duplicate retained messages. CASS reads only chat content. Raw mirroring and automatic fleet copying are disabled because the replica also holds secret and approval fields. This connector is separate from the Grok CLI connector and does not fetch cloud history.$KIMI_CODE_HOME/sessions/*/*/agents/*/wire.jsonl (default ~/.kimi-code; sub-agents index as <sessionId>:<agentId>), plus the legacy ~/.kimi/sessions/*/*/wire.jsonl layout (Session JSONL)~/.local/share/muse/sessions/<YYYY>/<MM>/<DD>/<session-id>/session.jsonl, including nested subagent/*/session.jsonl transcripts (override with CASS_MUSE_DATA_ROOT)~/.qwen/tmp/*/chats/session-*.json (Chat JSON)~/.factory/sessions (JSONL files organized by workspace slug)~/.gemini/antigravity/ and the CLI's ~/.gemini/antigravity-cli/ — each holding brain/<uuid>/.system_generated/logs/transcript.jsonl (clean JSONL transcript) with the durable per-conversation conversations/<uuid>.db (SQLite) mirrored alongside. IDE conversations are keyed ide/<uuid> so the two stores never collide; CASS_ANTIGRAVITY_DATA_ROOT replaces both with one explicit base. Resume with cass resume <transcript> --agent agy (agy --conversation <uuid>).~/.openhands/conversations/<id>/ — base_state.json metadata plus an events/event-NNNNN-<uuid>.json event stream (JSON)grok): ~/.grok/sessions/<percent-encoded-cwd>/<session-uuid>/ — updates.jsonl (authoritative ACP session-update stream) with summary.json metadata and chat_history.jsonl fallback (override the base dir with GROK_HOME). Resume with grok --resume <session-id>.codebuff): ~/.config/manicode/projects/<project>/chats/<chat-id>/chat-messages.json with its run-state.json (override with CASS_CODEBUFF_DATA_ROOT). Both products write the same Manicode store and no chat records which binary wrote it, so their sessions share one lineage identity, codebuff (filter with --agent codebuff). Messages are reconciled by their native IDs, so an edited message updates in place instead of duplicating.kiro): ~/.kiro/sessions/cli/<session-uuid>.jsonl (append-only event log: prompts, assistant messages, tool results) with the matching <session-uuid>.json snapshot read for session ID, working directory, title, timestamps and model.Claude Code Desktop sidecars preserve title, workspace, model, and session IDs, but not necessarily the full conversation body. If Claude Code has culled an old CLI JSONL body, cass can still index searchable sidecar metadata while reporting that the conversation body is unavailable.
Pi-Agent parses JSONL session files with rich event structure:
~/.pi/agent/sessions/ (override the agent home with PI_CODING_AGENT_DIR, or the sessions directory directly with PI_SESSIONS_DIR)session_start, message, model_change, thinking_level_change*_*.jsonl pattern in sessions directoryOh My Pi (omp) uses the same pi-family wire format but remains a separate
agent identity throughout search, analytics, resume, TUI, and HTML export:
~/.omp/agent/sessions/ and ~/.omp/profiles/<name>/agent/sessions/; OMP_PROFILE selects a profile and takes precedence over legacy PI_PROFILE$XDG_DATA_HOME/omp/sessions/ and $XDG_DATA_HOME/omp/profiles/<name>/sessions/ when the OMP XDG root existsPI_CODING_AGENT_SESSION_DIR names the exact OMP sessions directory. CASS_OMP_DATA_ROOT declares an OMP-only archive/store root and is the right choice for copied, mounted, or custom OMP data. PI_CODING_AGENT_DIR is shared by both pi-family programs, so CASS conservatively keeps otherwise-ambiguous paths under that root owned by Pi-Agent; use one of the OMP-specific variables when OMP identity matters. PI_CONFIG_DIR changes the home-relative .omp config directory name.omp [--profile <name>] --resume <id>; copied profiles, XDG archives, remote mirrors, and explicit roots also carry --session-dir <dir> so a canonical-looking archive cannot reopen a different live storepi_agent to omp using the same conservative canonical/XDG/remote-mirror ownership policy as live discovery, then the derived lexical index and analytics are rebuilt so a transcript cannot remain attributed to both agents. The conventional ~/.local/share/omp shape is durable path evidence; an arbitrary historical custom $XDG_DATA_HOME/omp path is reclassified only while that root is currently configured and resolvable. Without provider-qualified evidence, ambiguous historical paths fail closed as Pi-Agent rather than letting a generic .../omp/sessions directory steal ownership.OpenCode reads SQLite databases from workspace directories:
.opencode/ directories (scans recursively from home).opencode containing database filesSearch across agent sessions from multiple machines—your laptop, desktop, and remote servers—all from a single unified index. cass uses SSH/rsync to efficiently sync session data, tracking provenance so you know where each conversation originated.
The easiest way to configure multi-machine search is the interactive setup wizard:
cass sources setup
What the wizard does:
~/.ssh/configsources.toml with correct paths and mappingscass sources sync right after configuration (skipped with --skip-sync or --dry-run; --json setup defers it and reports the command to run)Wizard options:
| Flag | Purpose |
|---|---|
--hosts <names> | Configure only specific hosts (comma-separated) |
--dry-run | Preview changes without applying them |
--non-interactive | Use auto-detected defaults for scripting |
--skip-install | Don't install cass on remotes |
--skip-index | Don't run indexing on remotes |
--skip-sync | Skip the final cass sources sync. Interactive setup runs that sync after the hosts are configured and records it as complete only once it has actually finished; --json setup always defers it and reports sync.status = "pending" with the command to run |
--resume | Resume an interrupted setup |
--json | Output progress as JSON (for automation) |
Examples:
# Full interactive wizard
cass sources setup
# Configure specific hosts only
cass sources setup --hosts laptop,workstation,build-server
# Preview without making changes
cass sources setup --dry-run
# Resume interrupted setup
cass sources setup --resume
# Non-interactive for CI/CD
cass sources setup --non-interactive --hosts myserver --skip-install
Resumable state: If setup is interrupted (Ctrl+C, connection lost), state is saved to the cache directory (~/.cache/cass/setup_state.json on Linux). Resume with --resume.
Tailscale discovery is optional: cass sources discover --tailscale --json adds
online tailnet peers to SSH-config discovery, and cass sources setup --tailscale
offers them in setup. It reads local tailscale status --json with a five-second
deadline; a missing CLI, stopped daemon, or login failure produces a warning and
leaves SSH-config discovery available. Explicit setup --hosts skips discovery.
Connections use ordinary SSH over assigned Tailscale IPv4 addresses, so MagicDNS
is not required. Matching SSH aliases retain their user/key configuration;
otherwise SSH uses its normal defaults. IPv6-only peers are currently omitted.
Tailscale ACLs, SSH authorization and host-key checks still apply; discovery does
not log in, install Tailscale, or change either SSH or tailnet configuration.
The local fixture and Docker tests do not prove that your machines can sync and
search each other's sessions. The opt-in live harness uses actual SSH connections
and cass sources discover, sources add, sources sync, and search. It creates isolated synthetic
Codex sessions on each machine, checks source provenance and filters, repeats a
sync to detect duplicates, and appends messages. It checks both lexical and default
hybrid search, requires one JSON response per sync, holds the real indexing lock to
test busy refusal, and recovers transferred sessions through sources reingest.
A refused SSH connection must leave the other sources searchable.
Keep the inventory and SSH configuration outside this repository. For example, create a mode-0600 JSON file containing:
{
"ssh_config": "/private/path/to/ssh_config",
"hosts": [{"ssh": "workstation"}, {"ssh": "laptop"}]
}
Then run with an explicit binary:
python3 scripts/e2e/live_fleet_search.py \
--inventory /private/path/to/fleet.json \
--cass-bin /path/to/cass
Python 3 and authenticated SSH access are required on the remote machines.
The Unix runner needs Python 3.9+, rsync, and a CASS binary supporting the tested
commands. Each inventory alias must appear in the supplied SSH configuration;
included configuration files are supported. Host-key verification stays enabled.
To exercise actual tailnet discovery and transport, add --tailscale to the
harness command and use tailnet IPv4 addresses as the private inventory targets.
Keep any required SSH users, keys and trusted host-key aliases in the private SSH
configuration. For a discovery test independent of explicit aliases, use SSH
Match originalhost entries rather than literal Host entries for those addresses.
The harness retains fresh test directories and raw
receipts privately outside git; it never changes existing session archives or
deletes test data. Console results use ordinal labels. An unreachable machine
keeps the overall result failed, even if the other machines pass. Do not attach
raw receipts or inventories to public issues: they contain machine identities.
When the wizard installs cass on remote machines, it tries every viable method in this priority order, falling through to the next when one fails; setup fails only when all of them do, and the error lists each attempt:
| Priority | Method | Speed | Requirements |
|---|---|---|---|
| 1 | cargo-binstall | ~30s | cargo-binstall pre-installed, compatible release binary |
| 2 | Pre-built binary | ~10s | curl/wget, GitHub access, compatible release binary |
| 3 | cargo install | ~5min | Rust toolchain, 1GB disk, 2GB RAM |
| 4 | Full bootstrap | ~10min | curl, 1GB disk, 2GB RAM (installs rustup) |
crates.io publishing resumed at 0.7.0 (GH#416): the long-stale registry gap (0.6.13, published before the Quill/OMP era) is closed — the entire dependency chain now resolves from crates.io (
frankensearch 0.4.0, thefrankentorch-*family,frankenhnsw), socargo install coding-agent-searchbuilds the current line again. The installer and GitHub Release binaries remain the fastest paths.
Resource Requirements:
What Gets Installed:
cass binary (location depends on method: ~/.cargo/bin/cass for cargo-based, ~/.local/bin/cass for pre-built binary)Installation Progress: The wizard shows real-time progress for each stage:
Installing cass on laptop...
[1/4] Checking environment... ✓
[2/4] Downloading binary... ████████░░ 80%
[3/4] Verifying checksum... ✓
[4/4] Setting up PATH... ✓
Use --skip-install if you prefer to install manually on remotes.
The setup wizard automatically discovers SSH hosts from your configuration:
Discovery Sources:
~/.ssh/config (parses Host entries)*, ?) are automatically excludedProbe Results (for each discovered host):
| Check | Purpose |
|---|---|
| Connectivity | Can we establish SSH connection? |
| cass Version | Is cass already installed? What version? |
| Agent Data | Which agents have session data? |
| Session Count | How many conversations exist? |
| System Info | OS, architecture, disk space, memory |
Each setup run probes every selected host afresh; probe results are not cached between runs.
For manual configuration without the wizard:
# Add a remote machine using platform presets
cass sources add user@laptop.local --preset macos-defaults
# Or specify paths explicitly
cass sources add dev@workstation --path ~/.claude/projects --path ~/.codex/sessions
# Sync sessions from all configured sources
cass sources sync
# Check source health and connectivity
cass sources doctor
Remote source diagnostics are intentionally local-only. cass triage --json,
cass doctor --json, cass health --json, and cass status --json report the
remote_source_sync summary from cass-owned evidence: sources.toml,
sync_status.json, the local remotes/<source>/mirror/ copy, and archive DB
provenance rows. They do not open SSH sessions, mutate remote machines, or
rewrite provider session logs while classifying source gaps.
cass sources doctor is the explicit networked exception: it performs bounded,
read-only probes of configured source hosts. Its per-source human summary keeps
the same native reachability, binary-health, and mirror/sync state codes and
safe command as the JSON report. It intentionally does not claim local search
readiness, because a remote host probe cannot establish the controller's local
SQLite, lexical, or semantic asset state.
This matters because agent harnesses can prune their own logs. If a laptop is
retired, a remote path disappears, or a provider truncates older sessions, the
cass archive DB and cass-owned local mirror may be the only remaining evidence
for those conversations. Treat gap names such as remote_source_unavailable,
remote_source_pruned, local_archive_ahead_of_remote, and
remote_copy_ahead_verified as preservation signals first: keep the archive and
mirror intact, then run the recommended cass sources sync --json (all configured remote sources; --source <name> narrows it) or
source-specific sync command after reviewing the reported evidence.
Raw-mirror retention is explicit and audited. Use cass mirror prune --older-than 90d --json or cass mirror prune --max-size 100GB --json to get a
dry-run plan; add --apply only after reviewing the scope and totals. Preview
entries contain at most 1,000 manifest/blob details; omitted_entry_count
reports additional candidates. Planned counts and bytes cover the entire plan,
including omitted details. Use provider/path selectors to inspect a narrower
scope. Previews do not append audit records. Add
--keep-tag <tag> to pin captures linked to tagged conversations. prune
holds down blobs referenced by captures from the last 7 days by default, writes
complete intent/result records to raw-mirror/v1/pruned.jsonl for non-empty
applied plans, and refuses apply mode
while an index/watch job is active.
Applied pruning syncs the audit independently of the optional capture setting
CASS_RAW_MIRROR_FSYNC. Each completed result is recorded before the next
removal; a later failure preserves those earlier results. An abrupt crash
between a removal and its result record can still leave an intent without a
confirmed result.
Use --provider opencode and/or --source-path '*/opencode.db' with an age
or size rule to target one source without retiring unrelated captures.
Repeated providers are alternatives; a source-path glob further narrows them.
A pruned capture of a source that is still on disk is copied again by the next
index run. On a machine whose providers never delete their session files, set
CASS_RAW_MIRROR=0 (or false, no, off) to stop capturing altogether.
Indexing and search are unchanged, and existing captures stay until you prune
them. cass doctor reports raw_mirror_capture_disabled and warns about what is
given up: a session file its provider later deletes survives only in the
archive DB.
With a selector, --max-size measures unique blobs in that selection. Shared
blobs still referenced outside it and orphan blobs without source provenance
remain protected. The JSON plan records the selectors and scope_blob_bytes.
Large mutable sources are stored as 4 MiB content-addressed chunks. Growing
JSONL files reuse every unchanged complete chunk, and SQLite sources reuse
unchanged 4 MiB byte regions, so each historical snapshot remains byte-exact without
writing another full-file blob. Existing whole-blob manifests remain readable;
cass doctor --json reports storage_kind, chunk_count, the full-source
digest, and verifies every referenced chunk before treating a snapshot as
recovery authority.
Sources are configured in the platform config directory (Linux: ~/.config/cass/sources.toml, macOS: ~/Library/Application Support/cass/sources.toml):
[[sources]]
name = "laptop"
type = "ssh"
host = "user@laptop.local"
paths = ["~/.claude/projects", "~/.codex/sessions"]
sync_schedule = "manual"
[[sources]]
name = "workstation"
type = "ssh"
host = "dev@work.example.com"
paths = ["~/.claude/projects"]
sync_schedule = "daily"
# Path mappings rewrite remote paths to local equivalents
[[sources.path_mappings]]
from = "/home/dev/projects"
to = "/Users/me/projects"
# Agent-specific mappings
[[sources.path_mappings]]
from = "/opt/work"
to = "/Volumes/Work"
agents = ["claude_code"]
Configuration Fields:
| Field | Description |
|---|---|
name | Friendly identifier (becomes source_id) |
type | Connection type: ssh or local |
host | SSH host (user@hostname) |
paths | Paths to sync (supports ~ expansion) |
sync_schedule | manual, hourly, or daily. Only the jobs installed by cass schedule install run it; without them it is a label and syncs happen when you run cass sources sync |
path_mappings | Rewrite remote paths to local equivalents |
# List configured sources
cass sources list [--verbose] [--json]
# Add a new source
cass sources add <user@host> [--name <name>] [--preset macos-defaults|linux-defaults] [--path <path>...] [--no-test]
# Remove a source
cass sources remove <name> [--purge] [-y]
# Check connectivity and config
cass sources doctor [--source <name>] [--json]
# Sync sessions
cass sources sync [--source <name>] [--no-index] [--verbose] [--dry-run] [--json]
If one harness is generating mostly junk or looped output, you can disable it persistently even if its files remain on disk:
# Inspect current include/exclude state
cass sources agents list --json
# Stop indexing this harness in future runs
cass sources agents exclude openclaw
# Re-enable it later
cass sources agents include openclaw
cass stores this preference in sources.toml (~/.config/cass/sources.toml on Linux, ~/Library/Application Support/cass/sources.toml on macOS), so future scans, syncs, and watch-mode updates remember it automatically.
By default, cass sources agents exclude <agent> also removes already archived local data for that agent and rebuilds the lexical index so the exclusion frees space instead of only blocking future imports.
If you want to block future indexing but keep the data already archived:
cass sources agents exclude openclaw --keep-indexed-data
The sync engine uses rsync over SSH for efficient delta transfers and falls back to other transports when rsync is unavailable:
Transfer Methods (auto-detected):
| Method | When Used | Characteristics |
|---|---|---|
| rsync | rsync available on both ends | Delta transfers, compression, progress stats |
| WSL rsync | Windows without native rsync, WSL with rsync installed | Runs wsl rsync |
| scp | rsync unavailable | Full file copies through the system scp, inheriting the OpenSSH agent, keys and ~/.ssh/config |
| SFTP | the fallbacks above unavailable | Full file transfers via the SSH native protocol |
Safety Guarantees:
--delete, so remote deletions never propagate locally.-a and without -u, so a local mirror file that differs from the remote is overwritten, even when the local copy is newer. The mirror is a copy of the remote, not a place to edit sessions.--partial keeps a partly transferred file under its final name so the next sync continues it. A failed sync can therefore leave a truncated file until the next sync completes.Transfer Configuration:
| Setting | Default | Purpose |
|---|---|---|
| Connection timeout | 10s | Fail fast on unreachable hosts |
| Transfer timeout | 300 s of I/O inactivity | rsync --timeout aborts a transfer that stalls this long; there is no wall-clock limit on a transfer that keeps moving |
| Compression | Enabled | Reduce bandwidth for text-heavy sessions |
| Partial transfers | Enabled | Resume interrupted syncs |
rsync Flags Used:
-avz --links --safe-links --stats --partial [--protect-args | --secluded-args] --timeout 300 \
-e "ssh [-F $CASS_SSH_CONFIG] -o BatchMode=yes -o ConnectTimeout=10 -o ServerAliveInterval=15 -o ServerAliveCountMax=3 -o StrictHostKeyChecking=yes"
Where -avz = archive mode + verbose + compression. --protect-args/--secluded-args is auto-detected per remote rsync version (omitted when the remote rejects it), and --timeout carries the transfer timeout in seconds. StrictHostKeyChecking=yes means a host whose key is not already in known_hosts fails with "Host key verification failed". Connect once with plain ssh <host> and accept the key, or add it with ssh-keyscan, before the first sync.
Data Flow:
Remote: ~/.claude/projects/
↓ (rsync over SSH)
Local: ~/.local/share/coding-agent-search/remotes/<source>/mirror/<path>_<hash>/
↓ (connector scan)
Index: agent_search.db + index/v9-quill/
Where <path> is a filesystem-safe version of the remote path (e.g. .claude_projects), and <hash> is an FNV-1a hash of the original path in hex, so foo/bar and foo_bar never collide.
Sessions from remotes are indexed alongside local sessions, with provenance tracking to identify origin.
When viewing sessions from remote machines, workspace paths may not exist locally. Path mappings rewrite these paths so file links work on your local machine:
# List current mappings
cass sources mappings list laptop
# Add a mapping
cass sources mappings add laptop --from /home/user/projects --to /Users/me/projects
# Test how a path would be rewritten
cass sources mappings test laptop /home/user/projects/myapp/src/main.rs
# Output: /Users/me/projects/myapp/src/main.rs
# Agent-specific mappings (only apply for certain agents)
cass sources mappings add laptop --from /opt/work --to /Volumes/Work --agents claude_code,codex
# Remove a mapping by index
cass sources mappings remove laptop 0
In the TUI, filter sessions by origin:
Remote sessions display with a source indicator (e.g., [laptop]) in the results list.
Each conversation tracks its origin:
source_id: Machine identifier (e.g., "laptop", "workstation")origin_kind: local or remoteorigin_host: the remote host label, absent for local sessionsworkspace_original: Original path on the remote machine (before path mapping)--fields provenance selects exactly source_id, origin_kind and origin_host.
These fields appear in JSON/robot output and enable filtering:
cass search "auth error" --source laptop --json
cass timeline --since 7d --source remote
cass stats --by-source
cass is purpose-built for consumption by AI coding agents—not just as an afterthought, but as a first-class design goal. When you're an AI agent working on a codebase, your own session history and those of other agents become an invaluable knowledge base: solutions to similar problems, context about design decisions, debugging approaches that worked, and institutional memory that would otherwise be lost.
Imagine you're Claude Code working on a React authentication bug. With cass, you can instantly search across:
This cross-pollination of knowledge across different AI agents is transformative. Each agent has different strengths, different context windows, and encounters different problems. cass unifies all this collective intelligence into a single, searchable index.
cass teaches agents how to use it—no external documentation required:
# First-stop capability contract for agents
cass triage --json
cass capabilities --json
# → {"version": "...", "workflows": [...], "mistake_recoveries": [...], "commands": [...], "exit_codes": [...], "env_vars": [...]}
# Full API schema with argument types, defaults, and response shapes
cass introspect --json
# Topic-based help optimized for LLM consumption
cass robot-docs commands # All commands and flags
cass robot-docs schemas # Response JSON schemas
cass robot-docs examples # Copy-paste invocations
cass robot-docs exit-codes # Error handling guide
cass robot-docs guide # Quick-start walkthrough
AI agents sometimes make syntax mistakes. cass aggressively normalizes input to maximize acceptance when intent is clear:
| What you type | What cass understands | Correction note |
|---|---|---|
cass -robot --limit=5 | cass --robot --limit=5 | Single-dash long flags normalized |
cass --Robot --LIMIT 5 | cass --robot --limit 5 | Case normalized |
cass search "auth" --max_results 5 | cass search "auth" --limit 5 | Snake-case long flag normalized before alias recovery |
cass find "auth" | cass search "auth" | find/query/q → search via alias table |
cass --robot-docs | cass robot-docs | Flag-as-subcommand detected |
cass commands --json | cass robot-docs commands | Robot-docs topic shorthand detected |
cass schemas --json | cass robot-docs schemas | Robot-docs topic shorthand detected |
cass ready --json | cass triage --json | One-shot triage alias |
cass preflight --json | cass triage --json | One-shot triage alias |
cass --json | cass triage --json | Top-level robot request defaults to safe preflight |
cass --robot | cass triage --json | Top-level robot request defaults to safe preflight |
cass --json search "auth" | cass search "auth" --json | Leading structured flag moved to the robot-capable subcommand |
cass --robot status | cass status --json | Leading robot flag canonicalized to JSON output |
cass answer "auth" --json | cass pack "auth" --json | Cited-handoff aliases normalized to answer pack |
cass why auth failed --json --max-evidence 3 | cass pack "auth failed" --json --max-evidence 3 | Question/RC prompt aliases normalized to answer pack |
cass auth failed --json --max-evidence 3 | cass pack "auth failed" --json --max-evidence 3 | Bare robot queries with pack-only flags become answer packs |
cass search auth failed --json --max-evidence 3 | cass pack "auth failed" --json --max-evidence 3 | Explicit robot search with pack-only flags becomes an answer pack |
cass html-export session.jsonl --json | cass export-html session.jsonl --json | Reversed HTML export aliases normalized to the archive exporter |
cass current --json | cass sessions --current --json | Current-session shorthand normalized to session discovery |
cass sessions current --json | cass sessions --current --json | Positional current accepted as the sessions current flag |
cass search --query "auth" --json | cass search "auth" --json | Named query option converted to required positional query |
cass search --q "auth" --json | cass search "auth" --json | Short/familiar query aliases converted to required positional query |
cass search auth error --json | cass search "auth error" --json | Adjacent unquoted query words folded into one search |
cass auth error --json | cass search "auth error" --json | Unquoted robot-mode query words folded into search |
cass search --agent codex --limit 5 auth error --json | cass search "auth error" --agent codex --limit 5 --json | Query moved before leading search filters |
cass view --path session.jsonl --line 42 --json | cass view session.jsonl --line 42 --json | Named path option converted to required positional path |
cass view session.jsonl --line-number 42 --json | cass view session.jsonl --line 42 --json | Legacy alias for --line; still reads raw file line 42 |
cass view session.jsonl line_number=42 --json | cass view session.jsonl --message-index 42 --json | A pasted search-hit field selects canonical message 42, not raw line 42 |
cass view source_path=session.jsonl source_id=local line_number=42 --json | cass view session.jsonl --source local --message-index 42 --json | Search hit field bundle accepted as a follow-up command (add conversation_id when the file holds several conversations) |
cass search "auth" --format json | cass search "auth" --robot-format json | Familiar format spelling converted to robot format |
cass search "auth" --output json | cass search "auth" --robot-format json | Familiar output spelling converted to robot format |
cass help search --json | cass robot-docs commands | Structured help intent routed to the machine-readable command reference |
cass --format json status | cass status --robot-format json | Leading format request moved to the target subcommand |
cass search "auth" --max-results 5 | cass search "auth" --limit 5 | Result-count alias converted to canonical limit |
cass search "auth" -n 5 | cass search "auth" --limit 5 | Familiar short count flag converted to canonical limit |
cass search "auth" --last 7 --before now | cass search "auth" --since -7d --until now | Familiar time-window aliases converted to canonical filters |
cass search "auth" last=7d before=now | cass search "auth" --since -7d --until now | Bare time-window assignments converted to canonical filters |
cass search "auth" --provider codex | cass search "auth" --agent codex | Provider/tool/connector aliases converted to canonical agent filter |
cass search "auth" provider=codex | cass search "auth" --agent codex | Bare provider assignment converted to canonical agent filter |
cass search auth provider codex limit 5 | cass search auth --agent codex --limit 5 | Bare filter key/value pairs after a query converted to canonical flags |
cass search --limt 5 | cass search --limit 5 | Flag typos within Levenshtein distance ≤2 corrected |
The CLI applies multiple normalization layers:
--limt → --limit), and a first word within distance 2 of a subcommand is corrected to it (e.g. serach → search). A word that already names a subcommand is never changed, so cass status --jsn runs status --json. forget and upgrade are reached only by exact spelling.--Robot, --LIMIT → --robot, --limit--max_results, --data_dir, and other known snake_case long flags become canonical kebab-case before alias recovery runs-robot → --robot (common LLM mistake)ready/preflight → triage; find/query/q/grep/lookup → search; session → sessions; answer/evidence/bundle/handoff/why/explain/rca/root-cause/rootcause/summarize/summarise → pack; html-export/html_export/exporthtml → export-html; ls/list/info/summary → stats; st/state → status; reindex/idx/rebuild → index; show/get/read → view; diagnose/debug/check → diag; caps/cap → capabilities; inspect/intro → introspect; docs/help-robot/robotdocs → robot-docscommands, schemas, examples, exit-codes, and quickstart become robot-docs <topic> instead of falling through to search; command topics such as doctor and sources use structured help (cass help doctor --json, cass sources --help --json). Bare cass guide is reserved for the guided-operations planner; use cass robot-docs guide for the robot-docs walkthrough.cass --json, cass --robot, or cass --robot-format json with no subcommand runs read-only triage--json/--robot before a robot-capable subcommand is moved onto that subcommand--query/--q/--text/--pattern for search/pack and --path/--source-path/--file/--session for drill-down/export commands become the required positional argumentsearch/pack become one query positional--format json|jsonl|compact|sessions|toon, --output json|jsonl|compact|sessions|toon, and --output-format ... are accepted as --robot-format ... on robot-capable commands; export --format ... and export --output <file> keep their export meaningshelp --json, help commands --json, and search --help --json route to robot-docs guide / robot-docs commands; plain --help stays native clap help--max-results, --num-results, --results, --count, --top-k, and -n become --limit on commands with result limits--last 7, --before now, last=7d, and before=now become canonical --since/--until filters--provider, --tool, --connector, and matching assignments become canonical --agent filters on search-like commandsprovider codex, limit 5, and last 7d become canonical filter flags before the remaining words are folded into the querysearch with pack-only flags such as --max-evidence, --max-sessions, or --freshness-policy becomes pack, not implicit or explicit search--line-number, --line_number and line=42 become --line (a raw file line)line_number=42 pasted from a search hit becomes --message-index 42 (the canonical message ordinal the hit names), and a source_path=... source_id=... line_number=... bundle becomes the canonical path, --source and --message-index form for follow-up view/expand commandssearch query unless they look like a subcommand typocurrent, current-session, and sessions current become sessions --currentWhen corrections are applied, cass emits a teaching note to stderr so agents learn the canonical syntax. In robot/JSON mode the same information is emitted as one note: auto-corrected: <note> line per correction on stderr (at most two: the normalization note and the typo-recovery note), so stdout stays data-only. Robot-mode notes are printed only when the command succeeds; a failing command's stderr is its single JSON error envelope. The same notes appear in search output under _meta.effective.auto_corrections with --robot-meta.
Every command supports machine-readable output:
# Pretty-printed JSON (default robot mode)
cass search "error" --robot
# Streaming JSONL: one hit per line. Add --robot-meta to prepend a
# {budget, _meta} header line (elapsed_ms, next_cursor, state, index_freshness).
# The header also appears without --robot-meta when the search timed out
# (budget.timed_out), returned did-you-mean suggestions, --aggregate or --explain.
cass search "error" --robot-format jsonl # hits only
cass search "error" --robot-format jsonl --robot-meta # 1 _meta header + hits
# Compact single-line JSON (minimal bytes)
cass search "error" --robot-format compact
# Include performance metadata
cass search "error" --robot --robot-meta
# → { "hits": [...], "_meta": { "elapsed_ms": 12, "cache_hit": true, "wildcard_fallback": false, "lexical_degrade_reason": null, ... } }
# lexical_degrade_reason is "query_fuel_exhausted" when a hybrid search dropped its
# lexical leg because Quill's query fuel ran out (see CASS_QUILL_QUERY_FUEL_BUDGET)
# What the search actually ran (--robot-meta): check this instead of trusting the flags
cass search "error" --robot --robot-meta --days 7 | jq '._meta.effective'
# → { "command": "search", "query": "error",
# "query_structure": "error", // how the engine groups operands: `a OR b c` -> "a OR (b AND c)"
# "query_recoveries": [], // e.g. "1 unclosed '(' closed at the end of the query"
# "db_path": "/home/you/.local/share/coding-agent-search/agent_search.db",
# "db_path_source": "default", // --db | env:CASS_DB_PATH | --data-dir | env:CASS_DATA_DIR | env:XDG_DATA_HOME | default
# "time_window": { "since_ms": 1758067200000, "since_from": "--days 7", "until_ms": null, "until_from": null },
# "filters": { "agents": [], "workspaces": [], "source": "all", "sessions_from_paths": null },
# "auto_corrections": [] } // each argv correction, worded like its stderr note
# `cass pack "error" --json` carries the same object for the search it ran in
# its own `_meta.effective` ("command": "pack", and no search-only "daemon"),
# with home-directory paths, private hosts and secrets redacted like the rest
# of the pack (the db_path above reads "[REDACTED_PATH]/agent_search.db").
# Per-hit trust verdict (advisory; --robot-meta only)
cass search "error" --robot --robot-meta
# Each hit then carries a metadata-only `trust` block:
# "trust": {
# "schema_version": 1,
# "trust_tier": "unverified", // trusted | likely | unverified | stale | failed
# "confidence": "medium", // low | medium | high
# "provenance_refs": [], // e.g. ["commit:ab0d12ef90ab", "bead:xyz", "release:v0.6.15"]
# "stale_reason": "aged_out", // present only when not fully trusted
# "recommended_followup": "..." // advisory next step (never a destructive command)
# }
How agents should branch on trust_tier (relevance is not correctness — a
hit can be a landed fix or a failed attempt):
trust_tier | Meaning | What to do |
|---|---|---|
trusted | Landed, proof-backed, release/bead-contained | Safe to reuse |
likely | Has provenance (commit/closed bead) but not proof-pinned | Confirm via the cited ref first |
unverified | Relevant but no provenance link, or lexical-only corroboration | Corroborate before reuse |
stale | Aged out (aged_out) or superseded (superseded_by_newer) | Prefer a newer result |
failed | A failed/reverted attempt (failed_attempt) | Do not reuse |
The verdict is advisory metadata only — it never changes result ordering.
It is derived from metadata-only signals (recency, source health, realized
search mode, cwd-relative workspace match, and — opportunistically — linked
commit/bead/release provenance); it carries no raw session text. The same
trust block is attached to cass pack evidence. Branch on trust_tier and
stale_reason, not on confidence alone.
Provenance correlation is project-scoped and explicit-reference anchored:
for a hit from the project you are running cass in now, cass links it to a
closed bead or commit only when the hit's own indexed text references a known
identifier (bead:<id> or commit:<sha>), joined against that project's local
beads and git history. A linked commit's containing release is resolved from
Git. Release containment preserves provenance but does not establish proof of
the excerpt's claim: a landed commit remains proof_debt and cannot become
trusted from this correlation alone. A temporal or
workspace coincidence is never enough, so an unrelated conversation never
inherits another's trust. Off-project hits report workspace_mismatch, and a
hit whose local source file no longer exists on disk reports source_unhealthy
(archive-only) instead of overtrusting a dead path.
# Deterministic answer pack for handoff prompts
cass pack "why did checkout fail" --robot --max-tokens 12000 --limit 40
# Freshness-sensitive pack: fail if selected evidence is outside the window
cass pack "checkout timeout after redirect" --robot \
--freshness-policy strict --freshness-window-seconds 604800 \
--max-tokens 12000 --require-evidence
# Token-budgeted pack for pasting into another agent
cass pack "checkout timeout after redirect" --robot \
--max-tokens 4000 --max-evidence 8 --max-sessions 3 --max-excerpt-chars 600
# Pipeline from broad search to a bounded cited handoff
cass search "checkout timeout" --robot-format sessions \
| cass pack "checkout timeout root cause" --robot --sessions-from -
Design principle: stdout contains only parseable JSON data; all diagnostics, warnings, and progress go to stderr.
Use search when you are still exploring candidate sessions. Use pack when
you need a compact, cited, extractive artifact to hand to another agent or a
human operator. Use status/health before trusting freshness-sensitive output,
and use doctor only for diagnostics or safe repair workflows. Use
export-html when you need a full browsable session archive; packs are
token-budgeted evidence bundles, not full exports and not external
summarization.
Pack robot output includes health, freshness, privacy, and warnings.
Warnings such as privacy_redactions_applied, semantic_fallback_lexical,
or no_evidence_found are data, not prose; branch on the JSON fields before
copying the pack into another tool. Stale selected evidence is structural:
inspect freshness.stale_evidence_count.
Packs exclude injected skill payloads by default. Add --include-skill-content
to include them explicitly; credential redaction still applies.
privacy.skill_content_included reports whether the selected evidence includes
skill payloads, including after token-budget trimming.
Use the swarm surfaces when multiple agents are sharing one repo and you need a single read-only view before claiming work:
# Current shared-work snapshot; does not claim, reopen, release, or run builds
cass swarm status --json
# Advisory packet for one bead; still create real reservations and Beads updates yourself
cass swarm work-packet --json --bead coding_agent_session_search-example
# Coordination hygiene check before closeout or takeover review
cass swarm lint --json --bead coding_agent_session_search-example
# Read-only sibling dependency drift sentinel
cass swarm dependency-drift --json
swarm status and swarm work-packet collect bounded read-only Git state and
Beads exports when run from the repository root without a fixture. Git uses
porcelain-v2 with optional locks disabled. Beads uses br 0.6.x --no-db, so its
JSONL snapshot is explicitly partial: unexported database changes may exist.
Recheck Beads and reservations before claiming work. Child commands share a
single 15-second request budget and each has an 8 MiB output cap; Beads categories
cap at 512 issues.
Failures report unavailable providers and unknown summary counts, not zero work.
RCH contributes aggregate active/queued job counts, fleet slots and posture from
rch status --json (API 1.0, schema 1.0.0). Responses older than 60 seconds or
more than 5 seconds in the future are unavailable. Worker addresses, commands
and job details are omitted. This provider remains partial: local Cargo/CPU
state and build admission are unknown, even when RCH reports no active jobs.
Agent Mail roster and reservation reads are opt-in: set CASS_SWARM_AGENT_MAIL_URL
to the server's HTTP MCP endpoint and, if required, CASS_SWARM_AGENT_MAIL_TOKEN.
The reader uses only resources/read, never a local database fallback or inbox
read. Mail shares the total request budget with a 3-second cap of its own;
responses are capped at 8 MiB, rosters at 512 agents, and full 250-row reservation
pages are refused. The total reservation count remains unknown; source metadata
reports only the observed active count. Activity and expiry use
the observation time. Task descriptions, reservation reasons and message bodies
are omitted. These observations do not authorize claims or establish proof.
CASS evidence remains unwired. swarm lint still uses the placeholder
live snapshot. Fixture selection (--fixture <file> or --fixture-dir <dir> --fixture-id <id>) retains deterministic behavior; swarm dependency-drift
also has a live path.
swarm status composes Beads, Agent Mail metadata, git state, rch/build
pressure, cass health/status, and proof references. Stale candidates are
advisory only: coordinate through Beads and Agent Mail before reopening,
force-releasing, or taking over work. Suggested commands are robot-safe
templates, not automatic actions.
swarm dependency-drift reads Cargo.toml and optional sibling checkouts to
report manifest pins, local HEAD/dirty state, strict validation commands, and
release-risk recommendations. It does not fetch remotes, edit manifests, run
builds, update Beads, send Agent Mail, delete files, or mutate git state.
When status points at prior evidence, use cass pack "query" --robot to create
a bounded cited handoff for another agent. Packs complement the cockpit; they do
not replace Beads for ownership, Agent Mail for coordination, or rch for proof
commands.
LLMs have context limits. cass provides multiple levers to control output size:
| Flag | Effect |
|---|---|
--fields minimal | Only source_path, line_number, agent, source_id, conversation_id |
--fields summary | minimal plus title, score |
--fields score,title,snippet | Custom field selection |
--max-content-length 500 | Truncate long fields (UTF-8 safe, adds "...") |
--max-tokens 2000 | Soft budget (~4 chars/token); adjusts truncation dynamically |
--limit 5 | Cap number of results |
cass pack "query" --robot | Build a cited handoff pack from selected search evidence |
pack --max-tokens N | Set the pack planner's soft budget |
pack --max-evidence N | Cap evidence items selected into the pack |
pack --max-sessions N | Limit how many sessions can contribute evidence |
pack --max-excerpt-chars N | Shorten each cited excerpt before token estimation |
pack --fields summary | Return top-level summary fields for a smaller JSON envelope |
pack --field-mask minimal|standard|full | Select a documented pack projection; --fields accepts the same presets |
pack --freshness-policy strict --freshness-window-seconds N | Reject stale evidence instead of silently mixing it into a pack |
pack --sessions-from FILE | Restrict pack evidence to newline-delimited session paths; use - for stdin |
Truncated fields include a *_truncated: true indicator so agents know when they're seeing partial content.
Contributor verification for docs or contract changes should use rch, for example:
rch exec -- env CARGO_TARGET_DIR=${TMPDIR:-/tmp}/rch_target_cass_answer_pack_docs \
cargo test --test golden_robot_docs
Errors are structured, actionable, and include recovery hints. A real sample from cass search foo --robot against a fresh data dir:
{
"error": {
"code": 3,
"kind": "missing-index",
"message": "cass has not been initialized in <data_dir> yet, so search cannot run until the first index completes.",
"hint": "Run 'cass index --full' once to discover local sessions and build the initial archive.",
"retryable": true
}
}
Kind names are kebab-case (e.g. missing-index, missing-db, semantic-unavailable, embedder-unavailable, ambiguous-source, timeout, config, lock-busy). Agents that branch on err.kind should treat them as stable identifiers. The full set (about 90 kinds) is defined in src/model/cli_error_kind.rs; the canonical way to discover a kind programmatically is to trigger the condition and inspect err.kind from the JSON envelope.
Exit codes follow a semantic convention:
| Code | Meaning | Typical action |
|---|---|---|
| 0 | Success | Parse stdout |
| 1 | Health check failed | Run cass index --full |
| 2 | Usage error | Fix syntax (hint provided) |
| 3 | Index/DB missing | Run cass index --full (retryable: true) |
| 4 | I/O failure or unsafe operation refused (not a network code) | Branch on err.kind: fix path/permissions/space for io/output-not-writable; follow the hint for refused-unsafe |
| 5 | Data corruption, or maintenance required | Inspect cass health --json / cass status --json / cass doctor --json and follow recommended_action: usually rebuild derived assets (maintenance-required, checkpoint_incomplete start that rebuild themselves); only a canonical-archive failure needs repair or restore |
| 6 | Required input missing (password, resume command) | Supply the input (e.g. --password-stdin) and rerun |
| 7 | Lock/busy | Retry later |
| 8 | Partial result (sources sync only: some sources had path failures) | Inspect per-path errors in the JSON output and retry the failed sources |
| 9 | Unknown error | Check retryable flag |
| 10 | Config / timeout | Depends on err.kind |
| 11 | Config validation | Fix config |
| 12 | Source / SSH | Check remote host |
| 13 | Mapping / not-found | Depends on err.kind |
| 14 | I/O / mapping | Retry or inspect path |
| 15 | Semantic / embedder unavailable | Install model or --mode lexical |
| 20-21 | Model acquisition | Check err.kind, err.hint |
| 22 | I/O during model handling | Retry |
| 23 | Model download | Retry or use --from-file |
| 24 | I/O during model verify/install | Retry |
| 70 | cass index stalled and aborted (kind index-stalled envelope on stderr) | Inspect cass status --json, then rerun cass index |
| 130 | Interrupted (SIGINT) | Rerun; cass sources setup --resume continues an interrupted setup |
Search/pack timeouts are not exit 8: on expiry search and pack exit 0 with {"hits": [], "budget": {"timed_out": true, "skipped_sections": [...], "recommended_next_probe": "<command>", ...}}, and --robot-format sessions instead fails with exit 10, kind timeout. Explicit --mode semantic is the other exception: when the remaining budget cannot admit semantic setup or dispatch, search fails with exit 10, kind timeout, retryable: true, and a semantic_budget checkpoint=... message, rather than returning an empty or lexical result. Hybrid (explicit or default) instead falls back to lexical and reports semantic_budget_limited.
Codes ≥ 10 are domain-specific and the numeric value alone is ambiguous (e.g. code 10 maps to either config or timeout kinds depending on context). Agents should branch on err.kind from the JSON error envelope — not on the numeric code — when handling codes ≥ 10. See the Error Handling section above for the canonical kind list.
The retryable field tells agents whether a retry might succeed (e.g., transient I/O) vs. guaranteed failure (e.g., invalid path). A lexical query the engine refuses with posting cursor invariant failed (kind search, exit 9) is retryable: false: the same query fails the same way on the same index generation. The hint names the remedy, cass index --full --force-rebuild. Date-filtered searches over an index segment that holds deleted rows triggered it before frankensearch-quill 0.3.2 (GH #499).
Beyond search, cass provides commands for deep-diving into specific sessions:
# Discover the current session for this workspace
cass sessions --current --json
# List recent sessions for a specific project
cass sessions --workspace /path/to/project --json --limit 5
# Export full conversation to shareable format
cass export /path/to/session.jsonl --format markdown -o conversation.md
cass export /path/to/session.jsonl --format json --include-tools
# Export as self-contained HTML with encryption (recommended for sharing)
cass export-html /path/to/session.jsonl # To Downloads folder
printf '%s\n' "pwd" | cass export-html session.jsonl --encrypt --password-stdin
cass export-html session.jsonl --open --json # Open in browser, JSON output
# Common agent flow: find current session, then export it
cass export-html "$(cass sessions --current --json | jq -r '.sessions[0].path')" --json
# Expand context around a specific line (from search result)
cass expand /path/to/session.jsonl -n 42 -C 5 --json
# → Shows 5 messages before and after line 42
# Activity timeline: when were agents active?
cass timeline --today --json --group-by hour
cass timeline --since 7d --agent claude --json
# → Grouped activity counts, useful for understanding work patterns
Aggregate search results server-side to get counts and distributions without transferring full result data:
# Count results by agent
cass search "error" --robot --aggregate agent
# → { "aggregations": { "agent": { "buckets": [{"key": "claude_code", "count": 45}, ...] } } }
# Multi-field aggregation
cass search "bug" --robot --aggregate agent,workspace,date
# Combine with filters
cass search "TODO" --agent claude --robot --aggregate workspace
Aggregation Fields:
| Field | Description |
|---|---|
agent | Group by agent type (claude_code, codex, cursor, etc.) |
workspace | Group by workspace/project path |
date | Group by date (YYYY-MM-DD) |
match_type | Group by match type (exact, prefix, suffix, substring, wildcard, implicit_wildcard); one search has one type, so this shows a single bucket unless the wildcard fallback replaced the hits |
Response Format:
{
"aggregations": {
"agent": {
"buckets": [
{"key": "claude_code", "count": 120},
{"key": "codex", "count": 85}
],
"other_count": 15
}
}
}
Top 10 buckets are returned per field, with other_count for remaining items.
Mine recurrent CASS operational incidents from the canonical archive without dumping raw session text:
cass analytics incidents --limit 10 --json
# Tighten the bounded scan for automation or a very large archive
cass analytics incidents --max-sessions 500 --max-messages 50000 \
--max-bytes 67108864 --budget-ms 5000 --json
The response ranks top_sessions[] by hit count and category breadth and keeps
the exact conversation_id, agent, host, source_id, source_path,
live/archive state, dominant categories, and a structured cass view argv.
That argv carries the effective --db path plus --conversation-id, so it
opens the exact ranked archive row even when multiple sessions share a source
path or the report used a non-default database.
total_sessions, total_hits, and top_sessions_truncated distinguish the
bounded ranked result from the totals observed inside the scan scope.
discovery.partial and stop_reason explicitly distinguish a bounded partial
scan from a complete scan. Counts are scoped to scanned candidates whenever the
scan is partial. --budget-ms is a hard wall-clock result guard around the
independently row-bounded read-only worker. If it expires before a verified
result arrives, count_scope="no_verified_results_hard_timeout" returns an
empty partial report instead of overstating in-flight observations. Candidate
discovery is descending archive-row keyset paging;
--max-sessions bounds that newest-row window before dimensional filters, so a
selective filter can truthfully return a partial empty result instead of scanning
an unbounded archive. Individual messages are inspected through a bounded 4,096-char
fragment; an oversized message returns message-fragment-capped rather than
claiming a complete corpus scan. Raw prompt/tool content is always suppressed; evidence carries
only BLAKE3 fingerprints and basename-redacted paths. The actionable
source_path remains visible solely so the returned view command works.
Chain multiple searches together by piping session paths from one search to another:
# Find sessions mentioning "auth", then search within those for "token"
cass search "authentication" --robot-format sessions | \
cass search "refresh token" --sessions-from - --robot
# Build a filtered corpus from today's work
cass search --today --robot-format sessions > today_sessions.txt
cass search "bug fix" --sessions-from today_sessions.txt --robot
How It Works:
--robot-format sessions outputs one session path per line--sessions-from <file> restricts search to those sessions- to read from stdin for true pipingUse Cases:
Snippets always mark the terms the search engine matched with **bold**, in human-readable output and in robot/JSON output alike. The --highlight flag also marks the query terms' remaining literal occurrences and leaves already-marked text alone, so no term gets two pairs of marks. Only the snippet field is marked; content stays verbatim:
cass search "authentication error" --robot --highlight
# "snippet": "... **authentication** failed with **error** ..."
Highlighting is query-aware: quoted phrases like "auth error" highlight as a unit; individual terms highlight separately.
For large result sets, use cursor-based pagination:
# First page
cass search "TODO" --robot --robot-meta --limit 20
# → { "hits": [...], "_meta": { "next_cursor": "eyJ..." } }
# Next page
cass search "TODO" --robot --robot-meta --limit 20 --cursor "eyJ..."
A cursor is base64 JSON {"offset": N, "limit": M}: a plain page position, not a snapshot. Any index change between pages (a new session indexed, a rebuild, a forget) shifts the ranking, so the next page can skip or repeat hits. Page quickly, or fix the window with --until when you need stable pages.
total_matches answers "how many messages match?", but it is exact only when
cass can afford to count. Every search fetches one hit beyond --limit to learn
whether another page exists. When the page is full and the index is larger than
CASS_SEARCH_EXACT_TOTAL_COUNT_MAX_DOCS documents, cass skips the full count
and reports that limit + 1 as a lower bound. Below the threshold it counts
every match.
| Build | Threshold | stale lock with --limit 10 on a 1,034,219-document index |
|---|---|---|
| v0.9.0 and earlier | 50,000 documents | total_matches: 11 |
| Current (unreleased) | 5,000,000 documents | total_matches: 11915 |
--robot-meta says which kind of number you got:
cass search "stale lock" --robot --robot-meta --limit 10 \
| jq '{total_matches, precision: ._meta.cursor_manifest.count_precision, why: ._meta.cursor_manifest.count_reason}'
# → {"total_matches": 11915, "precision": "exact", "why": "total_matches is exact; no extra recount was needed"}
# A lower bound reads "precision": "lower_bound".
Why the threshold moved. The 50,000 cap dates from the Tantivy engine,
where counting a common term over millions of documents could dominate the
query. The Quill engine counts cheaply. Paired runs on that 1,034,219-document
archive, capped against exact (--limit 10, read-only, CPU time):
| Query | Exact total | Extra CPU for the exact count |
|---|---|---|
stale lock | 11,915 | ~0.00-0.04 s |
cargo build | 24,944 | ~0.01-0.03 s |
the | 439,461 | ~0.03-0.05 s |
AGENTS.md | 867,087 | ~0.06-0.11 s |
Each search cost about 0.8 s of CPU either way. The capped answer, meanwhile,
was wrong by up to five orders of magnitude, and agents read total_matches as
a count. The default now covers five times that archive; set
CASS_SEARCH_EXACT_TOTAL_COUNT_MAX_DOCS=0 to never count exactly, or raise it
for a larger archive.
Aggregations have their own window. --aggregate buckets are computed over
the top max(1000, limit + offset) hits, so bucket counts on a large archive
describe the best-ranked thousand matches, not the whole corpus. Use an exact
total_matches for "how many", and aggregations for "how are the top hits
distributed".
For debugging and logging, attach a request ID:
cass search "bug" --robot --request-id "req-12345"
# → { "request_id": "req-12345", "hits": [...], ... }
# (top level always; also under _meta.request_id with --robot-meta)
For safe retries (e.g., in CI pipelines or flaky networks):
cass index --full --idempotency-key "build-$(date +%Y%m%d)"
# If same key + params were used in last 24h, returns cached result
Debug why a search returned unexpected results:
cass search "auth*" --robot --explain
# → Adds "explanation": the sanitized query, flat lists of its terms, phrases and
# operators (not a tree), the query type and index strategy, a low/medium/high cost
# class, a filter summary and warnings. Wildcards are reported, not expanded.
cass search "auth error" --robot --dry-run
# → Validates query syntax without executing
For debugging agent pipelines:
cass search "error" --robot --trace-file /tmp/cass-trace.json
# Appends execution span with timing, exit code, and command details
cass index --full --json --robot-trace-ingest 2>/tmp/cass-ingest-trace.jsonl
# Streams one NDJSON record per ingest batch with wall_ms, batch_msgs,
# inserted_messages, and duplicate-lookup counters for perf bisects
| Flag | Purpose |
|---|---|
--robot / --json | JSON output (pretty-printed) |
--robot-format jsonl|compact | Streaming or single-line JSON |
--robot-meta | Include _meta block (elapsed_ms, cache stats, index freshness, lexical_degrade_reason: "query_fuel_exhausted" or null, wildcard_fallback_skipped: why a sparse result got no automatic wildcard retry, and effective: the database, time window, filters, auto-corrections and query grouping the search actually used) |
--fields minimal|summary|<list> | Reduce payload size |
--max-content-length N | Truncate content fields to N chars |
--max-tokens N | Apply an approximate token budget to robot output |
--timeout N | Timeout in milliseconds. On expiry search/pack still exit 0 and emit {"hits": [], "budget": {"timed_out": true, "skipped_sections": [...], "recommended_next_probe": "<command>", ...}}; --robot-format sessions fails with exit 10, kind timeout |
--cursor <token> | Cursor-based pagination (from _meta.next_cursor) |
--request-id ID | Echoed in response for correlation |
--aggregate agent,workspace,date | Server-side aggregations |
--explain | Include query analysis (parsed query, cost estimate) |
--dry-run | Validate query without executing |
--no-maintenance | Strict read-only search: never refresh, join, or spawn lexical maintenance, never auto-repair the archive while opening it, and never auto-spawn the daemon (conflicts with --refresh and --daemon) |
--source <source> | Filter by source: local, remote, all, or specific source ID |
--highlight | Also mark query-term occurrences the engine left unmarked (snippets always mark matched terms with **) |
| Flag | Purpose |
|---|---|
--idempotency-key KEY | Safe retries: same key + params returns cached result (24h TTL) |
--json | JSON output with stats |
--gc | Reclaim merge-retired lexical segment files and exit: runs the engine's grace-period garbage sweep (a folded segment file is unlinked only once no published MANIFEST generation has referenced it for 300 s) and reports files/bytes reclaimed. Every incremental cass index performs the same sweep at open; doctor --json reports the reclaimable bytes under storage_pressure.full_rebuild_readiness (GH #453) |
When health --json or status --json reports index.status: "hollow", the
live Quill generation serves fewer than half the documents certified by its
completed rebuild checkpoint. index.live_documents reports the served count.
Run cass index to let its pre-scan repair rebuild from the canonical archive;
cass index --full also rescans the session sources. A missing count provides
no hollow-generation verdict.
For machine-readable documentation, use cass robot-docs <topic>:
| Topic | Content |
|---|---|
commands | Full command reference with all flags |
env | Environment variables and defaults |
paths | Data directory locations per platform |
guide | Quick start guide for automation |
schemas | JSON response schemas |
exit-codes | Exit code meanings and retry guidance |
examples | Copy-paste usage examples |
contracts | API contract version and stability |
sources | Remote sources configuration guide |
# Get documentation programmatically
cass robot-docs guide
cass robot-docs schemas
cass robot-docs exit-codes
# Machine-first help (wide output, no TUI assumptions)
cass --robot-help
cass maintains a stable API contract for automation:
cass api-version --json
# → { "crate_version": "<cargo version>", "build_commit": "<sha or unknown>", "api_version": 1, "contract_version": "1" }
cass introspect --json
# → Full schema: all commands, arguments, response types
Contract Version: Currently 1. Increments only on breaking changes.
Guaranteed Stable:
--robot output_meta block format🔎 cass — Search All Your Agent History
What: cass indexes conversations from Claude Code, Codex, Cursor, Gemini, Aider, ChatGPT, and more into a unified, searchable index. Before solving a problem from scratch, check if any agent already solved something similar.
⚠️ NEVER run bare cass — it launches an interactive TUI. Always use --robot or --json.
Quick Start
# One-shot agent triage (read next_command when present)
cass triage --json
# Search across all agent histories
cass search "authentication error" --robot --limit 5
# Build a cited handoff pack from search evidence
cass pack "authentication error root cause" --robot --max-tokens 12000 --limit 40
# Tight handoff budget with freshness and privacy metadata
cass pack "authentication error root cause" --robot --max-tokens 4000 --max-evidence 8 --fields summary
# View a specific result (from search output)
cass view /path/to/session.jsonl -n 42 --json
# Expand context around a line
cass expand /path/to/session.jsonl -n 42 -C 3 --json
# Learn the full API
cass capabilities --json # Static agent self-description
cass robot-docs guide # LLM-optimized docs
Why Use It
- Cross-agent knowledge: Find solutions from Codex when using Claude, or vice versa
- Forgiving syntax: Typos and wrong flags are auto-corrected with teaching notes
- Token-efficient: --fields minimal returns only essential data; pack budgets cite only selected evidence
- Copy-safe handoffs: pack warnings include freshness and privacy/redaction status
Key Flags
| Flag | Purpose |
|------------------|--------------------------------------------------------|
| --robot / --json | Machine-readable JSON output (required!) |
| --fields minimal | Reduce payload: source_path, line_number, agent, source_id, conversation_id |
| pack --max-tokens N | Budget a cited handoff pack |
| --limit N | Cap result count |
| --agent NAME | Filter to specific agent (claude, codex, cursor, etc.) |
| --days N | Limit to recent N days |
stdout = data only, stderr = diagnostics. Exit 0 = success.
cass supports a rich query syntax designed for both humans and machines.
| Query | Matches |
|---|---|
error | Messages containing "error" (case-insensitive) |
python error | Messages containing both "python" AND "error" |
"authentication failed" | Exact phrase match |
auth fail | Both terms, in any order |
Combine terms with explicit operators for complex queries:
| Operator | Example | Meaning |
|---|---|---|
AND | python AND error | Both terms required (default) |
OR | error OR warning | Either term matches |
NOT | error NOT test | First term, excluding second |
- | error -test | Shorthand for NOT |
Operator Precedence: NOT binds tightest, then AND (explicit, &&, or implied between words), then OR (OR, ||). Parentheses group, in the TUI and robot mode alike: a OR b c means a OR (b AND c), while (a OR b) c needs the parentheses. A ( groups only at the start of a word, so code such as foo(bar) stays one term. NOT NOT x is x. Unbalanced parentheses are recovered rather than rejected. The SQLite fallback lanes, used while no lexical index is available, apply the same grammar.
# Complex boolean query
cass search "authentication AND (error OR failure) NOT test" --robot
# Exclude test files
cass search "bug fix -test -spec" --robot
# Either error type
cass search "TypeError OR ValueError" --robot
Wrap terms in double quotes for exact phrase matching:
| Query | Matches |
|---|---|
"file not found" | Exact sequence "file not found" |
"cannot read property" | Exact JavaScript error message |
"def test_" | Function definitions starting with test_ |
Phrases match their words adjacent and in order (no slop). Useful for error messages, code patterns, and specific terminology.
| Pattern | Type | Matches | Performance |
|---|---|---|---|
auth* | Prefix | "auth", "authentication", "authorize" | Fast (uses edge n-grams) |
*tion | Suffix | "authentication", "function", "exception" | Slower (term-dictionary expansion) |
*config* | Substring | "reconfigure", "config.json", "misconfigured" | Slowest (term-dictionary expansion) |
test_* | Prefix on test | anything whose token starts with "test" | Fast |
Tip: Prefix wildcards (foo*) use edge n-grams computed at index time: prefixes of 2–20 characters of every alphanumeric word, taken from titles and from the first 4 KiB of each message. Suffix and substring wildcards expand over the index's term dictionary, at most 16,384 terms per pattern; a pattern matching more terms fails rather than scanning. The tokenizer splits on anything that is not a letter or digit: test_* is a prefix match on test, c++ searches for c, and foo.bar means foo AND bar anywhere in the message, not the literal string. Phrases ("...") match adjacent words in order (slop 0).
# Field-specific search (in robot mode)
cass search "error" --agent claude --workspace /path/to/project
# Time-bounded search
cass search "bug" --since 2024-01-01 --until 2024-01-31
cass search "bug" --today
cass search "bug" --days 7
# Combined filters
cass search "authentication" --agent codex --workspace myproject --week
cass accepts a wide variety of time/date formats for filtering:
| Format | Examples | Description |
|---|---|---|
| Relative | -7d, -24h, -30m, -1w | Days, hours, minutes, weeks ago |
| Keywords | now, today, yesterday | Named reference points |
| ISO 8601 | 2024-11-25, 2024-11-25T14:30:00Z | Standard datetime |
| US Dates | 11/25/2024, 11-25-2024 | Month/Day/Year |
| Unix Timestamp | 1732579200 | Seconds since epoch |
| Unix Millis | 1732579200000 | Milliseconds (auto-detected) |
Intelligent Heuristics:
2024, not 24)today, yesterday) starts at local midnight as --since and runs through the day's last millisecond as --until, so --until 2024-01-31 includes January 31search and pack, a --since/--until value that cannot be parsed, or a --since later than --until, is a usage error (exit 2, kind usage); it is never silently ignored# All equivalent for "last week"
cass search "bug" --since -7d
cass search "bug" --since "-1w"
cass search "bug" --days 7
# Date range
cass search "feature" --since 2024-01-01 --until 2024-01-31
# Mix formats
cass search "error" --since yesterday --until now
Search results include a match_type indicator. It describes the query, not each hit: every hit of a search carries the type of the least precise pattern in the query (e.g. auth* *tion stamps suffix on all hits).
| Type | Meaning | Match Quality order |
|---|---|---|
exact | No wildcards: exact terms (and edge n-gram prefixes) | 1 |
prefix | Trailing wildcard (auth*) | 2 |
suffix | Leading wildcard (*tion) | 3 |
substring | Both sides (*config*) | 4 |
wildcard | Inner wildcard (f*o) | 5 |
implicit_wildcard | Automatic wildcard fallback on a sparse exact search | 6 |
Relevance scores carry no boost for the match type; only the TUI's Match Quality ranking mode (F12) orders by it.
When an exact query's first page returns fewer than 3 results (or fewer than a smaller --limit), cass retries with wildcard expansion:
auth → *auth*CASS_AUTOMATIC_WILDCARD_FALLBACK_MAX_DOCS; 0 disables it), so on a typical real archive it does not run._meta.wildcard_fallback: true. When a sparse result did not get the retry, _meta.wildcard_fallback_skipped says why: index_over_automatic_limit (the index is over the document cap), automatic_retry_disabled (the cap is 0) or long_query_term; add explicit wildcards to run it anyway| Key | Action |
|---|---|
Ctrl+C | Force quit |
Esc / F10 | Unwind: close the open modal or surface, otherwise quit |
F1 / Alt+? | Toggle help screen |
F2 / Alt+T | Next theme (cycles all 19 presets) |
Shift+F2 / Alt+Shift+T | Previous theme |
Ctrl+B | Toggle border style (rounded/square) |
Ctrl+P / Alt+P | Open the command palette |
Ctrl+S | Toggle the stats bar |
Ctrl+Shift+S | Open the sources management surface |
Alt+A | Open the analytics dashboard |
Alt+M | Toggle macro recording (replay with cass tui --play-macro FILE) |
Ctrl+Shift+I | Toggle the inspector overlay |
Ctrl+Shift+R | Force re-index |
Ctrl+Shift+Del | Reset all TUI state |
Ctrl+Z / Ctrl+Shift+Z | Undo / redo |
Launch-time flags: cass tui --refresh (alias --catch-up) runs an incremental index pass before opening; --record-macro FILE / --play-macro FILE record and replay input events.
| Key | Action |
|---|---|
| Type | Live search as you type; plain characters (including ?, y, o, c, 1-9, -, =) go into the query |
Enter | Open the selected hit; with no selected hit, submit the query (if the query is empty, edit the last filter chip) |
Backspace | Delete character; if the query is empty, remove the last filter chip |
Left/Right, Ctrl+Left/Ctrl+Right | Move the cursor by character / by word |
Home/End | Jump the cursor to the start / end of the query |
Ctrl+L | Clear the query |
Ctrl+U / Ctrl+K / Ctrl+W | Kill to line start / to line end / previous word |
Ctrl+R | Cycle through query history |
Ctrl+N / Ctrl+Shift+N | Next / previous query-history entry |
Ctrl+F | Toggle wildcard fallback |
Ctrl+Shift+Y | Copy the query |
| Key | Action |
|---|---|
Up/Down | Move selection in results list |
PageUp/PageDown | Scroll by page |
Tab / Shift+Tab | Toggle focus between results and detail pane / move focus left |
Alt+h/j/k/l | Vim-style directional focus (left/down/up/right) |
Alt+1..Alt+9 | Switch to pane N |
Alt+- / Alt+= | Shrink / grow the results pane |
Alt+D | Hide / show the detail pane |
Alt+[ / Alt+] | Timeline jump backward / forward |
| Key | Action |
|---|---|
F3 / Alt+G | Open agent filter palette |
Shift+F3 / Alt+Shift+G | Clear the agent filter |
F4 / Alt+W | Open workspace filter palette |
Shift+F4 / Alt+Shift+W / Ctrl+Del | Clear all active filters |
F5 | Set "from" time filter |
F6 | Set "to" time filter |
Shift+F5 | Cycle time presets: 24h → 7d → 30d → all |
F11 / Shift+F11 | Cycle the source filter / open the source filter menu |
Alt+/ | Open the pane filter |
| Key | Action |
|---|---|
F7 / Alt+C | Cycle context window size: S → M → L → XL |
Ctrl+Space | Momentary "peek" to XL context |
F9 | Toggle match mode: standard (default) ↔ prefix, where every bare word of 2+ characters also matches as a prefix (auth → auth*; phrases, operators and wildcards are left as typed) |
F12 / Alt+R | Cycle ranking: recent → balanced → relevance → quality → newest → oldest |
Alt+F | Cycle result grouping: agent → conversation → workspace → flat |
Alt+S | Cycle search mode (lexical / semantic / hybrid). Without an installed model (or vector index) the status line says results stay lexical and names the command: cass models install (offline --from-file <dir>) or cass index --semantic; nothing downloads on its own |
Ctrl+D | Cycle density: Compact → Cozy → Spacious |
Ctrl+1..Ctrl+9 | Save the current view to slot N |
Shift+1..Shift+9 | Load the view from slot N |
For one chronological list across agents, select flat grouping with Alt+F
and newest ranking with F12. Grouping is also available in the command
palette and is preserved with ranking and filters in saved views.
| Key | Action |
|---|---|
Enter / Ctrl+M | Open selected result in the detail modal (Messages tab by default) |
Ctrl+X | Toggle selection on current result |
Ctrl+A | Select/deselect all visible results |
Alt+B | Open bulk actions menu (when items selected) |
Ctrl+Enter | Add to multi-open queue |
Ctrl+O | Open all queued items in editor |
F8 / Alt+O | Open selected hit in $EDITOR |
Alt+V | View raw |
Alt+Shift+J | Toggle JSON view |
Ctrl+Y | Copy path |
Alt+Y | Copy snippet |
Ctrl+Shift+C | Copy content |
Ctrl+E | Open the export modal |
Ctrl+Shift+E | Export Markdown immediately |
Alt+U / Alt+N / Alt+I | Update banner: upgrade now / show release notes / skip this version |
These apply while the detail modal is open:
| Key | Action |
|---|---|
Esc | Close the detail modal |
Tab | Cycle detail tabs |
/ (or Ctrl+F, Alt+/) | Start find-in-detail; type to search, Enter advances to the next match |
n / N | Next / previous contextual search hit within this session |
Enter (Messages tab) | Next contextual search hit |
j / k, Up/Down | Scroll |
g / G, Home/End | Scroll to top / bottom |
{ / } | Jump to previous / next message |
[ / ] | Jump to previous / next user message |
w | Toggle line wrap |
e / c | Expand / collapse all tool and system messages |
e, h (Export tab) | Open the HTML export modal; m exports Markdown |
F7 | Cycle context window size |
Ctrl+Space | Momentary "peek" to XL context |
The detail pane has six tabs, cycled with Tab:
| Tab | Content | Best For |
|---|---|---|
| Messages | Full conversation with markdown rendering | Reading full context |
| Snippets | Keyword-extracted summaries | Quick scanning |
| Raw | Unformatted JSON/text | Debugging, copying exact content |
| Json | Syntax-highlighted, pretty-printed JSON (static; no collapsible tree) | Inspecting structured payloads |
| Analytics | Per-session token timeline, tool calls, message stats | Understanding one session |
| Export | Export actions and filename previews (HTML/Markdown) | Sharing a session |
Control how much content shows in the detail preview. Cycle with F7:
| Size | Characters | Use Case |
|---|---|---|
| Small | ~200 | Quick scanning, narrow terminals |
| Medium | ~400 | Default balanced view |
| Large | ~800 | Reading longer passages |
| XLarge | ~1600 | Full context, code review |
Peek Mode (Ctrl+Space): Temporarily expand to XL context. Press again to restore previous size. Useful for quick deep-dives without changing your preferred default.
Efficiently work with multiple search results at once:
Multi-Select Mode:
Ctrl+X to toggle selection on current result (checkbox appears)Ctrl+X againCtrl+A to select/deselect all visible resultsBulk Actions Menu (Alt+B when items selected):
| Action | Description |
|---|---|
| Open All | Open all selected files in editor |
| Copy Paths | Copy all file paths to clipboard |
| Export | Export selected results to file |
| Clear Selection | Deselect all items |
Multi-Open Queue: For opening many files without navigating away:
Ctrl+Enter to add current result to queueCtrl+O to open all queued itemsClipboard Operations:
Ctrl+Y - Copy the current item's pathAlt+Y - Copy the current item's snippetCtrl+Shift+C - Copy the current item's contentCycle through modes with F12 (or Alt+R) in the TUI. The search engine returns hits in its own relevance order; the mode then re-orders the results the TUI has loaded (the first page of up to 250 hits, plus every further page you load), so the whole loaded list always follows one order. Ranking modes are TUI-only; robot search returns engine order.
Recent Heavy: Score = relevance × 0.3 + recency × 0.7. Best for: "What was I working on?"
Balanced (default): Score = relevance × 0.5 + recency × 0.5. Best for general-purpose search.
Relevance: Score = relevance × 0.8 + recency × 0.2. Best for "find the best explanation of X".
Match Quality: exact matches first, then prefix, suffix, substring, wildcard, and finally automatic wildcard-fallback matches; within each class, the Relevance score. Best for precise technical searches.
Date Newest: newest first by message time. Best for "show me recent activity".
Date Oldest: oldest first. Best for "when did I first work on this?"
Undated hits sort last in the date modes. Hits with equal scores keep the engine's order. With an empty query, cass browses by date instead of searching; Date Oldest browses oldest first and every other mode newest first.
Relevance: the engine's score (BM25 for lexical search, fused rank for hybrid), min-max normalized over the loaded hits: the best loaded hit is 1.0 and the weakest 0.0 (all 1.0 when every score ties).
Recency: exponential decay from now with a 14-day half-life, 0.5 ^ (age_days / 14): 1.0 today, about 0.71 after a week, 0.5 after two weeks, about 0.23 after a month. Undated hits get 0.
The formulas live in src/ui/ranking.rs, and the tests in that file and in tests/ranking.rs exercise the same function the TUI calls.
Each connector transforms agent-specific formats into a unified schema:
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ Agent Files │ ──▶ │ Connector │ ──▶ │ Normalized │
│ (proprietary) │ │ (per-agent) │ │ Conversation │
└─────────────────┘ └──────────────────┘ └─────────────────┘
JSONL detect() agent_slug
SQLite scan() workspace
Markdown messages[]
JSON created_at
Different agents use different role names:
| Agent | Original | Normalized |
|---|---|---|
| Claude Code | human, assistant | user, assistant |
| Codex | user, assistant | user, assistant |
| ChatGPT | user, assistant, system | user, assistant, system |
| Cursor | user, assistant | user, assistant |
| Aider | (markdown headers) | user, assistant |
Agents store timestamps inconsistently:
| Format | Example | Handling |
|---|---|---|
| Unix milliseconds | 1699900000000 | Direct conversion |
| Unix seconds | 1699900000 | Multiply by 1000 |
| ISO 8601 | 2024-01-15T10:30:00Z | Parse with chrono |
| Missing | null | Use file modification time |
Tool calls, code blocks, and nested structures are flattened for searchability:
// Original (Claude Code)
{"type": "tool_use", "name": "Read", "input": {"path": "/foo/bar.rs"}}
// Flattened for indexing
"[Tool: Read] path=/foo/bar.rs"
The same conversation content can appear multiple times due to:
cass uses a multi-layer deduplication strategy:
Message identity: messages are keyed by UNIQUE(conversation_id, idx). Appends to a known conversation use INSERT OR IGNORE, and a new conversation's batched INSERT has the same unique index as its backstop, so re-indexing the same file never stores a message twice
Conversation identity: conversations are keyed by UNIQUE(source_id, agent_id, external_id)
Search-Time Dedup: hits are deduplicated on an exact key tuple — (source, source path, conversation id or title, line number, created_at, whitespace-invariant content hash) — keeping the highest-scored hit
Common low-value content is filtered from results:
# Find past solutions for similar errors
cass search "TypeError: Cannot read property" --days 30
# In TUI: F12 to switch to "relevance" mode for best matches
# What has ANY agent said about authentication in this project?
cass search "authentication" --workspace /path/to/project
# Export findings for a new agent's context
cass export /path/to/relevant/session.jsonl --format markdown
# What did I work on today?
cass timeline --today --json | jq '.groups[].conversations'
# TUI: Press Shift+F5 to cycle through time filters
# Find all debugging sessions for a specific file
cass search "debug src/auth/login.rs" --agent claude
# Expand context around a specific line in a session
cass expand /path/to/session.jsonl -n 150 -C 10
# Current agent searches what previous agents learned
cass search "database migration strategy" --robot --fields minimal
# Get full context for a relevant session
cass view /path/to/session.jsonl -n 42 --json
# Export high-quality problem-solving sessions
cass search "bug fix" --robot --limit 100 | \
jq '.hits[] | select(.score > 0.8)' > training_candidates.json
Press Ctrl+P to open the command palette—a fuzzy-searchable menu of all available actions.
| Command | Description |
|---|---|
| Toggle theme | Switch between dark/light mode |
| Toggle density | Cycle Compact → Cozy → Spacious |
| Toggle help strip | Pin/unpin the contextual help bar |
| Check updates | Show update assistant banner |
| Filter: agent | Open agent filter picker |
| Filter: workspace | Open workspace filter picker |
| Filter: today | Restrict results to today |
| Filter: last 7 days | Restrict results to past week |
| Filter: date range | Prompt for custom since/until |
| Saved views | List and manage saved view slots |
| Save view to slot N | Save current filters to slot 1-9 |
| Load view from slot N | Restore filters from slot 1-9 |
| Bulk actions | Open bulk menu (when items selected) |
| Reload index/view | Refresh the search reader |
Ctrl+P to openUp/Down to navigateEnter to executeEsc to closeSave your current filter configuration to one of 9 slots for instant recall.
| Key | Action |
|---|---|
Ctrl+1 through Ctrl+9 | Save current view to slot |
Shift+1 through Shift+9 | Load view from slot |
Ctrl+P → "Save view to slot N"Ctrl+P → "Load view from slot N"Ctrl+P → "Saved views" to list all slotsViews are stored in tui_state.json and persist across sessions. Clear all saved views with Ctrl+Shift+Del (resets all TUI state).
Control how many lines each search result occupies. Cycle with Ctrl+D or via the command palette.
| Mode | Lines per Result | Best For |
|---|---|---|
| Compact | 2 | Maximum results visible, scanning many items |
| Cozy (default) | 5 | Balanced view with context |
| Spacious | 6 | Detailed preview, fewer results |
The pane automatically adjusts how many results fit based on terminal height and density mode.
cass includes a sophisticated theming system with multiple presets, accessibility-aware color choices, and adaptive styling.
Cycle through 19 built-in theme presets with F2:
| Theme | Description | Best For |
|---|---|---|
| Tokyo Night (default) | Deep blues with restrained contrast | Low-light environments, extended sessions |
| Daylight | High-contrast light background | Bright environments, presentations |
| Catppuccin Mocha | Warm pastels, reduced eye strain | All-day coding, aesthetic preference |
| Dracula | Purple-accented dark theme | Popular among developers, familiar feel |
| Nord | Arctic-inspired cool tones | Calm, focused work sessions |
| Solarized Dark | Precisely tuned low-contrast palette | Long editing sessions, monitor-agnostic |
| Solarized Light | Solarized on a cream background | Paper-style readability in bright rooms |
| Monokai | Classic warm dark palette | Familiar Sublime/TextMate feel |
| Gruvbox Dark | Retro earth tones on dark | Warmer alternative to Tokyo Night |
| One Dark | Atom's signature balanced dark | Moderate contrast, friendly defaults |
| Rosé Pine | Soho-inspired muted roses | Gentle contrast, boutique look |
| Everforest | Forest-inspired green-brown palette | Calm, nature-adjacent mood |
| Kanagawa | Japanese ink-and-paper theme | Artistic, quietly distinctive |
| Ayu Mirage | Ayu's balanced muted dark | Blue-teal accents, relaxed contrast |
| Nightfox | Fox-inspired warm dark | Deep violets with orange highlights |
| Cyberpunk Aurora | Neon aurora on obsidian | Showy, high-saturation dark |
| Synthwave '84 | Retro neon magenta/cyan | 80s aesthetic, fun demos |
| High Contrast | Maximum readability | Accessibility needs, bright monitors |
| Colorblind | Deuteranopia/protanopia-safe palette | Color-vision-deficient users |
Every theme preset is checked in the test suite against WCAG contrast ratios, with these floors:
These floors are below WCAG AA's 4.5:1 for body text, so cass does not claim AA conformance. At runtime, role and status badges pick whichever candidate foreground has the highest contrast against their background.
Conversation messages are color-coded by role for quick visual parsing:
| Role | Visual Treatment | Purpose |
|---|---|---|
| User | Blue-tinted background, bold | Your input, easy to scan |
| Assistant | Green-tinted background | AI responses |
| System | Gray/muted background | Context, instructions |
| Tool | Orange-tinted background | Tool calls, file operations |
Each agent type (Claude, Codex, Cursor, etc.) also receives a subtle tint, making multi-agent result lists instantly scannable.
Border decorations adapt to terminal width and to render pressure:
| Condition | Style | Example |
|---|---|---|
| Narrow (<80 cols) | Square box-drawing | ┌─ content ─┐ |
| 80 cols and wider | Rounded corners | ╭─ content ─╮ |
| Frame budget under pressure | Square, then no borders | ┌─┐, then none |
There is no double-line tier. Ctrl+B toggles between rounded and square Unicode borders; both are box-drawing characters, not ASCII.
Bookmarks are a CLI feature: cass bookmarks add|list|remove|search|export|import --json manages user-authored annotations on search results (a source path, optional line number, note, and tags). The TUI has no bookmark keybindings today.
# Bookmark a search hit (source_path + line_number from search output)
cass bookmarks add /path/to/session.jsonl -n 42 --title "JWT refresh fix" \
--note "Good explanation of the refresh flow" --tags "auth,jwt" --json
# List (optionally by tag), search notes/titles/snippets, remove by id
cass bookmarks list --tag auth --json
cass bookmarks search "refresh" --json
cass bookmarks remove 1 --json # exit 13 (`bookmark-not-found`) if the id is unknown
# Back up and restore
cass bookmarks export -o bookmarks.json --json
cass bookmarks import bookmarks.json --json
bookmarks.db (SQLite), separate from the search index and never pruned by doctor/cleanup flowslist can filter by tag{
"id": 1,
"title": "Auth bug fix discussion",
"source_path": "/path/to/session.jsonl",
"line_number": 42,
"agent": "claude_code",
"workspace": "/projects/myapp",
"note": "Good explanation of JWT refresh flow",
"tags": "auth, jwt, important",
"snippet": "The token refresh logic should..."
}
Bookmarks are stored separately from the main index:
~/.local/share/coding-agent-search/bookmarks.db~/Library/Application Support/coding-agent-search/bookmarks.db%APPDATA%\coding-agent-search\bookmarks.dbcass uses a non-intrusive toast notification system for transient feedback—operations complete, errors occur, or state changes without modal dialogs interrupting your workflow.
| Type | Icon | Auto-Dismiss | Use Case |
|---|---|---|---|
| Info | i | 3 seconds | Status updates, tips |
| Success | * | 2 seconds | Operations completed |
| Warning | ! | 4 seconds | Non-critical issues |
| Error | x | 6 seconds | Failures requiring attention |
Toasts feature:
| Trigger | Toast |
|---|---|
| Bulk copy of selected paths | * "Copied 3 paths" |
| Bulk export | * "Exported 3 items as JSON" |
| Copy failure | x "Copy failed: ..." |
| Slow search | "Slow search: 1840ms" |
| Semantic refinement failure | "Refinement failed: ..." |
| Saved views | * "Renamed slot 2", ! "Slot 4 is empty" |
To achieve sub-60ms latency on large datasets, cass implements a multi-tier caching strategy in src/search/query.rs:
prefix_cache is organized into shards (default 256 entries each) that bound per-prefix memory; all shards sit behind one Mutex, so the sharding limits size, not lock contention. The cache lives in the searching process: it pays off in the TUI, where each keystroke re-queries, and not for one-shot CLI searches, which start with an empty cache.WarmJob thread watches the input. When the user pauses typing, it runs a lightweight query against the lexical reader to pre-load relevant index segments into the OS page cache. One-shot CLI searches run with warming disabled.The system is designed for extensibility via the Connector trait (src/connectors/mod.rs). This allows cass to treat disparate log formats as a uniform stream of events.
classDiagram
class Connector {
<<interface>>
+detect() DetectionResult
+scan(ScanContext) Vec~NormalizedConversation~
}
class NormalizedConversation {
+agent_slug String
+messages Vec~NormalizedMessage~
}
Connector <|-- CodexConnector
Connector <|-- ClineConnector
Connector <|-- ClaudeCodeConnector
Connector <|-- GeminiConnector
Connector <|-- ClawdbotConnector
Connector <|-- VibeConnector
Connector <|-- OpenCodeConnector
Connector <|-- AmpConnector
Connector <|-- CursorConnector
Connector <|-- ChatGptConnector
Connector <|-- AiderConnector
Connector <|-- PiAgentConnector
Connector <|-- FactoryConnector
Connector <|-- CopilotConnector
Connector <|-- CopilotCliConnector
Connector <|-- OpenClawConnector
Connector <|-- CrushConnector
Connector <|-- HermesConnector
Connector <|-- KimiConnector
Connector <|-- QwenConnector
CodexConnector ..> NormalizedConversation : emits
ClineConnector ..> NormalizedConversation : emits
ClaudeCodeConnector ..> NormalizedConversation : emits
GeminiConnector ..> NormalizedConversation : emits
ClawdbotConnector ..> NormalizedConversation : emits
VibeConnector ..> NormalizedConversation : emits
OpenCodeConnector ..> NormalizedConversation : emits
AmpConnector ..> NormalizedConversation : emits
CursorConnector ..> NormalizedConversation : emits
ChatGptConnector ..> NormalizedConversation : emits
AiderConnector ..> NormalizedConversation : emits
PiAgentConnector ..> NormalizedConversation : emits
FactoryConnector ..> NormalizedConversation : emits
CopilotConnector ..> NormalizedConversation : emits
CopilotCliConnector ..> NormalizedConversation : emits
OpenClawConnector ..> NormalizedConversation : emits
CrushConnector ..> NormalizedConversation : emits
HermesConnector ..> NormalizedConversation : emits
KimiConnector ..> NormalizedConversation : emits
QwenConnector ..> NormalizedConversation : emits
Box<dyn Connector> instances that are unaware of each other's underlying file formats (JSONL, SQLite, specialized JSON).cass uses frankensqlite as the durable source of truth and frankensearch as a derived speed layer, powered by a suite of integrated "franken" libraries.
cass capabilities --json lists them as connectors.messages, conversations, agents) via frankensqlite — a pure-Rust SQLite reimplementation. cass turns on the engine's concurrent mode (PRAGMA fsqlite.concurrent_mode = ON), so a plain BEGIN runs as BEGIN CONCURRENT. Indexing still has one writer at a time, because index-run.lock admits a single indexer. BEGIN IMMEDIATE appears only in the daemon job queue and in logical-archive import/migrate. An experimental opt-in parallel persist path (CASS_INDEXER_BEGIN_CONCURRENT=1, off by default) exists but is not the default.title, content, agent, workspace, created_at.title_prefix and content_prefix use Index-Time Edge N-Grams (not stored on disk to save space) for instant prefix matching.flowchart LR
classDef pastel fill:#f4f2ff,stroke:#c2b5ff,color:#2e2963;
classDef pastel2 fill:#e6f7ff,stroke:#9bd5f5,color:#0f3a4d;
classDef pastel3 fill:#e8fff3,stroke:#9fe3c5,color:#0f3d28;
classDef pastel4 fill:#fff7e6,stroke:#f2c27f,color:#4d350f;
classDef pastel5 fill:#ffeef2,stroke:#f5b0c2,color:#4d1f2c;
subgraph Sources["Local Sources"]
A1[Codex]:::pastel
A2[Cline]:::pastel
A3[Gemini]:::pastel
A4[Claude]:::pastel
A5[OpenCode]:::pastel
A6[Amp]:::pastel
A7[Cursor]:::pastel
A8[ChatGPT]:::pastel
A9[Aider]:::pastel
A10[Pi-Agent]:::pastel
A11[Factory]:::pastel
A12[Copilot Chat]:::pastel
A13[Copilot CLI]:::pastel
A14[OpenClaw]:::pastel
A15[Clawdbot]:::pastel
A16[Vibe]:::pastel
A17[Crush]:::pastel
A18[Hermes]:::pastel
A19[Kimi]:::pastel
A20[Qwen]:::pastel
end
subgraph Remote["Remote Sources"]
R1["sources.toml"]:::pastel
R2["SSH/rsync\nSync Engine"]:::pastel2
R3["remotes/\nSynced Data"]:::pastel3
end
subgraph "Ingestion Layer"
C1["franken_agent_detection\nAuto-Discover & Scan\nNormalize & Dedupe"]:::pastel2
end
subgraph "Storage + Search"
S1["frankensqlite (WAL)\nSource of Truth\nBEGIN CONCURRENT\nMigrations"]:::pastel3
T1["frankensearch\nBM25 + Semantic\nRRF Fusion\nReranking"]:::pastel4
end
subgraph "Presentation"
U1["TUI (FrankenTUI)\nElm Architecture\nAnalytics Dashboard\nAsync Search"]:::pastel5
U2["CLI / Robot\nJSON Output\nAutomation"]:::pastel5
end
A1 --> C1
A2 --> C1
A3 --> C1
A4 --> C1
A5 --> C1
A6 --> C1
A7 --> C1
A8 --> C1
A9 --> C1
A10 --> C1
A11 --> C1
A12 --> C1
A13 --> C1
A14 --> C1
A15 --> C1
A16 --> C1
A17 --> C1
A18 --> C1
A19 --> C1
A20 --> C1
R1 --> R2
R2 --> R3
R3 --> C1
C1 -->|Persist| S1
C1 -->|Index| T1
S1 -.->|Rebuild| T1
T1 -->|Query| U1
T1 -->|Query| U2
cass index --watch, foreground): Uses file system watchers (notify) to detect changes in agent logs. When you save a file or an agent replies, cass re-indexes just that conversation. The TUI does not start a watcher on its own; see Keeping the Index Fresh below for what runs automatically.An index that is always a little behind is the most common complaint about any local search tool, so cass has three cooperating mechanisms. None of them block a search; all of them run cass index --background, which lowers its own CPU (nice 15) and I/O (ionice idle on Linux) priority before touching anything, and all of them respect the single index-run.lock — two indexers never run at once.
| Layer | What | When it runs | Enable |
|---|---|---|---|
| Stale-on-read catch-up | search, pack, and TUI launch check index freshness. If the index is stale (> 30 min), partial, or has pending sessions, a detached incremental cass index --background is spawned in its own process group and the current results are returned immediately. The next search is fresh. | On demand, at most once per 5 min per data dir (CASS_AUTO_REFRESH_COOLDOWN_SECS). Never for data dirs under the OS temp dir, and never for search --no-maintenance. A catch-up that ends without advancing the index is not respawned blindly: 1 h, then 6 h between attempts, and three failures trip the breaker until any run completes. | On by default. CASS_AUTO_REFRESH=0 disables globally. --robot-meta reports index_freshness.auto_refresh.{outcome,trigger,pid,consecutive_failures,detail}. |
OS scheduler (cass schedule install) | launchd LaunchAgents (macOS) or systemd user timers (Linux): an incremental job every 15 min and a nightly job (03:00) that performs a full source census with conditional lexical rebuilding, then one bounded models backfill --scheduled worker per tier (fast/hash always; quality/MiniLM when installed). Due remote-source syncs run first. Priority is delegated to the OS: launchd jobs set Nice=15 only, without ProcessType=Background or LowPriorityIO (background I/O throttling starved scheduled indexing on macOS). systemd units set Nice=19, IOSchedulingClass=idle and CPUSchedulingPolicy=idle. | On the timer, even when no cass process is running; survives reboots (Persistent=true / launchd). | cass schedule install [--interval-mins 15] [--nightly-hour 3] [--no-nightly] [--no-semantic] [--dry-run]; cass schedule status; cass schedule uninstall. |
| Resident daemon timer | The warm-model daemon (cass daemon, started by hand or auto-spawned by a human-mode semantic/hybrid search with --daemon; robot searches never spawn it) can also kick an incremental background index while it is resident. | Every CASS_DAEMON_INDEX_INTERVAL_SECS seconds while the daemon lives (it exits after its idle timeout). | Off by default; CASS_DAEMON_INDEX_INTERVAL_SECS=900 recommended. |
Idle awareness: scheduled work skips a run when the machine is under severe load (Linux /proc/loadavg + PSI; macOS sysctl vm.loadavg). On macOS you can additionally require the console to have been idle — CASS_RESPONSIVENESS_MIN_USER_IDLE_SECS=600 makes the nightly job and scheduled semantic backfill wait until nobody has touched the keyboard for ten minutes (the gate fails open where idle time is unavailable). Foreground cass index is never gated.
After an upgrade, the storage engine repairs and migrates an existing archive once, on its first writable open; that pass copies and rewrites the whole archive. Background runs (stale-on-read catch-up and scheduled jobs) never start it on an archive larger than CASS_INDEX_INTEGRITY_PREFLIGHT_MAX_BYTES (default 2 GiB): they exit 7 with kind migration-repair-pending, touch nothing, and cass schedule status names the cause. Run cass index --full in the foreground at a quiet time to perform it once; it keeps the original as a .pre-migration-bak copy, so plan for that much free space.
For a slow hosted disk, start with cass schedule install --interval-mins 60 and measure before shortening the interval. On Linux, cass index --json reports indexing_stats.bytes_written: the process block-write counter increase during indexing, including final checkpointing. It measures physical writes across all indexing layers, not just new transcript bytes or lexical segments; a cache-backed filesystem can report zero. The field is omitted when the counter is unavailable, including on other platforms. Check this alongside elapsed_ms on both changed-source and unchanged-source runs.
Nightly indexing retains index --full source coverage because timestamp-only connectors can miss restored files with old modification times. Connectors with valid durable source observations can reuse unchanged sources. When a completed checkpoint matches the archive and the lexical index passes validation, new messages are indexed inline. Missing or invalid checkpoint evidence, sparse or corrupt lexical assets, deferred lexical updates, and provenance repairs retain authoritative rebuilding from SQLite. Explicit cass index --full and --full --force-rebuild keep their existing repair behavior.
The nightly census still pays for source discovery and archive integrity, salvage, analytics, and FTS maintenance where required. A run that resumes an interrupted lexical rebuild can finish canonical recovery before returning; source discovery resumes on a later indexing run. Disappearing source files do not erase the preserved canonical history.
Each semantic worker retains its loaded model across its admitted batches and releases the previous batch's messages, vectors, storage handle, and lock at every checkpoint. CASS_SCHEDULE_MAX_BACKFILL_BATCHES bounds total attempts across tiers. Standalone cass models backfill --max-batches N uses the same worker; its default remains one batch.
Everything a scheduled job did is recorded under <data_dir>/schedule/ (state.json, runs.jsonl, per-job logs) and the last stale-on-read spawn under <data_dir>/auto-refresh-state.json / auto-refresh.log; cass schedule status --json reads all of it.
# See what would be registered, then register it
cass schedule install --dry-run
cass schedule install
# Run a job by hand (what the units invoke); --force ignores load/idle gates
cass schedule run --job incremental --json
cass schedule run --job nightly --force
# Inspect
cass schedule status --json
cass search "auth" --robot --robot-meta | jq '._meta.index_freshness.auto_refresh'
Stale-on-read catch-up handles an index that is behind. A search can also find the lexical index missing or unusable: a first run, a rebuild that never finished, a schema change. The search then has three choices: answer from what exists, rebuild before answering, or refuse. cass picks by the size of the job and by who is asking.
| Situation | What cass search does |
|---|---|
| A readable index exists but its checkpoint metadata is stale | Searches the existing index and leaves the heavy repair to an index run |
No usable index, and the archive is within the inline repair budget (CASS_INCREMENTAL_AUTHORITATIVE_LEXICAL_REPAIR_MAX_DB_BYTES, default 1 GiB, database plus WAL) | Rebuilds from SQLite inline, then answers. Robot callers get a bounded refusal instead when the ingest-quarantine circuit breaker is active |
| No usable index, and the archive is over that budget | Refuses with exit 5 maintenance-required and starts a detached cass index --full --background; the error hint names its pid |
| Robot caller, existing index whose incomplete checkpoint the cheap metadata refresh cannot reconcile | Refuses with exit 5 checkpoint_incomplete and starts the matching background run (--full above the size budget, plain cass index below it) |
| A rebuild is already running and no searchable generation exists | Robot callers get exit 7 index-busy immediately, with N of M conversations processed when the rebuild has recorded progress. Human callers wait up to CASS_SEARCH_ACTIVE_REBUILD_WAIT_MS (30 s) for it to publish |
Why the search never runs a large rebuild itself. A rebuild inside the
search process lives only as long as that process, and an agent's search is
almost always wrapped in a timeout: cass's own robot budget, or the agent
harness's command limit. On a large archive the rebuild commits nothing until
its first batch completes, so a killed rebuild keeps no progress. On one real
11 GB archive, a search-driven rebuild reached 160 of 4,324 conversations in
30 seconds (19 of them spent waiting on the in-flight byte budget) with
committed_offset still 0 when the search's budget ended it. The next search
started again from zero. Telling the agent to run cass index --full itself
failed the same way, because that command ran under the same timeout. The index
never converged. A detached child in its own process group survives the search,
so the rebuild finishes and the next search answers.
What the caller sees. The hint says what cass did and what to do next, so an agent never has to guess whether to run maintenance itself:
{"error": {"code": 5, "kind": "maintenance-required",
"message": "Automatic lexical repair was not started after detecting searchable lexical metadata missing: ...",
"hint": "cass started `cass index --full --json --background` as a detached process (pid 62704) to rebuild the search index. Retry this search after it finishes; `cass status --json` shows its progress under .rebuild. Do not run `cass index --full --json` yourself meanwhile: it would exit 7 (index-busy).",
"retryable": true}}
When no child is started, the hint says why: a run already holds the index
lock; a recent spawn is still inside its cooldown (it may still be starting, or
it failed, with the log path); earlier spawns failed and the breaker backed off
or tripped (with the failure detail); CASS_AUTO_REFRESH=0; or the spawn itself
failed. Those hints name the foreground command and warn that it needs a process
that is not killed by a short timeout.
Guard rails. The handoff reuses the stale-on-read machinery, so the same
limits apply: one spawner at a time (a file lock), the 5-minute cooldown, the
failure breaker (1 h, then 6 h, tripped after three failures), and the
index-run.lock that keeps two indexers from ever running together. Data dirs
under the OS temp dir and TUI_HEADLESS harnesses never spawn, and
search --no-maintenance never spawns anything. Under --timeout, a wait for an
active rebuild stops at nine tenths of the time remaining, so the caller
receives the index-busy verdict rather than an empty timed-out result.
Measured end to end (release build, an isolated data dir with 30 sessions,
lexical index moved aside, inline budget forced to one byte): the search
answered in 0.09 s with the spawned pid in its hint, the background index
published about 10 s later, and the next search reported all 60 matching
messages (50 returned at --limit 50). With CASS_AUTO_REFRESH=0 the same sequence never recovered within
180 s, which is the behaviour every large archive had before.
The interactive interface (src/ui/app.rs) uses FrankenTUI (ftui), a Rust TUI framework implementing the Elm architecture (Model-View-Update). The runtime handles terminal lifecycle, event polling, rendering, and cleanup.
CassMsg variant. The update() function produces Cmd effects (async tasks, ticks, quit).view() function renders the current state to an ftui Frame. The runtime diff engine minimizes terminal writes using Bayesian strategy selection.Cmd::Task, with results delivered as messages.graph TD
Input([User Input]) -->|Key/Mouse/Tick| Runtime
Runtime -->|CassMsg| Update[Model::update]
Update -->|Cmd| Runtime
Update -->|State Change| View[Model::view]
View -->|Frame| DiffEngine[Bayesian Diff]
DiffEngine -->|Minimal Writes| Terminal
Update -->|Cmd::Task| Background[Background Thread]
Background -->|Result Msg| Runtime
Data integrity is paramount. cass treats the SQLite database (src/storage/sqlite.rs, powered by frankensqlite) as the source of truth for conversations. History grows by insertion, and rows change only where the source changed:
forget and purge delete rows.UNIQUE(conversation_id, idx). New conversations are written with batched plain INSERTs, with the unique index as the backstop. Appends to a known conversation use INSERT OR IGNORE, so an agent re-writing a file cannot store a message twice. BLAKE3 content hashes are used only in memory, as merge fingerprints._schema_migrations table and a strict migration path keep upgrades safe and atomic; see Database Schema Migrations. Production creates a fresh database with one combined full_schema_v13 step and then applies v14–v21, so a new database records versions 13–21. The v1–v12 SQL is compiled only into tests.cass treats search indexes as derived assets. The SQLite archive is authoritative; lexical and semantic search data can be rebuilt from it.
Every lexical generation stores a schema_hash.json file containing the schema fingerprint:
{"schema_hash":"quill-fslx-schema-v9-hyphen-cjk-bigrams-bounded-content-prefix-preview-stored-content-external"}
| Scenario | Detection | Recovery |
|---|---|---|
| First run | No SQLite archive and no lexical index | cass index --full discovers sessions and creates both |
| Missing lexical index | No readable lexical asset | Rebuild from SQLite into scratch space, then publish |
| Schema mismatch | Hash differs from current | Rebuild derived lexical asset from SQLite |
| Corrupted metadata | Invalid or missing lexical metadata | Ignore the broken derivative and rebuild from SQLite |
| Semantic not ready | Model/vector assets absent or still backfilling | Continue lexical search and report semantic fallback/readiness |
# Check the current truth surface first
cass triage --json
cass health --json
cass status --json
# If not ready, run the first targeted command from recommended_commands[].
# For a fresh data dir this is usually:
cass index --full --json --no-progress-events --data-dir <same-data-dir>
Manual rebuild commands are for first setup, explicit operator refresh, or cases where recommended_commands[] asks for them. A normal missing/stale lexical asset should be repaired as derived state from SQLite, not treated as lost user data.
cass only reads agent files, never modifies themcass maintains multiple layers of redundancy to recover from corruption or schema changes:
Schema Hash Versioning:
Each lexical generation stores a schema_hash.json file containing a hash of the current schema definition. On startup:
This ensures that version upgrades with schema changes can rebuild the lexical derivative without user intervention.
Automatic Rebuild Triggers:
| Condition | Detection | Action |
|---|---|---|
| Schema version change | Hash mismatch in schema_hash.json | Full rebuild |
| Missing Quill publication manifest | Quill can't open index | Rebuild and publish a fresh derivative |
| Corrupted index files | Lexical reader open fails | Rebuild and publish a fresh derivative |
| Explicit request | --force-rebuild flag | Rebuild derived search assets from the canonical SQLite archive |
SQLite as Ground Truth: The SQLite database serves as the authoritative data store. Lexical rebuilds reconstruct the Quill index from SQLite:
// Iterate all conversations from SQLite
// Re-index each message into a fresh Quill index
// Progress tracked via IndexingProgress for UI feedback
This means corrupted lexical data is a repairable derivative-state problem. Operators should start with cass triage --json for the exact next command, or read cass health --json / cass status --json for the narrower readiness snapshot.
The SQLite database uses 21 versioned schema migrations, tracked in the _schema_migrations table (CURRENT_SCHEMA_VERSION = 21 and MIGRATION_NAMES in src/storage/sqlite.rs):
| Version | Migration | Version | Migration |
|---|---|---|---|
| 1 | core_tables | 12 | model_dimensions |
| 2 | fts_messages | 13 | plan_token_rollups |
| 3 | fts_messages_rebuild | 14 | fts_contentless |
| 4 | sources | 15 | conversation_tail_state_cache |
| 5 | provenance_columns | 16 | drop_redundant_message_conv_idx |
| 6 | source_path_index | 17 | drop_message_created_idx |
| 7 | msgpack_columns | 18 | conversation_tail_state_hot_table |
| 8 | daily_stats | 19 | conversation_external_lookup |
| 9 | embedding_jobs | 20 | conversation_external_tail_lookup |
| 10 | token_analytics | 21 | conversation_context_index (current) |
| 11 | message_metrics |
Migration Process:
cass checks _schema_migrations in the database (older databases that still record schema_version in the meta table are transitioned automatically)Safe Files (never deleted during rebuild):
bookmarks.db - Your saved bookmarkstui_state.json - UI preferencessources.toml - Remote source configuration.env - Environment configurationBackup and Retention Policy: Migration/rebuild backups preserve user data and are not treated as disposable source evidence. Derived lexical publish backups use the bounded retention policy documented above, while quarantined artifacts and repair candidates persist until an operator runs an explicit, fingerprinted cleanup flow.
The --watch flag enables real-time index updates as agent files change.
File change detected
↓
[2 second debounce window] ← Accumulate more changes
↓
[5 second max wait] ← Force flush if changes keep coming
↓
Re-index affected files
Each file system event is routed to the appropriate connector:
~/.claude/projects/foo.jsonl → ClaudeCodeConnector
~/.codex/sessions/rollout-*.jsonl → CodexConnector
~/.aider.chat.history.md → AiderConnector
Watch mode maintains watch_state.json, one scan watermark (ms) per connector under a short connector code (cd Claude, cx Codex, gm Gemini, ...):
{"v":1,"m":{"cd":1699900000000,"cx":1699900000000}}
The global last_scan_ts and last_indexed_at watermarks live in the SQLite meta table, not in this file.
Codex event_msg token_count usage is attached to the nearest preceding assistant turn during indexing.
If you indexed Codex sessions before this behavior existed, backfill usage coverage with:
cass index --full
cass analytics rebuild --track a
cass analytics rebuild re-derives the Track A rollups (message_metrics,
usage_hourly, usage_daily, usage_models_daily) from messages already in
the archive; it never re-parses raw session files. On a large archive a full
rebuild is a long single-core job, so daily refreshes should be windowed:
# Full rebuild (every rollup row dropped and recomputed)
cass analytics rebuild
# Only recompute the last two UTC days; older rollups are left untouched
cass analytics rebuild --days 2
cass analytics rebuild --since -2d # same window, relative syntax
cass analytics rebuild --since 2026-08-20 # from a date
The window is widened to the start of the UTC day containing the cutoff,
because rollups are bucketed by day and hour. Progress is logged per 10k
messages (analytics_rebuild_progress). Across analytics commands, --days
and --since are mutually exclusive, and malformed or reversed time bounds
return a usage error instead of silently running an unfiltered query.
--until, --agent, --workspace
and --source are query-time filters and are rejected here rather than
silently ignored. --track b also rejects --since/--days; with --track all, the window applies to Track A while Track B still rebuilds the complete
token_usage ledger. cass analytics validate likewise rejects every query
filter because its invariant checks always cover the complete analytics
database.
The TUI analytics dashboard never rebuilds rollups in-process: when rollups are
missing it spawns a detached cass analytics rebuild child, logs it to
<data_dir>/analytics-rebuild.log, and reports the pid in the status line;
reopen the dashboard once the rebuild finishes.
Generate tab-completion scripts for your shell.
Bash:
cass completions bash > ~/.local/share/bash-completion/completions/cass
# Or: cass completions bash >> ~/.bashrc
Zsh:
cass completions zsh > "${fpath[1]}/_cass"
# Or add to ~/.zshrc: eval "$(cass completions zsh)"
Fish:
cass completions fish > ~/.config/fish/completions/cass.fish
PowerShell:
cass completions powershell >> $PROFILE
search, index, stats, etc.)--robot, --agent, --limit)SIGILL hazard (the historical ONNX Runtime dependency was removed in cass#308).cargo install --git https://github.com/Dicklesworthstone/coding_agent_session_search. This requirement exists because CI builds target ubuntu-24.04 to access newer kernel features used by the frankensqlite storage engine. The install script probes the host's glibc (ldd --version) before downloading a Linux prebuilt binary and falls back to build-from-source with a warning when it is older than 2.38; --from-source forces that route, and --artifact-url bypasses the probe for an explicitly chosen artifact.Recommended: Homebrew (Apple Silicon macOS + Linux)
brew install dicklesworthstone/tap/cass
# Update later
brew upgrade cass
The Homebrew tap installs prebuilt release tarballs (not bottles) for Linux and Apple Silicon macOS. On Intel macOS, use the install script with --from-source.
Windows: Scoop
scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cass
Alternative: Install Script
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/coding_agent_session_search/main/install.sh?$(date +%s)" \
| bash -s -- --easy-mode --verify
Alternative: GitHub Release Binaries
SHA256SUMS.txt against the downloaded archive.cass into your PATH.Example (Linux x86_64, replace VERSION with an explicit release tag):
VERSION=v0.2.0 # e.g. v0.2.0
curl -L -o cass-linux-amd64.tar.gz \
"https://github.com/Dicklesworthstone/coding_agent_session_search/releases/download/${VERSION}/cass-linux-amd64.tar.gz"
curl -L -o SHA256SUMS.txt \
"https://github.com/Dicklesworthstone/coding_agent_session_search/releases/download/${VERSION}/SHA256SUMS.txt"
sha256sum -c SHA256SUMS.txt
tar -xzf cass-linux-amd64.tar.gz
install -m 755 cass ~/.local/bin/cass
cass
On first run, cass starts a full index in the background (a detached, low-priority cass index --full that keeps going if you quit) and shows its progress in the status line. Search goes live, without a restart, as soon as the first index is published. Until then there are no results; if automatic indexing is off (CASS_AUTO_REFRESH=0) or cannot start, the status line says so and names cass index --full.
foo* (prefix), *foo (suffix), or *foo* (contains) for flexible matching.Up/Down to select, Tab (or Alt+l) to focus the detail pane. Ctrl+N/Ctrl+Shift+N step through query history; Ctrl+R cycles it.F3: Filter by Agent (e.g., "codex").F4: Filter by Workspace/Project.F5/F6: Time filters (Today, Week, etc.).F2: Next theme (Shift+F2 previous; 19 presets).F12: Cycle ranking mode (recent → balanced → relevance → quality → newest → oldest).Ctrl+B: Toggle rounded/square borders.Enter: Open selected result in contextual detail modal (defaults to Messages tab).Enter with no selected hit: submit query behavior (no-op if empty).F8: Open selected hit in $EDITOR.Ctrl+Enter: Add current result to queue (multi-open).Ctrl+O: Open all queued results in editor.Ctrl+X: Toggle selection on current item (Ctrl+M opens the detail modal, like Enter).Alt+B: Bulk actions menu (when items selected).Ctrl+Y / Alt+Y / Ctrl+Shift+C: Copy file path / snippet / content to clipboard./: Find text within detail pane; Enter advances matches; n/N cycle contextual session hits; Esc closes the modal.Ctrl+Shift+R: Trigger manual re-index (refresh search results).Ctrl+Shift+Del: Reset TUI state (clear history, filters, layout).Aggregate sessions from your other machines into a unified index:
# Add a remote machine
cass sources add user@laptop.local --preset macos-defaults
# Sync sessions from all sources
cass sources sync
# Check source health and connectivity
cass sources doctor
See Remote Sources (Multi-Machine Search) for full documentation.
The cass binary supports both interactive use and automation.
# Interactive
cass [tui] [--data-dir DIR] [--once] [--asciicast FILE]
# Indexing
cass index [--full] [--watch] [--background] [--data-dir DIR] [--idempotency-key KEY]
cass schedule install [--interval-mins 15] [--nightly-hour 3] [--no-semantic] [--dry-run]
cass schedule status --json
# Search
cass search "query" --robot --limit 5 [--timeout 5000] [--explain] [--dry-run]
cass search "error" --robot --aggregate agent,workspace --fields minimal
cass pack "query" --robot --max-tokens 12000 [--limit 40] [--sessions-from FILE|-]
cass pack "query" --robot --freshness-policy strict --freshness-window-seconds 604800 --require-evidence
cass pack "query" --robot --max-tokens 4000 --max-evidence 8 --max-sessions 3 --max-excerpt-chars 600
# Inspection & Health
cass triage --json # One-shot agent preflight with exact next command
cass status --json # Quick health snapshot
cass health # Minimal pre-flight check (<50ms)
cass capabilities --json # First-stop agent self-description
cass introspect --json # Full API schema
cass swarm status --json # Read-only Beads/Agent Mail/git/rch swarm snapshot
cass swarm work-packet --json # Advisory claim packet; no mutations
cass swarm lint --json # Coordination and proof-gap lint
cass context /path/to/session --json # Find related sessions
cass view /path/to/file -n 42 --json # View source at line
# Session Analysis
cass export /path/to/session --format markdown -o out.md # Export conversation
cass expand /path/to/session -n 42 -C 5 --json # Context around line
cass timeline --today --json # Activity timeline
# Remote Sources
cass sources add user@host --preset macos-defaults # Add machine
cass sources sync # Sync sessions
cass sources doctor # Check connectivity
cass sources mappings list laptop # View path mappings
# Utilities
cass stats --json
cass completions bash > ~/.bash_completion.d/cass
| Command | Purpose |
|---|---|
cass (default) | Start TUI (a stale index triggers a detached low-priority catch-up; see Keeping the Index Fresh) |
cass tui --asciicast FILE | Run TUI and save terminal output as asciicast v2 |
index --full | Discover sessions and refresh the canonical DB plus derived search assets |
index --background | Same as index, but lowers its own CPU/I/O priority first (used by auto-refresh, schedule, and the daemon timer) |
index --watch | Foreground watch loop: reindex automatically on file changes |
schedule install|uninstall|status|run | Register incremental (15 min) + nightly (full source census, conditional lexical rebuild, bounded semantic backfill) jobs with launchd / systemd user timers |
search --robot | JSON output for automation pipelines |
pack --robot | Deterministic cited answer packs for agent/human handoffs; reports health, freshness, privacy, and warnings |
triage / ready / preflight | One-shot agent preflight: readiness, exact next command, docs, schemas, workflows, and recoveries |
guide [INTENT] | Intent-to-command planner (fix-ci, investigate-search-miss, prepare-release, repair-assets, export-session, onboard-source, support-capsule); dry-run by default. With --apply, 8 of its 19 allowlisted operations have evaluating proof adapters; the other 11 are observation-only and are reported, not evaluated |
status / state | Health snapshot: index freshness, DB stats, recommended action |
health | Minimal health check (<50ms on a healthy archive; the strict, mutation-free owner-thread probe shared with status has a 30 s hard deadline and never checkpoints a dirty WAL), exit 0=healthy, 1=unhealthy |
selftest | Archive-independent executable probe for installers and binary-promotion gates; exercises an in-memory FrankenSQLite write/read round-trip |
capabilities | First-stop agent self-description: workflow recipes, mistake recoveries, commands, global flags, exit codes, env vars, and limits |
introspect | Full API schema: commands, arguments, response shapes |
swarm status --json | Read-only shared-repo operations snapshot across Beads, Agent Mail metadata, git, build pressure, cass readiness, and proof refs |
swarm work-packet --json | Advisory one-agent packet with readiness, suggested reservations, verification commands, and closeout checklist; it does not claim or reserve |
swarm lint --json | Read-only coordination protocol lint for missing mail, stale reservations, status mismatches, and proof gaps. Only fixture input (--fixture, --fixture-dir --fixture-id) is linted today; the live path reports every provider live-provider-unimplemented and finds nothing |
swarm dependency-drift --json | Read-only sibling dependency sentinel for Cargo.toml pins, optional local checkout HEAD/dirty state, strict validation commands, and release-risk recommendations |
sessions [--workspace DIR] [--current] | Discover recent session files for follow-up actions |
context <path> | Find related sessions by workspace, day, or agent |
view <path> -n N | View source file at specific line (follow-up on search) |
export <path> | Export conversation to markdown/JSON |
export-html <path> | Export as self-contained HTML with optional encryption |
expand <path> -n N | Show messages around a specific line number |
timeline | Activity timeline with grouping by hour/day |
sources | Manage remote sources: add/list/remove/doctor/sync/mappings |
doctor | Diagnose and repair installation issues (safe, never deletes data) |
Other subcommands (all present in the Commands enum in src/lib.rs):
| Command | Purpose |
|---|---|
pages | Export an encrypted, searchable static-site archive with GitHub Pages / Cloudflare Pages deploy; runs the interactive wizard by default, with --export-only DIR, --verify BUNDLE, --preview BUNDLE, and --scan-secrets as non-wizard modes. --share-profile public|team|personal (config: bundle.share_profile; wizard: step 4) redacts every exported text, including titles, paths, metadata values, message bodies, their search indexes and snippets. public covers home paths, usernames, project names, hostnames, emails, phone numbers, IP addresses, social-security and card numbers, and internal URLs; team covers home paths, emails, social-security and card numbers; personal rewrites nothing. No profile rewrites a credential (API keys, private keys, connection strings): the staged secret scan rejects an export that contains one, so it fails closed instead of publishing a rewritten secret. The default is public for a plaintext export and team when encrypted. The bundle's export_meta records the profile and per-kind redaction counts, never the values. For a plaintext bundle, --verify (which every export also runs before it reports success) reads the declared profile and rescans every exported text surface, including both search indexes, with that profile's rules. It fails if any value the profile removes remains, reporting counts per surface. The home-path and username rules use the verifying account's home directory. An encrypted bundle keeps its profile inside the payload, which --verify does not open. |
pages key list|add-password|add-recovery|revoke|rotate --archive BUNDLE | Manage the key slots of an exported encrypted bundle (LUKS-style: several independently wrapped copies of one data key). Passwords come from an interactive prompt or --password-stdin (current password on line 1, new password on line 2), never from argv; --json for automation; recovery secrets are printed once and never stored. See docs/RECOVERY.md |
upgrade | Check for a newer release and optionally run the same checksum-verified installer the TUI uses (--check, --yes, --force) |
man | Generate the man page to stdout |
storage | On-disk storage footprint by component (DB, WAL, lexical index, raw mirror, semantic, quarantine) |
dedup | Collapse pre-existing duplicate conversation rows (projects/<rel> vs <rel> external-id twins); dry-run unless --apply |
support-bundle | Assemble a redacted, share-safe recovery/support evidence bundle |
state | Quick state/health check (alias of status) |
onboarding | Read-only first-run source onboarding + readiness wizard; --json for scripts, never launches the TUI |
quarantine | Inspect and manage the conversation-ingest quarantine (list / clear) |
forget | Prune already-indexed conversations by source-path glob; dry-run by default, --apply to commit. Deletes the canonical rows, then rebuilds FTS, analytics and the lexical index. Semantic vectors are not rewritten, so explicit semantic search then reports semantic-unavailable and hybrid falls back to lexical until cass index --semantic re-embeds from the canonical rows; no surface returns forgotten messages. dedup --apply and sources agents exclude behave the same way. It removes indexed copies, not source files: forget records each source's size and modification time, so an unchanged source stays forgotten across cass index, --full and rescans triggered by other sessions, but if its agent appends to it the whole conversation is indexed again. A store that keeps many sessions in one file (a SQLite database) counts as changed when any of its sessions changes. Delete or move the file to keep it out for good. The raw mirror keeps its verbatim capture of the source; once the source is gone, cass mirror prune --older-than 0s --safety-hold-down 0s --source-path '<glob>' --apply removes that capture too (a source still on disk is captured again by the next scan). A cass serve session opened before the forget keeps its pinned snapshot until reload; the TUI drops forgotten hits on its next search. If the lexical rebuild fails, forget exits 5 (lexical-rebuild) and cass index --full finishes the purge |
fleet upgrade-rehearsal | Fleet-safe upgrade rehearsal (dry run) with bounded post-upgrade verification; --live opts in to SSH probes of configured remotes |
lessons list|search | Mine and query durable, redacted lessons from local evidence (commits, closed beads, proof manifests) |
import chatgpt | Split a ChatGPT web export (conversations.json) into files the ChatGPT connector can index |
release-verify | Verify release distribution channels (GitHub, Homebrew, Scoop, crates.io, installer) from a recorded observation (--from) or live (--live) |
sources discover | Auto-discover SSH hosts from ~/.ssh/config |
sources reingest | Re-ingest an already-synced mirror into the canonical archive without re-running rsync |
sources artifact-manifest | Build or verify a lexical-artifact evidence manifest for remote exchange |
Two more commands are dispatched before the main parser (so they are absent from cass --help, introspect and completions); use their own --help:
| Command | Purpose |
|---|---|
serve --stdio [--mcp] (--data-dir DIR | --index PATH) | Persistent search service: reuses one read-only lexical index reader across a stream of newline-delimited JSON requests (or MCP with --mcp), with per-request deadlines (--request-timeout-ms), cross-process reader admission (--admission-dir, --admission-slots) and a resident-memory cap (--max-resident-mib). Search, status and startup never open the canonical database; --db opts in to canonical view/view_batch requests. Operations: search, semantic_search, refine, view, view_batch, status, reload, unload, shutdown. search is lexical; semantic_search needs --semantic-embedder minilm|multilingual-minilm|hash plus --data-dir and --db, and refine needs --reranker-model PATH (installed local assets only, never downloaded). See docs/SEARCH_SERVICE.md |
archive export|verify|search|view|import | Bounded, versioned logical archive of the canonical rows: export --output FILE --archive-id ID --include-private streams a read-only snapshot to a new private JSONL file; verify checks framing, identities, counts and digest without a database; search/view read verified message bodies without restoring; import restores into a new database and never replaces an existing archive |
| Tool | Purpose |
|---|---|
cass tui --asciicast FILE | Record TUI output as an asciicast v2 artifact; there is no separate cass cast subcommand |
scripts/bakeoff/cass_validation_e2e.sh | Run the bake-off validation harness for lexical, semantic, hybrid, and reranked search scenarios |
scripts/bakeoff/cass_embedder_e2e.sh | Exercise embedder bake-off flows against a generated validation corpus |
scripts/bakeoff/cass_rerank_e2e.sh | Exercise reranker bake-off flows and append results to the bake-off log |
Commands for troubleshooting, debugging, and understanding system state:
# One-shot agent preflight
cass triage --json
# → { "surface": "triage", "status": "healthy", "next_command": null, ... }
# Health check (fast, <50ms; the archive probe is bounded by a 30 s hard deadline)
cass health --json
# → { "healthy": true, "index_age_seconds": 120, "message_count": 5000 }
# Detailed status with recommendations
cass status --json
# → Includes index freshness, staleness threshold, recommended action
# System diagnostics
cass diag --verbose --json
# → Database stats, index info, connector status, environment
# Query explanation (debug why results are what they are)
cass search "auth" --explain --dry-run --robot
# → Shows parsed query, index strategy, cost estimate without executing
# Find related sessions
cass context /path/to/session.jsonl --json
# → Sessions from same workspace, same day, or same agent
# Archive-first diagnostic check
cass doctor check --json
# → Read-only checks for archive coverage, source authority, locks, backups,
# storage pressure, semantic fallback, and recommended next action
# Fingerprinted repair plan and apply
cass doctor repair --dry-run --json
cass doctor repair --yes --plan-fingerprint <plan_fingerprint> --json
# → Builds candidates and applies only the inspected matching fingerprint
# Legacy safe auto-run for low-risk derived repairs
cass doctor --fix --json
# → Emits operation_outcome and receipts; fails closed on archive/source risk
cass doctor is a comprehensive diagnostic and repair tool designed for troubleshooting installation and data issues. Its current recovery model is archive-first: preserve cass-owned evidence, prove source authority and coverage, then repair through candidates and receipts. The full operator runbook is docs/planning/RECOVERY_RUNBOOK.md.
What it checks:
| Surface | Purpose | Mutation Policy |
|---|---|---|
cass doctor check --json | Read-only truth surface for archive coverage, source authority, locks, storage pressure, semantic fallback, and recommended action | Never mutates |
cass doctor archive-scan --json | Read-only source inventory, raw mirror, coverage, sole-copy, and remote sync gap inspection | Never mutates |
cass doctor repair --dry-run --json | Builds a fingerprinted repair plan and candidate/promotion gates | Read-only plan |
cass doctor repair --yes --plan-fingerprint <fp> --json | Applies exactly the inspected repair fingerprint | Candidate-based, receipt-backed |
cass doctor backups list/verify/restore ... --json | Lists backups, verifies manifests, rehearses restore, then applies by fingerprint | Restore apply requires a matching rehearsal fingerprint |
cass doctor --repair-leaked-pages --dry-run --json | Classifies the engine's integrity verdict: clean, leaked_pages (pages no table, index or freelist owns — page N is never used), or other_damage | Read-only |
cass doctor --repair-leaked-pages --yes --json | Frees leaked pages in place when they are the only damage; re-runs integrity_check and compares conversation/message counts. Any other damage exits 5 data-corruption untouched | Backs up the live bundle first as a leaked-pages-repair backup (cass doctor backups restore <id>) |
cass doctor cleanup --json | Plans cleanup for derived or explicitly reclaimable assets | Apply requires a matching fingerprint |
cass doctor support-bundle --json | Creates a scrubbed diagnostic handoff bundle | Redacted by default; not a backup |
Safety guarantees:
sources.toml are not cleanup targets.plan_fingerprint.Recommended support checklist:
cass doctor check --json
cass doctor baseline diff <baseline_id> --json
cass doctor support-bundle --json
cass doctor support-bundle verify <bundle_or_manifest_path> --json
Send the doctor JSON, latest failure_context.json if present, support-bundle
manifest.json, any baseline diff, relevant artifact_manifest_path and
event_log_path values, and the exact command/exit code. Do not attach raw
sessions, full SQLite archives, private source files, or encrypted payloads
unless the user explicitly opts into sensitive evidence attachment.
Diagnostic Flags:
| Flag | Available On | Effect |
|---|---|---|
--explain | search | Show query parsing and strategy |
--dry-run | search | Validate without executing |
--verbose | most commands | Extra detail in output |
--trace-file | all | Append execution trace to file |
--robot-trace-ingest | index | Emit per-ingest-batch NDJSON timing and lookup counters on stderr |
Commands for managing the semantic search ML model:
# Check current model status (abbreviated schema — real output also
# includes cache_lifecycle, files[], revision, license, and more):
cass models status --json
# → {
# "model_id": "all-minilm-l6-v2",
# "model_dir": "~/.local/share/coding-agent-search/models/all-MiniLM-L6-v2",
# "installed": false,
# "state": "not_acquired",
# "state_detail": "model not acquired (user consent required); missing ...",
# "next_step": "Run `cass models install`, or keep using lexical search.",
# "lexical_fail_open": true,
# "revision": "c9745ed1...",
# "license": "Apache-2.0",
# "total_size_bytes": 90872535,
# "installed_size_bytes": 0,
# "observed_file_bytes": 0,
# "policy_source": "semantic_policy"
# }
# Install model (downloads ~90MB from Hugging Face on explicit request)
cass models install
# → Downloads from Hugging Face, verifies checksum
# Install from local directory (air-gapped environments)
cass models install --from-file /path/to/model-dir
# Verify model integrity
cass models verify --json
# → all_valid bool + per-file SHA-256 checks (see `cass models verify --help`)
# Check for model updates
cass models check-update --json
# → { "update_available": bool, "reason": str,
# "current_revision": str|null, "latest_revision": str }
In cass status --json, semantic.preferred_backend is "fastembed" when the native MiniLM lane is selected and "hash" for the hash tier; fastembed is only the id of the native pure-Rust MiniLM lane — no ONNX runtime is involved.
Model Files (stored in $CASS_DATA_DIR/models/all-MiniLM-L6-v2/):
model.safetensors - The neural network weights (~90MB)tokenizer.json - Vocabulary and tokenization rulesconfig.json - Model configurationspecial_tokens_map.json - Special token definitionstokenizer_config.json - Tokenizer settingsVerified Install: The installer enforces SHA256 checksums.
Sandboxed Data: All indexes/DBs live in standard platform data directories (~/.local/share/coding-agent-search on Linux).
Read-Only Source: cass never modifies your agent log files. It only reads them.
cass uses crash-safe atomic write patterns throughout to prevent data corruption:
TUI State Persistence (tui_state.json):
1. Serialize state to JSON
2. Write to temporary file (tui_state.json.tmp)
3. Atomic rename: temp → final
If a crash occurs during step 2, the original file is untouched. The rename operation (step 3) is atomic on all modern filesystems—it either completes fully or not at all.
ML Model Installation (models/all-MiniLM-L6-v2/):
1. Download to temp directory (models/all-MiniLM-L6-v2.tmp/)
2. Verify all checksums
3. If existing model present: rename to backup (models/all-MiniLM-L6-v2.bak/)
4. Atomic rename: temp → final
5. On success: remove backup
6. On failure: restore from backup
This backup-rename-cleanup pattern ensures that either the old model or new model is always available—never a half-installed state.
Configuration Files (sources.toml, watch_state.json):
All configuration writes follow the same temp-file-then-rename pattern, ensuring consistency even during power loss or unexpected termination.
Why This Matters:
Agent transcripts are full of credentials: keys pasted into prompts, tokens
echoed by tool output, .env files read into context. With the default
CASS_INDEX_REDACTION=full, cass scrubs them from every persisted message,
title, snippet and metadata blob before anything reaches SQLite or the lexical
index, so search results, exports and robot output never repeat them. (The
original session files and the raw-mirror blobs keep the raw text on the same
disk; redaction protects the queryable surfaces, not disk-at-rest secrecy.)
What is recognized. Thirteen pattern families: AWS access key IDs, AWS
secret keys and session tokens in assignment context, GitHub tokens (classic and
fine-grained), OpenAI and Anthropic API keys, Bearer tokens, JWTs, PEM/OpenSSH/PGP
private-key blocks, database connection URLs (Postgres, MySQL, MongoDB, Redis,
AMQP; the whole URL, since it may carry a password), generic password= /
api_key: / secret= style assignments, Slack tokens, and Stripe live keys.
JSON metadata is also redacted by field name: values under keys such as
passphrase, authorization, cookie, or anything ending in password,
token, secret or apikey are replaced whatever they look like.
How it runs. Redaction sits on the ingest hot path, since it touches every message, so it is built in three layers:
RegexSet.
One scan per string reports which patterns could match; a string with no
candidates (the vast majority) is returned untouched without allocating.replace_all pass, in a fixed order, replacing each match with [REDACTED].CASS_REDACT_MEMO_CAPACITY sets the entry cap).A frozen copy of the original algorithm lives in the test suite, and a 512-case property test requires the plain path and the memoized path (on both a cache miss and a cache hit) to produce byte-identical output to it. The memoized JSON path has its own equivalence test against the uncached one over nested shapes.
Why the patterns use ASCII word boundaries. Rust's regex crate runs a
RegexSet on a fast lazy DFA, but a Unicode word boundary (\b) is something
that DFA cannot evaluate once the haystack contains a non-ASCII byte. It then
falls back to the PikeVM, a far slower NFA simulation. Real transcripts are full
of non-ASCII text (emoji, CJK, box-drawing characters in tool output), so every
such message paid the slow path. The patterns now use (?-u:\b), the ASCII
word boundary, which the DFA handles.
This cannot weaken redaction. Every boundary in these patterns sits next to an
ASCII word character, and ASCII word characters are a subset of Unicode word
characters. So wherever a Unicode boundary exists, an ASCII boundary exists too,
and the ASCII patterns match everything the Unicode ones did. The only
difference is a token glued directly to a non-ASCII letter (凭据ghp_…): a
Unicode boundary sees two word characters and no boundary, so the old patterns
missed that token, while the ASCII form redacts it.
Measured on 593 MiB of real session text (3.1 million strings, 79,870 of them containing non-ASCII), with the same regex version cass ships:
| Word boundary | Prefilter throughput | Pattern matches |
|---|---|---|
Unicode \b | 9.8-10.6 MiB/s | baseline |
ASCII (?-u:\b) | 261-284 MiB/s | identical on all 3.1M strings |
In an A/B incremental index on two identical clones of a 2.1-million-message
archive (10 minutes each, run one after the other), the ASCII build committed
53,155 new messages on 10.1 CPU-minutes, against 7,028 on 12.8 CPU-minutes for
the Unicode build: about 9.5 times the ingest throughput per CPU second. A unit
test rejects any secret pattern that reintroduces a Unicode \b, so the cliff
cannot quietly come back.
The project ships with a robust installer (install.sh / install.ps1) designed for CI/CD and local use:
Checksum Verification: Validates artifacts against a .sha256 file or explicit --checksum flag.
Rustup Bootstrap: Source installs use the dated nightly and components pinned by the release's rust-toolchain.toml. The installer bootstraps rustup without an unrelated default toolchain when needed.
Easy Mode: --easy-mode automates installation to ~/.local/bin without prompts.
Platform Agnostic: Detects OS/Arch (Linux/macOS/Windows, x86_64/arm64) and fetches the correct binary.
cass includes a built-in update checker that notifies you when new versions are available, without interrupting your workflow.
When a new version is available, a one-line banner appears at the top of the TUI:
Update v<current> -> v<latest> | Alt+U upgrade | Alt+N notes | Alt+I ignore | Esc dismiss
Confirming Alt+U runs the same verified installer used for initial installation:
macOS/Linux:
curl -fsSL https://...install.sh | bash -s -- --easy-mode --verify
Windows (PowerShell):
`
Truncated — view the full README on GitHub.
(top 24 of 30)
530 followers · starred Jul 2026
694 followers · starred Jan 2026
69 followers · starred Jan 2026
1,125 followers · starred Jan 2026
Rust
95.3%
JavaScript
1.5%
Shell
1.5%
Unified TUI and CLI to index and search your local coding agent session history across 11+ providers (Codex, Claude, Gemini, Cursor, Aider, etc.)
Rust
1,158
5,529 commits
updated Oct 2, 2026
Unified, high-performance TUI to index and search your local coding agent history. Aggregates sessions from Codex, Claude Code, Gemini CLI, Cline, OpenCode, Amp, Cursor, ChatGPT, Aider, Pi-Agent, Prime Agent, Oh My Pi, GitHub Copilot Chat, Copilot CLI, OpenClaw, Clawdbot, Vibe, Crush, Goose, Hermes, Kimi Code, Muse Code, Qwen Code, Factory (Droid), Antigravity, OpenHands, Grok Build, Grok Bot, Codebuff/Freebuff, Devin CLI, Shelley, and Kiro CLI into a single, searchable timeline.
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/coding_agent_session_search/main/install.sh?$(date +%s)" \
| bash -s -- --easy-mode --verify
# Windows (PowerShell)
& ([scriptblock]::Create((irm "https://raw.githubusercontent.com/Dicklesworthstone/coding_agent_session_search/main/install.ps1"))) -EasyMode -Verify
Installs the latest release by default. Pass --version <tag> / -Version <tag> to pin a specific version.
Or via package managers:
# Homebrew (Apple Silicon macOS + Linux)
brew install dicklesworthstone/tap/cass
# Windows (Scoop)
scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cass
The Homebrew tap installs prebuilt release tarballs (not bottles) for Linux and Apple Silicon macOS. On Intel macOS, use the install script with --from-source.
⚠️ Never run bare cass in an agent context — it launches the interactive TUI. Always use --robot or --json.
# 1) Check the installed interface once per version (recipe verified on 0.8.0).
cass --version
cass search --help
# Verify a newly installed executable without opening the configured archive.
cass selftest --json
# `health --binary-only` still reports (and therefore probes) archive readiness.
# 2) For a quick history question, start with scoped read-only lexical retrieval.
# Hybrid remains the product default; lexical is explicit for this workflow.
cass search "performance regression" --workspace /path/to/project --days 7 \
--mode lexical --no-maintenance --robot --robot-meta --fields minimal \
--limit 5 --max-tokens 2000 --timeout 2000
# 3) Find the current or recent session for this workspace
cass sessions --current --json
cass sessions --workspace "$(pwd)" --json --limit 5
# 4) View + expand a hit (use source_path/line_number from search output)
cass view /path/to/session.jsonl -n 42 -C 3 --json --timeout 2000
# 5) Discover the full machine API
cass capabilities --json
cass robot-docs guide
cass robot-docs schemas
# 6) Exclude a noisy agent harness from future indexing
cass sources agents list --json
cass sources agents exclude openclaw
cass sources agents include openclaw
The retrieval flags above are available in 0.8.0. On older builds, check help;
if --no-maintenance is absent, report the mismatch instead of dropping the
read-only constraint. --timeout is in milliseconds, while --max-tokens limits
approximate output size. Also set a caller-side deadline (for example, GNU
timeout 10s); an externally interrupted command may leave incomplete JSON.
Inspect budget.timed_out even after exit 0: timed-out empty hits are not proof
that no history exists. A maintenance-required response ends the retrieval
attempt; indexing or repair is a separate mutating task. Use triage/health/status
for readiness diagnosis, not as repeated prerequisites to a short summary.
Broaden scope deliberately, expand useful hits, and preserve source/line citations.
view -C bounds context lines, not bytes; check excerpt size before including
a long JSONL record in an agent prompt.
Output conventions
Search asset contract
--robot --robot-meta) reports the requested mode, realized mode, semantic refinement status, and any lexical fallback reason when semantic assets are not ready.cass models install downloads the default all-minilm-l6-v2 (alias minilm, ~90 MB) only on explicit request; --model multilingual-minilm selects the larger multilingual MiniLM L12 model (~480 MB) for CJK/mixed-language archives. Cass never auto-downloads or auto-selects the multilingual space. Air-gapped installs use --from-file <dir>. While the selected model is absent, hybrid search uses lexical-only and reports fallback_mode="lexical" in health/status.cass triage --json combines readiness, next_command, recommended_commands[], docs/schema pointers, starter workflows, and accepted recoveries for diagnosis. Review recommended mutations before executing them. cass health --json and cass status --json remain the narrower truth surfaces for readiness, active rebuilds, and recovery.Lexical publish durability (atomic-swap)
src/indexer/mod.rs::publish_staged_lexical_index.<data_dir>/index/.lexical-publish-backups/<dated>/ for a bounded retention window. Default cap is 1 (keep just the most-recent prior generation for one-step rollback); override via the CASS_LEXICAL_PUBLISH_BACKUP_RETENTION env var (0 disables retention entirely, higher N keeps deeper history). Pruning runs after every successful publish and emits structured tracing::info! events with freed_bytes + retention_limit for observability.recover_or_finalize_interrupted_lexical_publish_backup at the start of the next lexical publish or rebuild (not at process startup), which moves any orphaned canonical sidecar (.<name>.publish-in-progress.bak) into .lexical-publish-backups/ before the next publish lands.Quarantine, GC, and the doctor/diag surface
cass diag --json --quarantine enumerates every quarantined artifact (failed seed bundles, retained publish backups, quarantined lexical generations) with size_bytes, age_seconds, safe_to_gc, and a human-readable gc_reason. The safe_to_gc flag is advisory — it reflects retention policy + cleanup dry-run eligibility and is not wired to any automatic deletion path.cass doctor --json surfaces the same quarantine summary plus checks[] status for every diagnostic the tool runs. Without --fix, doctor is read-only (auto_fix_applied=false, auto_fix_actions=[], issues_fixed=0); with --fix it applies only the repairs whose dry-run plans are proven safe (currently: Track A analytics rebuild, Track B rollup rebuild via rebuild_token_daily_stats when the token_usage ledger is intact).cass doctor --fix never have a generation reclaimed silently — every quarantine stays on disk until an explicit derived-asset rebuild (cass models backfill or an index refresh recommended by cass health --json) supersedes it.cass index runs escalates from a warning to a non-zero exit (#434): the counter persists in <data_dir>/index/.fts-repair-failure-streak.json, watch daemons log the escalation instead of exiting, and any run whose repair succeeds — or fails differently — resets it. Canonical rows and the Tantivy index are unaffected; run cass doctor --rebuild-canonical-fts --yes --json for the explicit repair.Schema stability guarantees
tests/golden/robot/: capabilities, selftest, health, status, diag, models status/verify/check-update, introspect, doctor, api-version, stats, search, export-html, onboarding, the quarantine and dedup commands and analytics incidents, plus sessions and pack on their missing-database and error paths only. swarm status scenarios are pinned under tests/golden/swarm_status/. triage, swarm work-packet and swarm lint have no golden files; their shape is covered only by assertion tests. A change to any field name, type, or nullability fails the golden test suite and requires a deliberate regeneration pass (UPDATE_GOLDENS=1 rch exec -- env CARGO_TARGET_DIR=/data/tmp/cass-golden-target cargo test --test golden_robot_json --test golden_robot_docs).cass introspect --json's response_schemas block enumerates every schema in a stable alphabetical order (BTreeMap-backed — see bead coding_agent_session_search-8sl73).{error: {code, kind, message, hint, retryable}}) have a fixed shape. kind values are kebab-case; branch on err.kind, not on the numeric code, for codes ≥ 10 (see the Error Handling section below).If your runtime does not expose built-in mcp-agent-mail tools (for example, list_mcp_resources is empty), you can still coordinate via direct MCP HTTP calls.
~/.local/pipx/venvs/mcp-agent-mail/bin/python -m mcp_agent_mail.cli serve-http --host 127.0.0.1 --port 8765
/mcp)curl -sS -X POST http://127.0.0.1:8765/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":"health","method":"tools/call","params":{"name":"health_check","arguments":{}}}'
# Ensure project
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"ensure","method":"tools/call","params":{"name":"ensure_project","arguments":{"human_key":"/data/projects/coding_agent_session_search"}}}'
# Register agent
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"register","method":"tools/call","params":{"name":"register_agent","arguments":{"project_key":"/data/projects/coding_agent_session_search","program":"codex","model":"gpt-5","name":"YourAgentName"}}}'
# Send message
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"send","method":"tools/call","params":{"name":"send_message","arguments":{"project_key":"/data/projects/coding_agent_session_search","sender_name":"YourAgentName","to":["PeerAgent"],"subject":"[coord] hello","thread_id":"coord-2026-02-13","ack_required":true,"body_md":"Online and starting work."}}}'
# Fetch inbox
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"inbox","method":"tools/call","params":{"name":"fetch_inbox","arguments":{"project_key":"/data/projects/coding_agent_session_search","agent_name":"YourAgentName","limit":50,"include_bodies":true}}}'
# Acknowledge message id 42
curl -sS -X POST http://127.0.0.1:8765/mcp -H 'Content-Type: application/json' -d \
'{"jsonrpc":"2.0","id":"ack","method":"tools/call","params":{"name":"call_extended_tool","arguments":{"tool_name":"acknowledge_message","arguments":{"project_key":"/data/projects/coding_agent_session_search","agent_name":"YourAgentName","message_id":42}}}}'
mcp_agent_mail defaults to sqlite+aiosqlite:///./storage.sqlite3. That means the server working directory determines which mailbox database you are using. To avoid "project not found" confusion, start the server from the same directory your team expects for mailbox state.
Three-pane layout with semantic styling: filter bar with pills, results list with color-coded agents and score tiers, and syntax-highlighted detail preview with tab navigation
Full conversation rendering with markdown formatting, code blocks, headers, and structured content
Built-in help screen (press F1 or ?) with all shortcuts, filters, modes, and navigation tips
AI coding agents are transforming how we write software. Claude Code, Codex, Cursor, Copilot, Aider, Pi-Agent; each creates a trail of conversations, debugging sessions, and problem-solving attempts. But this wealth of knowledge is scattered and unsearchable:
cass treats your coding agent history as a unified knowledge base. It:
snake_case ("my_var" matches "my" and "var"), hyphenated terms, and code symbols (c++, foo.bar) correctly.reader.reload() ensures new messages appear in the search bar immediately without restarting.cass search --robot currently spends roughly a second in archive open and integrity preflight on a ~10 GB archive; --robot-meta reports that separately as _meta.timing.other_ms, while search_ms stays in the tens of milliseconds.Local inference: Uses frankensearch's pure-Rust native MiniLM implementation with local safetensors weights. Once MiniLM is installed, no network traffic is required to answer queries.
Warm-daemon reuse: Semantic and hybrid CLI searches automatically use an
already-running local embedding daemon (including a socket selected with
CASS_DAEMON_SOCKET) and only initialize the installed in-process model if
daemon inference fails. Pass --daemon to permit auto-spawning a missing
daemon in human-mode searches (robot/JSON searches never spawn one, even
with --daemon, because their bounded budget cannot wait for a daemon to
start; start cass daemon yourself first; _meta.effective.daemon shows
the request and what applied), or --no-daemon to force
direct inference. --fast-only stays in
the deterministic hash-vector space. Each data directory gets a distinct
default socket and owner-private pinned key; fresh handshake, health,
embedding, batch, and rerank challenges authenticate the exact response and
immutable Frankensearch embedding identity before any daemon output is used.
--two-tier progressive refinement (fast results refined in place by the
quality tier) is experimental and currently inactive: the one-shot CLI
collapses it to a single-tier quality search and the TUI's progressive lanes
are disabled at HEAD, so hybrid search today is lexical plus one MiniLM
refinement pass when the model is installed.
Opt-in acquisition: cass models install downloads all-minilm-l6-v2 from Hugging Face on explicit request and verifies SHA256 checksums. cass models install --model multilingual-minilm explicitly selects paraphrase-multilingual-MiniLM-L12-v2 for CJK and mixed-language retrieval. Nothing is fetched until an install command runs, and merely installing the multilingual model never changes the active space.
Air-gapped install: cass models install --model <minilm|multilingual-minilm> --from-file <dir> accepts a pre-downloaded model directory so you can bring the assets in yourself.
Switching spaces: both models output 384 values, but their identities and vectors are incompatible. Set CASS_SEMANTIC_EMBEDDER=multilingual-minilm, then run cass models backfill --tier quality --embedder multilingual-minilm; cass keeps lexical fail-open active until the complete new generation is atomically published.
Required files (all must be present after install; cass models verify --model <minilm|multilingual-minilm> checks the selected model):
model.safetensorstokenizer.jsonconfig.jsonspecial_tokens_map.jsontokenizer_config.jsonVector index: Stored as vector_index/index-<embedder>.fsvi in the data directory.
Lexical fail-open: While the model is absent, cass returns lexical-only results and reports fallback_mode="lexical" in health/status; search never blocks on semantic assets.
The deterministic hash embedder is available only when explicitly selected, such as with --fast-only, --embedder hash, or CASS_SEMANTIC_EMBEDDER=hash. It is a separate lexical-feature vector space, not a silent substitute for missing MiniLM vectors:
| Feature | ML Model (MiniLM) | Hash Embedder (FNV-1a) |
|---|---|---|
| Meaning Understanding | ✅ "car" ≈ "automobile" | ❌ Exact tokens only |
| Initialization Time | ~500ms (model loading) | <1ms (instant) |
| Network Dependency | None (after install) | None |
| Disk Footprint | ~90MB model files | 0 bytes |
| Deterministic | ✅ Same input = same output | ✅ Same input = same output |
Algorithm:
When to Use:
Override: Set CASS_SEMANTIC_EMBEDDER=hash to force hash mode even when ML model is available.
cass uses the frankensearch FSVI vector index format (.fsvi) for storing semantic embeddings.
Features:
f32 and f16 storage for smaller on-disk size--approximate is passed and the HNSW sidecar file exists. hnsw_ready in status --json means only that the sidecar file is present, not that ANN is in useIndex Location: ~/.local/share/coding-agent-search/vector_index/index-<embedder>.fsvi
cass supports three search modes, selectable via --mode flag or Alt+S in the TUI:
| Mode | Algorithm | Best For |
|---|---|---|
| Lexical | BM25 full-text | Exact term matching, code searches |
| Semantic | Vector similarity | Conceptual queries, "find similar" |
| Hybrid (default) | Lexical + single-tier semantic refinement fused with RRF; lexical fail-open | Balanced precision and recall |
Lexical Search: Uses Quill's BM25 implementation with prefix matching. Best when you know the exact terms you're looking for. The lexical index is derived from SQLite; if it is missing, stale, or incompatible, cass reports the state and rebuilds through the normal indexing path from the canonical database.
Semantic Search: Computes vector similarity between query and indexed MiniLM embeddings. Finds conceptually related content even without exact term overlap. Explicit semantic mode requires the MiniLM model and a compatible MiniLM vector index; it never substitutes same-dimensional hash vectors.
Hybrid Search: The default. It combines lexical and semantic results using Reciprocal Rank Fusion (RRF) when semantic assets are ready, and it fails open to lexical when semantic enrichment is still catching up or disabled:
RRF_score = Σ 1 / (K + rank_i)
Where K=60 (tuning constant) and rank_i is the position in each result list. This balances the precision of lexical search with the recall of semantic search. Semantic refinement is a single pass over the installed MiniLM index; progressive two-tier refinement (--two-tier) is experimental and currently inactive.
# CLI examples
cass search "authentication" --mode lexical --robot
cass search "how to handle user login" --mode semantic --robot
cass search "auth error handling" --mode hybrid --robot
foo* - Prefix match (finds "foobar", "foo123")*foo - Suffix match (finds "barfoo", "configfoo")*foo* - Substring match (finds "afoob", "configuration")*term* wildcards to broaden matches. A visual indicator shows when the fallback is active.Up/Down arrows.F12) that prioritizes exact matches over wildcard/fuzzy results.**bold**, in human-readable and robot/JSON output alike; --highlight also marks the query terms' other occurrences, never marking a term twice.Powered by FrankenTUI (ftui) — a high-performance Elm-architecture TUI framework with adaptive frame budgets, Bayesian diff selection, and spring-based animations.
Indexing 150/2000 (7%)—plus active filters.Ctrl+Enter, then open all in your editor with Ctrl+O. Confirmation prompt for large batches (≥12 items)./ to search within the detail pane; matches highlighted with n/N navigation.recent/balanced/relevance/quality with F12; quality mode penalizes fuzzy matches.Alt+A; Esc returns to search.cass tui --inline to keep terminal scrollback intact. The UI anchors to a region of the terminal while logs scroll normally. Configure with --ui-height <rows> and --anchor top|bottom.cass tui --record-macro session.macro for reproducible bug reports and workflow automation. Events are saved as human-readable JSONL with full timing data.cass tui --asciicast demo.cast.
Export conversations as styled, portable HTML files with optional encryption:
cdn.jsdelivr.net, pinned with SRI hashes.onerror="...no-prism" — code blocks remain readable offline in plain monospace, and the page layout never depends on a network resource.TUI Usage: Press Ctrl+E in the detail view to open the export modal, or Ctrl+Shift+E to export Markdown immediately with defaults. On the detail pane's Export tab, e/h open the HTML export modal and m runs the Markdown export.
CLI Usage:
# Basic export
cass export-html /path/to/session.jsonl
# With encryption
printf '%s\n' "secret" | cass export-html /path/to/session.jsonl --encrypt --password-stdin
# Custom output location
cass export-html session.jsonl --output-dir ~/exports --filename "my-session"
# Open in browser after export
cass export-html session.jsonl --open
# Robot mode (JSON output)
cass export-html session.jsonl --json
Ingests history from 32 local agent connectors, normalizing them into a unified Conversation -> Message -> Snippet model. cass capabilities --json | jq .connectors is the canonical machine-readable inventory (kept in lockstep with the runtime registry):
~/.codex/sessions (Rollout JSONL)~/.gemini/tmp (Chat JSON)~/.claude/projects (Session JSONL), plus macOS Desktop metadata sidecars under
~/Library/Application Support/Claude/claude-code-sessions and
~/Library/Application Support/Claude/local-agent-mode-sessions~/.clawdbot/sessions (Session JSONL)~/.vibe/logs/session/*/messages.jsonl (Session JSONL).opencode directories (SQLite)~/.local/share/amp & VS Code storage~/Library/Application Support/Cursor/User/ global + workspace storage (SQLite state.vscdb)~/Library/Application Support/com.openai.chat (v1 unencrypted JSON; v2/v3 encrypted—see Environment)~/.aider.chat.history.md and per-project .aider.chat.history.md files (Markdown)~/.pi/agent/sessions (Session JSONL with thinking content)prime_agent): ~/.prime/agent/sessions/<session-id>.jsonl (versions 1–3). Indexes the active branch with omission counts for abandoned siblings; preserves thinking, tool results and context summaries. Overrides, in precedence order: PRIME_AGENT_SESSION_DIR, legacy PRIME_AGENT_CODING_AGENT_SESSION_DIR, then PRIME_AGENT_CODING_AGENT_DIR (with /sessions appended). Prime retains its own agent identity.omp): OMP v18's default ~/.omp/agent/sessions, named profiles under ~/.omp/profiles/<name>/agent/sessions, XDG stores under $XDG_DATA_HOME/omp, and explicit OMP-only archive roots via CASS_OMP_DATA_ROOT (pi-family JSONL, including per-session sub-agent transcripts)github.copilot-chat (JSON)~/.copilot/session-state, legacy ~/.copilot/history-session-state, and gh copilot config paths (JSONL/JSON)~/.openclaw/agents/*/sessions (Session JSONL)~/.local/share/goose/sessions/sessions.db (SQLite, v1.20+), plus the earlier per-session *.jsonl layout under ~/.goose/sessions~/.crush/crush.db and per-project .crush/crush.db (SQLite)~/.hermes/state.db and project-local .hermes/state.db (SQLite)~/.local/share/devin/cli/sessions.db (SQLite; override with CASS_DEVIN_DATA_ROOT). Indexes visible local sessions along their active parent chain, preserving tool messages and excluding abandoned branches and inline image payloads. Cloud-only sessions are outside this connector's scope.CASS_SHELLEY_DB=/absolute/path/to/shelley.db, or add that file to the paths of a type = "local" source in sources.toml. Any filename is accepted after schema validation. Defaults include ~/.config/shelley/shelley.db and shelley.db in the current directory. Live indexing watches the database and its WAL/SHM sidecars; metadata changes refresh existing sessions. CASS_SKIP_SUBAGENTS=1 excludes conversations with a Shelley parent ID. The database can also contain credentials and application settings, so raw mirroring and remote database ingestion are disabled; keep the database on its original machine.CASS_GROK_BOT_DATA_ROOT to its persistence directory. Native message IDs preserve already indexed history as older messages leave the application's window; repeat scans do not duplicate retained messages. CASS reads only chat content. Raw mirroring and automatic fleet copying are disabled because the replica also holds secret and approval fields. This connector is separate from the Grok CLI connector and does not fetch cloud history.$KIMI_CODE_HOME/sessions/*/*/agents/*/wire.jsonl (default ~/.kimi-code; sub-agents index as <sessionId>:<agentId>), plus the legacy ~/.kimi/sessions/*/*/wire.jsonl layout (Session JSONL)~/.local/share/muse/sessions/<YYYY>/<MM>/<DD>/<session-id>/session.jsonl, including nested subagent/*/session.jsonl transcripts (override with CASS_MUSE_DATA_ROOT)~/.qwen/tmp/*/chats/session-*.json (Chat JSON)~/.factory/sessions (JSONL files organized by workspace slug)~/.gemini/antigravity/ and the CLI's ~/.gemini/antigravity-cli/ — each holding brain/<uuid>/.system_generated/logs/transcript.jsonl (clean JSONL transcript) with the durable per-conversation conversations/<uuid>.db (SQLite) mirrored alongside. IDE conversations are keyed ide/<uuid> so the two stores never collide; CASS_ANTIGRAVITY_DATA_ROOT replaces both with one explicit base. Resume with cass resume <transcript> --agent agy (agy --conversation <uuid>).~/.openhands/conversations/<id>/ — base_state.json metadata plus an events/event-NNNNN-<uuid>.json event stream (JSON)grok): ~/.grok/sessions/<percent-encoded-cwd>/<session-uuid>/ — updates.jsonl (authoritative ACP session-update stream) with summary.json metadata and chat_history.jsonl fallback (override the base dir with GROK_HOME). Resume with grok --resume <session-id>.codebuff): ~/.config/manicode/projects/<project>/chats/<chat-id>/chat-messages.json with its run-state.json (override with CASS_CODEBUFF_DATA_ROOT). Both products write the same Manicode store and no chat records which binary wrote it, so their sessions share one lineage identity, codebuff (filter with --agent codebuff). Messages are reconciled by their native IDs, so an edited message updates in place instead of duplicating.kiro): ~/.kiro/sessions/cli/<session-uuid>.jsonl (append-only event log: prompts, assistant messages, tool results) with the matching <session-uuid>.json snapshot read for session ID, working directory, title, timestamps and model.Claude Code Desktop sidecars preserve title, workspace, model, and session IDs, but not necessarily the full conversation body. If Claude Code has culled an old CLI JSONL body, cass can still index searchable sidecar metadata while reporting that the conversation body is unavailable.
Pi-Agent parses JSONL session files with rich event structure:
~/.pi/agent/sessions/ (override the agent home with PI_CODING_AGENT_DIR, or the sessions directory directly with PI_SESSIONS_DIR)session_start, message, model_change, thinking_level_change*_*.jsonl pattern in sessions directoryOh My Pi (omp) uses the same pi-family wire format but remains a separate
agent identity throughout search, analytics, resume, TUI, and HTML export:
~/.omp/agent/sessions/ and ~/.omp/profiles/<name>/agent/sessions/; OMP_PROFILE selects a profile and takes precedence over legacy PI_PROFILE$XDG_DATA_HOME/omp/sessions/ and $XDG_DATA_HOME/omp/profiles/<name>/sessions/ when the OMP XDG root existsPI_CODING_AGENT_SESSION_DIR names the exact OMP sessions directory. CASS_OMP_DATA_ROOT declares an OMP-only archive/store root and is the right choice for copied, mounted, or custom OMP data. PI_CODING_AGENT_DIR is shared by both pi-family programs, so CASS conservatively keeps otherwise-ambiguous paths under that root owned by Pi-Agent; use one of the OMP-specific variables when OMP identity matters. PI_CONFIG_DIR changes the home-relative .omp config directory name.omp [--profile <name>] --resume <id>; copied profiles, XDG archives, remote mirrors, and explicit roots also carry --session-dir <dir> so a canonical-looking archive cannot reopen a different live storepi_agent to omp using the same conservative canonical/XDG/remote-mirror ownership policy as live discovery, then the derived lexical index and analytics are rebuilt so a transcript cannot remain attributed to both agents. The conventional ~/.local/share/omp shape is durable path evidence; an arbitrary historical custom $XDG_DATA_HOME/omp path is reclassified only while that root is currently configured and resolvable. Without provider-qualified evidence, ambiguous historical paths fail closed as Pi-Agent rather than letting a generic .../omp/sessions directory steal ownership.OpenCode reads SQLite databases from workspace directories:
.opencode/ directories (scans recursively from home).opencode containing database filesSearch across agent sessions from multiple machines—your laptop, desktop, and remote servers—all from a single unified index. cass uses SSH/rsync to efficiently sync session data, tracking provenance so you know where each conversation originated.
The easiest way to configure multi-machine search is the interactive setup wizard:
cass sources setup
What the wizard does:
~/.ssh/configsources.toml with correct paths and mappingscass sources sync right after configuration (skipped with --skip-sync or --dry-run; --json setup defers it and reports the command to run)Wizard options:
| Flag | Purpose |
|---|---|
--hosts <names> | Configure only specific hosts (comma-separated) |
--dry-run | Preview changes without applying them |
--non-interactive | Use auto-detected defaults for scripting |
--skip-install | Don't install cass on remotes |
--skip-index | Don't run indexing on remotes |
--skip-sync | Skip the final cass sources sync. Interactive setup runs that sync after the hosts are configured and records it as complete only once it has actually finished; --json setup always defers it and reports sync.status = "pending" with the command to run |
--resume | Resume an interrupted setup |
--json | Output progress as JSON (for automation) |
Examples:
# Full interactive wizard
cass sources setup
# Configure specific hosts only
cass sources setup --hosts laptop,workstation,build-server
# Preview without making changes
cass sources setup --dry-run
# Resume interrupted setup
cass sources setup --resume
# Non-interactive for CI/CD
cass sources setup --non-interactive --hosts myserver --skip-install
Resumable state: If setup is interrupted (Ctrl+C, connection lost), state is saved to the cache directory (~/.cache/cass/setup_state.json on Linux). Resume with --resume.
Tailscale discovery is optional: cass sources discover --tailscale --json adds
online tailnet peers to SSH-config discovery, and cass sources setup --tailscale
offers them in setup. It reads local tailscale status --json with a five-second
deadline; a missing CLI, stopped daemon, or login failure produces a warning and
leaves SSH-config discovery available. Explicit setup --hosts skips discovery.
Connections use ordinary SSH over assigned Tailscale IPv4 addresses, so MagicDNS
is not required. Matching SSH aliases retain their user/key configuration;
otherwise SSH uses its normal defaults. IPv6-only peers are currently omitted.
Tailscale ACLs, SSH authorization and host-key checks still apply; discovery does
not log in, install Tailscale, or change either SSH or tailnet configuration.
The local fixture and Docker tests do not prove that your machines can sync and
search each other's sessions. The opt-in live harness uses actual SSH connections
and cass sources discover, sources add, sources sync, and search. It creates isolated synthetic
Codex sessions on each machine, checks source provenance and filters, repeats a
sync to detect duplicates, and appends messages. It checks both lexical and default
hybrid search, requires one JSON response per sync, holds the real indexing lock to
test busy refusal, and recovers transferred sessions through sources reingest.
A refused SSH connection must leave the other sources searchable.
Keep the inventory and SSH configuration outside this repository. For example, create a mode-0600 JSON file containing:
{
"ssh_config": "/private/path/to/ssh_config",
"hosts": [{"ssh": "workstation"}, {"ssh": "laptop"}]
}
Then run with an explicit binary:
python3 scripts/e2e/live_fleet_search.py \
--inventory /private/path/to/fleet.json \
--cass-bin /path/to/cass
Python 3 and authenticated SSH access are required on the remote machines.
The Unix runner needs Python 3.9+, rsync, and a CASS binary supporting the tested
commands. Each inventory alias must appear in the supplied SSH configuration;
included configuration files are supported. Host-key verification stays enabled.
To exercise actual tailnet discovery and transport, add --tailscale to the
harness command and use tailnet IPv4 addresses as the private inventory targets.
Keep any required SSH users, keys and trusted host-key aliases in the private SSH
configuration. For a discovery test independent of explicit aliases, use SSH
Match originalhost entries rather than literal Host entries for those addresses.
The harness retains fresh test directories and raw
receipts privately outside git; it never changes existing session archives or
deletes test data. Console results use ordinal labels. An unreachable machine
keeps the overall result failed, even if the other machines pass. Do not attach
raw receipts or inventories to public issues: they contain machine identities.
When the wizard installs cass on remote machines, it tries every viable method in this priority order, falling through to the next when one fails; setup fails only when all of them do, and the error lists each attempt:
| Priority | Method | Speed | Requirements |
|---|---|---|---|
| 1 | cargo-binstall | ~30s | cargo-binstall pre-installed, compatible release binary |
| 2 | Pre-built binary | ~10s | curl/wget, GitHub access, compatible release binary |
| 3 | cargo install | ~5min | Rust toolchain, 1GB disk, 2GB RAM |
| 4 | Full bootstrap | ~10min | curl, 1GB disk, 2GB RAM (installs rustup) |
crates.io publishing resumed at 0.7.0 (GH#416): the long-stale registry gap (0.6.13, published before the Quill/OMP era) is closed — the entire dependency chain now resolves from crates.io (
frankensearch 0.4.0, thefrankentorch-*family,frankenhnsw), socargo install coding-agent-searchbuilds the current line again. The installer and GitHub Release binaries remain the fastest paths.
Resource Requirements:
What Gets Installed:
cass binary (location depends on method: ~/.cargo/bin/cass for cargo-based, ~/.local/bin/cass for pre-built binary)Installation Progress: The wizard shows real-time progress for each stage:
Installing cass on laptop...
[1/4] Checking environment... ✓
[2/4] Downloading binary... ████████░░ 80%
[3/4] Verifying checksum... ✓
[4/4] Setting up PATH... ✓
Use --skip-install if you prefer to install manually on remotes.
The setup wizard automatically discovers SSH hosts from your configuration:
Discovery Sources:
~/.ssh/config (parses Host entries)*, ?) are automatically excludedProbe Results (for each discovered host):
| Check | Purpose |
|---|---|
| Connectivity | Can we establish SSH connection? |
| cass Version | Is cass already installed? What version? |
| Agent Data | Which agents have session data? |
| Session Count | How many conversations exist? |
| System Info | OS, architecture, disk space, memory |
Each setup run probes every selected host afresh; probe results are not cached between runs.
For manual configuration without the wizard:
# Add a remote machine using platform presets
cass sources add user@laptop.local --preset macos-defaults
# Or specify paths explicitly
cass sources add dev@workstation --path ~/.claude/projects --path ~/.codex/sessions
# Sync sessions from all configured sources
cass sources sync
# Check source health and connectivity
cass sources doctor
Remote source diagnostics are intentionally local-only. cass triage --json,
cass doctor --json, cass health --json, and cass status --json report the
remote_source_sync summary from cass-owned evidence: sources.toml,
sync_status.json, the local remotes/<source>/mirror/ copy, and archive DB
provenance rows. They do not open SSH sessions, mutate remote machines, or
rewrite provider session logs while classifying source gaps.
cass sources doctor is the explicit networked exception: it performs bounded,
read-only probes of configured source hosts. Its per-source human summary keeps
the same native reachability, binary-health, and mirror/sync state codes and
safe command as the JSON report. It intentionally does not claim local search
readiness, because a remote host probe cannot establish the controller's local
SQLite, lexical, or semantic asset state.
This matters because agent harnesses can prune their own logs. If a laptop is
retired, a remote path disappears, or a provider truncates older sessions, the
cass archive DB and cass-owned local mirror may be the only remaining evidence
for those conversations. Treat gap names such as remote_source_unavailable,
remote_source_pruned, local_archive_ahead_of_remote, and
remote_copy_ahead_verified as preservation signals first: keep the archive and
mirror intact, then run the recommended cass sources sync --json (all configured remote sources; --source <name> narrows it) or
source-specific sync command after reviewing the reported evidence.
Raw-mirror retention is explicit and audited. Use cass mirror prune --older-than 90d --json or cass mirror prune --max-size 100GB --json to get a
dry-run plan; add --apply only after reviewing the scope and totals. Preview
entries contain at most 1,000 manifest/blob details; omitted_entry_count
reports additional candidates. Planned counts and bytes cover the entire plan,
including omitted details. Use provider/path selectors to inspect a narrower
scope. Previews do not append audit records. Add
--keep-tag <tag> to pin captures linked to tagged conversations. prune
holds down blobs referenced by captures from the last 7 days by default, writes
complete intent/result records to raw-mirror/v1/pruned.jsonl for non-empty
applied plans, and refuses apply mode
while an index/watch job is active.
Applied pruning syncs the audit independently of the optional capture setting
CASS_RAW_MIRROR_FSYNC. Each completed result is recorded before the next
removal; a later failure preserves those earlier results. An abrupt crash
between a removal and its result record can still leave an intent without a
confirmed result.
Use --provider opencode and/or --source-path '*/opencode.db' with an age
or size rule to target one source without retiring unrelated captures.
Repeated providers are alternatives; a source-path glob further narrows them.
A pruned capture of a source that is still on disk is copied again by the next
index run. On a machine whose providers never delete their session files, set
CASS_RAW_MIRROR=0 (or false, no, off) to stop capturing altogether.
Indexing and search are unchanged, and existing captures stay until you prune
them. cass doctor reports raw_mirror_capture_disabled and warns about what is
given up: a session file its provider later deletes survives only in the
archive DB.
With a selector, --max-size measures unique blobs in that selection. Shared
blobs still referenced outside it and orphan blobs without source provenance
remain protected. The JSON plan records the selectors and scope_blob_bytes.
Large mutable sources are stored as 4 MiB content-addressed chunks. Growing
JSONL files reuse every unchanged complete chunk, and SQLite sources reuse
unchanged 4 MiB byte regions, so each historical snapshot remains byte-exact without
writing another full-file blob. Existing whole-blob manifests remain readable;
cass doctor --json reports storage_kind, chunk_count, the full-source
digest, and verifies every referenced chunk before treating a snapshot as
recovery authority.
Sources are configured in the platform config directory (Linux: ~/.config/cass/sources.toml, macOS: ~/Library/Application Support/cass/sources.toml):
[[sources]]
name = "laptop"
type = "ssh"
host = "user@laptop.local"
paths = ["~/.claude/projects", "~/.codex/sessions"]
sync_schedule = "manual"
[[sources]]
name = "workstation"
type = "ssh"
host = "dev@work.example.com"
paths = ["~/.claude/projects"]
sync_schedule = "daily"
# Path mappings rewrite remote paths to local equivalents
[[sources.path_mappings]]
from = "/home/dev/projects"
to = "/Users/me/projects"
# Agent-specific mappings
[[sources.path_mappings]]
from = "/opt/work"
to = "/Volumes/Work"
agents = ["claude_code"]
Configuration Fields:
| Field | Description |
|---|---|
name | Friendly identifier (becomes source_id) |
type | Connection type: ssh or local |
host | SSH host (user@hostname) |
paths | Paths to sync (supports ~ expansion) |
sync_schedule | manual, hourly, or daily. Only the jobs installed by cass schedule install run it; without them it is a label and syncs happen when you run cass sources sync |
path_mappings | Rewrite remote paths to local equivalents |
# List configured sources
cass sources list [--verbose] [--json]
# Add a new source
cass sources add <user@host> [--name <name>] [--preset macos-defaults|linux-defaults] [--path <path>...] [--no-test]
# Remove a source
cass sources remove <name> [--purge] [-y]
# Check connectivity and config
cass sources doctor [--source <name>] [--json]
# Sync sessions
cass sources sync [--source <name>] [--no-index] [--verbose] [--dry-run] [--json]
If one harness is generating mostly junk or looped output, you can disable it persistently even if its files remain on disk:
# Inspect current include/exclude state
cass sources agents list --json
# Stop indexing this harness in future runs
cass sources agents exclude openclaw
# Re-enable it later
cass sources agents include openclaw
cass stores this preference in sources.toml (~/.config/cass/sources.toml on Linux, ~/Library/Application Support/cass/sources.toml on macOS), so future scans, syncs, and watch-mode updates remember it automatically.
By default, cass sources agents exclude <agent> also removes already archived local data for that agent and rebuilds the lexical index so the exclusion frees space instead of only blocking future imports.
If you want to block future indexing but keep the data already archived:
cass sources agents exclude openclaw --keep-indexed-data
The sync engine uses rsync over SSH for efficient delta transfers and falls back to other transports when rsync is unavailable:
Transfer Methods (auto-detected):
| Method | When Used | Characteristics |
|---|---|---|
| rsync | rsync available on both ends | Delta transfers, compression, progress stats |
| WSL rsync | Windows without native rsync, WSL with rsync installed | Runs wsl rsync |
| scp | rsync unavailable | Full file copies through the system scp, inheriting the OpenSSH agent, keys and ~/.ssh/config |
| SFTP | the fallbacks above unavailable | Full file transfers via the SSH native protocol |
Safety Guarantees:
--delete, so remote deletions never propagate locally.-a and without -u, so a local mirror file that differs from the remote is overwritten, even when the local copy is newer. The mirror is a copy of the remote, not a place to edit sessions.--partial keeps a partly transferred file under its final name so the next sync continues it. A failed sync can therefore leave a truncated file until the next sync completes.Transfer Configuration:
| Setting | Default | Purpose |
|---|---|---|
| Connection timeout | 10s | Fail fast on unreachable hosts |
| Transfer timeout | 300 s of I/O inactivity | rsync --timeout aborts a transfer that stalls this long; there is no wall-clock limit on a transfer that keeps moving |
| Compression | Enabled | Reduce bandwidth for text-heavy sessions |
| Partial transfers | Enabled | Resume interrupted syncs |
rsync Flags Used:
-avz --links --safe-links --stats --partial [--protect-args | --secluded-args] --timeout 300 \
-e "ssh [-F $CASS_SSH_CONFIG] -o BatchMode=yes -o ConnectTimeout=10 -o ServerAliveInterval=15 -o ServerAliveCountMax=3 -o StrictHostKeyChecking=yes"
Where -avz = archive mode + verbose + compression. --protect-args/--secluded-args is auto-detected per remote rsync version (omitted when the remote rejects it), and --timeout carries the transfer timeout in seconds. StrictHostKeyChecking=yes means a host whose key is not already in known_hosts fails with "Host key verification failed". Connect once with plain ssh <host> and accept the key, or add it with ssh-keyscan, before the first sync.
Data Flow:
Remote: ~/.claude/projects/
↓ (rsync over SSH)
Local: ~/.local/share/coding-agent-search/remotes/<source>/mirror/<path>_<hash>/
↓ (connector scan)
Index: agent_search.db + index/v9-quill/
Where <path> is a filesystem-safe version of the remote path (e.g. .claude_projects), and <hash> is an FNV-1a hash of the original path in hex, so foo/bar and foo_bar never collide.
Sessions from remotes are indexed alongside local sessions, with provenance tracking to identify origin.
When viewing sessions from remote machines, workspace paths may not exist locally. Path mappings rewrite these paths so file links work on your local machine:
# List current mappings
cass sources mappings list laptop
# Add a mapping
cass sources mappings add laptop --from /home/user/projects --to /Users/me/projects
# Test how a path would be rewritten
cass sources mappings test laptop /home/user/projects/myapp/src/main.rs
# Output: /Users/me/projects/myapp/src/main.rs
# Agent-specific mappings (only apply for certain agents)
cass sources mappings add laptop --from /opt/work --to /Volumes/Work --agents claude_code,codex
# Remove a mapping by index
cass sources mappings remove laptop 0
In the TUI, filter sessions by origin:
Remote sessions display with a source indicator (e.g., [laptop]) in the results list.
Each conversation tracks its origin:
source_id: Machine identifier (e.g., "laptop", "workstation")origin_kind: local or remoteorigin_host: the remote host label, absent for local sessionsworkspace_original: Original path on the remote machine (before path mapping)--fields provenance selects exactly source_id, origin_kind and origin_host.
These fields appear in JSON/robot output and enable filtering:
cass search "auth error" --source laptop --json
cass timeline --since 7d --source remote
cass stats --by-source
cass is purpose-built for consumption by AI coding agents—not just as an afterthought, but as a first-class design goal. When you're an AI agent working on a codebase, your own session history and those of other agents become an invaluable knowledge base: solutions to similar problems, context about design decisions, debugging approaches that worked, and institutional memory that would otherwise be lost.
Imagine you're Claude Code working on a React authentication bug. With cass, you can instantly search across:
This cross-pollination of knowledge across different AI agents is transformative. Each agent has different strengths, different context windows, and encounters different problems. cass unifies all this collective intelligence into a single, searchable index.
cass teaches agents how to use it—no external documentation required:
# First-stop capability contract for agents
cass triage --json
cass capabilities --json
# → {"version": "...", "workflows": [...], "mistake_recoveries": [...], "commands": [...], "exit_codes": [...], "env_vars": [...]}
# Full API schema with argument types, defaults, and response shapes
cass introspect --json
# Topic-based help optimized for LLM consumption
cass robot-docs commands # All commands and flags
cass robot-docs schemas # Response JSON schemas
cass robot-docs examples # Copy-paste invocations
cass robot-docs exit-codes # Error handling guide
cass robot-docs guide # Quick-start walkthrough
AI agents sometimes make syntax mistakes. cass aggressively normalizes input to maximize acceptance when intent is clear:
| What you type | What cass understands | Correction note |
|---|---|---|
cass -robot --limit=5 | cass --robot --limit=5 | Single-dash long flags normalized |
cass --Robot --LIMIT 5 | cass --robot --limit 5 | Case normalized |
cass search "auth" --max_results 5 | cass search "auth" --limit 5 | Snake-case long flag normalized before alias recovery |
cass find "auth" | cass search "auth" | find/query/q → search via alias table |
cass --robot-docs | cass robot-docs | Flag-as-subcommand detected |
cass commands --json | cass robot-docs commands | Robot-docs topic shorthand detected |
cass schemas --json | cass robot-docs schemas | Robot-docs topic shorthand detected |
cass ready --json | cass triage --json | One-shot triage alias |
cass preflight --json | cass triage --json | One-shot triage alias |
cass --json | cass triage --json | Top-level robot request defaults to safe preflight |
cass --robot | cass triage --json | Top-level robot request defaults to safe preflight |
cass --json search "auth" | cass search "auth" --json | Leading structured flag moved to the robot-capable subcommand |
cass --robot status | cass status --json | Leading robot flag canonicalized to JSON output |
cass answer "auth" --json | cass pack "auth" --json | Cited-handoff aliases normalized to answer pack |
cass why auth failed --json --max-evidence 3 | cass pack "auth failed" --json --max-evidence 3 | Question/RC prompt aliases normalized to answer pack |
cass auth failed --json --max-evidence 3 | cass pack "auth failed" --json --max-evidence 3 | Bare robot queries with pack-only flags become answer packs |
cass search auth failed --json --max-evidence 3 | cass pack "auth failed" --json --max-evidence 3 | Explicit robot search with pack-only flags becomes an answer pack |
cass html-export session.jsonl --json | cass export-html session.jsonl --json | Reversed HTML export aliases normalized to the archive exporter |
cass current --json | cass sessions --current --json | Current-session shorthand normalized to session discovery |
cass sessions current --json | cass sessions --current --json | Positional current accepted as the sessions current flag |
cass search --query "auth" --json | cass search "auth" --json | Named query option converted to required positional query |
cass search --q "auth" --json | cass search "auth" --json | Short/familiar query aliases converted to required positional query |
cass search auth error --json | cass search "auth error" --json | Adjacent unquoted query words folded into one search |
cass auth error --json | cass search "auth error" --json | Unquoted robot-mode query words folded into search |
cass search --agent codex --limit 5 auth error --json | cass search "auth error" --agent codex --limit 5 --json | Query moved before leading search filters |
cass view --path session.jsonl --line 42 --json | cass view session.jsonl --line 42 --json | Named path option converted to required positional path |
cass view session.jsonl --line-number 42 --json | cass view session.jsonl --line 42 --json | Legacy alias for --line; still reads raw file line 42 |
cass view session.jsonl line_number=42 --json | cass view session.jsonl --message-index 42 --json | A pasted search-hit field selects canonical message 42, not raw line 42 |
cass view source_path=session.jsonl source_id=local line_number=42 --json | cass view session.jsonl --source local --message-index 42 --json | Search hit field bundle accepted as a follow-up command (add conversation_id when the file holds several conversations) |
cass search "auth" --format json | cass search "auth" --robot-format json | Familiar format spelling converted to robot format |
cass search "auth" --output json | cass search "auth" --robot-format json | Familiar output spelling converted to robot format |
cass help search --json | cass robot-docs commands | Structured help intent routed to the machine-readable command reference |
cass --format json status | cass status --robot-format json | Leading format request moved to the target subcommand |
cass search "auth" --max-results 5 | cass search "auth" --limit 5 | Result-count alias converted to canonical limit |
cass search "auth" -n 5 | cass search "auth" --limit 5 | Familiar short count flag converted to canonical limit |
cass search "auth" --last 7 --before now | cass search "auth" --since -7d --until now | Familiar time-window aliases converted to canonical filters |
cass search "auth" last=7d before=now | cass search "auth" --since -7d --until now | Bare time-window assignments converted to canonical filters |
cass search "auth" --provider codex | cass search "auth" --agent codex | Provider/tool/connector aliases converted to canonical agent filter |
cass search "auth" provider=codex | cass search "auth" --agent codex | Bare provider assignment converted to canonical agent filter |
cass search auth provider codex limit 5 | cass search auth --agent codex --limit 5 | Bare filter key/value pairs after a query converted to canonical flags |
cass search --limt 5 | cass search --limit 5 | Flag typos within Levenshtein distance ≤2 corrected |
The CLI applies multiple normalization layers:
--limt → --limit), and a first word within distance 2 of a subcommand is corrected to it (e.g. serach → search). A word that already names a subcommand is never changed, so cass status --jsn runs status --json. forget and upgrade are reached only by exact spelling.--Robot, --LIMIT → --robot, --limit--max_results, --data_dir, and other known snake_case long flags become canonical kebab-case before alias recovery runs-robot → --robot (common LLM mistake)ready/preflight → triage; find/query/q/grep/lookup → search; session → sessions; answer/evidence/bundle/handoff/why/explain/rca/root-cause/rootcause/summarize/summarise → pack; html-export/html_export/exporthtml → export-html; ls/list/info/summary → stats; st/state → status; reindex/idx/rebuild → index; show/get/read → view; diagnose/debug/check → diag; caps/cap → capabilities; inspect/intro → introspect; docs/help-robot/robotdocs → robot-docscommands, schemas, examples, exit-codes, and quickstart become robot-docs <topic> instead of falling through to search; command topics such as doctor and sources use structured help (cass help doctor --json, cass sources --help --json). Bare cass guide is reserved for the guided-operations planner; use cass robot-docs guide for the robot-docs walkthrough.cass --json, cass --robot, or cass --robot-format json with no subcommand runs read-only triage--json/--robot before a robot-capable subcommand is moved onto that subcommand--query/--q/--text/--pattern for search/pack and --path/--source-path/--file/--session for drill-down/export commands become the required positional argumentsearch/pack become one query positional--format json|jsonl|compact|sessions|toon, --output json|jsonl|compact|sessions|toon, and --output-format ... are accepted as --robot-format ... on robot-capable commands; export --format ... and export --output <file> keep their export meaningshelp --json, help commands --json, and search --help --json route to robot-docs guide / robot-docs commands; plain --help stays native clap help--max-results, --num-results, --results, --count, --top-k, and -n become --limit on commands with result limits--last 7, --before now, last=7d, and before=now become canonical --since/--until filters--provider, --tool, --connector, and matching assignments become canonical --agent filters on search-like commandsprovider codex, limit 5, and last 7d become canonical filter flags before the remaining words are folded into the querysearch with pack-only flags such as --max-evidence, --max-sessions, or --freshness-policy becomes pack, not implicit or explicit search--line-number, --line_number and line=42 become --line (a raw file line)line_number=42 pasted from a search hit becomes --message-index 42 (the canonical message ordinal the hit names), and a source_path=... source_id=... line_number=... bundle becomes the canonical path, --source and --message-index form for follow-up view/expand commandssearch query unless they look like a subcommand typocurrent, current-session, and sessions current become sessions --currentWhen corrections are applied, cass emits a teaching note to stderr so agents learn the canonical syntax. In robot/JSON mode the same information is emitted as one note: auto-corrected: <note> line per correction on stderr (at most two: the normalization note and the typo-recovery note), so stdout stays data-only. Robot-mode notes are printed only when the command succeeds; a failing command's stderr is its single JSON error envelope. The same notes appear in search output under _meta.effective.auto_corrections with --robot-meta.
Every command supports machine-readable output:
# Pretty-printed JSON (default robot mode)
cass search "error" --robot
# Streaming JSONL: one hit per line. Add --robot-meta to prepend a
# {budget, _meta} header line (elapsed_ms, next_cursor, state, index_freshness).
# The header also appears without --robot-meta when the search timed out
# (budget.timed_out), returned did-you-mean suggestions, --aggregate or --explain.
cass search "error" --robot-format jsonl # hits only
cass search "error" --robot-format jsonl --robot-meta # 1 _meta header + hits
# Compact single-line JSON (minimal bytes)
cass search "error" --robot-format compact
# Include performance metadata
cass search "error" --robot --robot-meta
# → { "hits": [...], "_meta": { "elapsed_ms": 12, "cache_hit": true, "wildcard_fallback": false, "lexical_degrade_reason": null, ... } }
# lexical_degrade_reason is "query_fuel_exhausted" when a hybrid search dropped its
# lexical leg because Quill's query fuel ran out (see CASS_QUILL_QUERY_FUEL_BUDGET)
# What the search actually ran (--robot-meta): check this instead of trusting the flags
cass search "error" --robot --robot-meta --days 7 | jq '._meta.effective'
# → { "command": "search", "query": "error",
# "query_structure": "error", // how the engine groups operands: `a OR b c` -> "a OR (b AND c)"
# "query_recoveries": [], // e.g. "1 unclosed '(' closed at the end of the query"
# "db_path": "/home/you/.local/share/coding-agent-search/agent_search.db",
# "db_path_source": "default", // --db | env:CASS_DB_PATH | --data-dir | env:CASS_DATA_DIR | env:XDG_DATA_HOME | default
# "time_window": { "since_ms": 1758067200000, "since_from": "--days 7", "until_ms": null, "until_from": null },
# "filters": { "agents": [], "workspaces": [], "source": "all", "sessions_from_paths": null },
# "auto_corrections": [] } // each argv correction, worded like its stderr note
# `cass pack "error" --json` carries the same object for the search it ran in
# its own `_meta.effective` ("command": "pack", and no search-only "daemon"),
# with home-directory paths, private hosts and secrets redacted like the rest
# of the pack (the db_path above reads "[REDACTED_PATH]/agent_search.db").
# Per-hit trust verdict (advisory; --robot-meta only)
cass search "error" --robot --robot-meta
# Each hit then carries a metadata-only `trust` block:
# "trust": {
# "schema_version": 1,
# "trust_tier": "unverified", // trusted | likely | unverified | stale | failed
# "confidence": "medium", // low | medium | high
# "provenance_refs": [], // e.g. ["commit:ab0d12ef90ab", "bead:xyz", "release:v0.6.15"]
# "stale_reason": "aged_out", // present only when not fully trusted
# "recommended_followup": "..." // advisory next step (never a destructive command)
# }
How agents should branch on trust_tier (relevance is not correctness — a
hit can be a landed fix or a failed attempt):
trust_tier | Meaning | What to do |
|---|---|---|
trusted | Landed, proof-backed, release/bead-contained | Safe to reuse |
likely | Has provenance (commit/closed bead) but not proof-pinned | Confirm via the cited ref first |
unverified | Relevant but no provenance link, or lexical-only corroboration | Corroborate before reuse |
stale | Aged out (aged_out) or superseded (superseded_by_newer) | Prefer a newer result |
failed | A failed/reverted attempt (failed_attempt) | Do not reuse |
The verdict is advisory metadata only — it never changes result ordering.
It is derived from metadata-only signals (recency, source health, realized
search mode, cwd-relative workspace match, and — opportunistically — linked
commit/bead/release provenance); it carries no raw session text. The same
trust block is attached to cass pack evidence. Branch on trust_tier and
stale_reason, not on confidence alone.
Provenance correlation is project-scoped and explicit-reference anchored:
for a hit from the project you are running cass in now, cass links it to a
closed bead or commit only when the hit's own indexed text references a known
identifier (bead:<id> or commit:<sha>), joined against that project's local
beads and git history. A linked commit's containing release is resolved from
Git. Release containment preserves provenance but does not establish proof of
the excerpt's claim: a landed commit remains proof_debt and cannot become
trusted from this correlation alone. A temporal or
workspace coincidence is never enough, so an unrelated conversation never
inherits another's trust. Off-project hits report workspace_mismatch, and a
hit whose local source file no longer exists on disk reports source_unhealthy
(archive-only) instead of overtrusting a dead path.
# Deterministic answer pack for handoff prompts
cass pack "why did checkout fail" --robot --max-tokens 12000 --limit 40
# Freshness-sensitive pack: fail if selected evidence is outside the window
cass pack "checkout timeout after redirect" --robot \
--freshness-policy strict --freshness-window-seconds 604800 \
--max-tokens 12000 --require-evidence
# Token-budgeted pack for pasting into another agent
cass pack "checkout timeout after redirect" --robot \
--max-tokens 4000 --max-evidence 8 --max-sessions 3 --max-excerpt-chars 600
# Pipeline from broad search to a bounded cited handoff
cass search "checkout timeout" --robot-format sessions \
| cass pack "checkout timeout root cause" --robot --sessions-from -
Design principle: stdout contains only parseable JSON data; all diagnostics, warnings, and progress go to stderr.
Use search when you are still exploring candidate sessions. Use pack when
you need a compact, cited, extractive artifact to hand to another agent or a
human operator. Use status/health before trusting freshness-sensitive output,
and use doctor only for diagnostics or safe repair workflows. Use
export-html when you need a full browsable session archive; packs are
token-budgeted evidence bundles, not full exports and not external
summarization.
Pack robot output includes health, freshness, privacy, and warnings.
Warnings such as privacy_redactions_applied, semantic_fallback_lexical,
or no_evidence_found are data, not prose; branch on the JSON fields before
copying the pack into another tool. Stale selected evidence is structural:
inspect freshness.stale_evidence_count.
Packs exclude injected skill payloads by default. Add --include-skill-content
to include them explicitly; credential redaction still applies.
privacy.skill_content_included reports whether the selected evidence includes
skill payloads, including after token-budget trimming.
Use the swarm surfaces when multiple agents are sharing one repo and you need a single read-only view before claiming work:
# Current shared-work snapshot; does not claim, reopen, release, or run builds
cass swarm status --json
# Advisory packet for one bead; still create real reservations and Beads updates yourself
cass swarm work-packet --json --bead coding_agent_session_search-example
# Coordination hygiene check before closeout or takeover review
cass swarm lint --json --bead coding_agent_session_search-example
# Read-only sibling dependency drift sentinel
cass swarm dependency-drift --json
swarm status and swarm work-packet collect bounded read-only Git state and
Beads exports when run from the repository root without a fixture. Git uses
porcelain-v2 with optional locks disabled. Beads uses br 0.6.x --no-db, so its
JSONL snapshot is explicitly partial: unexported database changes may exist.
Recheck Beads and reservations before claiming work. Child commands share a
single 15-second request budget and each has an 8 MiB output cap; Beads categories
cap at 512 issues.
Failures report unavailable providers and unknown summary counts, not zero work.
RCH contributes aggregate active/queued job counts, fleet slots and posture from
rch status --json (API 1.0, schema 1.0.0). Responses older than 60 seconds or
more than 5 seconds in the future are unavailable. Worker addresses, commands
and job details are omitted. This provider remains partial: local Cargo/CPU
state and build admission are unknown, even when RCH reports no active jobs.
Agent Mail roster and reservation reads are opt-in: set CASS_SWARM_AGENT_MAIL_URL
to the server's HTTP MCP endpoint and, if required, CASS_SWARM_AGENT_MAIL_TOKEN.
The reader uses only resources/read, never a local database fallback or inbox
read. Mail shares the total request budget with a 3-second cap of its own;
responses are capped at 8 MiB, rosters at 512 agents, and full 250-row reservation
pages are refused. The total reservation count remains unknown; source metadata
reports only the observed active count. Activity and expiry use
the observation time. Task descriptions, reservation reasons and message bodies
are omitted. These observations do not authorize claims or establish proof.
CASS evidence remains unwired. swarm lint still uses the placeholder
live snapshot. Fixture selection (--fixture <file> or --fixture-dir <dir> --fixture-id <id>) retains deterministic behavior; swarm dependency-drift
also has a live path.
swarm status composes Beads, Agent Mail metadata, git state, rch/build
pressure, cass health/status, and proof references. Stale candidates are
advisory only: coordinate through Beads and Agent Mail before reopening,
force-releasing, or taking over work. Suggested commands are robot-safe
templates, not automatic actions.
swarm dependency-drift reads Cargo.toml and optional sibling checkouts to
report manifest pins, local HEAD/dirty state, strict validation commands, and
release-risk recommendations. It does not fetch remotes, edit manifests, run
builds, update Beads, send Agent Mail, delete files, or mutate git state.
When status points at prior evidence, use cass pack "query" --robot to create
a bounded cited handoff for another agent. Packs complement the cockpit; they do
not replace Beads for ownership, Agent Mail for coordination, or rch for proof
commands.
LLMs have context limits. cass provides multiple levers to control output size:
| Flag | Effect |
|---|---|
--fields minimal | Only source_path, line_number, agent, source_id, conversation_id |
--fields summary | minimal plus title, score |
--fields score,title,snippet | Custom field selection |
--max-content-length 500 | Truncate long fields (UTF-8 safe, adds "...") |
--max-tokens 2000 | Soft budget (~4 chars/token); adjusts truncation dynamically |
--limit 5 | Cap number of results |
cass pack "query" --robot | Build a cited handoff pack from selected search evidence |
pack --max-tokens N | Set the pack planner's soft budget |
pack --max-evidence N | Cap evidence items selected into the pack |
pack --max-sessions N | Limit how many sessions can contribute evidence |
pack --max-excerpt-chars N | Shorten each cited excerpt before token estimation |
pack --fields summary | Return top-level summary fields for a smaller JSON envelope |
pack --field-mask minimal|standard|full | Select a documented pack projection; --fields accepts the same presets |
pack --freshness-policy strict --freshness-window-seconds N | Reject stale evidence instead of silently mixing it into a pack |
pack --sessions-from FILE | Restrict pack evidence to newline-delimited session paths; use - for stdin |
Truncated fields include a *_truncated: true indicator so agents know when they're seeing partial content.
Contributor verification for docs or contract changes should use rch, for example:
rch exec -- env CARGO_TARGET_DIR=${TMPDIR:-/tmp}/rch_target_cass_answer_pack_docs \
cargo test --test golden_robot_docs
Errors are structured, actionable, and include recovery hints. A real sample from cass search foo --robot against a fresh data dir:
{
"error": {
"code": 3,
"kind": "missing-index",
"message": "cass has not been initialized in <data_dir> yet, so search cannot run until the first index completes.",
"hint": "Run 'cass index --full' once to discover local sessions and build the initial archive.",
"retryable": true
}
}
Kind names are kebab-case (e.g. missing-index, missing-db, semantic-unavailable, embedder-unavailable, ambiguous-source, timeout, config, lock-busy). Agents that branch on err.kind should treat them as stable identifiers. The full set (about 90 kinds) is defined in src/model/cli_error_kind.rs; the canonical way to discover a kind programmatically is to trigger the condition and inspect err.kind from the JSON envelope.
Exit codes follow a semantic convention:
| Code | Meaning | Typical action |
|---|---|---|
| 0 | Success | Parse stdout |
| 1 | Health check failed | Run cass index --full |
| 2 | Usage error | Fix syntax (hint provided) |
| 3 | Index/DB missing | Run cass index --full (retryable: true) |
| 4 | I/O failure or unsafe operation refused (not a network code) | Branch on err.kind: fix path/permissions/space for io/output-not-writable; follow the hint for refused-unsafe |
| 5 | Data corruption, or maintenance required | Inspect cass health --json / cass status --json / cass doctor --json and follow recommended_action: usually rebuild derived assets (maintenance-required, checkpoint_incomplete start that rebuild themselves); only a canonical-archive failure needs repair or restore |
| 6 | Required input missing (password, resume command) | Supply the input (e.g. --password-stdin) and rerun |
| 7 | Lock/busy | Retry later |
| 8 | Partial result (sources sync only: some sources had path failures) | Inspect per-path errors in the JSON output and retry the failed sources |
| 9 | Unknown error | Check retryable flag |
| 10 | Config / timeout | Depends on err.kind |
| 11 | Config validation | Fix config |
| 12 | Source / SSH | Check remote host |
| 13 | Mapping / not-found | Depends on err.kind |
| 14 | I/O / mapping | Retry or inspect path |
| 15 | Semantic / embedder unavailable | Install model or --mode lexical |
| 20-21 | Model acquisition | Check err.kind, err.hint |
| 22 | I/O during model handling | Retry |
| 23 | Model download | Retry or use --from-file |
| 24 | I/O during model verify/install | Retry |
| 70 | cass index stalled and aborted (kind index-stalled envelope on stderr) | Inspect cass status --json, then rerun cass index |
| 130 | Interrupted (SIGINT) | Rerun; cass sources setup --resume continues an interrupted setup |
Search/pack timeouts are not exit 8: on expiry search and pack exit 0 with {"hits": [], "budget": {"timed_out": true, "skipped_sections": [...], "recommended_next_probe": "<command>", ...}}, and --robot-format sessions instead fails with exit 10, kind timeout. Explicit --mode semantic is the other exception: when the remaining budget cannot admit semantic setup or dispatch, search fails with exit 10, kind timeout, retryable: true, and a semantic_budget checkpoint=... message, rather than returning an empty or lexical result. Hybrid (explicit or default) instead falls back to lexical and reports semantic_budget_limited.
Codes ≥ 10 are domain-specific and the numeric value alone is ambiguous (e.g. code 10 maps to either config or timeout kinds depending on context). Agents should branch on err.kind from the JSON error envelope — not on the numeric code — when handling codes ≥ 10. See the Error Handling section above for the canonical kind list.
The retryable field tells agents whether a retry might succeed (e.g., transient I/O) vs. guaranteed failure (e.g., invalid path). A lexical query the engine refuses with posting cursor invariant failed (kind search, exit 9) is retryable: false: the same query fails the same way on the same index generation. The hint names the remedy, cass index --full --force-rebuild. Date-filtered searches over an index segment that holds deleted rows triggered it before frankensearch-quill 0.3.2 (GH #499).
Beyond search, cass provides commands for deep-diving into specific sessions:
# Discover the current session for this workspace
cass sessions --current --json
# List recent sessions for a specific project
cass sessions --workspace /path/to/project --json --limit 5
# Export full conversation to shareable format
cass export /path/to/session.jsonl --format markdown -o conversation.md
cass export /path/to/session.jsonl --format json --include-tools
# Export as self-contained HTML with encryption (recommended for sharing)
cass export-html /path/to/session.jsonl # To Downloads folder
printf '%s\n' "pwd" | cass export-html session.jsonl --encrypt --password-stdin
cass export-html session.jsonl --open --json # Open in browser, JSON output
# Common agent flow: find current session, then export it
cass export-html "$(cass sessions --current --json | jq -r '.sessions[0].path')" --json
# Expand context around a specific line (from search result)
cass expand /path/to/session.jsonl -n 42 -C 5 --json
# → Shows 5 messages before and after line 42
# Activity timeline: when were agents active?
cass timeline --today --json --group-by hour
cass timeline --since 7d --agent claude --json
# → Grouped activity counts, useful for understanding work patterns
Aggregate search results server-side to get counts and distributions without transferring full result data:
# Count results by agent
cass search "error" --robot --aggregate agent
# → { "aggregations": { "agent": { "buckets": [{"key": "claude_code", "count": 45}, ...] } } }
# Multi-field aggregation
cass search "bug" --robot --aggregate agent,workspace,date
# Combine with filters
cass search "TODO" --agent claude --robot --aggregate workspace
Aggregation Fields:
| Field | Description |
|---|---|
agent | Group by agent type (claude_code, codex, cursor, etc.) |
workspace | Group by workspace/project path |
date | Group by date (YYYY-MM-DD) |
match_type | Group by match type (exact, prefix, suffix, substring, wildcard, implicit_wildcard); one search has one type, so this shows a single bucket unless the wildcard fallback replaced the hits |
Response Format:
{
"aggregations": {
"agent": {
"buckets": [
{"key": "claude_code", "count": 120},
{"key": "codex", "count": 85}
],
"other_count": 15
}
}
}
Top 10 buckets are returned per field, with other_count for remaining items.
Mine recurrent CASS operational incidents from the canonical archive without dumping raw session text:
cass analytics incidents --limit 10 --json
# Tighten the bounded scan for automation or a very large archive
cass analytics incidents --max-sessions 500 --max-messages 50000 \
--max-bytes 67108864 --budget-ms 5000 --json
The response ranks top_sessions[] by hit count and category breadth and keeps
the exact conversation_id, agent, host, source_id, source_path,
live/archive state, dominant categories, and a structured cass view argv.
That argv carries the effective --db path plus --conversation-id, so it
opens the exact ranked archive row even when multiple sessions share a source
path or the report used a non-default database.
total_sessions, total_hits, and top_sessions_truncated distinguish the
bounded ranked result from the totals observed inside the scan scope.
discovery.partial and stop_reason explicitly distinguish a bounded partial
scan from a complete scan. Counts are scoped to scanned candidates whenever the
scan is partial. --budget-ms is a hard wall-clock result guard around the
independently row-bounded read-only worker. If it expires before a verified
result arrives, count_scope="no_verified_results_hard_timeout" returns an
empty partial report instead of overstating in-flight observations. Candidate
discovery is descending archive-row keyset paging;
--max-sessions bounds that newest-row window before dimensional filters, so a
selective filter can truthfully return a partial empty result instead of scanning
an unbounded archive. Individual messages are inspected through a bounded 4,096-char
fragment; an oversized message returns message-fragment-capped rather than
claiming a complete corpus scan. Raw prompt/tool content is always suppressed; evidence carries
only BLAKE3 fingerprints and basename-redacted paths. The actionable
source_path remains visible solely so the returned view command works.
Chain multiple searches together by piping session paths from one search to another:
# Find sessions mentioning "auth", then search within those for "token"
cass search "authentication" --robot-format sessions | \
cass search "refresh token" --sessions-from - --robot
# Build a filtered corpus from today's work
cass search --today --robot-format sessions > today_sessions.txt
cass search "bug fix" --sessions-from today_sessions.txt --robot
How It Works:
--robot-format sessions outputs one session path per line--sessions-from <file> restricts search to those sessions- to read from stdin for true pipingUse Cases:
Snippets always mark the terms the search engine matched with **bold**, in human-readable output and in robot/JSON output alike. The --highlight flag also marks the query terms' remaining literal occurrences and leaves already-marked text alone, so no term gets two pairs of marks. Only the snippet field is marked; content stays verbatim:
cass search "authentication error" --robot --highlight
# "snippet": "... **authentication** failed with **error** ..."
Highlighting is query-aware: quoted phrases like "auth error" highlight as a unit; individual terms highlight separately.
For large result sets, use cursor-based pagination:
# First page
cass search "TODO" --robot --robot-meta --limit 20
# → { "hits": [...], "_meta": { "next_cursor": "eyJ..." } }
# Next page
cass search "TODO" --robot --robot-meta --limit 20 --cursor "eyJ..."
A cursor is base64 JSON {"offset": N, "limit": M}: a plain page position, not a snapshot. Any index change between pages (a new session indexed, a rebuild, a forget) shifts the ranking, so the next page can skip or repeat hits. Page quickly, or fix the window with --until when you need stable pages.
total_matches answers "how many messages match?", but it is exact only when
cass can afford to count. Every search fetches one hit beyond --limit to learn
whether another page exists. When the page is full and the index is larger than
CASS_SEARCH_EXACT_TOTAL_COUNT_MAX_DOCS documents, cass skips the full count
and reports that limit + 1 as a lower bound. Below the threshold it counts
every match.
| Build | Threshold | stale lock with --limit 10 on a 1,034,219-document index |
|---|---|---|
| v0.9.0 and earlier | 50,000 documents | total_matches: 11 |
| Current (unreleased) | 5,000,000 documents | total_matches: 11915 |
--robot-meta says which kind of number you got:
cass search "stale lock" --robot --robot-meta --limit 10 \
| jq '{total_matches, precision: ._meta.cursor_manifest.count_precision, why: ._meta.cursor_manifest.count_reason}'
# → {"total_matches": 11915, "precision": "exact", "why": "total_matches is exact; no extra recount was needed"}
# A lower bound reads "precision": "lower_bound".
Why the threshold moved. The 50,000 cap dates from the Tantivy engine,
where counting a common term over millions of documents could dominate the
query. The Quill engine counts cheaply. Paired runs on that 1,034,219-document
archive, capped against exact (--limit 10, read-only, CPU time):
| Query | Exact total | Extra CPU for the exact count |
|---|---|---|
stale lock | 11,915 | ~0.00-0.04 s |
cargo build | 24,944 | ~0.01-0.03 s |
the | 439,461 | ~0.03-0.05 s |
AGENTS.md | 867,087 | ~0.06-0.11 s |
Each search cost about 0.8 s of CPU either way. The capped answer, meanwhile,
was wrong by up to five orders of magnitude, and agents read total_matches as
a count. The default now covers five times that archive; set
CASS_SEARCH_EXACT_TOTAL_COUNT_MAX_DOCS=0 to never count exactly, or raise it
for a larger archive.
Aggregations have their own window. --aggregate buckets are computed over
the top max(1000, limit + offset) hits, so bucket counts on a large archive
describe the best-ranked thousand matches, not the whole corpus. Use an exact
total_matches for "how many", and aggregations for "how are the top hits
distributed".
For debugging and logging, attach a request ID:
cass search "bug" --robot --request-id "req-12345"
# → { "request_id": "req-12345", "hits": [...], ... }
# (top level always; also under _meta.request_id with --robot-meta)
For safe retries (e.g., in CI pipelines or flaky networks):
cass index --full --idempotency-key "build-$(date +%Y%m%d)"
# If same key + params were used in last 24h, returns cached result
Debug why a search returned unexpected results:
cass search "auth*" --robot --explain
# → Adds "explanation": the sanitized query, flat lists of its terms, phrases and
# operators (not a tree), the query type and index strategy, a low/medium/high cost
# class, a filter summary and warnings. Wildcards are reported, not expanded.
cass search "auth error" --robot --dry-run
# → Validates query syntax without executing
For debugging agent pipelines:
cass search "error" --robot --trace-file /tmp/cass-trace.json
# Appends execution span with timing, exit code, and command details
cass index --full --json --robot-trace-ingest 2>/tmp/cass-ingest-trace.jsonl
# Streams one NDJSON record per ingest batch with wall_ms, batch_msgs,
# inserted_messages, and duplicate-lookup counters for perf bisects
| Flag | Purpose |
|---|---|
--robot / --json | JSON output (pretty-printed) |
--robot-format jsonl|compact | Streaming or single-line JSON |
--robot-meta | Include _meta block (elapsed_ms, cache stats, index freshness, lexical_degrade_reason: "query_fuel_exhausted" or null, wildcard_fallback_skipped: why a sparse result got no automatic wildcard retry, and effective: the database, time window, filters, auto-corrections and query grouping the search actually used) |
--fields minimal|summary|<list> | Reduce payload size |
--max-content-length N | Truncate content fields to N chars |
--max-tokens N | Apply an approximate token budget to robot output |
--timeout N | Timeout in milliseconds. On expiry search/pack still exit 0 and emit {"hits": [], "budget": {"timed_out": true, "skipped_sections": [...], "recommended_next_probe": "<command>", ...}}; --robot-format sessions fails with exit 10, kind timeout |
--cursor <token> | Cursor-based pagination (from _meta.next_cursor) |
--request-id ID | Echoed in response for correlation |
--aggregate agent,workspace,date | Server-side aggregations |
--explain | Include query analysis (parsed query, cost estimate) |
--dry-run | Validate query without executing |
--no-maintenance | Strict read-only search: never refresh, join, or spawn lexical maintenance, never auto-repair the archive while opening it, and never auto-spawn the daemon (conflicts with --refresh and --daemon) |
--source <source> | Filter by source: local, remote, all, or specific source ID |
--highlight | Also mark query-term occurrences the engine left unmarked (snippets always mark matched terms with **) |
| Flag | Purpose |
|---|---|
--idempotency-key KEY | Safe retries: same key + params returns cached result (24h TTL) |
--json | JSON output with stats |
--gc | Reclaim merge-retired lexical segment files and exit: runs the engine's grace-period garbage sweep (a folded segment file is unlinked only once no published MANIFEST generation has referenced it for 300 s) and reports files/bytes reclaimed. Every incremental cass index performs the same sweep at open; doctor --json reports the reclaimable bytes under storage_pressure.full_rebuild_readiness (GH #453) |
When health --json or status --json reports index.status: "hollow", the
live Quill generation serves fewer than half the documents certified by its
completed rebuild checkpoint. index.live_documents reports the served count.
Run cass index to let its pre-scan repair rebuild from the canonical archive;
cass index --full also rescans the session sources. A missing count provides
no hollow-generation verdict.
For machine-readable documentation, use cass robot-docs <topic>:
| Topic | Content |
|---|---|
commands | Full command reference with all flags |
env | Environment variables and defaults |
paths | Data directory locations per platform |
guide | Quick start guide for automation |
schemas | JSON response schemas |
exit-codes | Exit code meanings and retry guidance |
examples | Copy-paste usage examples |
contracts | API contract version and stability |
sources | Remote sources configuration guide |
# Get documentation programmatically
cass robot-docs guide
cass robot-docs schemas
cass robot-docs exit-codes
# Machine-first help (wide output, no TUI assumptions)
cass --robot-help
cass maintains a stable API contract for automation:
cass api-version --json
# → { "crate_version": "<cargo version>", "build_commit": "<sha or unknown>", "api_version": 1, "contract_version": "1" }
cass introspect --json
# → Full schema: all commands, arguments, response types
Contract Version: Currently 1. Increments only on breaking changes.
Guaranteed Stable:
--robot output_meta block format🔎 cass — Search All Your Agent History
What: cass indexes conversations from Claude Code, Codex, Cursor, Gemini, Aider, ChatGPT, and more into a unified, searchable index. Before solving a problem from scratch, check if any agent already solved something similar.
⚠️ NEVER run bare cass — it launches an interactive TUI. Always use --robot or --json.
Quick Start
# One-shot agent triage (read next_command when present)
cass triage --json
# Search across all agent histories
cass search "authentication error" --robot --limit 5
# Build a cited handoff pack from search evidence
cass pack "authentication error root cause" --robot --max-tokens 12000 --limit 40
# Tight handoff budget with freshness and privacy metadata
cass pack "authentication error root cause" --robot --max-tokens 4000 --max-evidence 8 --fields summary
# View a specific result (from search output)
cass view /path/to/session.jsonl -n 42 --json
# Expand context around a line
cass expand /path/to/session.jsonl -n 42 -C 3 --json
# Learn the full API
cass capabilities --json # Static agent self-description
cass robot-docs guide # LLM-optimized docs
Why Use It
- Cross-agent knowledge: Find solutions from Codex when using Claude, or vice versa
- Forgiving syntax: Typos and wrong flags are auto-corrected with teaching notes
- Token-efficient: --fields minimal returns only essential data; pack budgets cite only selected evidence
- Copy-safe handoffs: pack warnings include freshness and privacy/redaction status
Key Flags
| Flag | Purpose |
|------------------|--------------------------------------------------------|
| --robot / --json | Machine-readable JSON output (required!) |
| --fields minimal | Reduce payload: source_path, line_number, agent, source_id, conversation_id |
| pack --max-tokens N | Budget a cited handoff pack |
| --limit N | Cap result count |
| --agent NAME | Filter to specific agent (claude, codex, cursor, etc.) |
| --days N | Limit to recent N days |
stdout = data only, stderr = diagnostics. Exit 0 = success.
cass supports a rich query syntax designed for both humans and machines.
| Query | Matches |
|---|---|
error | Messages containing "error" (case-insensitive) |
python error | Messages containing both "python" AND "error" |
"authentication failed" | Exact phrase match |
auth fail | Both terms, in any order |
Combine terms with explicit operators for complex queries:
| Operator | Example | Meaning |
|---|---|---|
AND | python AND error | Both terms required (default) |
OR | error OR warning | Either term matches |
NOT | error NOT test | First term, excluding second |
- | error -test | Shorthand for NOT |
Operator Precedence: NOT binds tightest, then AND (explicit, &&, or implied between words), then OR (OR, ||). Parentheses group, in the TUI and robot mode alike: a OR b c means a OR (b AND c), while (a OR b) c needs the parentheses. A ( groups only at the start of a word, so code such as foo(bar) stays one term. NOT NOT x is x. Unbalanced parentheses are recovered rather than rejected. The SQLite fallback lanes, used while no lexical index is available, apply the same grammar.
# Complex boolean query
cass search "authentication AND (error OR failure) NOT test" --robot
# Exclude test files
cass search "bug fix -test -spec" --robot
# Either error type
cass search "TypeError OR ValueError" --robot
Wrap terms in double quotes for exact phrase matching:
| Query | Matches |
|---|---|
"file not found" | Exact sequence "file not found" |
"cannot read property" | Exact JavaScript error message |
"def test_" | Function definitions starting with test_ |
Phrases match their words adjacent and in order (no slop). Useful for error messages, code patterns, and specific terminology.
| Pattern | Type | Matches | Performance |
|---|---|---|---|
auth* | Prefix | "auth", "authentication", "authorize" | Fast (uses edge n-grams) |
*tion | Suffix | "authentication", "function", "exception" | Slower (term-dictionary expansion) |
*config* | Substring | "reconfigure", "config.json", "misconfigured" | Slowest (term-dictionary expansion) |
test_* | Prefix on test | anything whose token starts with "test" | Fast |
Tip: Prefix wildcards (foo*) use edge n-grams computed at index time: prefixes of 2–20 characters of every alphanumeric word, taken from titles and from the first 4 KiB of each message. Suffix and substring wildcards expand over the index's term dictionary, at most 16,384 terms per pattern; a pattern matching more terms fails rather than scanning. The tokenizer splits on anything that is not a letter or digit: test_* is a prefix match on test, c++ searches for c, and foo.bar means foo AND bar anywhere in the message, not the literal string. Phrases ("...") match adjacent words in order (slop 0).
# Field-specific search (in robot mode)
cass search "error" --agent claude --workspace /path/to/project
# Time-bounded search
cass search "bug" --since 2024-01-01 --until 2024-01-31
cass search "bug" --today
cass search "bug" --days 7
# Combined filters
cass search "authentication" --agent codex --workspace myproject --week
cass accepts a wide variety of time/date formats for filtering:
| Format | Examples | Description |
|---|---|---|
| Relative | -7d, -24h, -30m, -1w | Days, hours, minutes, weeks ago |
| Keywords | now, today, yesterday | Named reference points |
| ISO 8601 | 2024-11-25, 2024-11-25T14:30:00Z | Standard datetime |
| US Dates | 11/25/2024, 11-25-2024 | Month/Day/Year |
| Unix Timestamp | 1732579200 | Seconds since epoch |
| Unix Millis | 1732579200000 | Milliseconds (auto-detected) |
Intelligent Heuristics:
2024, not 24)today, yesterday) starts at local midnight as --since and runs through the day's last millisecond as --until, so --until 2024-01-31 includes January 31search and pack, a --since/--until value that cannot be parsed, or a --since later than --until, is a usage error (exit 2, kind usage); it is never silently ignored# All equivalent for "last week"
cass search "bug" --since -7d
cass search "bug" --since "-1w"
cass search "bug" --days 7
# Date range
cass search "feature" --since 2024-01-01 --until 2024-01-31
# Mix formats
cass search "error" --since yesterday --until now
Search results include a match_type indicator. It describes the query, not each hit: every hit of a search carries the type of the least precise pattern in the query (e.g. auth* *tion stamps suffix on all hits).
| Type | Meaning | Match Quality order |
|---|---|---|
exact | No wildcards: exact terms (and edge n-gram prefixes) | 1 |
prefix | Trailing wildcard (auth*) | 2 |
suffix | Leading wildcard (*tion) | 3 |
substring | Both sides (*config*) | 4 |
wildcard | Inner wildcard (f*o) | 5 |
implicit_wildcard | Automatic wildcard fallback on a sparse exact search | 6 |
Relevance scores carry no boost for the match type; only the TUI's Match Quality ranking mode (F12) orders by it.
When an exact query's first page returns fewer than 3 results (or fewer than a smaller --limit), cass retries with wildcard expansion:
auth → *auth*CASS_AUTOMATIC_WILDCARD_FALLBACK_MAX_DOCS; 0 disables it), so on a typical real archive it does not run._meta.wildcard_fallback: true. When a sparse result did not get the retry, _meta.wildcard_fallback_skipped says why: index_over_automatic_limit (the index is over the document cap), automatic_retry_disabled (the cap is 0) or long_query_term; add explicit wildcards to run it anyway| Key | Action |
|---|---|
Ctrl+C | Force quit |
Esc / F10 | Unwind: close the open modal or surface, otherwise quit |
F1 / Alt+? | Toggle help screen |
F2 / Alt+T | Next theme (cycles all 19 presets) |
Shift+F2 / Alt+Shift+T | Previous theme |
Ctrl+B | Toggle border style (rounded/square) |
Ctrl+P / Alt+P | Open the command palette |
Ctrl+S | Toggle the stats bar |
Ctrl+Shift+S | Open the sources management surface |
Alt+A | Open the analytics dashboard |
Alt+M | Toggle macro recording (replay with cass tui --play-macro FILE) |
Ctrl+Shift+I | Toggle the inspector overlay |
Ctrl+Shift+R | Force re-index |
Ctrl+Shift+Del | Reset all TUI state |
Ctrl+Z / Ctrl+Shift+Z | Undo / redo |
Launch-time flags: cass tui --refresh (alias --catch-up) runs an incremental index pass before opening; --record-macro FILE / --play-macro FILE record and replay input events.
| Key | Action |
|---|---|
| Type | Live search as you type; plain characters (including ?, y, o, c, 1-9, -, =) go into the query |
Enter | Open the selected hit; with no selected hit, submit the query (if the query is empty, edit the last filter chip) |
Backspace | Delete character; if the query is empty, remove the last filter chip |
Left/Right, Ctrl+Left/Ctrl+Right | Move the cursor by character / by word |
Home/End | Jump the cursor to the start / end of the query |
Ctrl+L | Clear the query |
Ctrl+U / Ctrl+K / Ctrl+W | Kill to line start / to line end / previous word |
Ctrl+R | Cycle through query history |
Ctrl+N / Ctrl+Shift+N | Next / previous query-history entry |
Ctrl+F | Toggle wildcard fallback |
Ctrl+Shift+Y | Copy the query |
| Key | Action |
|---|---|
Up/Down | Move selection in results list |
PageUp/PageDown | Scroll by page |
Tab / Shift+Tab | Toggle focus between results and detail pane / move focus left |
Alt+h/j/k/l | Vim-style directional focus (left/down/up/right) |
Alt+1..Alt+9 | Switch to pane N |
Alt+- / Alt+= | Shrink / grow the results pane |
Alt+D | Hide / show the detail pane |
Alt+[ / Alt+] | Timeline jump backward / forward |
| Key | Action |
|---|---|
F3 / Alt+G | Open agent filter palette |
Shift+F3 / Alt+Shift+G | Clear the agent filter |
F4 / Alt+W | Open workspace filter palette |
Shift+F4 / Alt+Shift+W / Ctrl+Del | Clear all active filters |
F5 | Set "from" time filter |
F6 | Set "to" time filter |
Shift+F5 | Cycle time presets: 24h → 7d → 30d → all |
F11 / Shift+F11 | Cycle the source filter / open the source filter menu |
Alt+/ | Open the pane filter |
| Key | Action |
|---|---|
F7 / Alt+C | Cycle context window size: S → M → L → XL |
Ctrl+Space | Momentary "peek" to XL context |
F9 | Toggle match mode: standard (default) ↔ prefix, where every bare word of 2+ characters also matches as a prefix (auth → auth*; phrases, operators and wildcards are left as typed) |
F12 / Alt+R | Cycle ranking: recent → balanced → relevance → quality → newest → oldest |
Alt+F | Cycle result grouping: agent → conversation → workspace → flat |
Alt+S | Cycle search mode (lexical / semantic / hybrid). Without an installed model (or vector index) the status line says results stay lexical and names the command: cass models install (offline --from-file <dir>) or cass index --semantic; nothing downloads on its own |
Ctrl+D | Cycle density: Compact → Cozy → Spacious |
Ctrl+1..Ctrl+9 | Save the current view to slot N |
Shift+1..Shift+9 | Load the view from slot N |
For one chronological list across agents, select flat grouping with Alt+F
and newest ranking with F12. Grouping is also available in the command
palette and is preserved with ranking and filters in saved views.
| Key | Action |
|---|---|
Enter / Ctrl+M | Open selected result in the detail modal (Messages tab by default) |
Ctrl+X | Toggle selection on current result |
Ctrl+A | Select/deselect all visible results |
Alt+B | Open bulk actions menu (when items selected) |
Ctrl+Enter | Add to multi-open queue |
Ctrl+O | Open all queued items in editor |
F8 / Alt+O | Open selected hit in $EDITOR |
Alt+V | View raw |
Alt+Shift+J | Toggle JSON view |
Ctrl+Y | Copy path |
Alt+Y | Copy snippet |
Ctrl+Shift+C | Copy content |
Ctrl+E | Open the export modal |
Ctrl+Shift+E | Export Markdown immediately |
Alt+U / Alt+N / Alt+I | Update banner: upgrade now / show release notes / skip this version |
These apply while the detail modal is open:
| Key | Action |
|---|---|
Esc | Close the detail modal |
Tab | Cycle detail tabs |
/ (or Ctrl+F, Alt+/) | Start find-in-detail; type to search, Enter advances to the next match |
n / N | Next / previous contextual search hit within this session |
Enter (Messages tab) | Next contextual search hit |
j / k, Up/Down | Scroll |
g / G, Home/End | Scroll to top / bottom |
{ / } | Jump to previous / next message |
[ / ] | Jump to previous / next user message |
w | Toggle line wrap |
e / c | Expand / collapse all tool and system messages |
e, h (Export tab) | Open the HTML export modal; m exports Markdown |
F7 | Cycle context window size |
Ctrl+Space | Momentary "peek" to XL context |
The detail pane has six tabs, cycled with Tab:
| Tab | Content | Best For |
|---|---|---|
| Messages | Full conversation with markdown rendering | Reading full context |
| Snippets | Keyword-extracted summaries | Quick scanning |
| Raw | Unformatted JSON/text | Debugging, copying exact content |
| Json | Syntax-highlighted, pretty-printed JSON (static; no collapsible tree) | Inspecting structured payloads |
| Analytics | Per-session token timeline, tool calls, message stats | Understanding one session |
| Export | Export actions and filename previews (HTML/Markdown) | Sharing a session |
Control how much content shows in the detail preview. Cycle with F7:
| Size | Characters | Use Case |
|---|---|---|
| Small | ~200 | Quick scanning, narrow terminals |
| Medium | ~400 | Default balanced view |
| Large | ~800 | Reading longer passages |
| XLarge | ~1600 | Full context, code review |
Peek Mode (Ctrl+Space): Temporarily expand to XL context. Press again to restore previous size. Useful for quick deep-dives without changing your preferred default.
Efficiently work with multiple search results at once:
Multi-Select Mode:
Ctrl+X to toggle selection on current result (checkbox appears)Ctrl+X againCtrl+A to select/deselect all visible resultsBulk Actions Menu (Alt+B when items selected):
| Action | Description |
|---|---|
| Open All | Open all selected files in editor |
| Copy Paths | Copy all file paths to clipboard |
| Export | Export selected results to file |
| Clear Selection | Deselect all items |
Multi-Open Queue: For opening many files without navigating away:
Ctrl+Enter to add current result to queueCtrl+O to open all queued itemsClipboard Operations:
Ctrl+Y - Copy the current item's pathAlt+Y - Copy the current item's snippetCtrl+Shift+C - Copy the current item's contentCycle through modes with F12 (or Alt+R) in the TUI. The search engine returns hits in its own relevance order; the mode then re-orders the results the TUI has loaded (the first page of up to 250 hits, plus every further page you load), so the whole loaded list always follows one order. Ranking modes are TUI-only; robot search returns engine order.
Recent Heavy: Score = relevance × 0.3 + recency × 0.7. Best for: "What was I working on?"
Balanced (default): Score = relevance × 0.5 + recency × 0.5. Best for general-purpose search.
Relevance: Score = relevance × 0.8 + recency × 0.2. Best for "find the best explanation of X".
Match Quality: exact matches first, then prefix, suffix, substring, wildcard, and finally automatic wildcard-fallback matches; within each class, the Relevance score. Best for precise technical searches.
Date Newest: newest first by message time. Best for "show me recent activity".
Date Oldest: oldest first. Best for "when did I first work on this?"
Undated hits sort last in the date modes. Hits with equal scores keep the engine's order. With an empty query, cass browses by date instead of searching; Date Oldest browses oldest first and every other mode newest first.
Relevance: the engine's score (BM25 for lexical search, fused rank for hybrid), min-max normalized over the loaded hits: the best loaded hit is 1.0 and the weakest 0.0 (all 1.0 when every score ties).
Recency: exponential decay from now with a 14-day half-life, 0.5 ^ (age_days / 14): 1.0 today, about 0.71 after a week, 0.5 after two weeks, about 0.23 after a month. Undated hits get 0.
The formulas live in src/ui/ranking.rs, and the tests in that file and in tests/ranking.rs exercise the same function the TUI calls.
Each connector transforms agent-specific formats into a unified schema:
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ Agent Files │ ──▶ │ Connector │ ──▶ │ Normalized │
│ (proprietary) │ │ (per-agent) │ │ Conversation │
└─────────────────┘ └──────────────────┘ └─────────────────┘
JSONL detect() agent_slug
SQLite scan() workspace
Markdown messages[]
JSON created_at
Different agents use different role names:
| Agent | Original | Normalized |
|---|---|---|
| Claude Code | human, assistant | user, assistant |
| Codex | user, assistant | user, assistant |
| ChatGPT | user, assistant, system | user, assistant, system |
| Cursor | user, assistant | user, assistant |
| Aider | (markdown headers) | user, assistant |
Agents store timestamps inconsistently:
| Format | Example | Handling |
|---|---|---|
| Unix milliseconds | 1699900000000 | Direct conversion |
| Unix seconds | 1699900000 | Multiply by 1000 |
| ISO 8601 | 2024-01-15T10:30:00Z | Parse with chrono |
| Missing | null | Use file modification time |
Tool calls, code blocks, and nested structures are flattened for searchability:
// Original (Claude Code)
{"type": "tool_use", "name": "Read", "input": {"path": "/foo/bar.rs"}}
// Flattened for indexing
"[Tool: Read] path=/foo/bar.rs"
The same conversation content can appear multiple times due to:
cass uses a multi-layer deduplication strategy:
Message identity: messages are keyed by UNIQUE(conversation_id, idx). Appends to a known conversation use INSERT OR IGNORE, and a new conversation's batched INSERT has the same unique index as its backstop, so re-indexing the same file never stores a message twice
Conversation identity: conversations are keyed by UNIQUE(source_id, agent_id, external_id)
Search-Time Dedup: hits are deduplicated on an exact key tuple — (source, source path, conversation id or title, line number, created_at, whitespace-invariant content hash) — keeping the highest-scored hit
Common low-value content is filtered from results:
# Find past solutions for similar errors
cass search "TypeError: Cannot read property" --days 30
# In TUI: F12 to switch to "relevance" mode for best matches
# What has ANY agent said about authentication in this project?
cass search "authentication" --workspace /path/to/project
# Export findings for a new agent's context
cass export /path/to/relevant/session.jsonl --format markdown
# What did I work on today?
cass timeline --today --json | jq '.groups[].conversations'
# TUI: Press Shift+F5 to cycle through time filters
# Find all debugging sessions for a specific file
cass search "debug src/auth/login.rs" --agent claude
# Expand context around a specific line in a session
cass expand /path/to/session.jsonl -n 150 -C 10
# Current agent searches what previous agents learned
cass search "database migration strategy" --robot --fields minimal
# Get full context for a relevant session
cass view /path/to/session.jsonl -n 42 --json
# Export high-quality problem-solving sessions
cass search "bug fix" --robot --limit 100 | \
jq '.hits[] | select(.score > 0.8)' > training_candidates.json
Press Ctrl+P to open the command palette—a fuzzy-searchable menu of all available actions.
| Command | Description |
|---|---|
| Toggle theme | Switch between dark/light mode |
| Toggle density | Cycle Compact → Cozy → Spacious |
| Toggle help strip | Pin/unpin the contextual help bar |
| Check updates | Show update assistant banner |
| Filter: agent | Open agent filter picker |
| Filter: workspace | Open workspace filter picker |
| Filter: today | Restrict results to today |
| Filter: last 7 days | Restrict results to past week |
| Filter: date range | Prompt for custom since/until |
| Saved views | List and manage saved view slots |
| Save view to slot N | Save current filters to slot 1-9 |
| Load view from slot N | Restore filters from slot 1-9 |
| Bulk actions | Open bulk menu (when items selected) |
| Reload index/view | Refresh the search reader |
Ctrl+P to openUp/Down to navigateEnter to executeEsc to closeSave your current filter configuration to one of 9 slots for instant recall.
| Key | Action |
|---|---|
Ctrl+1 through Ctrl+9 | Save current view to slot |
Shift+1 through Shift+9 | Load view from slot |
Ctrl+P → "Save view to slot N"Ctrl+P → "Load view from slot N"Ctrl+P → "Saved views" to list all slotsViews are stored in tui_state.json and persist across sessions. Clear all saved views with Ctrl+Shift+Del (resets all TUI state).
Control how many lines each search result occupies. Cycle with Ctrl+D or via the command palette.
| Mode | Lines per Result | Best For |
|---|---|---|
| Compact | 2 | Maximum results visible, scanning many items |
| Cozy (default) | 5 | Balanced view with context |
| Spacious | 6 | Detailed preview, fewer results |
The pane automatically adjusts how many results fit based on terminal height and density mode.
cass includes a sophisticated theming system with multiple presets, accessibility-aware color choices, and adaptive styling.
Cycle through 19 built-in theme presets with F2:
| Theme | Description | Best For |
|---|---|---|
| Tokyo Night (default) | Deep blues with restrained contrast | Low-light environments, extended sessions |
| Daylight | High-contrast light background | Bright environments, presentations |
| Catppuccin Mocha | Warm pastels, reduced eye strain | All-day coding, aesthetic preference |
| Dracula | Purple-accented dark theme | Popular among developers, familiar feel |
| Nord | Arctic-inspired cool tones | Calm, focused work sessions |
| Solarized Dark | Precisely tuned low-contrast palette | Long editing sessions, monitor-agnostic |
| Solarized Light | Solarized on a cream background | Paper-style readability in bright rooms |
| Monokai | Classic warm dark palette | Familiar Sublime/TextMate feel |
| Gruvbox Dark | Retro earth tones on dark | Warmer alternative to Tokyo Night |
| One Dark | Atom's signature balanced dark | Moderate contrast, friendly defaults |
| Rosé Pine | Soho-inspired muted roses | Gentle contrast, boutique look |
| Everforest | Forest-inspired green-brown palette | Calm, nature-adjacent mood |
| Kanagawa | Japanese ink-and-paper theme | Artistic, quietly distinctive |
| Ayu Mirage | Ayu's balanced muted dark | Blue-teal accents, relaxed contrast |
| Nightfox | Fox-inspired warm dark | Deep violets with orange highlights |
| Cyberpunk Aurora | Neon aurora on obsidian | Showy, high-saturation dark |
| Synthwave '84 | Retro neon magenta/cyan | 80s aesthetic, fun demos |
| High Contrast | Maximum readability | Accessibility needs, bright monitors |
| Colorblind | Deuteranopia/protanopia-safe palette | Color-vision-deficient users |
Every theme preset is checked in the test suite against WCAG contrast ratios, with these floors:
These floors are below WCAG AA's 4.5:1 for body text, so cass does not claim AA conformance. At runtime, role and status badges pick whichever candidate foreground has the highest contrast against their background.
Conversation messages are color-coded by role for quick visual parsing:
| Role | Visual Treatment | Purpose |
|---|---|---|
| User | Blue-tinted background, bold | Your input, easy to scan |
| Assistant | Green-tinted background | AI responses |
| System | Gray/muted background | Context, instructions |
| Tool | Orange-tinted background | Tool calls, file operations |
Each agent type (Claude, Codex, Cursor, etc.) also receives a subtle tint, making multi-agent result lists instantly scannable.
Border decorations adapt to terminal width and to render pressure:
| Condition | Style | Example |
|---|---|---|
| Narrow (<80 cols) | Square box-drawing | ┌─ content ─┐ |
| 80 cols and wider | Rounded corners | ╭─ content ─╮ |
| Frame budget under pressure | Square, then no borders | ┌─┐, then none |
There is no double-line tier. Ctrl+B toggles between rounded and square Unicode borders; both are box-drawing characters, not ASCII.
Bookmarks are a CLI feature: cass bookmarks add|list|remove|search|export|import --json manages user-authored annotations on search results (a source path, optional line number, note, and tags). The TUI has no bookmark keybindings today.
# Bookmark a search hit (source_path + line_number from search output)
cass bookmarks add /path/to/session.jsonl -n 42 --title "JWT refresh fix" \
--note "Good explanation of the refresh flow" --tags "auth,jwt" --json
# List (optionally by tag), search notes/titles/snippets, remove by id
cass bookmarks list --tag auth --json
cass bookmarks search "refresh" --json
cass bookmarks remove 1 --json # exit 13 (`bookmark-not-found`) if the id is unknown
# Back up and restore
cass bookmarks export -o bookmarks.json --json
cass bookmarks import bookmarks.json --json
bookmarks.db (SQLite), separate from the search index and never pruned by doctor/cleanup flowslist can filter by tag{
"id": 1,
"title": "Auth bug fix discussion",
"source_path": "/path/to/session.jsonl",
"line_number": 42,
"agent": "claude_code",
"workspace": "/projects/myapp",
"note": "Good explanation of JWT refresh flow",
"tags": "auth, jwt, important",
"snippet": "The token refresh logic should..."
}
Bookmarks are stored separately from the main index:
~/.local/share/coding-agent-search/bookmarks.db~/Library/Application Support/coding-agent-search/bookmarks.db%APPDATA%\coding-agent-search\bookmarks.dbcass uses a non-intrusive toast notification system for transient feedback—operations complete, errors occur, or state changes without modal dialogs interrupting your workflow.
| Type | Icon | Auto-Dismiss | Use Case |
|---|---|---|---|
| Info | i | 3 seconds | Status updates, tips |
| Success | * | 2 seconds | Operations completed |
| Warning | ! | 4 seconds | Non-critical issues |
| Error | x | 6 seconds | Failures requiring attention |
Toasts feature:
| Trigger | Toast |
|---|---|
| Bulk copy of selected paths | * "Copied 3 paths" |
| Bulk export | * "Exported 3 items as JSON" |
| Copy failure | x "Copy failed: ..." |
| Slow search | "Slow search: 1840ms" |
| Semantic refinement failure | "Refinement failed: ..." |
| Saved views | * "Renamed slot 2", ! "Slot 4 is empty" |
To achieve sub-60ms latency on large datasets, cass implements a multi-tier caching strategy in src/search/query.rs:
prefix_cache is organized into shards (default 256 entries each) that bound per-prefix memory; all shards sit behind one Mutex, so the sharding limits size, not lock contention. The cache lives in the searching process: it pays off in the TUI, where each keystroke re-queries, and not for one-shot CLI searches, which start with an empty cache.WarmJob thread watches the input. When the user pauses typing, it runs a lightweight query against the lexical reader to pre-load relevant index segments into the OS page cache. One-shot CLI searches run with warming disabled.The system is designed for extensibility via the Connector trait (src/connectors/mod.rs). This allows cass to treat disparate log formats as a uniform stream of events.
classDiagram
class Connector {
<<interface>>
+detect() DetectionResult
+scan(ScanContext) Vec~NormalizedConversation~
}
class NormalizedConversation {
+agent_slug String
+messages Vec~NormalizedMessage~
}
Connector <|-- CodexConnector
Connector <|-- ClineConnector
Connector <|-- ClaudeCodeConnector
Connector <|-- GeminiConnector
Connector <|-- ClawdbotConnector
Connector <|-- VibeConnector
Connector <|-- OpenCodeConnector
Connector <|-- AmpConnector
Connector <|-- CursorConnector
Connector <|-- ChatGptConnector
Connector <|-- AiderConnector
Connector <|-- PiAgentConnector
Connector <|-- FactoryConnector
Connector <|-- CopilotConnector
Connector <|-- CopilotCliConnector
Connector <|-- OpenClawConnector
Connector <|-- CrushConnector
Connector <|-- HermesConnector
Connector <|-- KimiConnector
Connector <|-- QwenConnector
CodexConnector ..> NormalizedConversation : emits
ClineConnector ..> NormalizedConversation : emits
ClaudeCodeConnector ..> NormalizedConversation : emits
GeminiConnector ..> NormalizedConversation : emits
ClawdbotConnector ..> NormalizedConversation : emits
VibeConnector ..> NormalizedConversation : emits
OpenCodeConnector ..> NormalizedConversation : emits
AmpConnector ..> NormalizedConversation : emits
CursorConnector ..> NormalizedConversation : emits
ChatGptConnector ..> NormalizedConversation : emits
AiderConnector ..> NormalizedConversation : emits
PiAgentConnector ..> NormalizedConversation : emits
FactoryConnector ..> NormalizedConversation : emits
CopilotConnector ..> NormalizedConversation : emits
CopilotCliConnector ..> NormalizedConversation : emits
OpenClawConnector ..> NormalizedConversation : emits
CrushConnector ..> NormalizedConversation : emits
HermesConnector ..> NormalizedConversation : emits
KimiConnector ..> NormalizedConversation : emits
QwenConnector ..> NormalizedConversation : emits
Box<dyn Connector> instances that are unaware of each other's underlying file formats (JSONL, SQLite, specialized JSON).cass uses frankensqlite as the durable source of truth and frankensearch as a derived speed layer, powered by a suite of integrated "franken" libraries.
cass capabilities --json lists them as connectors.messages, conversations, agents) via frankensqlite — a pure-Rust SQLite reimplementation. cass turns on the engine's concurrent mode (PRAGMA fsqlite.concurrent_mode = ON), so a plain BEGIN runs as BEGIN CONCURRENT. Indexing still has one writer at a time, because index-run.lock admits a single indexer. BEGIN IMMEDIATE appears only in the daemon job queue and in logical-archive import/migrate. An experimental opt-in parallel persist path (CASS_INDEXER_BEGIN_CONCURRENT=1, off by default) exists but is not the default.title, content, agent, workspace, created_at.title_prefix and content_prefix use Index-Time Edge N-Grams (not stored on disk to save space) for instant prefix matching.flowchart LR
classDef pastel fill:#f4f2ff,stroke:#c2b5ff,color:#2e2963;
classDef pastel2 fill:#e6f7ff,stroke:#9bd5f5,color:#0f3a4d;
classDef pastel3 fill:#e8fff3,stroke:#9fe3c5,color:#0f3d28;
classDef pastel4 fill:#fff7e6,stroke:#f2c27f,color:#4d350f;
classDef pastel5 fill:#ffeef2,stroke:#f5b0c2,color:#4d1f2c;
subgraph Sources["Local Sources"]
A1[Codex]:::pastel
A2[Cline]:::pastel
A3[Gemini]:::pastel
A4[Claude]:::pastel
A5[OpenCode]:::pastel
A6[Amp]:::pastel
A7[Cursor]:::pastel
A8[ChatGPT]:::pastel
A9[Aider]:::pastel
A10[Pi-Agent]:::pastel
A11[Factory]:::pastel
A12[Copilot Chat]:::pastel
A13[Copilot CLI]:::pastel
A14[OpenClaw]:::pastel
A15[Clawdbot]:::pastel
A16[Vibe]:::pastel
A17[Crush]:::pastel
A18[Hermes]:::pastel
A19[Kimi]:::pastel
A20[Qwen]:::pastel
end
subgraph Remote["Remote Sources"]
R1["sources.toml"]:::pastel
R2["SSH/rsync\nSync Engine"]:::pastel2
R3["remotes/\nSynced Data"]:::pastel3
end
subgraph "Ingestion Layer"
C1["franken_agent_detection\nAuto-Discover & Scan\nNormalize & Dedupe"]:::pastel2
end
subgraph "Storage + Search"
S1["frankensqlite (WAL)\nSource of Truth\nBEGIN CONCURRENT\nMigrations"]:::pastel3
T1["frankensearch\nBM25 + Semantic\nRRF Fusion\nReranking"]:::pastel4
end
subgraph "Presentation"
U1["TUI (FrankenTUI)\nElm Architecture\nAnalytics Dashboard\nAsync Search"]:::pastel5
U2["CLI / Robot\nJSON Output\nAutomation"]:::pastel5
end
A1 --> C1
A2 --> C1
A3 --> C1
A4 --> C1
A5 --> C1
A6 --> C1
A7 --> C1
A8 --> C1
A9 --> C1
A10 --> C1
A11 --> C1
A12 --> C1
A13 --> C1
A14 --> C1
A15 --> C1
A16 --> C1
A17 --> C1
A18 --> C1
A19 --> C1
A20 --> C1
R1 --> R2
R2 --> R3
R3 --> C1
C1 -->|Persist| S1
C1 -->|Index| T1
S1 -.->|Rebuild| T1
T1 -->|Query| U1
T1 -->|Query| U2
cass index --watch, foreground): Uses file system watchers (notify) to detect changes in agent logs. When you save a file or an agent replies, cass re-indexes just that conversation. The TUI does not start a watcher on its own; see Keeping the Index Fresh below for what runs automatically.An index that is always a little behind is the most common complaint about any local search tool, so cass has three cooperating mechanisms. None of them block a search; all of them run cass index --background, which lowers its own CPU (nice 15) and I/O (ionice idle on Linux) priority before touching anything, and all of them respect the single index-run.lock — two indexers never run at once.
| Layer | What | When it runs | Enable |
|---|---|---|---|
| Stale-on-read catch-up | search, pack, and TUI launch check index freshness. If the index is stale (> 30 min), partial, or has pending sessions, a detached incremental cass index --background is spawned in its own process group and the current results are returned immediately. The next search is fresh. | On demand, at most once per 5 min per data dir (CASS_AUTO_REFRESH_COOLDOWN_SECS). Never for data dirs under the OS temp dir, and never for search --no-maintenance. A catch-up that ends without advancing the index is not respawned blindly: 1 h, then 6 h between attempts, and three failures trip the breaker until any run completes. | On by default. CASS_AUTO_REFRESH=0 disables globally. --robot-meta reports index_freshness.auto_refresh.{outcome,trigger,pid,consecutive_failures,detail}. |
OS scheduler (cass schedule install) | launchd LaunchAgents (macOS) or systemd user timers (Linux): an incremental job every 15 min and a nightly job (03:00) that performs a full source census with conditional lexical rebuilding, then one bounded models backfill --scheduled worker per tier (fast/hash always; quality/MiniLM when installed). Due remote-source syncs run first. Priority is delegated to the OS: launchd jobs set Nice=15 only, without ProcessType=Background or LowPriorityIO (background I/O throttling starved scheduled indexing on macOS). systemd units set Nice=19, IOSchedulingClass=idle and CPUSchedulingPolicy=idle. | On the timer, even when no cass process is running; survives reboots (Persistent=true / launchd). | cass schedule install [--interval-mins 15] [--nightly-hour 3] [--no-nightly] [--no-semantic] [--dry-run]; cass schedule status; cass schedule uninstall. |
| Resident daemon timer | The warm-model daemon (cass daemon, started by hand or auto-spawned by a human-mode semantic/hybrid search with --daemon; robot searches never spawn it) can also kick an incremental background index while it is resident. | Every CASS_DAEMON_INDEX_INTERVAL_SECS seconds while the daemon lives (it exits after its idle timeout). | Off by default; CASS_DAEMON_INDEX_INTERVAL_SECS=900 recommended. |
Idle awareness: scheduled work skips a run when the machine is under severe load (Linux /proc/loadavg + PSI; macOS sysctl vm.loadavg). On macOS you can additionally require the console to have been idle — CASS_RESPONSIVENESS_MIN_USER_IDLE_SECS=600 makes the nightly job and scheduled semantic backfill wait until nobody has touched the keyboard for ten minutes (the gate fails open where idle time is unavailable). Foreground cass index is never gated.
After an upgrade, the storage engine repairs and migrates an existing archive once, on its first writable open; that pass copies and rewrites the whole archive. Background runs (stale-on-read catch-up and scheduled jobs) never start it on an archive larger than CASS_INDEX_INTEGRITY_PREFLIGHT_MAX_BYTES (default 2 GiB): they exit 7 with kind migration-repair-pending, touch nothing, and cass schedule status names the cause. Run cass index --full in the foreground at a quiet time to perform it once; it keeps the original as a .pre-migration-bak copy, so plan for that much free space.
For a slow hosted disk, start with cass schedule install --interval-mins 60 and measure before shortening the interval. On Linux, cass index --json reports indexing_stats.bytes_written: the process block-write counter increase during indexing, including final checkpointing. It measures physical writes across all indexing layers, not just new transcript bytes or lexical segments; a cache-backed filesystem can report zero. The field is omitted when the counter is unavailable, including on other platforms. Check this alongside elapsed_ms on both changed-source and unchanged-source runs.
Nightly indexing retains index --full source coverage because timestamp-only connectors can miss restored files with old modification times. Connectors with valid durable source observations can reuse unchanged sources. When a completed checkpoint matches the archive and the lexical index passes validation, new messages are indexed inline. Missing or invalid checkpoint evidence, sparse or corrupt lexical assets, deferred lexical updates, and provenance repairs retain authoritative rebuilding from SQLite. Explicit cass index --full and --full --force-rebuild keep their existing repair behavior.
The nightly census still pays for source discovery and archive integrity, salvage, analytics, and FTS maintenance where required. A run that resumes an interrupted lexical rebuild can finish canonical recovery before returning; source discovery resumes on a later indexing run. Disappearing source files do not erase the preserved canonical history.
Each semantic worker retains its loaded model across its admitted batches and releases the previous batch's messages, vectors, storage handle, and lock at every checkpoint. CASS_SCHEDULE_MAX_BACKFILL_BATCHES bounds total attempts across tiers. Standalone cass models backfill --max-batches N uses the same worker; its default remains one batch.
Everything a scheduled job did is recorded under <data_dir>/schedule/ (state.json, runs.jsonl, per-job logs) and the last stale-on-read spawn under <data_dir>/auto-refresh-state.json / auto-refresh.log; cass schedule status --json reads all of it.
# See what would be registered, then register it
cass schedule install --dry-run
cass schedule install
# Run a job by hand (what the units invoke); --force ignores load/idle gates
cass schedule run --job incremental --json
cass schedule run --job nightly --force
# Inspect
cass schedule status --json
cass search "auth" --robot --robot-meta | jq '._meta.index_freshness.auto_refresh'
Stale-on-read catch-up handles an index that is behind. A search can also find the lexical index missing or unusable: a first run, a rebuild that never finished, a schema change. The search then has three choices: answer from what exists, rebuild before answering, or refuse. cass picks by the size of the job and by who is asking.
| Situation | What cass search does |
|---|---|
| A readable index exists but its checkpoint metadata is stale | Searches the existing index and leaves the heavy repair to an index run |
No usable index, and the archive is within the inline repair budget (CASS_INCREMENTAL_AUTHORITATIVE_LEXICAL_REPAIR_MAX_DB_BYTES, default 1 GiB, database plus WAL) | Rebuilds from SQLite inline, then answers. Robot callers get a bounded refusal instead when the ingest-quarantine circuit breaker is active |
| No usable index, and the archive is over that budget | Refuses with exit 5 maintenance-required and starts a detached cass index --full --background; the error hint names its pid |
| Robot caller, existing index whose incomplete checkpoint the cheap metadata refresh cannot reconcile | Refuses with exit 5 checkpoint_incomplete and starts the matching background run (--full above the size budget, plain cass index below it) |
| A rebuild is already running and no searchable generation exists | Robot callers get exit 7 index-busy immediately, with N of M conversations processed when the rebuild has recorded progress. Human callers wait up to CASS_SEARCH_ACTIVE_REBUILD_WAIT_MS (30 s) for it to publish |
Why the search never runs a large rebuild itself. A rebuild inside the
search process lives only as long as that process, and an agent's search is
almost always wrapped in a timeout: cass's own robot budget, or the agent
harness's command limit. On a large archive the rebuild commits nothing until
its first batch completes, so a killed rebuild keeps no progress. On one real
11 GB archive, a search-driven rebuild reached 160 of 4,324 conversations in
30 seconds (19 of them spent waiting on the in-flight byte budget) with
committed_offset still 0 when the search's budget ended it. The next search
started again from zero. Telling the agent to run cass index --full itself
failed the same way, because that command ran under the same timeout. The index
never converged. A detached child in its own process group survives the search,
so the rebuild finishes and the next search answers.
What the caller sees. The hint says what cass did and what to do next, so an agent never has to guess whether to run maintenance itself:
{"error": {"code": 5, "kind": "maintenance-required",
"message": "Automatic lexical repair was not started after detecting searchable lexical metadata missing: ...",
"hint": "cass started `cass index --full --json --background` as a detached process (pid 62704) to rebuild the search index. Retry this search after it finishes; `cass status --json` shows its progress under .rebuild. Do not run `cass index --full --json` yourself meanwhile: it would exit 7 (index-busy).",
"retryable": true}}
When no child is started, the hint says why: a run already holds the index
lock; a recent spawn is still inside its cooldown (it may still be starting, or
it failed, with the log path); earlier spawns failed and the breaker backed off
or tripped (with the failure detail); CASS_AUTO_REFRESH=0; or the spawn itself
failed. Those hints name the foreground command and warn that it needs a process
that is not killed by a short timeout.
Guard rails. The handoff reuses the stale-on-read machinery, so the same
limits apply: one spawner at a time (a file lock), the 5-minute cooldown, the
failure breaker (1 h, then 6 h, tripped after three failures), and the
index-run.lock that keeps two indexers from ever running together. Data dirs
under the OS temp dir and TUI_HEADLESS harnesses never spawn, and
search --no-maintenance never spawns anything. Under --timeout, a wait for an
active rebuild stops at nine tenths of the time remaining, so the caller
receives the index-busy verdict rather than an empty timed-out result.
Measured end to end (release build, an isolated data dir with 30 sessions,
lexical index moved aside, inline budget forced to one byte): the search
answered in 0.09 s with the spawned pid in its hint, the background index
published about 10 s later, and the next search reported all 60 matching
messages (50 returned at --limit 50). With CASS_AUTO_REFRESH=0 the same sequence never recovered within
180 s, which is the behaviour every large archive had before.
The interactive interface (src/ui/app.rs) uses FrankenTUI (ftui), a Rust TUI framework implementing the Elm architecture (Model-View-Update). The runtime handles terminal lifecycle, event polling, rendering, and cleanup.
CassMsg variant. The update() function produces Cmd effects (async tasks, ticks, quit).view() function renders the current state to an ftui Frame. The runtime diff engine minimizes terminal writes using Bayesian strategy selection.Cmd::Task, with results delivered as messages.graph TD
Input([User Input]) -->|Key/Mouse/Tick| Runtime
Runtime -->|CassMsg| Update[Model::update]
Update -->|Cmd| Runtime
Update -->|State Change| View[Model::view]
View -->|Frame| DiffEngine[Bayesian Diff]
DiffEngine -->|Minimal Writes| Terminal
Update -->|Cmd::Task| Background[Background Thread]
Background -->|Result Msg| Runtime
Data integrity is paramount. cass treats the SQLite database (src/storage/sqlite.rs, powered by frankensqlite) as the source of truth for conversations. History grows by insertion, and rows change only where the source changed:
forget and purge delete rows.UNIQUE(conversation_id, idx). New conversations are written with batched plain INSERTs, with the unique index as the backstop. Appends to a known conversation use INSERT OR IGNORE, so an agent re-writing a file cannot store a message twice. BLAKE3 content hashes are used only in memory, as merge fingerprints._schema_migrations table and a strict migration path keep upgrades safe and atomic; see Database Schema Migrations. Production creates a fresh database with one combined full_schema_v13 step and then applies v14–v21, so a new database records versions 13–21. The v1–v12 SQL is compiled only into tests.cass treats search indexes as derived assets. The SQLite archive is authoritative; lexical and semantic search data can be rebuilt from it.
Every lexical generation stores a schema_hash.json file containing the schema fingerprint:
{"schema_hash":"quill-fslx-schema-v9-hyphen-cjk-bigrams-bounded-content-prefix-preview-stored-content-external"}
| Scenario | Detection | Recovery |
|---|---|---|
| First run | No SQLite archive and no lexical index | cass index --full discovers sessions and creates both |
| Missing lexical index | No readable lexical asset | Rebuild from SQLite into scratch space, then publish |
| Schema mismatch | Hash differs from current | Rebuild derived lexical asset from SQLite |
| Corrupted metadata | Invalid or missing lexical metadata | Ignore the broken derivative and rebuild from SQLite |
| Semantic not ready | Model/vector assets absent or still backfilling | Continue lexical search and report semantic fallback/readiness |
# Check the current truth surface first
cass triage --json
cass health --json
cass status --json
# If not ready, run the first targeted command from recommended_commands[].
# For a fresh data dir this is usually:
cass index --full --json --no-progress-events --data-dir <same-data-dir>
Manual rebuild commands are for first setup, explicit operator refresh, or cases where recommended_commands[] asks for them. A normal missing/stale lexical asset should be repaired as derived state from SQLite, not treated as lost user data.
cass only reads agent files, never modifies themcass maintains multiple layers of redundancy to recover from corruption or schema changes:
Schema Hash Versioning:
Each lexical generation stores a schema_hash.json file containing a hash of the current schema definition. On startup:
This ensures that version upgrades with schema changes can rebuild the lexical derivative without user intervention.
Automatic Rebuild Triggers:
| Condition | Detection | Action |
|---|---|---|
| Schema version change | Hash mismatch in schema_hash.json | Full rebuild |
| Missing Quill publication manifest | Quill can't open index | Rebuild and publish a fresh derivative |
| Corrupted index files | Lexical reader open fails | Rebuild and publish a fresh derivative |
| Explicit request | --force-rebuild flag | Rebuild derived search assets from the canonical SQLite archive |
SQLite as Ground Truth: The SQLite database serves as the authoritative data store. Lexical rebuilds reconstruct the Quill index from SQLite:
// Iterate all conversations from SQLite
// Re-index each message into a fresh Quill index
// Progress tracked via IndexingProgress for UI feedback
This means corrupted lexical data is a repairable derivative-state problem. Operators should start with cass triage --json for the exact next command, or read cass health --json / cass status --json for the narrower readiness snapshot.
The SQLite database uses 21 versioned schema migrations, tracked in the _schema_migrations table (CURRENT_SCHEMA_VERSION = 21 and MIGRATION_NAMES in src/storage/sqlite.rs):
| Version | Migration | Version | Migration |
|---|---|---|---|
| 1 | core_tables | 12 | model_dimensions |
| 2 | fts_messages | 13 | plan_token_rollups |
| 3 | fts_messages_rebuild | 14 | fts_contentless |
| 4 | sources | 15 | conversation_tail_state_cache |
| 5 | provenance_columns | 16 | drop_redundant_message_conv_idx |
| 6 | source_path_index | 17 | drop_message_created_idx |
| 7 | msgpack_columns | 18 | conversation_tail_state_hot_table |
| 8 | daily_stats | 19 | conversation_external_lookup |
| 9 | embedding_jobs | 20 | conversation_external_tail_lookup |
| 10 | token_analytics | 21 | conversation_context_index (current) |
| 11 | message_metrics |
Migration Process:
cass checks _schema_migrations in the database (older databases that still record schema_version in the meta table are transitioned automatically)Safe Files (never deleted during rebuild):
bookmarks.db - Your saved bookmarkstui_state.json - UI preferencessources.toml - Remote source configuration.env - Environment configurationBackup and Retention Policy: Migration/rebuild backups preserve user data and are not treated as disposable source evidence. Derived lexical publish backups use the bounded retention policy documented above, while quarantined artifacts and repair candidates persist until an operator runs an explicit, fingerprinted cleanup flow.
The --watch flag enables real-time index updates as agent files change.
File change detected
↓
[2 second debounce window] ← Accumulate more changes
↓
[5 second max wait] ← Force flush if changes keep coming
↓
Re-index affected files
Each file system event is routed to the appropriate connector:
~/.claude/projects/foo.jsonl → ClaudeCodeConnector
~/.codex/sessions/rollout-*.jsonl → CodexConnector
~/.aider.chat.history.md → AiderConnector
Watch mode maintains watch_state.json, one scan watermark (ms) per connector under a short connector code (cd Claude, cx Codex, gm Gemini, ...):
{"v":1,"m":{"cd":1699900000000,"cx":1699900000000}}
The global last_scan_ts and last_indexed_at watermarks live in the SQLite meta table, not in this file.
Codex event_msg token_count usage is attached to the nearest preceding assistant turn during indexing.
If you indexed Codex sessions before this behavior existed, backfill usage coverage with:
cass index --full
cass analytics rebuild --track a
cass analytics rebuild re-derives the Track A rollups (message_metrics,
usage_hourly, usage_daily, usage_models_daily) from messages already in
the archive; it never re-parses raw session files. On a large archive a full
rebuild is a long single-core job, so daily refreshes should be windowed:
# Full rebuild (every rollup row dropped and recomputed)
cass analytics rebuild
# Only recompute the last two UTC days; older rollups are left untouched
cass analytics rebuild --days 2
cass analytics rebuild --since -2d # same window, relative syntax
cass analytics rebuild --since 2026-08-20 # from a date
The window is widened to the start of the UTC day containing the cutoff,
because rollups are bucketed by day and hour. Progress is logged per 10k
messages (analytics_rebuild_progress). Across analytics commands, --days
and --since are mutually exclusive, and malformed or reversed time bounds
return a usage error instead of silently running an unfiltered query.
--until, --agent, --workspace
and --source are query-time filters and are rejected here rather than
silently ignored. --track b also rejects --since/--days; with --track all, the window applies to Track A while Track B still rebuilds the complete
token_usage ledger. cass analytics validate likewise rejects every query
filter because its invariant checks always cover the complete analytics
database.
The TUI analytics dashboard never rebuilds rollups in-process: when rollups are
missing it spawns a detached cass analytics rebuild child, logs it to
<data_dir>/analytics-rebuild.log, and reports the pid in the status line;
reopen the dashboard once the rebuild finishes.
Generate tab-completion scripts for your shell.
Bash:
cass completions bash > ~/.local/share/bash-completion/completions/cass
# Or: cass completions bash >> ~/.bashrc
Zsh:
cass completions zsh > "${fpath[1]}/_cass"
# Or add to ~/.zshrc: eval "$(cass completions zsh)"
Fish:
cass completions fish > ~/.config/fish/completions/cass.fish
PowerShell:
cass completions powershell >> $PROFILE
search, index, stats, etc.)--robot, --agent, --limit)SIGILL hazard (the historical ONNX Runtime dependency was removed in cass#308).cargo install --git https://github.com/Dicklesworthstone/coding_agent_session_search. This requirement exists because CI builds target ubuntu-24.04 to access newer kernel features used by the frankensqlite storage engine. The install script probes the host's glibc (ldd --version) before downloading a Linux prebuilt binary and falls back to build-from-source with a warning when it is older than 2.38; --from-source forces that route, and --artifact-url bypasses the probe for an explicitly chosen artifact.Recommended: Homebrew (Apple Silicon macOS + Linux)
brew install dicklesworthstone/tap/cass
# Update later
brew upgrade cass
The Homebrew tap installs prebuilt release tarballs (not bottles) for Linux and Apple Silicon macOS. On Intel macOS, use the install script with --from-source.
Windows: Scoop
scoop bucket add dicklesworthstone https://github.com/Dicklesworthstone/scoop-bucket
scoop install dicklesworthstone/cass
Alternative: Install Script
curl -fsSL "https://raw.githubusercontent.com/Dicklesworthstone/coding_agent_session_search/main/install.sh?$(date +%s)" \
| bash -s -- --easy-mode --verify
Alternative: GitHub Release Binaries
SHA256SUMS.txt against the downloaded archive.cass into your PATH.Example (Linux x86_64, replace VERSION with an explicit release tag):
VERSION=v0.2.0 # e.g. v0.2.0
curl -L -o cass-linux-amd64.tar.gz \
"https://github.com/Dicklesworthstone/coding_agent_session_search/releases/download/${VERSION}/cass-linux-amd64.tar.gz"
curl -L -o SHA256SUMS.txt \
"https://github.com/Dicklesworthstone/coding_agent_session_search/releases/download/${VERSION}/SHA256SUMS.txt"
sha256sum -c SHA256SUMS.txt
tar -xzf cass-linux-amd64.tar.gz
install -m 755 cass ~/.local/bin/cass
cass
On first run, cass starts a full index in the background (a detached, low-priority cass index --full that keeps going if you quit) and shows its progress in the status line. Search goes live, without a restart, as soon as the first index is published. Until then there are no results; if automatic indexing is off (CASS_AUTO_REFRESH=0) or cannot start, the status line says so and names cass index --full.
foo* (prefix), *foo (suffix), or *foo* (contains) for flexible matching.Up/Down to select, Tab (or Alt+l) to focus the detail pane. Ctrl+N/Ctrl+Shift+N step through query history; Ctrl+R cycles it.F3: Filter by Agent (e.g., "codex").F4: Filter by Workspace/Project.F5/F6: Time filters (Today, Week, etc.).F2: Next theme (Shift+F2 previous; 19 presets).F12: Cycle ranking mode (recent → balanced → relevance → quality → newest → oldest).Ctrl+B: Toggle rounded/square borders.Enter: Open selected result in contextual detail modal (defaults to Messages tab).Enter with no selected hit: submit query behavior (no-op if empty).F8: Open selected hit in $EDITOR.Ctrl+Enter: Add current result to queue (multi-open).Ctrl+O: Open all queued results in editor.Ctrl+X: Toggle selection on current item (Ctrl+M opens the detail modal, like Enter).Alt+B: Bulk actions menu (when items selected).Ctrl+Y / Alt+Y / Ctrl+Shift+C: Copy file path / snippet / content to clipboard./: Find text within detail pane; Enter advances matches; n/N cycle contextual session hits; Esc closes the modal.Ctrl+Shift+R: Trigger manual re-index (refresh search results).Ctrl+Shift+Del: Reset TUI state (clear history, filters, layout).Aggregate sessions from your other machines into a unified index:
# Add a remote machine
cass sources add user@laptop.local --preset macos-defaults
# Sync sessions from all sources
cass sources sync
# Check source health and connectivity
cass sources doctor
See Remote Sources (Multi-Machine Search) for full documentation.
The cass binary supports both interactive use and automation.
# Interactive
cass [tui] [--data-dir DIR] [--once] [--asciicast FILE]
# Indexing
cass index [--full] [--watch] [--background] [--data-dir DIR] [--idempotency-key KEY]
cass schedule install [--interval-mins 15] [--nightly-hour 3] [--no-semantic] [--dry-run]
cass schedule status --json
# Search
cass search "query" --robot --limit 5 [--timeout 5000] [--explain] [--dry-run]
cass search "error" --robot --aggregate agent,workspace --fields minimal
cass pack "query" --robot --max-tokens 12000 [--limit 40] [--sessions-from FILE|-]
cass pack "query" --robot --freshness-policy strict --freshness-window-seconds 604800 --require-evidence
cass pack "query" --robot --max-tokens 4000 --max-evidence 8 --max-sessions 3 --max-excerpt-chars 600
# Inspection & Health
cass triage --json # One-shot agent preflight with exact next command
cass status --json # Quick health snapshot
cass health # Minimal pre-flight check (<50ms)
cass capabilities --json # First-stop agent self-description
cass introspect --json # Full API schema
cass swarm status --json # Read-only Beads/Agent Mail/git/rch swarm snapshot
cass swarm work-packet --json # Advisory claim packet; no mutations
cass swarm lint --json # Coordination and proof-gap lint
cass context /path/to/session --json # Find related sessions
cass view /path/to/file -n 42 --json # View source at line
# Session Analysis
cass export /path/to/session --format markdown -o out.md # Export conversation
cass expand /path/to/session -n 42 -C 5 --json # Context around line
cass timeline --today --json # Activity timeline
# Remote Sources
cass sources add user@host --preset macos-defaults # Add machine
cass sources sync # Sync sessions
cass sources doctor # Check connectivity
cass sources mappings list laptop # View path mappings
# Utilities
cass stats --json
cass completions bash > ~/.bash_completion.d/cass
| Command | Purpose |
|---|---|
cass (default) | Start TUI (a stale index triggers a detached low-priority catch-up; see Keeping the Index Fresh) |
cass tui --asciicast FILE | Run TUI and save terminal output as asciicast v2 |
index --full | Discover sessions and refresh the canonical DB plus derived search assets |
index --background | Same as index, but lowers its own CPU/I/O priority first (used by auto-refresh, schedule, and the daemon timer) |
index --watch | Foreground watch loop: reindex automatically on file changes |
schedule install|uninstall|status|run | Register incremental (15 min) + nightly (full source census, conditional lexical rebuild, bounded semantic backfill) jobs with launchd / systemd user timers |
search --robot | JSON output for automation pipelines |
pack --robot | Deterministic cited answer packs for agent/human handoffs; reports health, freshness, privacy, and warnings |
triage / ready / preflight | One-shot agent preflight: readiness, exact next command, docs, schemas, workflows, and recoveries |
guide [INTENT] | Intent-to-command planner (fix-ci, investigate-search-miss, prepare-release, repair-assets, export-session, onboard-source, support-capsule); dry-run by default. With --apply, 8 of its 19 allowlisted operations have evaluating proof adapters; the other 11 are observation-only and are reported, not evaluated |
status / state | Health snapshot: index freshness, DB stats, recommended action |
health | Minimal health check (<50ms on a healthy archive; the strict, mutation-free owner-thread probe shared with status has a 30 s hard deadline and never checkpoints a dirty WAL), exit 0=healthy, 1=unhealthy |
selftest | Archive-independent executable probe for installers and binary-promotion gates; exercises an in-memory FrankenSQLite write/read round-trip |
capabilities | First-stop agent self-description: workflow recipes, mistake recoveries, commands, global flags, exit codes, env vars, and limits |
introspect | Full API schema: commands, arguments, response shapes |
swarm status --json | Read-only shared-repo operations snapshot across Beads, Agent Mail metadata, git, build pressure, cass readiness, and proof refs |
swarm work-packet --json | Advisory one-agent packet with readiness, suggested reservations, verification commands, and closeout checklist; it does not claim or reserve |
swarm lint --json | Read-only coordination protocol lint for missing mail, stale reservations, status mismatches, and proof gaps. Only fixture input (--fixture, --fixture-dir --fixture-id) is linted today; the live path reports every provider live-provider-unimplemented and finds nothing |
swarm dependency-drift --json | Read-only sibling dependency sentinel for Cargo.toml pins, optional local checkout HEAD/dirty state, strict validation commands, and release-risk recommendations |
sessions [--workspace DIR] [--current] | Discover recent session files for follow-up actions |
context <path> | Find related sessions by workspace, day, or agent |
view <path> -n N | View source file at specific line (follow-up on search) |
export <path> | Export conversation to markdown/JSON |
export-html <path> | Export as self-contained HTML with optional encryption |
expand <path> -n N | Show messages around a specific line number |
timeline | Activity timeline with grouping by hour/day |
sources | Manage remote sources: add/list/remove/doctor/sync/mappings |
doctor | Diagnose and repair installation issues (safe, never deletes data) |
Other subcommands (all present in the Commands enum in src/lib.rs):
| Command | Purpose |
|---|---|
pages | Export an encrypted, searchable static-site archive with GitHub Pages / Cloudflare Pages deploy; runs the interactive wizard by default, with --export-only DIR, --verify BUNDLE, --preview BUNDLE, and --scan-secrets as non-wizard modes. --share-profile public|team|personal (config: bundle.share_profile; wizard: step 4) redacts every exported text, including titles, paths, metadata values, message bodies, their search indexes and snippets. public covers home paths, usernames, project names, hostnames, emails, phone numbers, IP addresses, social-security and card numbers, and internal URLs; team covers home paths, emails, social-security and card numbers; personal rewrites nothing. No profile rewrites a credential (API keys, private keys, connection strings): the staged secret scan rejects an export that contains one, so it fails closed instead of publishing a rewritten secret. The default is public for a plaintext export and team when encrypted. The bundle's export_meta records the profile and per-kind redaction counts, never the values. For a plaintext bundle, --verify (which every export also runs before it reports success) reads the declared profile and rescans every exported text surface, including both search indexes, with that profile's rules. It fails if any value the profile removes remains, reporting counts per surface. The home-path and username rules use the verifying account's home directory. An encrypted bundle keeps its profile inside the payload, which --verify does not open. |
pages key list|add-password|add-recovery|revoke|rotate --archive BUNDLE | Manage the key slots of an exported encrypted bundle (LUKS-style: several independently wrapped copies of one data key). Passwords come from an interactive prompt or --password-stdin (current password on line 1, new password on line 2), never from argv; --json for automation; recovery secrets are printed once and never stored. See docs/RECOVERY.md |
upgrade | Check for a newer release and optionally run the same checksum-verified installer the TUI uses (--check, --yes, --force) |
man | Generate the man page to stdout |
storage | On-disk storage footprint by component (DB, WAL, lexical index, raw mirror, semantic, quarantine) |
dedup | Collapse pre-existing duplicate conversation rows (projects/<rel> vs <rel> external-id twins); dry-run unless --apply |
support-bundle | Assemble a redacted, share-safe recovery/support evidence bundle |
state | Quick state/health check (alias of status) |
onboarding | Read-only first-run source onboarding + readiness wizard; --json for scripts, never launches the TUI |
quarantine | Inspect and manage the conversation-ingest quarantine (list / clear) |
forget | Prune already-indexed conversations by source-path glob; dry-run by default, --apply to commit. Deletes the canonical rows, then rebuilds FTS, analytics and the lexical index. Semantic vectors are not rewritten, so explicit semantic search then reports semantic-unavailable and hybrid falls back to lexical until cass index --semantic re-embeds from the canonical rows; no surface returns forgotten messages. dedup --apply and sources agents exclude behave the same way. It removes indexed copies, not source files: forget records each source's size and modification time, so an unchanged source stays forgotten across cass index, --full and rescans triggered by other sessions, but if its agent appends to it the whole conversation is indexed again. A store that keeps many sessions in one file (a SQLite database) counts as changed when any of its sessions changes. Delete or move the file to keep it out for good. The raw mirror keeps its verbatim capture of the source; once the source is gone, cass mirror prune --older-than 0s --safety-hold-down 0s --source-path '<glob>' --apply removes that capture too (a source still on disk is captured again by the next scan). A cass serve session opened before the forget keeps its pinned snapshot until reload; the TUI drops forgotten hits on its next search. If the lexical rebuild fails, forget exits 5 (lexical-rebuild) and cass index --full finishes the purge |
fleet upgrade-rehearsal | Fleet-safe upgrade rehearsal (dry run) with bounded post-upgrade verification; --live opts in to SSH probes of configured remotes |
lessons list|search | Mine and query durable, redacted lessons from local evidence (commits, closed beads, proof manifests) |
import chatgpt | Split a ChatGPT web export (conversations.json) into files the ChatGPT connector can index |
release-verify | Verify release distribution channels (GitHub, Homebrew, Scoop, crates.io, installer) from a recorded observation (--from) or live (--live) |
sources discover | Auto-discover SSH hosts from ~/.ssh/config |
sources reingest | Re-ingest an already-synced mirror into the canonical archive without re-running rsync |
sources artifact-manifest | Build or verify a lexical-artifact evidence manifest for remote exchange |
Two more commands are dispatched before the main parser (so they are absent from cass --help, introspect and completions); use their own --help:
| Command | Purpose |
|---|---|
serve --stdio [--mcp] (--data-dir DIR | --index PATH) | Persistent search service: reuses one read-only lexical index reader across a stream of newline-delimited JSON requests (or MCP with --mcp), with per-request deadlines (--request-timeout-ms), cross-process reader admission (--admission-dir, --admission-slots) and a resident-memory cap (--max-resident-mib). Search, status and startup never open the canonical database; --db opts in to canonical view/view_batch requests. Operations: search, semantic_search, refine, view, view_batch, status, reload, unload, shutdown. search is lexical; semantic_search needs --semantic-embedder minilm|multilingual-minilm|hash plus --data-dir and --db, and refine needs --reranker-model PATH (installed local assets only, never downloaded). See docs/SEARCH_SERVICE.md |
archive export|verify|search|view|import | Bounded, versioned logical archive of the canonical rows: export --output FILE --archive-id ID --include-private streams a read-only snapshot to a new private JSONL file; verify checks framing, identities, counts and digest without a database; search/view read verified message bodies without restoring; import restores into a new database and never replaces an existing archive |
| Tool | Purpose |
|---|---|
cass tui --asciicast FILE | Record TUI output as an asciicast v2 artifact; there is no separate cass cast subcommand |
scripts/bakeoff/cass_validation_e2e.sh | Run the bake-off validation harness for lexical, semantic, hybrid, and reranked search scenarios |
scripts/bakeoff/cass_embedder_e2e.sh | Exercise embedder bake-off flows against a generated validation corpus |
scripts/bakeoff/cass_rerank_e2e.sh | Exercise reranker bake-off flows and append results to the bake-off log |
Commands for troubleshooting, debugging, and understanding system state:
# One-shot agent preflight
cass triage --json
# → { "surface": "triage", "status": "healthy", "next_command": null, ... }
# Health check (fast, <50ms; the archive probe is bounded by a 30 s hard deadline)
cass health --json
# → { "healthy": true, "index_age_seconds": 120, "message_count": 5000 }
# Detailed status with recommendations
cass status --json
# → Includes index freshness, staleness threshold, recommended action
# System diagnostics
cass diag --verbose --json
# → Database stats, index info, connector status, environment
# Query explanation (debug why results are what they are)
cass search "auth" --explain --dry-run --robot
# → Shows parsed query, index strategy, cost estimate without executing
# Find related sessions
cass context /path/to/session.jsonl --json
# → Sessions from same workspace, same day, or same agent
# Archive-first diagnostic check
cass doctor check --json
# → Read-only checks for archive coverage, source authority, locks, backups,
# storage pressure, semantic fallback, and recommended next action
# Fingerprinted repair plan and apply
cass doctor repair --dry-run --json
cass doctor repair --yes --plan-fingerprint <plan_fingerprint> --json
# → Builds candidates and applies only the inspected matching fingerprint
# Legacy safe auto-run for low-risk derived repairs
cass doctor --fix --json
# → Emits operation_outcome and receipts; fails closed on archive/source risk
cass doctor is a comprehensive diagnostic and repair tool designed for troubleshooting installation and data issues. Its current recovery model is archive-first: preserve cass-owned evidence, prove source authority and coverage, then repair through candidates and receipts. The full operator runbook is docs/planning/RECOVERY_RUNBOOK.md.
What it checks:
| Surface | Purpose | Mutation Policy |
|---|---|---|
cass doctor check --json | Read-only truth surface for archive coverage, source authority, locks, storage pressure, semantic fallback, and recommended action | Never mutates |
cass doctor archive-scan --json | Read-only source inventory, raw mirror, coverage, sole-copy, and remote sync gap inspection | Never mutates |
cass doctor repair --dry-run --json | Builds a fingerprinted repair plan and candidate/promotion gates | Read-only plan |
cass doctor repair --yes --plan-fingerprint <fp> --json | Applies exactly the inspected repair fingerprint | Candidate-based, receipt-backed |
cass doctor backups list/verify/restore ... --json | Lists backups, verifies manifests, rehearses restore, then applies by fingerprint | Restore apply requires a matching rehearsal fingerprint |
cass doctor --repair-leaked-pages --dry-run --json | Classifies the engine's integrity verdict: clean, leaked_pages (pages no table, index or freelist owns — page N is never used), or other_damage | Read-only |
cass doctor --repair-leaked-pages --yes --json | Frees leaked pages in place when they are the only damage; re-runs integrity_check and compares conversation/message counts. Any other damage exits 5 data-corruption untouched | Backs up the live bundle first as a leaked-pages-repair backup (cass doctor backups restore <id>) |
cass doctor cleanup --json | Plans cleanup for derived or explicitly reclaimable assets | Apply requires a matching fingerprint |
cass doctor support-bundle --json | Creates a scrubbed diagnostic handoff bundle | Redacted by default; not a backup |
Safety guarantees:
sources.toml are not cleanup targets.plan_fingerprint.Recommended support checklist:
cass doctor check --json
cass doctor baseline diff <baseline_id> --json
cass doctor support-bundle --json
cass doctor support-bundle verify <bundle_or_manifest_path> --json
Send the doctor JSON, latest failure_context.json if present, support-bundle
manifest.json, any baseline diff, relevant artifact_manifest_path and
event_log_path values, and the exact command/exit code. Do not attach raw
sessions, full SQLite archives, private source files, or encrypted payloads
unless the user explicitly opts into sensitive evidence attachment.
Diagnostic Flags:
| Flag | Available On | Effect |
|---|---|---|
--explain | search | Show query parsing and strategy |
--dry-run | search | Validate without executing |
--verbose | most commands | Extra detail in output |
--trace-file | all | Append execution trace to file |
--robot-trace-ingest | index | Emit per-ingest-batch NDJSON timing and lookup counters on stderr |
Commands for managing the semantic search ML model:
# Check current model status (abbreviated schema — real output also
# includes cache_lifecycle, files[], revision, license, and more):
cass models status --json
# → {
# "model_id": "all-minilm-l6-v2",
# "model_dir": "~/.local/share/coding-agent-search/models/all-MiniLM-L6-v2",
# "installed": false,
# "state": "not_acquired",
# "state_detail": "model not acquired (user consent required); missing ...",
# "next_step": "Run `cass models install`, or keep using lexical search.",
# "lexical_fail_open": true,
# "revision": "c9745ed1...",
# "license": "Apache-2.0",
# "total_size_bytes": 90872535,
# "installed_size_bytes": 0,
# "observed_file_bytes": 0,
# "policy_source": "semantic_policy"
# }
# Install model (downloads ~90MB from Hugging Face on explicit request)
cass models install
# → Downloads from Hugging Face, verifies checksum
# Install from local directory (air-gapped environments)
cass models install --from-file /path/to/model-dir
# Verify model integrity
cass models verify --json
# → all_valid bool + per-file SHA-256 checks (see `cass models verify --help`)
# Check for model updates
cass models check-update --json
# → { "update_available": bool, "reason": str,
# "current_revision": str|null, "latest_revision": str }
In cass status --json, semantic.preferred_backend is "fastembed" when the native MiniLM lane is selected and "hash" for the hash tier; fastembed is only the id of the native pure-Rust MiniLM lane — no ONNX runtime is involved.
Model Files (stored in $CASS_DATA_DIR/models/all-MiniLM-L6-v2/):
model.safetensors - The neural network weights (~90MB)tokenizer.json - Vocabulary and tokenization rulesconfig.json - Model configurationspecial_tokens_map.json - Special token definitionstokenizer_config.json - Tokenizer settingsVerified Install: The installer enforces SHA256 checksums.
Sandboxed Data: All indexes/DBs live in standard platform data directories (~/.local/share/coding-agent-search on Linux).
Read-Only Source: cass never modifies your agent log files. It only reads them.
cass uses crash-safe atomic write patterns throughout to prevent data corruption:
TUI State Persistence (tui_state.json):
1. Serialize state to JSON
2. Write to temporary file (tui_state.json.tmp)
3. Atomic rename: temp → final
If a crash occurs during step 2, the original file is untouched. The rename operation (step 3) is atomic on all modern filesystems—it either completes fully or not at all.
ML Model Installation (models/all-MiniLM-L6-v2/):
1. Download to temp directory (models/all-MiniLM-L6-v2.tmp/)
2. Verify all checksums
3. If existing model present: rename to backup (models/all-MiniLM-L6-v2.bak/)
4. Atomic rename: temp → final
5. On success: remove backup
6. On failure: restore from backup
This backup-rename-cleanup pattern ensures that either the old model or new model is always available—never a half-installed state.
Configuration Files (sources.toml, watch_state.json):
All configuration writes follow the same temp-file-then-rename pattern, ensuring consistency even during power loss or unexpected termination.
Why This Matters:
Agent transcripts are full of credentials: keys pasted into prompts, tokens
echoed by tool output, .env files read into context. With the default
CASS_INDEX_REDACTION=full, cass scrubs them from every persisted message,
title, snippet and metadata blob before anything reaches SQLite or the lexical
index, so search results, exports and robot output never repeat them. (The
original session files and the raw-mirror blobs keep the raw text on the same
disk; redaction protects the queryable surfaces, not disk-at-rest secrecy.)
What is recognized. Thirteen pattern families: AWS access key IDs, AWS
secret keys and session tokens in assignment context, GitHub tokens (classic and
fine-grained), OpenAI and Anthropic API keys, Bearer tokens, JWTs, PEM/OpenSSH/PGP
private-key blocks, database connection URLs (Postgres, MySQL, MongoDB, Redis,
AMQP; the whole URL, since it may carry a password), generic password= /
api_key: / secret= style assignments, Slack tokens, and Stripe live keys.
JSON metadata is also redacted by field name: values under keys such as
passphrase, authorization, cookie, or anything ending in password,
token, secret or apikey are replaced whatever they look like.
How it runs. Redaction sits on the ingest hot path, since it touches every message, so it is built in three layers:
RegexSet.
One scan per string reports which patterns could match; a string with no
candidates (the vast majority) is returned untouched without allocating.replace_all pass, in a fixed order, replacing each match with [REDACTED].CASS_REDACT_MEMO_CAPACITY sets the entry cap).A frozen copy of the original algorithm lives in the test suite, and a 512-case property test requires the plain path and the memoized path (on both a cache miss and a cache hit) to produce byte-identical output to it. The memoized JSON path has its own equivalence test against the uncached one over nested shapes.
Why the patterns use ASCII word boundaries. Rust's regex crate runs a
RegexSet on a fast lazy DFA, but a Unicode word boundary (\b) is something
that DFA cannot evaluate once the haystack contains a non-ASCII byte. It then
falls back to the PikeVM, a far slower NFA simulation. Real transcripts are full
of non-ASCII text (emoji, CJK, box-drawing characters in tool output), so every
such message paid the slow path. The patterns now use (?-u:\b), the ASCII
word boundary, which the DFA handles.
This cannot weaken redaction. Every boundary in these patterns sits next to an
ASCII word character, and ASCII word characters are a subset of Unicode word
characters. So wherever a Unicode boundary exists, an ASCII boundary exists too,
and the ASCII patterns match everything the Unicode ones did. The only
difference is a token glued directly to a non-ASCII letter (凭据ghp_…): a
Unicode boundary sees two word characters and no boundary, so the old patterns
missed that token, while the ASCII form redacts it.
Measured on 593 MiB of real session text (3.1 million strings, 79,870 of them containing non-ASCII), with the same regex version cass ships:
| Word boundary | Prefilter throughput | Pattern matches |
|---|---|---|
Unicode \b | 9.8-10.6 MiB/s | baseline |
ASCII (?-u:\b) | 261-284 MiB/s | identical on all 3.1M strings |
In an A/B incremental index on two identical clones of a 2.1-million-message
archive (10 minutes each, run one after the other), the ASCII build committed
53,155 new messages on 10.1 CPU-minutes, against 7,028 on 12.8 CPU-minutes for
the Unicode build: about 9.5 times the ingest throughput per CPU second. A unit
test rejects any secret pattern that reintroduces a Unicode \b, so the cliff
cannot quietly come back.
The project ships with a robust installer (install.sh / install.ps1) designed for CI/CD and local use:
Checksum Verification: Validates artifacts against a .sha256 file or explicit --checksum flag.
Rustup Bootstrap: Source installs use the dated nightly and components pinned by the release's rust-toolchain.toml. The installer bootstraps rustup without an unrelated default toolchain when needed.
Easy Mode: --easy-mode automates installation to ~/.local/bin without prompts.
Platform Agnostic: Detects OS/Arch (Linux/macOS/Windows, x86_64/arm64) and fetches the correct binary.
cass includes a built-in update checker that notifies you when new versions are available, without interrupting your workflow.
When a new version is available, a one-line banner appears at the top of the TUI:
Update v<current> -> v<latest> | Alt+U upgrade | Alt+N notes | Alt+I ignore | Esc dismiss
Confirming Alt+U runs the same verified installer used for initial installation:
macOS/Linux:
curl -fsSL https://...install.sh | bash -s -- --easy-mode --verify
Windows (PowerShell):
`
Truncated — view the full README on GitHub.
(top 24 of 30)
530 followers · starred Jul 2026
694 followers · starred Jan 2026
69 followers · starred Jan 2026
1,125 followers · starred Jan 2026
Rust
95.3%
JavaScript
1.5%
Shell
1.5%