Acronym extraction and expansion — local, offline.
ae reads the text that flows past it and sorts every acronym into three buckets at once:
KPI (Key Performance Indicator)); the new term is pulled out, and the dictionary grows as it reads.$ ae "The OKR review needs a TPS (Transaction Processing System) sign-off before the XYZ ships."
OKR Objectives and Key Results expansion v1.00 c0.50
TPS Transaction Processing System extraction 0.95
XYZ (no expansion) candidate
Unlike a static glossary, ae grows its own dictionary by watching the stream — extracting inline definitions, and mining speculative expansions from prose where word initials spell a watched acronym (no parentheses needed).
Internally, every (acronym, expansion) pair is judged on two independent axes — the idea ae is built on:
1.0 when a human verified it, 0.9 for an inline definition, 0.0 for a speculative mined guess.0.5 with no evidence yet.A pair can be rock-solid valid (PT → Part Time) yet a poor fit in a physical-therapy paragraph. The structured output folds the two into one number — confidence = validity × context fit — so a consumer reads a single "trust this expansion here" score instead of guessing which axis to threshold. Both axes still drive ranking and pruning underneath, along a provenance continuum by source — user (verified) > inline > mined (speculative).
ae's real interface is a pipe: send text on stdin, read structured JSON on stdout. The output is a flat list of findings — one per expansion match, inline extraction, or unresolved candidate — identical whether the input was one blob or a stream: -j is a pretty array, -J is one finding per line.
$ printf 'ship the OKR review this sprint' | ae -j
[
{
"kind": "expansion",
"acronym": "OKR",
"expansion": "Objectives and Key Results",
"confidence": 0.5
}
]
An acronym with several candidate expansions yields several expansion findings, each with its own confidence; group by acronym to rank them. A candidate is just {"kind": "candidate", "acronym": "…"}; an extraction adds pattern_type. Each finding carries a single confidence — provenance (validity) is an internal prior that feeds it, surfaced on the dictionary commands (list/show) rather than per finding.
Agents and tools hit org-specific acronyms constantly and can't phone a server for them — ae resolves them locally, in-process, in milliseconds.
No network calls, ever. The dictionary is a bundled SQLite database
($XDG_DATA_HOME/ae/acronyms.db); embeddings run locally via a quantized ONNX
model, with a deterministic hash-embedder fallback so offline builds still work.
Across concurrent callers, ae elects one in-process Leader that holds the
warm state behind a Unix socket while the rest proxy to it — no daemon to manage,
and an idle Leader cleans itself up.
brew install dpep/tools/ae # binary `ae`
cargo install acronym-engine # same binary, from crates.io
make install # from a source checkout → ~/.cargo/bin/ae
The embedding model is fetched on first use from the HuggingFace Hub into the
shared cache (~/.cache/huggingface/hub, honoring $HF_HOME) — never committed,
never downloaded at build time, and reused across any tool that pulls the same
model. Point ae elsewhere with --model <dir | .onnx | org/name> or pin a local
copy via $AE_MODEL_DIR. If nothing loads (offline and uncached), ae falls back
to a deterministic hash embedder so it still runs.
ae "text to scan" # analyze one string → findings array
cat access.log | ae -J # stream stdin line by line, findings as NDJSON
ae -f notes.md # stream a file line by line
ae -r "…" # read-only: expand known acronyms, never learn
Full flags and subcommands: ae --help. The learned dictionary persists in
SQLite ($XDG_DATA_HOME/ae/acronyms.db); the daemon and the in-process fallback
share it, so what's learned in one call is there for the next.
Subcommands curate the dictionary directly (no flags needed — they're distinct from analysis input, which arrives as a quoted argument or via stdin):
ae add MVP "Minimum Viable Product" "Most Valuable Player" # add (one or more)
ae list # list everything
ae list perf # filter by substring of acronym or expansion
ae show KPI # expansions of one acronym
ae candidates # acronyms seen but undefined, by frequency
ae add PB&J # declare a token as an acronym (ae mines its expansion later)
ae suggest MVP # speculative expansions, --limit N / --min-confidence
ae define MVP # promote interactively (fzf), or pass expansions
ae prune # GC: spell-fix + dedup (prefix+fuzzy) + drop noise
ae ignore IOS # mute: keep it but make it inert (unignore reverses)
ae ignore # list muted acronyms
ae rm MVP # delete it outright (ignore mutes; rm removes)
ae rm --all # wipe the whole dictionary (backed up first)
ae rm --restore # restore the most recent backup (or pass a path)
-q/--quiet suppresses normal output everywhere (e.g. ae "…" -q silently
learns; ae add … -q adds without printing).
rm --all (no acronym) clears the entire dictionary and resets to the built-in
defaults. It isn't gated by a prompt — instead it first writes a timestamped
snapshot to /tmp/ae/backup_<ts>.db, and ae rm --restore [PATH] brings it back
(with no path, the most recent backup; the pre-restore state is itself snapshotted,
so a restore is reversible).
ignore (alias mute) is for acronym-shaped tokens you never want surfaced or
mined — e.g. an all-caps word a colleague keeps shouting. The entry stays in the
DB but is left out of expansion, mining, suggestions, and candidate flagging
until unignore. That's the difference from rm, which deletes outright.
Fully-capitalized lines are also skipped automatically: with no lowercase to
contrast against, an all-caps sentence would otherwise flag every short word.
Each acronym has a provenance: declared (you said it's an acronym, via
ae add ACR with no expansion) or seen (ae noticed it). An acronym joins the
watch list — where we hunt its expansions in later text — once it's declared
or has been seen enough times (default 3); below that it's noise and ae prune
drops it. Punctuated acronyms (PB&J, R&D, U.S.A) are
detected and mined too (the &/. maps to a skipped filler word), and a longer
match wins over its parts (PB&J beats PB, maximal munch).
ae prune also spell-corrects mined expansions against the system word list
(/usr/share/dict/words, if present — nothing bundled), so "minimum viabel
product" converges to "minimum viable product" before dedup.
You rarely need to run it by hand: the same pass runs automatically on a
cadence after a write — at most once per AE_CONSOLIDATE_SECS (default daily; a
negative value disables it). The point is the quality half — spell-fix and
dedup pool evidence and lift confidence — so it's worth running regularly; the
deletion half is gentle (a candidate seen within AE_PRUNE_GRACE_SECS, default
~30 days, is spared, so an infrequent token can recur weeks later before it's
ever considered noise). The warm daemon amortizes it across requests.
ae define MVP with no expansions picks interactively from the mined
suggestions — via fzf (multi-select) if installed, else a numbered prompt. An
acronym can hold several expansions, so multi-select is first-class.
ae prune is the occasional GC: it merges prefix-duplicate expansions
("min viable product" → "minimum viable product"), drops ones below a confidence
floor (default 0.15), and removes seen-once noise candidates. ae suggest keeps
a higher bar by default (0.30) since speculation is noisy; both take
--min-confidence to override, and suggest takes -l/--limit N.
ae suggest is the payoff of tracking candidates. When ae analyzes text it
mines word-sequences whose initials spell a watched candidate acronym — no
parentheses needed, tolerant of skipped filler words (OKR = Objectives and
Key Results), and across subsequent inputs too (a definition mentioned in a
later sentence that never repeats the acronym still accrues). It counts how often
each phrase recurs and how well its context fits where the acronym is used, then
blends both into a confidence:
$ ae add FOO # declare it — start watching
ae: now watching FOO for expansions
$ ae "the Foundations Of Onboarding workshop was great" # initials spell FOO
No acronyms found.
$ ae suggest FOO --min-confidence 0 # ae mined the phrase
FOO foundation of onboarding 0.50 (1)
Confirm one with ae define MVP (interactive) or ae add MVP "Minimum Viable Product" (which clears it from the candidate/suggestion lists). It's heuristic —
short acronyms attract noise, which sinks to the bottom on low confidence and is
hidden by the default threshold.
Removal disambiguates when an acronym has several expansions:
ae rm MVP # removes it if there's one expansion; else lists them and stops
ae rm MVP valuable # substring picks one ("Most Valuable Player")
ae rm MVP --all # removes every expansion
ae candidates is fed automatically: every analysis records the acronym-shaped
tokens it couldn't resolve and counts how often you've used them, so you can see
what's worth defining. (Defining one clears it from the list.) All of these
honor -j/-J too.
Default output is human-readable; -j/-J switch to JSON/NDJSON. A positional
TEXT is analyzed as one blob; piped stdin and --file are streamed line by
line — but either way the structured output is the same flat list of findings,
so a consumer parses one shape. With nothing to do, ae prints help. stdout
carries only data — all logs go to stderr, so ae … | jq is always safe. Every
command is machine-friendly: -j/-J work everywhere, and --daemon/--stop
emit a {"status": …} object in those modes. ae --status reports a running
daemon's version, embedder, and uptime (read-only — it never starts one), and
exits non-zero when none is up, so --status -q is a silent health check.
Under -J each finding is one compact object, flushed as it's produced (so
tail -f … | ae -J streams live); -j collects them into one pretty array.
--read-only is the safe path for untrusted or high-volume input — it expands
known acronyms without ever writing to the dictionary.
--model lets you point at any compatible model — an absolute/relative path to a
model directory or .onnx file, or a HuggingFace org/name repo id (fetched into
the shared Hub cache). With no flag, ae uses the default model from the Hub
($AE_MODEL_DIR overrides with a local copy), and falls back to the hash embedder
if none loads.
input ─▶ file lock ─▶ Leader (UDS server, warm trie + dictionary + embedder)
└─▶ Follower (forwards text, pipes back JSON)
evaluation = STAGE 1 expansion (trie scan → dictionary → 64-d MRL vector match)
+ STAGE 2 learning (rule-based extraction of inline definitions)
Embeddings come from all-MiniLM-L6-v2 (int8-quantized ONNX) run locally via ONNX Runtime — tokenize, mean-pool, then compress with Matryoshka Representation Learning: the 384-d vector is truncated to its first 64 coordinates and L2-normalized, shrinking the vector store ~6×. See docs/SPEC.md for the full design and docs/ROADMAP.md for status and deliberate deviations (notably the model choice and the runtime model-fetch strategy).
cargo test # unit + integration tests
cargo clippy --all-targets
cargo fmt
See CLAUDE.md for conventions.
MIT — see LICENSE.txt.
70 commits
1 commits
Rust
99.3%
Acronym extraction and expansion — local, offline.
ae reads the text that flows past it and sorts every acronym into three buckets at once:
KPI (Key Performance Indicator)); the new term is pulled out, and the dictionary grows as it reads.$ ae "The OKR review needs a TPS (Transaction Processing System) sign-off before the XYZ ships."
OKR Objectives and Key Results expansion v1.00 c0.50
TPS Transaction Processing System extraction 0.95
XYZ (no expansion) candidate
Unlike a static glossary, ae grows its own dictionary by watching the stream — extracting inline definitions, and mining speculative expansions from prose where word initials spell a watched acronym (no parentheses needed).
Internally, every (acronym, expansion) pair is judged on two independent axes — the idea ae is built on:
1.0 when a human verified it, 0.9 for an inline definition, 0.0 for a speculative mined guess.0.5 with no evidence yet.A pair can be rock-solid valid (PT → Part Time) yet a poor fit in a physical-therapy paragraph. The structured output folds the two into one number — confidence = validity × context fit — so a consumer reads a single "trust this expansion here" score instead of guessing which axis to threshold. Both axes still drive ranking and pruning underneath, along a provenance continuum by source — user (verified) > inline > mined (speculative).
ae's real interface is a pipe: send text on stdin, read structured JSON on stdout. The output is a flat list of findings — one per expansion match, inline extraction, or unresolved candidate — identical whether the input was one blob or a stream: -j is a pretty array, -J is one finding per line.
$ printf 'ship the OKR review this sprint' | ae -j
[
{
"kind": "expansion",
"acronym": "OKR",
"expansion": "Objectives and Key Results",
"confidence": 0.5
}
]
An acronym with several candidate expansions yields several expansion findings, each with its own confidence; group by acronym to rank them. A candidate is just {"kind": "candidate", "acronym": "…"}; an extraction adds pattern_type. Each finding carries a single confidence — provenance (validity) is an internal prior that feeds it, surfaced on the dictionary commands (list/show) rather than per finding.
Agents and tools hit org-specific acronyms constantly and can't phone a server for them — ae resolves them locally, in-process, in milliseconds.
No network calls, ever. The dictionary is a bundled SQLite database
($XDG_DATA_HOME/ae/acronyms.db); embeddings run locally via a quantized ONNX
model, with a deterministic hash-embedder fallback so offline builds still work.
Across concurrent callers, ae elects one in-process Leader that holds the
warm state behind a Unix socket while the rest proxy to it — no daemon to manage,
and an idle Leader cleans itself up.
brew install dpep/tools/ae # binary `ae`
cargo install acronym-engine # same binary, from crates.io
make install # from a source checkout → ~/.cargo/bin/ae
The embedding model is fetched on first use from the HuggingFace Hub into the
shared cache (~/.cache/huggingface/hub, honoring $HF_HOME) — never committed,
never downloaded at build time, and reused across any tool that pulls the same
model. Point ae elsewhere with --model <dir | .onnx | org/name> or pin a local
copy via $AE_MODEL_DIR. If nothing loads (offline and uncached), ae falls back
to a deterministic hash embedder so it still runs.
ae "text to scan" # analyze one string → findings array
cat access.log | ae -J # stream stdin line by line, findings as NDJSON
ae -f notes.md # stream a file line by line
ae -r "…" # read-only: expand known acronyms, never learn
Full flags and subcommands: ae --help. The learned dictionary persists in
SQLite ($XDG_DATA_HOME/ae/acronyms.db); the daemon and the in-process fallback
share it, so what's learned in one call is there for the next.
Subcommands curate the dictionary directly (no flags needed — they're distinct from analysis input, which arrives as a quoted argument or via stdin):
ae add MVP "Minimum Viable Product" "Most Valuable Player" # add (one or more)
ae list # list everything
ae list perf # filter by substring of acronym or expansion
ae show KPI # expansions of one acronym
ae candidates # acronyms seen but undefined, by frequency
ae add PB&J # declare a token as an acronym (ae mines its expansion later)
ae suggest MVP # speculative expansions, --limit N / --min-confidence
ae define MVP # promote interactively (fzf), or pass expansions
ae prune # GC: spell-fix + dedup (prefix+fuzzy) + drop noise
ae ignore IOS # mute: keep it but make it inert (unignore reverses)
ae ignore # list muted acronyms
ae rm MVP # delete it outright (ignore mutes; rm removes)
ae rm --all # wipe the whole dictionary (backed up first)
ae rm --restore # restore the most recent backup (or pass a path)
-q/--quiet suppresses normal output everywhere (e.g. ae "…" -q silently
learns; ae add … -q adds without printing).
rm --all (no acronym) clears the entire dictionary and resets to the built-in
defaults. It isn't gated by a prompt — instead it first writes a timestamped
snapshot to /tmp/ae/backup_<ts>.db, and ae rm --restore [PATH] brings it back
(with no path, the most recent backup; the pre-restore state is itself snapshotted,
so a restore is reversible).
ignore (alias mute) is for acronym-shaped tokens you never want surfaced or
mined — e.g. an all-caps word a colleague keeps shouting. The entry stays in the
DB but is left out of expansion, mining, suggestions, and candidate flagging
until unignore. That's the difference from rm, which deletes outright.
Fully-capitalized lines are also skipped automatically: with no lowercase to
contrast against, an all-caps sentence would otherwise flag every short word.
Each acronym has a provenance: declared (you said it's an acronym, via
ae add ACR with no expansion) or seen (ae noticed it). An acronym joins the
watch list — where we hunt its expansions in later text — once it's declared
or has been seen enough times (default 3); below that it's noise and ae prune
drops it. Punctuated acronyms (PB&J, R&D, U.S.A) are
detected and mined too (the &/. maps to a skipped filler word), and a longer
match wins over its parts (PB&J beats PB, maximal munch).
ae prune also spell-corrects mined expansions against the system word list
(/usr/share/dict/words, if present — nothing bundled), so "minimum viabel
product" converges to "minimum viable product" before dedup.
You rarely need to run it by hand: the same pass runs automatically on a
cadence after a write — at most once per AE_CONSOLIDATE_SECS (default daily; a
negative value disables it). The point is the quality half — spell-fix and
dedup pool evidence and lift confidence — so it's worth running regularly; the
deletion half is gentle (a candidate seen within AE_PRUNE_GRACE_SECS, default
~30 days, is spared, so an infrequent token can recur weeks later before it's
ever considered noise). The warm daemon amortizes it across requests.
ae define MVP with no expansions picks interactively from the mined
suggestions — via fzf (multi-select) if installed, else a numbered prompt. An
acronym can hold several expansions, so multi-select is first-class.
ae prune is the occasional GC: it merges prefix-duplicate expansions
("min viable product" → "minimum viable product"), drops ones below a confidence
floor (default 0.15), and removes seen-once noise candidates. ae suggest keeps
a higher bar by default (0.30) since speculation is noisy; both take
--min-confidence to override, and suggest takes -l/--limit N.
ae suggest is the payoff of tracking candidates. When ae analyzes text it
mines word-sequences whose initials spell a watched candidate acronym — no
parentheses needed, tolerant of skipped filler words (OKR = Objectives and
Key Results), and across subsequent inputs too (a definition mentioned in a
later sentence that never repeats the acronym still accrues). It counts how often
each phrase recurs and how well its context fits where the acronym is used, then
blends both into a confidence:
$ ae add FOO # declare it — start watching
ae: now watching FOO for expansions
$ ae "the Foundations Of Onboarding workshop was great" # initials spell FOO
No acronyms found.
$ ae suggest FOO --min-confidence 0 # ae mined the phrase
FOO foundation of onboarding 0.50 (1)
Confirm one with ae define MVP (interactive) or ae add MVP "Minimum Viable Product" (which clears it from the candidate/suggestion lists). It's heuristic —
short acronyms attract noise, which sinks to the bottom on low confidence and is
hidden by the default threshold.
Removal disambiguates when an acronym has several expansions:
ae rm MVP # removes it if there's one expansion; else lists them and stops
ae rm MVP valuable # substring picks one ("Most Valuable Player")
ae rm MVP --all # removes every expansion
ae candidates is fed automatically: every analysis records the acronym-shaped
tokens it couldn't resolve and counts how often you've used them, so you can see
what's worth defining. (Defining one clears it from the list.) All of these
honor -j/-J too.
Default output is human-readable; -j/-J switch to JSON/NDJSON. A positional
TEXT is analyzed as one blob; piped stdin and --file are streamed line by
line — but either way the structured output is the same flat list of findings,
so a consumer parses one shape. With nothing to do, ae prints help. stdout
carries only data — all logs go to stderr, so ae … | jq is always safe. Every
command is machine-friendly: -j/-J work everywhere, and --daemon/--stop
emit a {"status": …} object in those modes. ae --status reports a running
daemon's version, embedder, and uptime (read-only — it never starts one), and
exits non-zero when none is up, so --status -q is a silent health check.
Under -J each finding is one compact object, flushed as it's produced (so
tail -f … | ae -J streams live); -j collects them into one pretty array.
--read-only is the safe path for untrusted or high-volume input — it expands
known acronyms without ever writing to the dictionary.
--model lets you point at any compatible model — an absolute/relative path to a
model directory or .onnx file, or a HuggingFace org/name repo id (fetched into
the shared Hub cache). With no flag, ae uses the default model from the Hub
($AE_MODEL_DIR overrides with a local copy), and falls back to the hash embedder
if none loads.
input ─▶ file lock ─▶ Leader (UDS server, warm trie + dictionary + embedder)
└─▶ Follower (forwards text, pipes back JSON)
evaluation = STAGE 1 expansion (trie scan → dictionary → 64-d MRL vector match)
+ STAGE 2 learning (rule-based extraction of inline definitions)
Embeddings come from all-MiniLM-L6-v2 (int8-quantized ONNX) run locally via ONNX Runtime — tokenize, mean-pool, then compress with Matryoshka Representation Learning: the 384-d vector is truncated to its first 64 coordinates and L2-normalized, shrinking the vector store ~6×. See docs/SPEC.md for the full design and docs/ROADMAP.md for status and deliberate deviations (notably the model choice and the runtime model-fetch strategy).
cargo test # unit + integration tests
cargo clippy --all-targets
cargo fmt
See CLAUDE.md for conventions.
MIT — see LICENSE.txt.
70 commits
1 commits
Rust
99.3%