A software Turing–Welchman Bombe that breaks real Enigma traffic, with a probabilistic analyst deciding what counts as German.

enigma-jev rebuilds Bletchley Park's pipeline in software, then measures it. The pipeline is a verified Enigma I, M3 and M4, crib dragging, a Bombe with Welchman's diagonal board, and a ciphertext-only hill-climber. Two steps in that pipeline were human judgement calls: which probable phrase to try, and whether a trial decryption is German. Here Jev, a model that answers only typed probabilistic questions, makes those two calls. The code runs a tiered backtest against historical intercepts with published keys. A second study compares Jev against n-gram judges and XGBoost on the same candidates.
It comes with three pages:
The machine (/) | A 3D Enigma I you can key, type on and transmit from. Bletchley then breaks your message live: crib dragging, Jev's crib ranking, the Bombe across a worker pool, and Jev's verdict. |
Research & history (/research) | A long-form paper covering Rejewski to the M4 Project, how the system is built, and its evaluation. Five figures are computed live in the browser, including a running Bombe. |
Jev performance (/jev) | A second paper evaluating Jev as crib selector and plaintext judge. It covers discrimination, calibration, decisions, reliability, cost, and a head-to-head against XGBoost. |
You need Bun 1.3 or later, and a TypeSafe API key for the Jev steps.
bun install
bun run web
Open http://localhost:5199. The machine stays closed until you enter a TypeSafe key. The server checks it with one Jev question, then seals it into an encrypted, HttpOnly session cookie that the page's scripts can't read and only the server can open. The key is never written to disk and never logged. The Research and Jev pages need no key.
On the command line, put the key in .env (see .env.example):
bun run cli doctor
bun run cli break "GCDSE AHUGW TQGRK VLFGX UCALX VYMIG MMNMF DXTGN VHVRM MEVOU YFZSL RHDRR XFJWC FHUHM UNZEF RDISI KBGPM YVXUZ" --machine I --date 1930 --crib FEINDLIQEINFANTERIE
That is the test message from the 1930 Enigma I manual. The Bombe runs your crib first and recovers the key (UKW A, II-I-III, plugs AM FI NV PS TU WZ). Add --no-jev to judge with n-gram statistics only; drop --crib to let Jev rank the built-in crib list.
bun run cli encrypt "ANGRIFF IM MORGENGRAUEN" --rotors II,IV,V --rings BUL --start BLA --plugs "AV BS"
The repository deploys to Vercel as it is:
bun run build → dist/)./api/* route runs in one Bun function (api/index.ts), with up to 300 seconds for a break.vercel.json holds the configuration. The project needs two settings:
vercel.json sets the build.SESSION_SECRET: any long random string, for example from openssl rand -base64 32. It encrypts the session cookie, so every function instance can read it. Without it, unlocking is refused.A break on Vercel runs on the function's CPU, and its duration is capped at 300 seconds on the Hobby plan. That is ample for a break from your own crib; a long search through the crib list can hit the cap. For those, run the server locally.
These results are from the main run (seed 1941): 9 historical messages with published keys and 16 synthetic messages under seeded random keys. A tier breaks a message when its best candidate matches at least 90% of the published plaintext.
| Tier | What is known | Broken |
|---|---|---|
verify | the full key; the machine must reproduce the published plaintext | 25/25 |
key | the daily key; the message key is searched | 25/25 |
crib | the message's true first 14 letters | 22/25 |
bombe | machine and service; cribs are chosen from a fixed list | 0/25 |
climb | the machine only (ciphertext only) | 0/25 |
Ten-plug traffic falls to a Bombe with the right crib and not to ciphertext-only statistics. That was also the position in 1940.
Jev as judge (478 candidates from 160 judge calls, labelled against the known plaintext):
Against the standard judges (leave-one-text-out cross-validation, cluster bootstrap):
| Judge | AUC | Right decisions | Trained on the evaluation traffic? |
|---|---|---|---|
| Index of coincidence | 0.988 | 88.8% | yes |
| Trigram German-ness | 0.993 | 93.8% | yes |
| Quadgram fitness | 0.9995 | 98.8% | yes |
| Kneser–Ney 5-gram | 0.9985 | 98.8% | yes |
| Logistic regression, 12 features | 0.9996 | 97.5% | yes |
| XGBoost, 12 features | 0.993 | 98.8% | yes |
| Jev, zero-shot | 0.9999 | 98.8% | no |
On the same kind of traffic, trained judges tie with Jev. Under distribution shift they do not. Trained on synthetic traffic and tested on the real intercepts, XGBoost decides 53% of calls correctly, while zero-shot Jev decides 97%.
The Bombe against Weinbaum (2025). The test register reproduces the stop counts published for one- and two-loop menus, within their confidence intervals. It also reproduces the share of settings that can produce a Zygalski female: 0.4059 measured over all 60 wheel orders, against 0.4052 in Weinbaum's model.
The full tables, intervals and caveats are in the two papers. Every number there is read from reports/ at page load.
flowchart LR
I[Intercept] --> D[Crib dragging<br/>no letter enciphers to itself]
D --> R{{Jev: which crib?}}
R --> B[Bombe<br/>menu, diagonal board,<br/>every wheel order × 17,576]
B --> S[Stop ranking<br/>ring-independent trigram score]
S --> H[Plugboard hill-climb<br/>+ ring refinement]
H --> J{{Jev: is this German,<br/>and which one?}}
J -->|accept| K[Key + plaintext]
J -->|none| B
B -.->|no crib breaks it| C[Ciphertext-only<br/>IoC scan + climb]
C --> J
choice question over the surviving cribs, plus "none".choice over up to four distinct candidates plus "none", and a noul P(correct) for each. Jev sees only the texts, never the search's scores. On the web page, German letter statistics must also agree before a break is shown.More detail is in docs/architecture.md, the web API and event stream, and how Jev is used.
bun run reproduce
This runs offline. It runs the tests, rebuilds the Jev analysis from the local audit log (if there is one), recounts the Bombe stops and reruns the judge comparison. --backtests and --experiments rerun the steps that make Jev calls; those need a key.
| Result | Command | Output | Shown in |
|---|---|---|---|
| Tier table, main run | bun run backtest all --seed 1941 | reports/backtest-*.{md,json} | Research C.3, Jev §7 |
| Tier table, holdout | bun run backtest synthetic --seed 2024 | reports/backtest-*.{md,json} | Research C.3 |
| Jev discrimination, calibration, decisions | bun run analyze:jev | reports/jev-analysis.json | Jev §§1–7 |
| Test–retest and position experiments | bun run src/analysis/jev-experiments.ts | reports/jev-experiments.json | Jev §4.6 |
| Judge comparison, learning curve, shift | bun run analyze:judges | reports/judge-comparison.json | Jev §4.8 |
| Bombe stops against Weinbaum | bun run analyze:stops | reports/bombe-stops.json | Research B.5, Table B1 |
The Python step needs a virtual environment:
python3 -m venv .venv && .venv/bin/pip install -r analysis/requirements.txt
See docs/evaluation.md for the protocol, the scoring rules and how each report is built.
![]() | ![]() |
src/
enigma/ the machine: rotors I–VIII, Beta/Gamma, reflectors, rings, plugboard, double step
break/ Bombe, test register, ciphertext-only climb, the shared search engine
lang/ German, English and Spanish n-gram models; operator-style text normalisation
jev/ the Jev client, crib ranking, the judge, response schemas
pipeline/ tier runners and the acceptance rules shared by the CLI and the web page
backtest/ historical and synthetic cases, the runner, report writer
analysis/ Jev evaluation, reliability experiments, Bombe stop counts
web/ server, session gate, streamed break (SSE), worker pool
cli.ts break | backtest | encrypt | decrypt | doctor
api/ the Vercel function: every /api/* route (shares src/web/api.ts)
web/
pages/ index, research, jev
app/ the machine page: steps, gate, Bletchley panel
machine/ the CSS 3D Enigma
research/ live figures and the Bombe demo
jev/ charts and tables for the Jev paper
shared/ DOM helpers, glossary, citation previews, operator habits
analysis/ Python: n-gram and XGBoost judges, metrics, bootstrap
data/ historical intercepts, synthetic plaintexts, language corpora
reports/ every number the papers show
scripts/ reproduce.ts
test/ 52 tests
Jev is a hosted model served by TypeSafe's System One API. It answers typed questions about a text: noul (a probability), choice (a distribution over named options) and score. It never returns prose, so it cannot write out a decryption. All the cryptanalysis runs locally. Jev only makes the two calls an analyst made.
~/.enigma-jev/jev-calls.jsonl. The Jev paper is rebuilt from that file.TYPESAFE_API_KEY, then ./.env, then ~/.enigma-jev/.env. JEV_MODEL overrides the pinned default jev-1.13.0, the model behind every published number.Independence. enigma-jev is an independent open-source project. It is not affiliated with, endorsed by or sponsored by TypeSafe. "TypeSafe", "Jev" and "System One" are used only to name the service this code calls. You need your own key, and your use of the API is governed by TypeSafe's terms.
bombe tier.jev-1.13.0 and the raw calls are logged.Contributions are welcome, and historical intercepts with sound provenance most of all. CONTRIBUTING.md covers setup and the checks every change must pass:
bun run check
data/README.md explains how to add a message. Report security issues as described in SECURITY.md.
If you use this code or its results, please cite it as described in CITATION.cff.
The historical messages and their keys come from the work of Frode Weierud and Geoff Sullivan (cryptocellar), Stefan Krah and the M4 Message Breaking Project (bytereef), Enigma-Hörenberg, the Crypto Museum, the Franklin Heath Enigma wiki, the ringstellung Enigma collection and German Wikipedia. Each message's own sources are listed in data/historical/messages.json. The Bombe validation follows Jonah Weinbaum, Action This Day (Dartmouth College MS thesis, 2025).
MIT © 2026 Andres Godoy
TypeScript
65.2%
HTML
18.2%
CSS
12.9%
Python
3.8%
A software Turing–Welchman Bombe that breaks real Enigma traffic, with a probabilistic analyst deciding what counts as German.

enigma-jev rebuilds Bletchley Park's pipeline in software, then measures it. The pipeline is a verified Enigma I, M3 and M4, crib dragging, a Bombe with Welchman's diagonal board, and a ciphertext-only hill-climber. Two steps in that pipeline were human judgement calls: which probable phrase to try, and whether a trial decryption is German. Here Jev, a model that answers only typed probabilistic questions, makes those two calls. The code runs a tiered backtest against historical intercepts with published keys. A second study compares Jev against n-gram judges and XGBoost on the same candidates.
It comes with three pages:
The machine (/) | A 3D Enigma I you can key, type on and transmit from. Bletchley then breaks your message live: crib dragging, Jev's crib ranking, the Bombe across a worker pool, and Jev's verdict. |
Research & history (/research) | A long-form paper covering Rejewski to the M4 Project, how the system is built, and its evaluation. Five figures are computed live in the browser, including a running Bombe. |
Jev performance (/jev) | A second paper evaluating Jev as crib selector and plaintext judge. It covers discrimination, calibration, decisions, reliability, cost, and a head-to-head against XGBoost. |
You need Bun 1.3 or later, and a TypeSafe API key for the Jev steps.
bun install
bun run web
Open http://localhost:5199. The machine stays closed until you enter a TypeSafe key. The server checks it with one Jev question, then seals it into an encrypted, HttpOnly session cookie that the page's scripts can't read and only the server can open. The key is never written to disk and never logged. The Research and Jev pages need no key.
On the command line, put the key in .env (see .env.example):
bun run cli doctor
bun run cli break "GCDSE AHUGW TQGRK VLFGX UCALX VYMIG MMNMF DXTGN VHVRM MEVOU YFZSL RHDRR XFJWC FHUHM UNZEF RDISI KBGPM YVXUZ" --machine I --date 1930 --crib FEINDLIQEINFANTERIE
That is the test message from the 1930 Enigma I manual. The Bombe runs your crib first and recovers the key (UKW A, II-I-III, plugs AM FI NV PS TU WZ). Add --no-jev to judge with n-gram statistics only; drop --crib to let Jev rank the built-in crib list.
bun run cli encrypt "ANGRIFF IM MORGENGRAUEN" --rotors II,IV,V --rings BUL --start BLA --plugs "AV BS"
The repository deploys to Vercel as it is:
bun run build → dist/)./api/* route runs in one Bun function (api/index.ts), with up to 300 seconds for a break.vercel.json holds the configuration. The project needs two settings:
vercel.json sets the build.SESSION_SECRET: any long random string, for example from openssl rand -base64 32. It encrypts the session cookie, so every function instance can read it. Without it, unlocking is refused.A break on Vercel runs on the function's CPU, and its duration is capped at 300 seconds on the Hobby plan. That is ample for a break from your own crib; a long search through the crib list can hit the cap. For those, run the server locally.
These results are from the main run (seed 1941): 9 historical messages with published keys and 16 synthetic messages under seeded random keys. A tier breaks a message when its best candidate matches at least 90% of the published plaintext.
| Tier | What is known | Broken |
|---|---|---|
verify | the full key; the machine must reproduce the published plaintext | 25/25 |
key | the daily key; the message key is searched | 25/25 |
crib | the message's true first 14 letters | 22/25 |
bombe | machine and service; cribs are chosen from a fixed list | 0/25 |
climb | the machine only (ciphertext only) | 0/25 |
Ten-plug traffic falls to a Bombe with the right crib and not to ciphertext-only statistics. That was also the position in 1940.
Jev as judge (478 candidates from 160 judge calls, labelled against the known plaintext):
Against the standard judges (leave-one-text-out cross-validation, cluster bootstrap):
| Judge | AUC | Right decisions | Trained on the evaluation traffic? |
|---|---|---|---|
| Index of coincidence | 0.988 | 88.8% | yes |
| Trigram German-ness | 0.993 | 93.8% | yes |
| Quadgram fitness | 0.9995 | 98.8% | yes |
| Kneser–Ney 5-gram | 0.9985 | 98.8% | yes |
| Logistic regression, 12 features | 0.9996 | 97.5% | yes |
| XGBoost, 12 features | 0.993 | 98.8% | yes |
| Jev, zero-shot | 0.9999 | 98.8% | no |
On the same kind of traffic, trained judges tie with Jev. Under distribution shift they do not. Trained on synthetic traffic and tested on the real intercepts, XGBoost decides 53% of calls correctly, while zero-shot Jev decides 97%.
The Bombe against Weinbaum (2025). The test register reproduces the stop counts published for one- and two-loop menus, within their confidence intervals. It also reproduces the share of settings that can produce a Zygalski female: 0.4059 measured over all 60 wheel orders, against 0.4052 in Weinbaum's model.
The full tables, intervals and caveats are in the two papers. Every number there is read from reports/ at page load.
flowchart LR
I[Intercept] --> D[Crib dragging<br/>no letter enciphers to itself]
D --> R{{Jev: which crib?}}
R --> B[Bombe<br/>menu, diagonal board,<br/>every wheel order × 17,576]
B --> S[Stop ranking<br/>ring-independent trigram score]
S --> H[Plugboard hill-climb<br/>+ ring refinement]
H --> J{{Jev: is this German,<br/>and which one?}}
J -->|accept| K[Key + plaintext]
J -->|none| B
B -.->|no crib breaks it| C[Ciphertext-only<br/>IoC scan + climb]
C --> J
choice question over the surviving cribs, plus "none".choice over up to four distinct candidates plus "none", and a noul P(correct) for each. Jev sees only the texts, never the search's scores. On the web page, German letter statistics must also agree before a break is shown.More detail is in docs/architecture.md, the web API and event stream, and how Jev is used.
bun run reproduce
This runs offline. It runs the tests, rebuilds the Jev analysis from the local audit log (if there is one), recounts the Bombe stops and reruns the judge comparison. --backtests and --experiments rerun the steps that make Jev calls; those need a key.
| Result | Command | Output | Shown in |
|---|---|---|---|
| Tier table, main run | bun run backtest all --seed 1941 | reports/backtest-*.{md,json} | Research C.3, Jev §7 |
| Tier table, holdout | bun run backtest synthetic --seed 2024 | reports/backtest-*.{md,json} | Research C.3 |
| Jev discrimination, calibration, decisions | bun run analyze:jev | reports/jev-analysis.json | Jev §§1–7 |
| Test–retest and position experiments | bun run src/analysis/jev-experiments.ts | reports/jev-experiments.json | Jev §4.6 |
| Judge comparison, learning curve, shift | bun run analyze:judges | reports/judge-comparison.json | Jev §4.8 |
| Bombe stops against Weinbaum | bun run analyze:stops | reports/bombe-stops.json | Research B.5, Table B1 |
The Python step needs a virtual environment:
python3 -m venv .venv && .venv/bin/pip install -r analysis/requirements.txt
See docs/evaluation.md for the protocol, the scoring rules and how each report is built.
![]() | ![]() |
src/
enigma/ the machine: rotors I–VIII, Beta/Gamma, reflectors, rings, plugboard, double step
break/ Bombe, test register, ciphertext-only climb, the shared search engine
lang/ German, English and Spanish n-gram models; operator-style text normalisation
jev/ the Jev client, crib ranking, the judge, response schemas
pipeline/ tier runners and the acceptance rules shared by the CLI and the web page
backtest/ historical and synthetic cases, the runner, report writer
analysis/ Jev evaluation, reliability experiments, Bombe stop counts
web/ server, session gate, streamed break (SSE), worker pool
cli.ts break | backtest | encrypt | decrypt | doctor
api/ the Vercel function: every /api/* route (shares src/web/api.ts)
web/
pages/ index, research, jev
app/ the machine page: steps, gate, Bletchley panel
machine/ the CSS 3D Enigma
research/ live figures and the Bombe demo
jev/ charts and tables for the Jev paper
shared/ DOM helpers, glossary, citation previews, operator habits
analysis/ Python: n-gram and XGBoost judges, metrics, bootstrap
data/ historical intercepts, synthetic plaintexts, language corpora
reports/ every number the papers show
scripts/ reproduce.ts
test/ 52 tests
Jev is a hosted model served by TypeSafe's System One API. It answers typed questions about a text: noul (a probability), choice (a distribution over named options) and score. It never returns prose, so it cannot write out a decryption. All the cryptanalysis runs locally. Jev only makes the two calls an analyst made.
~/.enigma-jev/jev-calls.jsonl. The Jev paper is rebuilt from that file.TYPESAFE_API_KEY, then ./.env, then ~/.enigma-jev/.env. JEV_MODEL overrides the pinned default jev-1.13.0, the model behind every published number.Independence. enigma-jev is an independent open-source project. It is not affiliated with, endorsed by or sponsored by TypeSafe. "TypeSafe", "Jev" and "System One" are used only to name the service this code calls. You need your own key, and your use of the API is governed by TypeSafe's terms.
bombe tier.jev-1.13.0 and the raw calls are logged.Contributions are welcome, and historical intercepts with sound provenance most of all. CONTRIBUTING.md covers setup and the checks every change must pass:
bun run check
data/README.md explains how to add a message. Report security issues as described in SECURITY.md.
If you use this code or its results, please cite it as described in CITATION.cff.
The historical messages and their keys come from the work of Frode Weierud and Geoff Sullivan (cryptocellar), Stefan Krah and the M4 Message Breaking Project (bytereef), Enigma-Hörenberg, the Crypto Museum, the Franklin Heath Enigma wiki, the ringstellung Enigma collection and German Wikipedia. Each message's own sources are listed in data/historical/messages.json. The Bombe validation follows Jonah Weinbaum, Action This Day (Dartmouth College MS thesis, 2025).
MIT © 2026 Andres Godoy
TypeScript
65.2%
HTML
18.2%
CSS
12.9%
Python
3.8%