A command generator selected for its balance of ALFA and internal benchmark results. This is a trained EasyCommand checkpoint derived from
Qwen/Qwen2.5-Coder-1.5B-Instruct
at revision 2e1fd397ee46e1388853d2af2c993145b0f1098a. It generates GNU/Linux Bash commands from
English requests as {"kind":"COMMAND","value":"<command>"}.
All GGUFs are complete merged models with tokenizer and chat-template metadata. They require no separate adapter or upstream weight download for inference. The selected research checkpoint is A3; all quants here share that trained state.
| File | Quant | Size | ALFA-updated /300 | Internal /1320 |
|---|---|---|---|---|
| ec-1.5b.Q4_K_M.gguf | Q4_K_M | 940.4 MiB | 212/300 | 1119/1320 |
| Model/checkpoint | Quant | ALFA-updated /300 | Pass rate | Internal /1320 |
|---|---|---|---|---|
| EC 1.5B A3 (this release) | Q4_K_M | 212/300 | 70.67% | 1119/1320 |
| whatisit / nl2sh-1.5b | Q4_K_M | 191/300 | 63.67% | 82/1320 |
| EC 0.6B A1 | Q8_0 | 174/300 | 58.00% | 922/1320 |
| EC 0.6B A1 | Q4_K_M | 165/300 | 55.00% | 911/1320 |
Benchmark note: This comparison uses an updated version of the ALFA benchmark, ALFA-updated, with documented environment and correctness repairs. The benchmark source is public. These are locally measured results under the same frozen grading definition, not scores from the original ALFA grader. Benchmark errors remain in the 300-task denominator: one for each EC Q4 run, two for EC Q8, and five for whatisit.
Each model uses its own serving profile: EC uses the JSON system prompt below and a 256-token limit; whatisit uses its native plain-command prompt, a 64-token limit and its published repetition settings. Its saved predictions were regraded without changing the commands. This compares deployed configurations, not an experiment holding prompt and decoding settings constant. The complete whatisit internal measurement is 82/1320 (6.21%) across all six current panels. We reused 1,160 saved responses after exact ID/request and native extraction checks, generated the missing transfer160 responses, and graded all 1,320 with the current accepted fixture policies. No older panel scores were pooled. The internal run had 1 benchmark error, retained as failed cases in the 1,320 denominator. The native plaintext responses were admitted through a model-specific adapter; they were not required to satisfy EC's JSON generation protocol.
These author-reported results provide context only. They are not directly comparable with the ALFA-updated scores above. Sources were checked on 2026-10-05; none of these rows was regraded with ALFA-updated here.
| Model/configuration | Quant | Reported download size | Original-ALFA pass rate | Source |
|---|---|---|---|---|
| GPT-4o (cloud API; parsed output) | Hosted | Not downloadable | 73.0% | whatisit HF results, paper, Table 5 |
| nl2sh-3b | Q4_K_M | 1.9 GB | 65.7% | whatisit benchmarks |
| nl2sh-qwen25-coder-1.5b (120k-pool build) | Q4_K_M | 941 MB | 65.67% | Model card |
| nl2sh-qwen25-coder-1.5b (TPU build) | Q4_K_M | 0.99 GB | 63.67% | Model card |
| whatisit / nl2sh-1.5b | Q4_K_M | 941 MB | 62.0% | whatisit HF results |
| Qwen2.5-Coder-7B-Instruct, untuned | Not specified in source table | 4.4 GB | 61.3% | whatisit HF results |
| Qwen2.5-Coder-1.5B-Instruct, untuned | Not specified in source table | 941 MB | 54.0% | whatisit HF results |
Sizes are quoted in each source's units and rounding; they are not standardized
byte measurements. GPT-4o's 73% is the reference cited by whatisit; Table 5 of the
original paper reports 73% with parsing (also 73% with in-context learning),
versus 74% for its baseline. It used gpt-4o-2024-08-06. This row is a paper
measurement, not a fresh run by us or by whatisit, and does not inherit their
64-token serving profile.
The whatisit and 120k-pool sources report the unmodified upstream scorer, greedy generation, a 64-token limit and embedding threshold 0.75. The TPU card reports 300 tasks and temperature zero but does not specify every grading option. The published 62.0% for whatisit and our local 191/300 ALFA-updated result are separate measurements and should not be substituted for one another.
The community author also remeasured whatisit at 59.0% on their own rig, compared with its upstream published 62.0%. That is an external original-ALFA measurement, not our local remeasurement and not an ALFA-updated result.
This is the A3 repair checkpoint, not the earlier parent or the find specialist. A3 Q4 scored 212/300 ALFA and 1119/1320 internal. The parent's latest repeat scored 208/1115; its earlier 210 ALFA and 1127 internal results are separate measurements, with 1127 using an older prompt. The find specialist scored 217/1056 and was not selected for this release. The differences over the parent are modest; these runs do not establish a statistically reliable improvement.
| Model / quant | Commands /244 | Quoting /256 | Operands /512 | Time /84 | Transfers /160 | English /64 |
|---|---|---|---|---|---|---|
| EC 1.5B A3 / Q4_K_M | 232 | 237 | 417 | 80 | 89 | 64 |
| whatisit 1.5B / Q4_K_M | 80 | 0 | 0 | 2 | 0 | 0 |
The transfer panel measures existing operation/source-versus-runtime diagnostics.
It is the current transfer160 panel, not a sum of historical pair64/fresh96 panels.
Whatisit inference used the pinned Q4 model and llama-server 0.4.1-dev
(161755f29), eight CPU threads, context 4096 and one slot for the new responses,
matching the recorded historical serving profile. The time observer error was
caught at the observer boundary so remaining cases could be measured; its case
remains a failure. Frozen fixture functions and expected outputs were unchanged.
Suite composition: Quoting256, operands512 and English64 account for 832/1320 (63.0%) cases and test closely related literal-search/output-mode contracts. Transfer160 covers ten text-processing families in source-text and runtime-output variants; time84 tests age predicates. These are concentrated retention/stress diagnostics, not broad shell-task accuracy. Separate known-correct native plaintext controls passed 12/12 across the four zero panels, confirming that their commands can pass the same extraction and frozen fixtures without generating JSON. An actual native-template check preserved all 1,320 requests. Four of 14 separate native replays differed from the saved commands; saved and newly generated responses are separately bound, and deterministic replay equivalence is not assumed.
ALFA scores use ALFA-updated, a documented
variant with grading definition SHA256 6f50e4fea37c125f73f93e5f785e95e91fda448b4bf347a9a2986db4d83b48dc. They are not original ALFA
scores and should not be compared directly with published results from another protocol.
The ALFA run had one benchmark error, retained in the 300-task denominator. The internal panels contain related templates and were consulted during
development; 1320 cases are not 1320 independent unseen tasks. ALFA also informed
repair strategy and checkpoint selection. These are development measurements.
Only the published quants' measured results are shown; they do not imply BF16 performance.
evaluation.json contains the per-panel counts and source-summary hashes.
Install the EasyCommand application, then download the Q4 model from that checkout using its checksum-verifying helper:
python3 scripts/download-model \
'https://huggingface.co/dirac-run/ec-1.5b-gguf/resolve/main/ec-1.5b.Q4_K_M.gguf' \
--sha256 135c8ec1a7ad6c24ba25971ee01cb3659130676031e45d876bd1822185c9e044 \
--output ec-1.5b.Q4_K_M.gguf
ec --model ec-1.5b.Q4_K_M.gguf --preview print the system uptime
ec --model ec-1.5b.Q4_K_M.gguf list the last five commits in this repository
--preview prints COMMAND JSON without executing. Ordinary ec usage extracts
the command, displays it and asks for confirmation before running it.
Use the exact system message in system-prompt.txt:
You are a GNU/Linux shell command generator. Produce the simplest Bash command that fulfills the entire request. Return only valid JSON: {"kind":"COMMAND","value":"<command>"}.
It is 177 bytes, SHA256 3a9028d5aebb73c3ed7363e63eb3e751689218775ea5572ab0dc6807e522238b. Evaluations used greedy decoding, no
generation grammar, a 256-token output limit, EOS 151645 and the native CPU
decoder based on llama.cpp revision 1af554f8fc78ba029665a47b839484d9763e2a75.
Use a system message and one user message, then the assistant generation prefix.
Use the upstream Qwen2.5 ChatML format with an assistant generation prefix. Do not add a Qwen3 thinking prefix. The ec application applies the correct format. See inference.json for the profile. Other templates,
prompts, runtime builds or CPU kernels can change outputs and should be evaluated separately.
The model started from the pinned upstream weights, received one weighted epoch of LoRA training (474,635 presentations; 14,833 updates), and then 200 incremental repair/replay updates from that trained parent. This is not a stock model and not a fresh 200-step fine-tune from stock weights.
Both stages used rank 32, alpha 64, dropout 0.05, effective batch 32 and adapters
on q/k/v/o attention projections and gate/up/down MLP projections. The base stayed
BF16 with FP32 trainable adapters. AdamW used linear warmup and cosine decay.
The parent peak learning rate was 0.0001; the continuation used
5e-06, ten warmup updates and 50% repair / 50% replay sampling from
2,820 source rows. Only assistant answer/EOS tokens were supervised.
The parent epoch ran on H100 and the continuation on A40.
The parent used the older prompt in training-system-prompt.txt. Continuation training and reported release evaluations used the shorter serving prompt. training.json records the actual recipe and lineage.
The released dataset has 401,975 deduplicated request/answer pairs, including original and concise descriptions, corrected/simplified commands and incremental repair data. It also includes rows from other repair experiments. A single pass over it does not reproduce these models' weighted exposure or sampling history. Merge used the original base and trained adapter in FP32, followed by F16 GGUF conversion and direct quantization; no requantization or importance matrix was used.
Designed for English requests targeting Bash and GNU/Linux utilities. The model does not inspect the live filesystem or know which tools are installed. It gives a best-effort command rather than clarification or inability responses. Generated commands can be wrong, incomplete or destructive; review them before execution.
This repository contains GGUF inference exports. The trainable BF16 model and original LoRA adapter are released separately for further fine-tuning. They identify the same trained checkpoint; numeric formats and kernels can yield different outputs. To train from stock weights, use the released dataset and record your own exposure and validation.
Weights and documentation are released under Apache-2.0, retaining
the upstream attribution in NOTICE. The application has its own license.
Verify downloads from this directory with sha256sum --check SHA256SUMS.
manifest.json lists the weight files and hashes.
A command generator selected for its balance of ALFA and internal benchmark results. This is a trained EasyCommand checkpoint derived from
Qwen/Qwen2.5-Coder-1.5B-Instruct
at revision 2e1fd397ee46e1388853d2af2c993145b0f1098a. It generates GNU/Linux Bash commands from
English requests as {"kind":"COMMAND","value":"<command>"}.
All GGUFs are complete merged models with tokenizer and chat-template metadata. They require no separate adapter or upstream weight download for inference. The selected research checkpoint is A3; all quants here share that trained state.
| File | Quant | Size | ALFA-updated /300 | Internal /1320 |
|---|---|---|---|---|
| ec-1.5b.Q4_K_M.gguf | Q4_K_M | 940.4 MiB | 212/300 | 1119/1320 |
| Model/checkpoint | Quant | ALFA-updated /300 | Pass rate | Internal /1320 |
|---|---|---|---|---|
| EC 1.5B A3 (this release) | Q4_K_M | 212/300 | 70.67% | 1119/1320 |
| whatisit / nl2sh-1.5b | Q4_K_M | 191/300 | 63.67% | 82/1320 |
| EC 0.6B A1 | Q8_0 | 174/300 | 58.00% | 922/1320 |
| EC 0.6B A1 | Q4_K_M | 165/300 | 55.00% | 911/1320 |
Benchmark note: This comparison uses an updated version of the ALFA benchmark, ALFA-updated, with documented environment and correctness repairs. The benchmark source is public. These are locally measured results under the same frozen grading definition, not scores from the original ALFA grader. Benchmark errors remain in the 300-task denominator: one for each EC Q4 run, two for EC Q8, and five for whatisit.
Each model uses its own serving profile: EC uses the JSON system prompt below and a 256-token limit; whatisit uses its native plain-command prompt, a 64-token limit and its published repetition settings. Its saved predictions were regraded without changing the commands. This compares deployed configurations, not an experiment holding prompt and decoding settings constant. The complete whatisit internal measurement is 82/1320 (6.21%) across all six current panels. We reused 1,160 saved responses after exact ID/request and native extraction checks, generated the missing transfer160 responses, and graded all 1,320 with the current accepted fixture policies. No older panel scores were pooled. The internal run had 1 benchmark error, retained as failed cases in the 1,320 denominator. The native plaintext responses were admitted through a model-specific adapter; they were not required to satisfy EC's JSON generation protocol.
These author-reported results provide context only. They are not directly comparable with the ALFA-updated scores above. Sources were checked on 2026-10-05; none of these rows was regraded with ALFA-updated here.
| Model/configuration | Quant | Reported download size | Original-ALFA pass rate | Source |
|---|---|---|---|---|
| GPT-4o (cloud API; parsed output) | Hosted | Not downloadable | 73.0% | whatisit HF results, paper, Table 5 |
| nl2sh-3b | Q4_K_M | 1.9 GB | 65.7% | whatisit benchmarks |
| nl2sh-qwen25-coder-1.5b (120k-pool build) | Q4_K_M | 941 MB | 65.67% | Model card |
| nl2sh-qwen25-coder-1.5b (TPU build) | Q4_K_M | 0.99 GB | 63.67% | Model card |
| whatisit / nl2sh-1.5b | Q4_K_M | 941 MB | 62.0% | whatisit HF results |
| Qwen2.5-Coder-7B-Instruct, untuned | Not specified in source table | 4.4 GB | 61.3% | whatisit HF results |
| Qwen2.5-Coder-1.5B-Instruct, untuned | Not specified in source table | 941 MB | 54.0% | whatisit HF results |
Sizes are quoted in each source's units and rounding; they are not standardized
byte measurements. GPT-4o's 73% is the reference cited by whatisit; Table 5 of the
original paper reports 73% with parsing (also 73% with in-context learning),
versus 74% for its baseline. It used gpt-4o-2024-08-06. This row is a paper
measurement, not a fresh run by us or by whatisit, and does not inherit their
64-token serving profile.
The whatisit and 120k-pool sources report the unmodified upstream scorer, greedy generation, a 64-token limit and embedding threshold 0.75. The TPU card reports 300 tasks and temperature zero but does not specify every grading option. The published 62.0% for whatisit and our local 191/300 ALFA-updated result are separate measurements and should not be substituted for one another.
The community author also remeasured whatisit at 59.0% on their own rig, compared with its upstream published 62.0%. That is an external original-ALFA measurement, not our local remeasurement and not an ALFA-updated result.
This is the A3 repair checkpoint, not the earlier parent or the find specialist. A3 Q4 scored 212/300 ALFA and 1119/1320 internal. The parent's latest repeat scored 208/1115; its earlier 210 ALFA and 1127 internal results are separate measurements, with 1127 using an older prompt. The find specialist scored 217/1056 and was not selected for this release. The differences over the parent are modest; these runs do not establish a statistically reliable improvement.
| Model / quant | Commands /244 | Quoting /256 | Operands /512 | Time /84 | Transfers /160 | English /64 |
|---|---|---|---|---|---|---|
| EC 1.5B A3 / Q4_K_M | 232 | 237 | 417 | 80 | 89 | 64 |
| whatisit 1.5B / Q4_K_M | 80 | 0 | 0 | 2 | 0 | 0 |
The transfer panel measures existing operation/source-versus-runtime diagnostics.
It is the current transfer160 panel, not a sum of historical pair64/fresh96 panels.
Whatisit inference used the pinned Q4 model and llama-server 0.4.1-dev
(161755f29), eight CPU threads, context 4096 and one slot for the new responses,
matching the recorded historical serving profile. The time observer error was
caught at the observer boundary so remaining cases could be measured; its case
remains a failure. Frozen fixture functions and expected outputs were unchanged.
Suite composition: Quoting256, operands512 and English64 account for 832/1320 (63.0%) cases and test closely related literal-search/output-mode contracts. Transfer160 covers ten text-processing families in source-text and runtime-output variants; time84 tests age predicates. These are concentrated retention/stress diagnostics, not broad shell-task accuracy. Separate known-correct native plaintext controls passed 12/12 across the four zero panels, confirming that their commands can pass the same extraction and frozen fixtures without generating JSON. An actual native-template check preserved all 1,320 requests. Four of 14 separate native replays differed from the saved commands; saved and newly generated responses are separately bound, and deterministic replay equivalence is not assumed.
ALFA scores use ALFA-updated, a documented
variant with grading definition SHA256 6f50e4fea37c125f73f93e5f785e95e91fda448b4bf347a9a2986db4d83b48dc. They are not original ALFA
scores and should not be compared directly with published results from another protocol.
The ALFA run had one benchmark error, retained in the 300-task denominator. The internal panels contain related templates and were consulted during
development; 1320 cases are not 1320 independent unseen tasks. ALFA also informed
repair strategy and checkpoint selection. These are development measurements.
Only the published quants' measured results are shown; they do not imply BF16 performance.
evaluation.json contains the per-panel counts and source-summary hashes.
Install the EasyCommand application, then download the Q4 model from that checkout using its checksum-verifying helper:
python3 scripts/download-model \
'https://huggingface.co/dirac-run/ec-1.5b-gguf/resolve/main/ec-1.5b.Q4_K_M.gguf' \
--sha256 135c8ec1a7ad6c24ba25971ee01cb3659130676031e45d876bd1822185c9e044 \
--output ec-1.5b.Q4_K_M.gguf
ec --model ec-1.5b.Q4_K_M.gguf --preview print the system uptime
ec --model ec-1.5b.Q4_K_M.gguf list the last five commits in this repository
--preview prints COMMAND JSON without executing. Ordinary ec usage extracts
the command, displays it and asks for confirmation before running it.
Use the exact system message in system-prompt.txt:
You are a GNU/Linux shell command generator. Produce the simplest Bash command that fulfills the entire request. Return only valid JSON: {"kind":"COMMAND","value":"<command>"}.
It is 177 bytes, SHA256 3a9028d5aebb73c3ed7363e63eb3e751689218775ea5572ab0dc6807e522238b. Evaluations used greedy decoding, no
generation grammar, a 256-token output limit, EOS 151645 and the native CPU
decoder based on llama.cpp revision 1af554f8fc78ba029665a47b839484d9763e2a75.
Use a system message and one user message, then the assistant generation prefix.
Use the upstream Qwen2.5 ChatML format with an assistant generation prefix. Do not add a Qwen3 thinking prefix. The ec application applies the correct format. See inference.json for the profile. Other templates,
prompts, runtime builds or CPU kernels can change outputs and should be evaluated separately.
The model started from the pinned upstream weights, received one weighted epoch of LoRA training (474,635 presentations; 14,833 updates), and then 200 incremental repair/replay updates from that trained parent. This is not a stock model and not a fresh 200-step fine-tune from stock weights.
Both stages used rank 32, alpha 64, dropout 0.05, effective batch 32 and adapters
on q/k/v/o attention projections and gate/up/down MLP projections. The base stayed
BF16 with FP32 trainable adapters. AdamW used linear warmup and cosine decay.
The parent peak learning rate was 0.0001; the continuation used
5e-06, ten warmup updates and 50% repair / 50% replay sampling from
2,820 source rows. Only assistant answer/EOS tokens were supervised.
The parent epoch ran on H100 and the continuation on A40.
The parent used the older prompt in training-system-prompt.txt. Continuation training and reported release evaluations used the shorter serving prompt. training.json records the actual recipe and lineage.
The released dataset has 401,975 deduplicated request/answer pairs, including original and concise descriptions, corrected/simplified commands and incremental repair data. It also includes rows from other repair experiments. A single pass over it does not reproduce these models' weighted exposure or sampling history. Merge used the original base and trained adapter in FP32, followed by F16 GGUF conversion and direct quantization; no requantization or importance matrix was used.
Designed for English requests targeting Bash and GNU/Linux utilities. The model does not inspect the live filesystem or know which tools are installed. It gives a best-effort command rather than clarification or inability responses. Generated commands can be wrong, incomplete or destructive; review them before execution.
This repository contains GGUF inference exports. The trainable BF16 model and original LoRA adapter are released separately for further fine-tuning. They identify the same trained checkpoint; numeric formats and kernels can yield different outputs. To train from stock weights, use the released dataset and record your own exposure and validation.
Weights and documentation are released under Apache-2.0, retaining
the upstream attribution in NOTICE. The application has its own license.
Verify downloads from this directory with sha256sum --check SHA256SUMS.
manifest.json lists the weight files and hashes.