Generate GNU/Linux shell commands locally with small trained models and an embedded llama.cpp runtime
C++
0
0 commits
updated Oct 5, 2026
ec)Turn an English request into a Bash command, entirely on your own machine. EasyCommand prints the command and asks before executing it.
Models and data · Research story · ALFA-updated benchmark
ec list the last five commits in this repository
The C++ application embeds llama.cpp. Inference needs a compatible GGUF model and ordinary Linux system libraries; it does not need Python, PyTorch, a GPU, Ollama or a network connection. Python is used only by optional setup, testing and training utilities.
Locally measured results for the released EC models and an external reference:
| Model | Quant | ALFA-updated /300 | Pass rate | Internal /1320 |
|---|---|---|---|---|
| EC 1.5B A3 | Q4_K_M | 212/300 | 70.67% | 1119/1320 |
| whatisit / nl2sh-1.5b | Q4_K_M | 191/300 | 63.67% | 82/1320 |
| EC 0.6B A1 | Q8_0 | 174/300 | 58.00% | 922/1320 |
| EC 0.6B A1 | Q4_K_M | 165/300 | 55.00% | 911/1320 |
These use ALFA-updated, with documented environment and correctness repairs. They are separate from published original-ALFA scores. Benchmark errors remain in the 300-task denominator: one for each EC Q4 run, two for EC Q8, and five for whatisit.
Internal suite breakdown: ordinary commands 244, quoting 256, exact operands 512, time predicates 84, text-processing transfer 160, and English wording 64. Quoting, operands and English contribute 832/1320 (63.0%) closely related literal-search/output-mode stress tests. Cases share templates, and these panels guided EC development; the total is not broad shell accuracy on 1,320 independent unseen tasks. The whatisit internal result retains one observer error as a failure.
The comparison keeps each model's serving profile: EC uses COMMAND JSON and a 256-token limit; whatisit uses its native plain-command prompt and a 64-token limit. This compares deployed configurations with the same graders, rather than holding inference settings constant. No fresh full-suite result is assigned to the merged BF16 exports.
See per-panel results and methodology, machine-readable measurements, and the research story for controls, historical comparisons, and limitations.
These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.
| Model/configuration | Reported download size | Original-ALFA pass rate | Source |
|---|---|---|---|
| GPT-4o, cloud API (published reference) | — | 73.0% | whatisit benchmarks |
| nl2sh-3b Q4_K_M | 1.9 GB | 65.7% | whatisit benchmarks |
| Community nl2sh-qwen25-coder-1.5b Q4_K_M | 941 MB | 65.67% | Community model card |
| whatisit / nl2sh-1.5b Q4_K_M | 941 MB | 62.0% | whatisit benchmarks |
| Qwen2.5-Coder-7B, untuned | 4.4 GB | 61.3% | whatisit benchmarks |
| Qwen2.5-Coder-1.5B, untuned | 941 MB | 54.0% | whatisit benchmarks |
The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.
The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.
Requirements: Linux, Bash, a C/C++ compiler, CMake 3.20+, and network access for the first dependency download. Build dependencies have fixed versions and checksums. Python 3.11+ is needed for the installer and tests.
A prebuilt Linux x86_64 bundle is also available, with its checksum, license notices and exact CPU/system requirements. Build from source if your machine does not meet those requirements.
git clone https://github.com/dirac-run/ec.git
cd ec
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --parallel
ctest --test-dir build --output-on-failure
python3 scripts/install
The installer defaults to ~/.local/bin/ec and refuses an existing unrelated
command, including one elsewhere on PATH. If ec is already used, install
with python3 scripts/install --name easycommand. Direct CMake installation
also supports cmake --install build --prefix ~/.local, but does not perform
the installer's ownership/collision checks.
build/easycommand is a compatibility symlink to build/ec. For a build
optimized for your own CPU, add -DEC_NATIVE=ON; that binary may not run on
another CPU. Prebuilt binaries require their stated operating-system/CPU profile.
Use an EasyCommand GGUF based on Qwen3-0.6B or Qwen2.5-Coder-1.5B-Instruct. These are trained checkpoints; the untouched upstream models are not equivalent. Weights are kept separately from this source repository.
| Release | Model repository | Published quants |
|---|---|---|
| 0.6B, selected A1 checkpoint | ec-0.6b-gguf | Q4_K_M, Q8_0 |
| 1.5B, selected A3 checkpoint | ec-1.5b-gguf | Q4_K_M |
For the 1.5B release, download and verify the model from this checkout:
python3 scripts/download-model \
'https://huggingface.co/dirac-run/ec-1.5b-gguf/resolve/main/ec-1.5b.Q4_K_M.gguf' \
--sha256 135c8ec1a7ad6c24ba25971ee01cb3659130676031e45d876bd1822185c9e044
ec print the system uptime
ec --model /path/to/model.gguf print the system uptime
ec --model /path/to/model.gguf --preview list the files in this directory
Without --model, the default is $XDG_DATA_HOME/easycommand/model.gguf, or
~/.local/share/easycommand/model.gguf. To install a downloaded model, put it
there or create a symlink to it. scripts/download-model accepts an HTTPS model
URL and its published SHA256 checksum; use --help for its arguments.
The runtime uses the exact prompt in config/system-prompt.txt, greedy decoding, a 256-token output limit, and thinking disabled for Qwen3. It expects COMMAND JSON and extracts the command. No JSON generation grammar is applied. Unsupported architectures are rejected explicitly.
ec --help lists the options. --preview prints validated JSON and executes
nothing. Ordinary requests print the proposed command; Enter, y or Y
confirms execution, and other answers cancel it. Execution requires interactive
input and output terminals. Commands are checked for Bash syntax and explicitly
named tools before the prompt. These checks do not establish that a generated
command fulfills the request; review the command before accepting it.
For repeated requests, a resident worker can keep the model in memory:
ec --serve --socket /tmp/ec.sock --model /path/to/model.gguf --threads 8
# In another terminal:
ec --socket /tmp/ec.sock --expect-model /path/to/model.gguf print the system uptime
The worker resets inference state between requests and verifies the requested model identity. Remove its socket after the worker exits before starting another. No systemd service or machine-specific shell configuration is required.
See the training guide for dataset format, LoRA continuation and GGUF conversion, and measured results for checkpoint and precision comparisons. ALFA-updated is a separately documented benchmark variant; its scores are not interchangeable with original ALFA scores. The project post draft covers the data expansion, training experiments, negative results, release choices and future directions.
Code is MIT licensed. Dependency licenses and credits are in THIRD_PARTY.md. Model and dataset terms are supplied separately with their respective distributions.
Generate GNU/Linux shell commands locally with small trained models and an embedded llama.cpp runtime
C++
0
0 commits
updated Oct 5, 2026
ec)Turn an English request into a Bash command, entirely on your own machine. EasyCommand prints the command and asks before executing it.
Models and data · Research story · ALFA-updated benchmark
ec list the last five commits in this repository
The C++ application embeds llama.cpp. Inference needs a compatible GGUF model and ordinary Linux system libraries; it does not need Python, PyTorch, a GPU, Ollama or a network connection. Python is used only by optional setup, testing and training utilities.
Locally measured results for the released EC models and an external reference:
| Model | Quant | ALFA-updated /300 | Pass rate | Internal /1320 |
|---|---|---|---|---|
| EC 1.5B A3 | Q4_K_M | 212/300 | 70.67% | 1119/1320 |
| whatisit / nl2sh-1.5b | Q4_K_M | 191/300 | 63.67% | 82/1320 |
| EC 0.6B A1 | Q8_0 | 174/300 | 58.00% | 922/1320 |
| EC 0.6B A1 | Q4_K_M | 165/300 | 55.00% | 911/1320 |
These use ALFA-updated, with documented environment and correctness repairs. They are separate from published original-ALFA scores. Benchmark errors remain in the 300-task denominator: one for each EC Q4 run, two for EC Q8, and five for whatisit.
Internal suite breakdown: ordinary commands 244, quoting 256, exact operands 512, time predicates 84, text-processing transfer 160, and English wording 64. Quoting, operands and English contribute 832/1320 (63.0%) closely related literal-search/output-mode stress tests. Cases share templates, and these panels guided EC development; the total is not broad shell accuracy on 1,320 independent unseen tasks. The whatisit internal result retains one observer error as a failure.
The comparison keeps each model's serving profile: EC uses COMMAND JSON and a 256-token limit; whatisit uses its native plain-command prompt and a 64-token limit. This compares deployed configurations with the same graders, rather than holding inference settings constant. No fresh full-suite result is assigned to the merged BF16 exports.
See per-panel results and methodology, machine-readable measurements, and the research story for controls, historical comparisons, and limitations.
These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.
| Model/configuration | Reported download size | Original-ALFA pass rate | Source |
|---|---|---|---|
| GPT-4o, cloud API (published reference) | — | 73.0% | whatisit benchmarks |
| nl2sh-3b Q4_K_M | 1.9 GB | 65.7% | whatisit benchmarks |
| Community nl2sh-qwen25-coder-1.5b Q4_K_M | 941 MB | 65.67% | Community model card |
| whatisit / nl2sh-1.5b Q4_K_M | 941 MB | 62.0% | whatisit benchmarks |
| Qwen2.5-Coder-7B, untuned | 4.4 GB | 61.3% | whatisit benchmarks |
| Qwen2.5-Coder-1.5B, untuned | 941 MB | 54.0% | whatisit benchmarks |
The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.
The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.
Requirements: Linux, Bash, a C/C++ compiler, CMake 3.20+, and network access for the first dependency download. Build dependencies have fixed versions and checksums. Python 3.11+ is needed for the installer and tests.
A prebuilt Linux x86_64 bundle is also available, with its checksum, license notices and exact CPU/system requirements. Build from source if your machine does not meet those requirements.
git clone https://github.com/dirac-run/ec.git
cd ec
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --parallel
ctest --test-dir build --output-on-failure
python3 scripts/install
The installer defaults to ~/.local/bin/ec and refuses an existing unrelated
command, including one elsewhere on PATH. If ec is already used, install
with python3 scripts/install --name easycommand. Direct CMake installation
also supports cmake --install build --prefix ~/.local, but does not perform
the installer's ownership/collision checks.
build/easycommand is a compatibility symlink to build/ec. For a build
optimized for your own CPU, add -DEC_NATIVE=ON; that binary may not run on
another CPU. Prebuilt binaries require their stated operating-system/CPU profile.
Use an EasyCommand GGUF based on Qwen3-0.6B or Qwen2.5-Coder-1.5B-Instruct. These are trained checkpoints; the untouched upstream models are not equivalent. Weights are kept separately from this source repository.
| Release | Model repository | Published quants |
|---|---|---|
| 0.6B, selected A1 checkpoint | ec-0.6b-gguf | Q4_K_M, Q8_0 |
| 1.5B, selected A3 checkpoint | ec-1.5b-gguf | Q4_K_M |
For the 1.5B release, download and verify the model from this checkout:
python3 scripts/download-model \
'https://huggingface.co/dirac-run/ec-1.5b-gguf/resolve/main/ec-1.5b.Q4_K_M.gguf' \
--sha256 135c8ec1a7ad6c24ba25971ee01cb3659130676031e45d876bd1822185c9e044
ec print the system uptime
ec --model /path/to/model.gguf print the system uptime
ec --model /path/to/model.gguf --preview list the files in this directory
Without --model, the default is $XDG_DATA_HOME/easycommand/model.gguf, or
~/.local/share/easycommand/model.gguf. To install a downloaded model, put it
there or create a symlink to it. scripts/download-model accepts an HTTPS model
URL and its published SHA256 checksum; use --help for its arguments.
The runtime uses the exact prompt in config/system-prompt.txt, greedy decoding, a 256-token output limit, and thinking disabled for Qwen3. It expects COMMAND JSON and extracts the command. No JSON generation grammar is applied. Unsupported architectures are rejected explicitly.
ec --help lists the options. --preview prints validated JSON and executes
nothing. Ordinary requests print the proposed command; Enter, y or Y
confirms execution, and other answers cancel it. Execution requires interactive
input and output terminals. Commands are checked for Bash syntax and explicitly
named tools before the prompt. These checks do not establish that a generated
command fulfills the request; review the command before accepting it.
For repeated requests, a resident worker can keep the model in memory:
ec --serve --socket /tmp/ec.sock --model /path/to/model.gguf --threads 8
# In another terminal:
ec --socket /tmp/ec.sock --expect-model /path/to/model.gguf print the system uptime
The worker resets inference state between requests and verifies the requested model identity. Remove its socket after the worker exits before starting another. No systemd service or machine-specific shell configuration is required.
See the training guide for dataset format, LoRA continuation and GGUF conversion, and measured results for checkpoint and precision comparisons. ALFA-updated is a separately documented benchmark variant; its scores are not interchangeable with original ALFA scores. The project post draft covers the data expansion, training experiments, negative results, release choices and future directions.
Code is MIT licensed. Dependency licenses and credits are in THIRD_PARTY.md. Model and dataset terms are supplied separately with their respective distributions.