dirac-run/ec

Generate GNU/Linux shell commands locally with small trained models and an embedded llama.cpp runtime

C++

0

0 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset (r/LocalLLaMA)

Purely a hobby side project to see how far I can push a really small model, using (mostly) automated training pipelines Full synthetic data: [https://huggingface.co/datasets/dirac-run/ec-training-data](https://huggingface.co/datasets/dirac-run/ec-training-data) Models:…

10

Oct 5, 2026

Show HN: I finetuned 1.5B Qwen to near GPT-4o level bash generation perf

1

Oct 5, 2026

README

EasyCommand (ec)

Turn an English request into a Bash command, entirely on your own machine. EasyCommand prints the command and asks before executing it.

Models and data · Research story · ALFA-updated benchmark

ec list the last five commits in this repository

The C++ application embeds llama.cpp. Inference needs a compatible GGUF model and ordinary Linux system libraries; it does not need Python, PyTorch, a GPU, Ollama or a network connection. Python is used only by optional setup, testing and training utilities.

Benchmark results

Locally measured results for the released EC models and an external reference:

ModelQuantALFA-updated /300Pass rateInternal /1320
EC 1.5B A3Q4_K_M212/30070.67%1119/1320
whatisit / nl2sh-1.5bQ4_K_M191/30063.67%82/1320
EC 0.6B A1Q8_0174/30058.00%922/1320
EC 0.6B A1Q4_K_M165/30055.00%911/1320

These use ALFA-updated, with documented environment and correctness repairs. They are separate from published original-ALFA scores. Benchmark errors remain in the 300-task denominator: one for each EC Q4 run, two for EC Q8, and five for whatisit.

Internal suite breakdown: ordinary commands 244, quoting 256, exact operands 512, time predicates 84, text-processing transfer 160, and English wording 64. Quoting, operands and English contribute 832/1320 (63.0%) closely related literal-search/output-mode stress tests. Cases share templates, and these panels guided EC development; the total is not broad shell accuracy on 1,320 independent unseen tasks. The whatisit internal result retains one observer error as a failure.

The comparison keeps each model's serving profile: EC uses COMMAND JSON and a 256-token limit; whatisit uses its native plain-command prompt and a 64-token limit. This compares deployed configurations with the same graders, rather than holding inference settings constant. No fresh full-suite result is assigned to the merged BF16 exports.

See per-panel results and methodology, machine-readable measurements, and the research story for controls, historical comparisons, and limitations.

Published original-ALFA results (different protocol)

These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.

Model/configurationReported download sizeOriginal-ALFA pass rateSource
GPT-4o, cloud API (published reference)—73.0%whatisit benchmarks
nl2sh-3b Q4_K_M1.9 GB65.7%whatisit benchmarks
Community nl2sh-qwen25-coder-1.5b Q4_K_M941 MB65.67%Community model card
whatisit / nl2sh-1.5b Q4_K_M941 MB62.0%whatisit benchmarks
Qwen2.5-Coder-7B, untuned4.4 GB61.3%whatisit benchmarks
Qwen2.5-Coder-1.5B, untuned941 MB54.0%whatisit benchmarks

The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.

The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.

Build and install

Requirements: Linux, Bash, a C/C++ compiler, CMake 3.20+, and network access for the first dependency download. Build dependencies have fixed versions and checksums. Python 3.11+ is needed for the installer and tests.

A prebuilt Linux x86_64 bundle is also available, with its checksum, license notices and exact CPU/system requirements. Build from source if your machine does not meet those requirements.

git clone https://github.com/dirac-run/ec.git
cd ec
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --parallel
ctest --test-dir build --output-on-failure
python3 scripts/install

The installer defaults to ~/.local/bin/ec and refuses an existing unrelated command, including one elsewhere on PATH. If ec is already used, install with python3 scripts/install --name easycommand. Direct CMake installation also supports cmake --install build --prefix ~/.local, but does not perform the installer's ownership/collision checks.

build/easycommand is a compatibility symlink to build/ec. For a build optimized for your own CPU, add -DEC_NATIVE=ON; that binary may not run on another CPU. Prebuilt binaries require their stated operating-system/CPU profile.

Choose a model

Use an EasyCommand GGUF based on Qwen3-0.6B or Qwen2.5-Coder-1.5B-Instruct. These are trained checkpoints; the untouched upstream models are not equivalent. Weights are kept separately from this source repository.

ReleaseModel repositoryPublished quants
0.6B, selected A1 checkpointec-0.6b-ggufQ4_K_M, Q8_0
1.5B, selected A3 checkpointec-1.5b-ggufQ4_K_M

For the 1.5B release, download and verify the model from this checkout:

python3 scripts/download-model \
  'https://huggingface.co/dirac-run/ec-1.5b-gguf/resolve/main/ec-1.5b.Q4_K_M.gguf' \
  --sha256 135c8ec1a7ad6c24ba25971ee01cb3659130676031e45d876bd1822185c9e044
ec print the system uptime
ec --model /path/to/model.gguf print the system uptime
ec --model /path/to/model.gguf --preview list the files in this directory

Without --model, the default is $XDG_DATA_HOME/easycommand/model.gguf, or ~/.local/share/easycommand/model.gguf. To install a downloaded model, put it there or create a symlink to it. scripts/download-model accepts an HTTPS model URL and its published SHA256 checksum; use --help for its arguments.

The runtime uses the exact prompt in config/system-prompt.txt, greedy decoding, a 256-token output limit, and thinking disabled for Qwen3. It expects COMMAND JSON and extracts the command. No JSON generation grammar is applied. Unsupported architectures are rejected explicitly.

Usage

ec --help lists the options. --preview prints validated JSON and executes nothing. Ordinary requests print the proposed command; Enter, y or Y confirms execution, and other answers cancel it. Execution requires interactive input and output terminals. Commands are checked for Bash syntax and explicitly named tools before the prompt. These checks do not establish that a generated command fulfills the request; review the command before accepting it.

For repeated requests, a resident worker can keep the model in memory:

ec --serve --socket /tmp/ec.sock --model /path/to/model.gguf --threads 8
# In another terminal:
ec --socket /tmp/ec.sock --expect-model /path/to/model.gguf print the system uptime

The worker resets inference state between requests and verifies the requested model identity. Remove its socket after the worker exits before starting another. No systemd service or machine-specific shell configuration is required.

Training and results

See the training guide for dataset format, LoRA continuation and GGUF conversion, and measured results for checkpoint and precision comparisons. ALFA-updated is a separately documented benchmark variant; its scores are not interchangeable with original ALFA scores. The project post draft covers the data expansion, training experiments, negative results, release choices and future directions.

Code is MIT licensed. Dependency licenses and credits are in THIRD_PARTY.md. Model and dataset terms are supplied separately with their respective distributions.

dirac-run/ec

Generate GNU/Linux shell commands locally with small trained models and an embedded llama.cpp runtime

C++

0

0 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset (r/LocalLLaMA)

Purely a hobby side project to see how far I can push a really small model, using (mostly) automated training pipelines Full synthetic data: [https://huggingface.co/datasets/dirac-run/ec-training-data](https://huggingface.co/datasets/dirac-run/ec-training-data) Models:…

10

Oct 5, 2026

Show HN: I finetuned 1.5B Qwen to near GPT-4o level bash generation perf

1

Oct 5, 2026

README

EasyCommand (ec)

Turn an English request into a Bash command, entirely on your own machine. EasyCommand prints the command and asks before executing it.

Models and data · Research story · ALFA-updated benchmark

ec list the last five commits in this repository

The C++ application embeds llama.cpp. Inference needs a compatible GGUF model and ordinary Linux system libraries; it does not need Python, PyTorch, a GPU, Ollama or a network connection. Python is used only by optional setup, testing and training utilities.

Benchmark results

Locally measured results for the released EC models and an external reference:

ModelQuantALFA-updated /300Pass rateInternal /1320
EC 1.5B A3Q4_K_M212/30070.67%1119/1320
whatisit / nl2sh-1.5bQ4_K_M191/30063.67%82/1320
EC 0.6B A1Q8_0174/30058.00%922/1320
EC 0.6B A1Q4_K_M165/30055.00%911/1320

These use ALFA-updated, with documented environment and correctness repairs. They are separate from published original-ALFA scores. Benchmark errors remain in the 300-task denominator: one for each EC Q4 run, two for EC Q8, and five for whatisit.

Internal suite breakdown: ordinary commands 244, quoting 256, exact operands 512, time predicates 84, text-processing transfer 160, and English wording 64. Quoting, operands and English contribute 832/1320 (63.0%) closely related literal-search/output-mode stress tests. Cases share templates, and these panels guided EC development; the total is not broad shell accuracy on 1,320 independent unseen tasks. The whatisit internal result retains one observer error as a failure.

The comparison keeps each model's serving profile: EC uses COMMAND JSON and a 256-token limit; whatisit uses its native plain-command prompt and a 64-token limit. This compares deployed configurations with the same graders, rather than holding inference settings constant. No fresh full-suite result is assigned to the merged BF16 exports.

See per-panel results and methodology, machine-readable measurements, and the research story for controls, historical comparisons, and limitations.

Published original-ALFA results (different protocol)

These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.

Model/configurationReported download sizeOriginal-ALFA pass rateSource
GPT-4o, cloud API (published reference)—73.0%whatisit benchmarks
nl2sh-3b Q4_K_M1.9 GB65.7%whatisit benchmarks
Community nl2sh-qwen25-coder-1.5b Q4_K_M941 MB65.67%Community model card
whatisit / nl2sh-1.5b Q4_K_M941 MB62.0%whatisit benchmarks
Qwen2.5-Coder-7B, untuned4.4 GB61.3%whatisit benchmarks
Qwen2.5-Coder-1.5B, untuned941 MB54.0%whatisit benchmarks

The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.

The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.

Build and install

Requirements: Linux, Bash, a C/C++ compiler, CMake 3.20+, and network access for the first dependency download. Build dependencies have fixed versions and checksums. Python 3.11+ is needed for the installer and tests.

A prebuilt Linux x86_64 bundle is also available, with its checksum, license notices and exact CPU/system requirements. Build from source if your machine does not meet those requirements.

git clone https://github.com/dirac-run/ec.git
cd ec
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --parallel
ctest --test-dir build --output-on-failure
python3 scripts/install

The installer defaults to ~/.local/bin/ec and refuses an existing unrelated command, including one elsewhere on PATH. If ec is already used, install with python3 scripts/install --name easycommand. Direct CMake installation also supports cmake --install build --prefix ~/.local, but does not perform the installer's ownership/collision checks.

build/easycommand is a compatibility symlink to build/ec. For a build optimized for your own CPU, add -DEC_NATIVE=ON; that binary may not run on another CPU. Prebuilt binaries require their stated operating-system/CPU profile.

Choose a model

Use an EasyCommand GGUF based on Qwen3-0.6B or Qwen2.5-Coder-1.5B-Instruct. These are trained checkpoints; the untouched upstream models are not equivalent. Weights are kept separately from this source repository.

ReleaseModel repositoryPublished quants
0.6B, selected A1 checkpointec-0.6b-ggufQ4_K_M, Q8_0
1.5B, selected A3 checkpointec-1.5b-ggufQ4_K_M

For the 1.5B release, download and verify the model from this checkout:

python3 scripts/download-model \
  'https://huggingface.co/dirac-run/ec-1.5b-gguf/resolve/main/ec-1.5b.Q4_K_M.gguf' \
  --sha256 135c8ec1a7ad6c24ba25971ee01cb3659130676031e45d876bd1822185c9e044
ec print the system uptime
ec --model /path/to/model.gguf print the system uptime
ec --model /path/to/model.gguf --preview list the files in this directory

Without --model, the default is $XDG_DATA_HOME/easycommand/model.gguf, or ~/.local/share/easycommand/model.gguf. To install a downloaded model, put it there or create a symlink to it. scripts/download-model accepts an HTTPS model URL and its published SHA256 checksum; use --help for its arguments.

The runtime uses the exact prompt in config/system-prompt.txt, greedy decoding, a 256-token output limit, and thinking disabled for Qwen3. It expects COMMAND JSON and extracts the command. No JSON generation grammar is applied. Unsupported architectures are rejected explicitly.

Usage

ec --help lists the options. --preview prints validated JSON and executes nothing. Ordinary requests print the proposed command; Enter, y or Y confirms execution, and other answers cancel it. Execution requires interactive input and output terminals. Commands are checked for Bash syntax and explicitly named tools before the prompt. These checks do not establish that a generated command fulfills the request; review the command before accepting it.

For repeated requests, a resident worker can keep the model in memory:

ec --serve --socket /tmp/ec.sock --model /path/to/model.gguf --threads 8
# In another terminal:
ec --socket /tmp/ec.sock --expect-model /path/to/model.gguf print the system uptime

The worker resets inference state between requests and verifies the requested model identity. Remove its socket after the worker exits before starting another. No systemd service or machine-specific shell configuration is required.

Training and results

See the training guide for dataset format, LoRA continuation and GGUF conversion, and measured results for checkpoint and precision comparisons. ALFA-updated is a separately documented benchmark variant; its scores are not interchangeable with original ALFA scores. The project post draft covers the data expansion, training experiments, negative results, release choices and future directions.

Code is MIT licensed. Dependency licenses and credits are in THIRD_PARTY.md. Model and dataset terms are supplied separately with their respective distributions.