dirac-run/ec-0.6b-gguf

Model

EasyCommand 0.6B

0

4 commits

3 linked in READMEs

updated Oct 5, 2026

See the code

README

EasyCommand 0.6B

A compact command generator, released in Q4_K_M and Q8_0. This is a trained EasyCommand checkpoint derived from Qwen/Qwen3-0.6B at revision c1899de289a04d12100db370d81485cdf75e47ca. It generates GNU/Linux Bash commands from English requests as {"kind":"COMMAND","value":"<command>"}.

All GGUFs are complete merged models with tokenizer and chat-template metadata. They require no separate adapter or upstream weight download for inference. The selected research checkpoint is A1; all quants here share that trained state.

Downloads and measured results

FileQuantSizeALFA-updated /300Internal /1320
ec-0.6b.Q4_K_M.ggufQ4_K_M461.8 MiB165/300911/1320
ec-0.6b.Q8_0.ggufQ8_0767.5 MiB174/300922/1320

A1 Q8 scored 174/300 ALFA and 922/1320 internal; Q4 scored 165/300 and 911/1320. These are separate measurements of two quants of the same checkpoint.

QuantCommands /244Quoting /256Operands /512Time /84Transfers /160English /64
Q4_K_M217202284727264
Q8_0219217272737764

ALFA scores use ALFA-updated, a documented variant with grading definition SHA256 6f50e4fea37c125f73f93e5f785e95e91fda448b4bf347a9a2986db4d83b48dc. They are not original ALFA scores and should not be compared directly with published results from another protocol. The Q4 run had one benchmark error; Q8 had two. Errors remain in the 300-task denominator. The internal panels contain related templates and were consulted during development; 1320 cases are not 1320 independent unseen tasks. ALFA also informed repair strategy and checkpoint selection. These are development measurements. Only the published quants' measured results are shown; they do not imply BF16 performance. evaluation.json contains the per-panel counts and source-summary hashes.

Published original-ALFA results (different protocol)

These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.

Model/configurationReported download sizeOriginal-ALFA pass rateSource
GPT-4o, cloud API (published reference)—73.0%whatisit benchmarks
nl2sh-3b Q4_K_M1.9 GB65.7%whatisit benchmarks
Community nl2sh-qwen25-coder-1.5b Q4_K_M941 MB65.67%Community model card
whatisit / nl2sh-1.5b Q4_K_M941 MB62.0%whatisit benchmarks
Qwen2.5-Coder-7B, untuned4.4 GB61.3%whatisit benchmarks
Qwen2.5-Coder-1.5B, untuned941 MB54.0%whatisit benchmarks

The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.

The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.

Use with ec

Install the EasyCommand application, then download the Q4 model from that checkout using its checksum-verifying helper:

python3 scripts/download-model \
  'https://huggingface.co/dirac-run/ec-0.6b-gguf/resolve/main/ec-0.6b.Q4_K_M.gguf' \
  --sha256 614925fb3df0e457ff74dfc52e8f5b86bb5aeeb6213f71de6cb2e2decfabc025 \
  --output ec-0.6b.Q4_K_M.gguf
ec --model ec-0.6b.Q4_K_M.gguf --preview print the system uptime
ec --model ec-0.6b.Q4_K_M.gguf list the last five commits in this repository

--preview prints COMMAND JSON without executing. Ordinary ec usage extracts the command, displays it and asks for confirmation before running it. The download paths point to the published model files in this repository.

Exact inference profile

Use the exact system message in system-prompt.txt:

You are a GNU/Linux shell command generator. Produce the simplest Bash command that fulfills the entire request. Return only valid JSON: {"kind":"COMMAND","value":"<command>"}.

It is 177 bytes, SHA256 3a9028d5aebb73c3ed7363e63eb3e751689218775ea5572ab0dc6807e522238b. Evaluations used greedy decoding, no generation grammar, a 256-token output limit, EOS 151645 and the native CPU decoder based on llama.cpp revision 1af554f8fc78ba029665a47b839484d9763e2a75. Use a system message and one user message, then the assistant generation prefix. For Qwen3, disable thinking (enable_thinking=False). The evaluated assistant prefix includes <think>\n\n</think>\n\n before the generated JSON. The ec application applies this automatically. See inference.json for the profile. Other templates, prompts, runtime builds or CPU kernels can change outputs and should be evaluated separately.

Training

The model started from the pinned upstream weights, received one weighted epoch of LoRA training (474,635 presentations; 14,833 updates), and then 200 incremental repair/replay updates from that trained parent. This is not a stock model and not a fresh 200-step fine-tune from stock weights.

Both stages used rank 32, alpha 64, dropout 0.05, effective batch 32 and adapters on q/k/v/o attention projections and gate/up/down MLP projections. The base stayed BF16 with FP32 trainable adapters. AdamW used linear warmup and cosine decay. The parent peak learning rate was 0.0002; the continuation used 5e-06, ten warmup updates and 50% repair / 50% replay sampling from 2,564 source rows. Only assistant answer/EOS tokens were supervised. The parent epoch ran on H100 and the continuation on A40.

The parent used the older prompt in training-system-prompt.txt. Continuation training and reported release evaluations used the shorter serving prompt. training.json records the actual recipe and lineage.

The released dataset has 401,975 deduplicated request/answer pairs, including original and concise descriptions, corrected/simplified commands and incremental repair data. It also includes rows from other repair experiments. A single pass over it does not reproduce these models' weighted exposure or sampling history. Merge used the original base and trained adapter in FP32, followed by F16 GGUF conversion and direct quantization; no requantization or importance matrix was used.

Scope and limitations

Designed for English requests targeting Bash and GNU/Linux utilities. The model does not inspect the live filesystem or know which tools are installed. It gives a best-effort command rather than clarification or inability responses. Generated commands can be wrong, incomplete or destructive; review them before execution.

This repository contains GGUF inference exports. The trainable BF16 model and original LoRA adapter are released separately for further fine-tuning. They identify the same trained checkpoint; numeric formats and kernels can yield different outputs. To train from stock weights, use the released dataset and record your own exposure and validation.

License and integrity

Weights and documentation are released under Apache-2.0, retaining the upstream attribution in NOTICE. The application has its own license. Verify downloads from this directory with sha256sum --check SHA256SUMS. manifest.json lists the weight files and hashes.

bash
command-generation
conversational
easycommand
endpoints_compatible
gguf
lora
shell
text-generation

dirac-run/ec-0.6b-gguf

Model

EasyCommand 0.6B

0

4 commits

3 linked in READMEs

updated Oct 5, 2026

See the code

README

EasyCommand 0.6B

A compact command generator, released in Q4_K_M and Q8_0. This is a trained EasyCommand checkpoint derived from Qwen/Qwen3-0.6B at revision c1899de289a04d12100db370d81485cdf75e47ca. It generates GNU/Linux Bash commands from English requests as {"kind":"COMMAND","value":"<command>"}.

All GGUFs are complete merged models with tokenizer and chat-template metadata. They require no separate adapter or upstream weight download for inference. The selected research checkpoint is A1; all quants here share that trained state.

Downloads and measured results

FileQuantSizeALFA-updated /300Internal /1320
ec-0.6b.Q4_K_M.ggufQ4_K_M461.8 MiB165/300911/1320
ec-0.6b.Q8_0.ggufQ8_0767.5 MiB174/300922/1320

A1 Q8 scored 174/300 ALFA and 922/1320 internal; Q4 scored 165/300 and 911/1320. These are separate measurements of two quants of the same checkpoint.

QuantCommands /244Quoting /256Operands /512Time /84Transfers /160English /64
Q4_K_M217202284727264
Q8_0219217272737764

ALFA scores use ALFA-updated, a documented variant with grading definition SHA256 6f50e4fea37c125f73f93e5f785e95e91fda448b4bf347a9a2986db4d83b48dc. They are not original ALFA scores and should not be compared directly with published results from another protocol. The Q4 run had one benchmark error; Q8 had two. Errors remain in the 300-task denominator. The internal panels contain related templates and were consulted during development; 1320 cases are not 1320 independent unseen tasks. ALFA also informed repair strategy and checkpoint selection. These are development measurements. Only the published quants' measured results are shown; they do not imply BF16 performance. evaluation.json contains the per-panel counts and source-summary hashes.

Published original-ALFA results (different protocol)

These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.

Model/configurationReported download sizeOriginal-ALFA pass rateSource
GPT-4o, cloud API (published reference)—73.0%whatisit benchmarks
nl2sh-3b Q4_K_M1.9 GB65.7%whatisit benchmarks
Community nl2sh-qwen25-coder-1.5b Q4_K_M941 MB65.67%Community model card
whatisit / nl2sh-1.5b Q4_K_M941 MB62.0%whatisit benchmarks
Qwen2.5-Coder-7B, untuned4.4 GB61.3%whatisit benchmarks
Qwen2.5-Coder-1.5B, untuned941 MB54.0%whatisit benchmarks

The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.

The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.

Use with ec

Install the EasyCommand application, then download the Q4 model from that checkout using its checksum-verifying helper:

python3 scripts/download-model \
  'https://huggingface.co/dirac-run/ec-0.6b-gguf/resolve/main/ec-0.6b.Q4_K_M.gguf' \
  --sha256 614925fb3df0e457ff74dfc52e8f5b86bb5aeeb6213f71de6cb2e2decfabc025 \
  --output ec-0.6b.Q4_K_M.gguf
ec --model ec-0.6b.Q4_K_M.gguf --preview print the system uptime
ec --model ec-0.6b.Q4_K_M.gguf list the last five commits in this repository

--preview prints COMMAND JSON without executing. Ordinary ec usage extracts the command, displays it and asks for confirmation before running it. The download paths point to the published model files in this repository.

Exact inference profile

Use the exact system message in system-prompt.txt:

You are a GNU/Linux shell command generator. Produce the simplest Bash command that fulfills the entire request. Return only valid JSON: {"kind":"COMMAND","value":"<command>"}.

It is 177 bytes, SHA256 3a9028d5aebb73c3ed7363e63eb3e751689218775ea5572ab0dc6807e522238b. Evaluations used greedy decoding, no generation grammar, a 256-token output limit, EOS 151645 and the native CPU decoder based on llama.cpp revision 1af554f8fc78ba029665a47b839484d9763e2a75. Use a system message and one user message, then the assistant generation prefix. For Qwen3, disable thinking (enable_thinking=False). The evaluated assistant prefix includes <think>\n\n</think>\n\n before the generated JSON. The ec application applies this automatically. See inference.json for the profile. Other templates, prompts, runtime builds or CPU kernels can change outputs and should be evaluated separately.

Training

The model started from the pinned upstream weights, received one weighted epoch of LoRA training (474,635 presentations; 14,833 updates), and then 200 incremental repair/replay updates from that trained parent. This is not a stock model and not a fresh 200-step fine-tune from stock weights.

Both stages used rank 32, alpha 64, dropout 0.05, effective batch 32 and adapters on q/k/v/o attention projections and gate/up/down MLP projections. The base stayed BF16 with FP32 trainable adapters. AdamW used linear warmup and cosine decay. The parent peak learning rate was 0.0002; the continuation used 5e-06, ten warmup updates and 50% repair / 50% replay sampling from 2,564 source rows. Only assistant answer/EOS tokens were supervised. The parent epoch ran on H100 and the continuation on A40.

The parent used the older prompt in training-system-prompt.txt. Continuation training and reported release evaluations used the shorter serving prompt. training.json records the actual recipe and lineage.

The released dataset has 401,975 deduplicated request/answer pairs, including original and concise descriptions, corrected/simplified commands and incremental repair data. It also includes rows from other repair experiments. A single pass over it does not reproduce these models' weighted exposure or sampling history. Merge used the original base and trained adapter in FP32, followed by F16 GGUF conversion and direct quantization; no requantization or importance matrix was used.

Scope and limitations

Designed for English requests targeting Bash and GNU/Linux utilities. The model does not inspect the live filesystem or know which tools are installed. It gives a best-effort command rather than clarification or inability responses. Generated commands can be wrong, incomplete or destructive; review them before execution.

This repository contains GGUF inference exports. The trainable BF16 model and original LoRA adapter are released separately for further fine-tuning. They identify the same trained checkpoint; numeric formats and kernels can yield different outputs. To train from stock weights, use the released dataset and record your own exposure and validation.

License and integrity

Weights and documentation are released under Apache-2.0, retaining the upstream attribution in NOTICE. The application has its own license. Verify downloads from this directory with sha256sum --check SHA256SUMS. manifest.json lists the weight files and hashes.

bash
command-generation
conversational
easycommand
endpoints_compatible
gguf
lora
shell
text-generation