Reproducible ALFA shell-command benchmark with documented environment and correctness repairs
Python
0
0 commits
updated Oct 5, 2026
ALFA-updated is a repaired execution environment and grader for the 300 public requests in InterCode-ALFA. It preserves task requirements and submitted model commands, while correcting documented transport, reset, fixture, reference and output-checking defects. Its scores must be identified as ALFA-updated, separately from original ALFA.
The complete change log lists every repair and its reason. REASONING.md explains the correctness boundaries, and known limitations states remaining weaknesses.
The pinned release assets support Linux x86_64, Docker and Python 3.11+. Docker must be running and accessible to your user. The embedding service runs on CPU. It is an embedding comparator, not a new generative LLM judge.
git clone https://github.com/dirac-run/ALFA-updated.git
cd ALFA-updated
python3 -m venv .venv
.venv/bin/pip install -r requirements.lock
.venv/bin/pip install --no-build-isolation .
.venv/bin/alfa-prepare
From a GitHub checkout, preparation derives the asset download URL from its
origin remote and the single release tag v0.1.0. It downloads only missing
assets, verifies SHA256 checksums, loads the exact five Docker image IDs, and
prepares the locked embedding runtime/model. It never rebuilds fixtures,
updates the lock, or pulls a mutable model tag.
The release downloads
contain all seven archives and their checksum file.
For a source archive or installed wheel without Git metadata, pass the asset
base URL with alfa-prepare --release-url URL, or download the release's seven
.tar files into one directory and use alfa-prepare --asset-dir /path/to/assets.
Cache files go under ~/.cache/alfa-updated/0.1.0; ALFA_CACHE_DIR can choose
another directory.
Start the embedding service in a separate terminal:
.venv/bin/alfa-embedding
This launches the exact locked Ollama runtime and mxbai-embed-large model on
localhost. It refuses occupied ports and does not replace other services.
For another port, use alfa-embedding --port 11435 and set
ALFA_EMBEDDING_URL=http://127.0.0.1:11435 when grading. This changes only the
service address, not its model or scoring parameters.
Generate a command for each task using your model's own prompt/inference setup. Submit each command without repairing or rewriting it:
from icalfa import submit_command, close
try:
score = submit_command(index=0, command="ls")
finally:
close()
The compatible import remains icalfa; install this package in a separate
environment from upstream icalfa. Task indices are zero-based, 0 through
299. Public task requests are in data/tasks.jsonl.
For a full batch, save exactly 300 unique rows in JSONL form:
{"index":0,"command":"ls"}
.venv/bin/alfa-grade --predictions commands.jsonl --output results/model-run
The output directory must be new. A run freezes commands and grader inputs, records per-task execution/checking evidence, and refuses changes to its definition during grading. Each fixture uses one owned container, resetting state between candidate/reference executions and subsequent tasks. The evaluator removes its containers and network on close.
Output scoring is a documented hybrid: supported task-information contracts
check correctness directly; remaining cases retain the locked embedding model,
the original 1000-character prefix and strict similarity threshold > 0.75.
Filesystem tasks check actual types, contents, permissions and links. No new
generative LLM judging is added.
Repairs are applied equally to all models. Candidate outputs, requested scope,
ordering and destructive effects are not relaxed to favor a model. Fixture and
reference changes are explicit in data/fixture_errata.jsonl
and data/reference_errata.jsonl. Positive and
deliberately incorrect controls are included in tests/.
Report passes / 300 together with model failures and benchmark errors.
Infrastructure errors have a null score and are disclosed; they do not justify
shrinking the denominator. Live host/network observations can vary, and a
passing embedding comparison is not proof of universal semantic correctness.
Report the model, quantization, prompt, decoding settings and benchmark
definition recorded in the run manifest.
| ID | What changed | Why |
|---|---|---|
| R01 | Exact argv transport to Bash | Preserve nested quotes and the actual generated program. |
| R02 | Normal compound-CD interpretation | Avoid the old interceptor rewriting valid commands. |
| R03 | Sequential reuse of one owned container per fixture | Identical candidate/reference initial environment; no per-question container pairs. |
| R04 | Verified PAX snapshot reset, including mtimes and ignored mutable roots | Prevent deletion/checkout timing, /tmp loss and residual files from contaminating later questions. |
| R05 | Real type/content/mode/owner/link filesystem checks | Avoid whitespace Git parsing, unchecked modifications, bad directory hashes and hardlink/symlink confusion. |
| R06 | Separate stdout/stderr/status and explicit benchmark errors | Error strings are not requested information or successful checksums. |
| R07 | Symmetric ten-second deadlines and process cleanup | Stop timed-out children affecting later tasks. |
| R08 | Task-information validators with positive/negative controls | Equivalent counts/layouts may pass; wrong values, missing interfaces and wrong execution-local PID/UID must fail. |
| R09 | AES-256-CBC plaintext validation for task 104 | Random salted ciphertext differs legitimately; incorrect plaintext/password or extra plaintext remains wrong. |
| R10 | Explicit reference errata for 176/182/210 | Fix nonempty-directory removal, rounded-size threshold, and owner/group permission confusion. |
| R11 | Uniform standard-tool additions and Bash/GNU fixture 5 | Support legitimate tools and upstream Bash-specific references; disclosed environment extension. |
| R12 | Dependency/data/image/embedding locks and local archives | Reproduction requires exact artifacts, not mutable latest names. |
| R13 | Retained public/internal query differences for 150/190/284 | Expose the actual requests models received rather than hide provenance differences. |
| R14 | Frozen per-run sources and inputs; no overwrite/resume patching | Repair only between runs and retain every earlier failure receipt. |
| R15 | Recorded relative-age fixtures; frozen calendar tasks only | Stable age predicates without falsifying live kernel boot/uptime facts. |
| R16 | Explicit task-local prerequisites for 6/11/52/60/62/63/64 | Missing targets and crontabs, or searching live /proc, did not provide valid independent tests. Sparse size-boundary files also reject incorrect thresholds. |
| R17 | Correct symlink references/destinations for 150/190/258 | Follow the public request's direction and link type. Hard links and reversed links must fail. |
| R18 | Task-local path-list and multiline sorting inputs for 248/284 | The original NUL-filled path list and single-line sorting source could not test the stated operation. |
| R19 | Explicit failure for unavailable live download resources | HTTP error pages are not text-file downloads. Reference 198 uses curl --fail/--location; no synthetic service or hosts substitution. Unavailable resources remain unscored. |
| R20 | Correct references for 139/189/294; complete archive checks | Prevent directory recursion from selecting out-of-scope files and prevent hashing a newly created output file itself. Verify complete archive contents, compression and requested metadata instead of incidental compressed bytes. |
| R21 | Diff/ping status handling and graceful deadline interrupt | diff status 1 means differences; ping status 1 means no replies. Correct diagnostics can succeed without pretending the address responded. |
| R22 | Physical/usable RAM, rotating DNS, ASCII-space, timing/hash/path checks | Check actual information. Preserve spaces and complete sets; physical RAM and usable RAM are distinct observations. |
| R23 | Requested timestamp preservation and system-file content checks | A copy that discards timestamps, or wrong system-file content, must fail even if filenames match. |
| R24 | Non-mutating inspection and access-time restoration | The checker changed find -amin results through its own reads. Restore access times and inspect with Linux O_NOATIME. |
| R25 | Reviewer-driven negative controls and task 264 reference correction | Retain archive-root metadata, bind embedding digest to its tag, check actual ping transmission, and reject incomplete/fabricated process reports. Docker's separate execution roots require a complete process-table tree for 264. |
| R26 | Crypto parameter verification, API freezing, and aborted-run summaries | Explicit KDF, salt and base64 options are verified by actual decryption. Opaque settings remain unscored. Both APIs freeze definitions; failures retain evidence and never produce numeric model rewards. |
| R27 | Discriminating age fixtures, observed process facts, immutable bootstrap | Both old/recent files must be handled correctly. Process names, arguments, parent relationships and stable completeness are checked; unsupported formats remain explicit. Ordinary setup loads only locked assets. |
| R28 | Explicit bounded hybrid scoring and scope safeguards | Correctness checks are documented changes to output equivalence, not an unchanged embedding-only judge. Headerless ps, numbered ip-link names and elapsed-only timing are valid layouts. DNS269 keeps the original comparator because RRset membership cannot establish first-answer/variable requirements. |
| R29 | Precise uptime duration parsing and disk-usage layouts | Weeks, scalar/one-column/full df reports preserve requested values. Uptime tolerance follows the displayed unit and observed interval, rather than accepting wrong exact-second values. |
| R30 | Complete ordinary IPv4 route records against captured kernel facts | Check destination/prefix/gateway/interface, printed metric and supported ip-route metadata. Independently observed resolver facts handle legitimate hostname labels; missing/extra/duplicate/wrong routes fail. Unsupported layouts remain explicit errors. |
| R31 | Complete difference and per-file MD5 records | Diff status1 is normal. Compare complete difference multisets and complete file/hash associations or unlabelled per-file hash counts where format is unspecified. No embedding prefix or aggregate checksum can substitute for omitted records. |
| R32 | Supply only task243’s named destination directory; normalize resolved link targets | The unchanged request requires links inside ~/newlinks. Original source list is unchanged; absolute and equivalent relative symlinks may pass, while wrong target/type/location and extra links fail. |
| R33 | Complete group/count/layout checks for 50/65/187/216/289 | Verify actual memberships, all requested counts, mountpoint and every disk-usage row. Optional filename labels and truthful report granularity do not change the answer. |
| R34 | Task129 DIRECTORY prerequisite and discriminating basenames | Supply the request's unset variable equally; reject case-sensitive-only or full-path vowel matches. The original aeiouy convention remains. |
| R35 | Correct content-only concatenation reference192 | Original -print contaminates the output with filenames. Verify complete source bytes in any whole-file order; reject added labels, omitted or repeated data. |
| R36 | Actual filesystem creation-time reference/check228 | Modification time is not birth time. Independently capture each restored inode's birth timestamp and permissions; unsupported platforms are explicit errors. |
| R37 | Preserve duplicate selected-path records and classify input mutation | Duplicate membership/path rows must fail. Candidate checksum input mutation is a model failure; missing independent baseline facts remain an infrastructure error. |
With prepared assets and the embedding service running:
.venv/bin/pytest -q -p no:cacheprovider tests
The lock pins task/errata/rule files, dependencies, prepared Docker images, embedding runtime and model. Docker build recipes are included for inspection; rebuilding with today's package repositories does not reproduce the locked image IDs. Use the accompanying verified archives for this release.
Based on westenfelder/InterCode-ALFA commit
2d3a69473a68569828ab0b4859073ef4a0ae482c, derived from Princeton NLP's InterCode.
Canonical requests come from westenfelder/NL2SH-ALFA, revision
a99cb5784cf5c2a42b1cc26c1903d9c3b35206ba.
The upstream MIT license is retained in LICENSE.txt.
Embedding/runtime archives retain their own license notices. The repaired
public API is submit_command or Evaluator; unsupported historical Gym and
external judging interfaces are not part of this package.
Reproducible ALFA shell-command benchmark with documented environment and correctness repairs
Python
0
0 commits
updated Oct 5, 2026
ALFA-updated is a repaired execution environment and grader for the 300 public requests in InterCode-ALFA. It preserves task requirements and submitted model commands, while correcting documented transport, reset, fixture, reference and output-checking defects. Its scores must be identified as ALFA-updated, separately from original ALFA.
The complete change log lists every repair and its reason. REASONING.md explains the correctness boundaries, and known limitations states remaining weaknesses.
The pinned release assets support Linux x86_64, Docker and Python 3.11+. Docker must be running and accessible to your user. The embedding service runs on CPU. It is an embedding comparator, not a new generative LLM judge.
git clone https://github.com/dirac-run/ALFA-updated.git
cd ALFA-updated
python3 -m venv .venv
.venv/bin/pip install -r requirements.lock
.venv/bin/pip install --no-build-isolation .
.venv/bin/alfa-prepare
From a GitHub checkout, preparation derives the asset download URL from its
origin remote and the single release tag v0.1.0. It downloads only missing
assets, verifies SHA256 checksums, loads the exact five Docker image IDs, and
prepares the locked embedding runtime/model. It never rebuilds fixtures,
updates the lock, or pulls a mutable model tag.
The release downloads
contain all seven archives and their checksum file.
For a source archive or installed wheel without Git metadata, pass the asset
base URL with alfa-prepare --release-url URL, or download the release's seven
.tar files into one directory and use alfa-prepare --asset-dir /path/to/assets.
Cache files go under ~/.cache/alfa-updated/0.1.0; ALFA_CACHE_DIR can choose
another directory.
Start the embedding service in a separate terminal:
.venv/bin/alfa-embedding
This launches the exact locked Ollama runtime and mxbai-embed-large model on
localhost. It refuses occupied ports and does not replace other services.
For another port, use alfa-embedding --port 11435 and set
ALFA_EMBEDDING_URL=http://127.0.0.1:11435 when grading. This changes only the
service address, not its model or scoring parameters.
Generate a command for each task using your model's own prompt/inference setup. Submit each command without repairing or rewriting it:
from icalfa import submit_command, close
try:
score = submit_command(index=0, command="ls")
finally:
close()
The compatible import remains icalfa; install this package in a separate
environment from upstream icalfa. Task indices are zero-based, 0 through
299. Public task requests are in data/tasks.jsonl.
For a full batch, save exactly 300 unique rows in JSONL form:
{"index":0,"command":"ls"}
.venv/bin/alfa-grade --predictions commands.jsonl --output results/model-run
The output directory must be new. A run freezes commands and grader inputs, records per-task execution/checking evidence, and refuses changes to its definition during grading. Each fixture uses one owned container, resetting state between candidate/reference executions and subsequent tasks. The evaluator removes its containers and network on close.
Output scoring is a documented hybrid: supported task-information contracts
check correctness directly; remaining cases retain the locked embedding model,
the original 1000-character prefix and strict similarity threshold > 0.75.
Filesystem tasks check actual types, contents, permissions and links. No new
generative LLM judging is added.
Repairs are applied equally to all models. Candidate outputs, requested scope,
ordering and destructive effects are not relaxed to favor a model. Fixture and
reference changes are explicit in data/fixture_errata.jsonl
and data/reference_errata.jsonl. Positive and
deliberately incorrect controls are included in tests/.
Report passes / 300 together with model failures and benchmark errors.
Infrastructure errors have a null score and are disclosed; they do not justify
shrinking the denominator. Live host/network observations can vary, and a
passing embedding comparison is not proof of universal semantic correctness.
Report the model, quantization, prompt, decoding settings and benchmark
definition recorded in the run manifest.
| ID | What changed | Why |
|---|---|---|
| R01 | Exact argv transport to Bash | Preserve nested quotes and the actual generated program. |
| R02 | Normal compound-CD interpretation | Avoid the old interceptor rewriting valid commands. |
| R03 | Sequential reuse of one owned container per fixture | Identical candidate/reference initial environment; no per-question container pairs. |
| R04 | Verified PAX snapshot reset, including mtimes and ignored mutable roots | Prevent deletion/checkout timing, /tmp loss and residual files from contaminating later questions. |
| R05 | Real type/content/mode/owner/link filesystem checks | Avoid whitespace Git parsing, unchecked modifications, bad directory hashes and hardlink/symlink confusion. |
| R06 | Separate stdout/stderr/status and explicit benchmark errors | Error strings are not requested information or successful checksums. |
| R07 | Symmetric ten-second deadlines and process cleanup | Stop timed-out children affecting later tasks. |
| R08 | Task-information validators with positive/negative controls | Equivalent counts/layouts may pass; wrong values, missing interfaces and wrong execution-local PID/UID must fail. |
| R09 | AES-256-CBC plaintext validation for task 104 | Random salted ciphertext differs legitimately; incorrect plaintext/password or extra plaintext remains wrong. |
| R10 | Explicit reference errata for 176/182/210 | Fix nonempty-directory removal, rounded-size threshold, and owner/group permission confusion. |
| R11 | Uniform standard-tool additions and Bash/GNU fixture 5 | Support legitimate tools and upstream Bash-specific references; disclosed environment extension. |
| R12 | Dependency/data/image/embedding locks and local archives | Reproduction requires exact artifacts, not mutable latest names. |
| R13 | Retained public/internal query differences for 150/190/284 | Expose the actual requests models received rather than hide provenance differences. |
| R14 | Frozen per-run sources and inputs; no overwrite/resume patching | Repair only between runs and retain every earlier failure receipt. |
| R15 | Recorded relative-age fixtures; frozen calendar tasks only | Stable age predicates without falsifying live kernel boot/uptime facts. |
| R16 | Explicit task-local prerequisites for 6/11/52/60/62/63/64 | Missing targets and crontabs, or searching live /proc, did not provide valid independent tests. Sparse size-boundary files also reject incorrect thresholds. |
| R17 | Correct symlink references/destinations for 150/190/258 | Follow the public request's direction and link type. Hard links and reversed links must fail. |
| R18 | Task-local path-list and multiline sorting inputs for 248/284 | The original NUL-filled path list and single-line sorting source could not test the stated operation. |
| R19 | Explicit failure for unavailable live download resources | HTTP error pages are not text-file downloads. Reference 198 uses curl --fail/--location; no synthetic service or hosts substitution. Unavailable resources remain unscored. |
| R20 | Correct references for 139/189/294; complete archive checks | Prevent directory recursion from selecting out-of-scope files and prevent hashing a newly created output file itself. Verify complete archive contents, compression and requested metadata instead of incidental compressed bytes. |
| R21 | Diff/ping status handling and graceful deadline interrupt | diff status 1 means differences; ping status 1 means no replies. Correct diagnostics can succeed without pretending the address responded. |
| R22 | Physical/usable RAM, rotating DNS, ASCII-space, timing/hash/path checks | Check actual information. Preserve spaces and complete sets; physical RAM and usable RAM are distinct observations. |
| R23 | Requested timestamp preservation and system-file content checks | A copy that discards timestamps, or wrong system-file content, must fail even if filenames match. |
| R24 | Non-mutating inspection and access-time restoration | The checker changed find -amin results through its own reads. Restore access times and inspect with Linux O_NOATIME. |
| R25 | Reviewer-driven negative controls and task 264 reference correction | Retain archive-root metadata, bind embedding digest to its tag, check actual ping transmission, and reject incomplete/fabricated process reports. Docker's separate execution roots require a complete process-table tree for 264. |
| R26 | Crypto parameter verification, API freezing, and aborted-run summaries | Explicit KDF, salt and base64 options are verified by actual decryption. Opaque settings remain unscored. Both APIs freeze definitions; failures retain evidence and never produce numeric model rewards. |
| R27 | Discriminating age fixtures, observed process facts, immutable bootstrap | Both old/recent files must be handled correctly. Process names, arguments, parent relationships and stable completeness are checked; unsupported formats remain explicit. Ordinary setup loads only locked assets. |
| R28 | Explicit bounded hybrid scoring and scope safeguards | Correctness checks are documented changes to output equivalence, not an unchanged embedding-only judge. Headerless ps, numbered ip-link names and elapsed-only timing are valid layouts. DNS269 keeps the original comparator because RRset membership cannot establish first-answer/variable requirements. |
| R29 | Precise uptime duration parsing and disk-usage layouts | Weeks, scalar/one-column/full df reports preserve requested values. Uptime tolerance follows the displayed unit and observed interval, rather than accepting wrong exact-second values. |
| R30 | Complete ordinary IPv4 route records against captured kernel facts | Check destination/prefix/gateway/interface, printed metric and supported ip-route metadata. Independently observed resolver facts handle legitimate hostname labels; missing/extra/duplicate/wrong routes fail. Unsupported layouts remain explicit errors. |
| R31 | Complete difference and per-file MD5 records | Diff status1 is normal. Compare complete difference multisets and complete file/hash associations or unlabelled per-file hash counts where format is unspecified. No embedding prefix or aggregate checksum can substitute for omitted records. |
| R32 | Supply only task243’s named destination directory; normalize resolved link targets | The unchanged request requires links inside ~/newlinks. Original source list is unchanged; absolute and equivalent relative symlinks may pass, while wrong target/type/location and extra links fail. |
| R33 | Complete group/count/layout checks for 50/65/187/216/289 | Verify actual memberships, all requested counts, mountpoint and every disk-usage row. Optional filename labels and truthful report granularity do not change the answer. |
| R34 | Task129 DIRECTORY prerequisite and discriminating basenames | Supply the request's unset variable equally; reject case-sensitive-only or full-path vowel matches. The original aeiouy convention remains. |
| R35 | Correct content-only concatenation reference192 | Original -print contaminates the output with filenames. Verify complete source bytes in any whole-file order; reject added labels, omitted or repeated data. |
| R36 | Actual filesystem creation-time reference/check228 | Modification time is not birth time. Independently capture each restored inode's birth timestamp and permissions; unsupported platforms are explicit errors. |
| R37 | Preserve duplicate selected-path records and classify input mutation | Duplicate membership/path rows must fail. Candidate checksum input mutation is a model failure; missing independent baseline facts remain an infrastructure error. |
With prepared assets and the embedding service running:
.venv/bin/pytest -q -p no:cacheprovider tests
The lock pins task/errata/rule files, dependencies, prepared Docker images, embedding runtime and model. Docker build recipes are included for inspection; rebuilding with today's package repositories does not reproduce the locked image IDs. Use the accompanying verified archives for this release.
Based on westenfelder/InterCode-ALFA commit
2d3a69473a68569828ab0b4859073ef4a0ae482c, derived from Princeton NLP's InterCode.
Canonical requests come from westenfelder/NL2SH-ALFA, revision
a99cb5784cf5c2a42b1cc26c1903d9c3b35206ba.
The upstream MIT license is retained in LICENSE.txt.
Embedding/runtime archives retain their own license notices. The repaired
public API is submit_command or Evaluator; unsupported historical Gym and
external judging interfaces are not part of this package.