dirac-run/ALFA-updated

Reproducible ALFA shell-command benchmark with documented environment and correctness repairs

Python

0

0 commits

updated Oct 5, 2026

See the code

README

ALFA-updated

ALFA-updated is a repaired execution environment and grader for the 300 public requests in InterCode-ALFA. It preserves task requirements and submitted model commands, while correcting documented transport, reset, fixture, reference and output-checking defects. Its scores must be identified as ALFA-updated, separately from original ALFA.

The complete change log lists every repair and its reason. REASONING.md explains the correctness boundaries, and known limitations states remaining weaknesses.

Install and prepare

The pinned release assets support Linux x86_64, Docker and Python 3.11+. Docker must be running and accessible to your user. The embedding service runs on CPU. It is an embedding comparator, not a new generative LLM judge.

git clone https://github.com/dirac-run/ALFA-updated.git
cd ALFA-updated
python3 -m venv .venv
.venv/bin/pip install -r requirements.lock
.venv/bin/pip install --no-build-isolation .
.venv/bin/alfa-prepare

From a GitHub checkout, preparation derives the asset download URL from its origin remote and the single release tag v0.1.0. It downloads only missing assets, verifies SHA256 checksums, loads the exact five Docker image IDs, and prepares the locked embedding runtime/model. It never rebuilds fixtures, updates the lock, or pulls a mutable model tag. The release downloads contain all seven archives and their checksum file.

For a source archive or installed wheel without Git metadata, pass the asset base URL with alfa-prepare --release-url URL, or download the release's seven .tar files into one directory and use alfa-prepare --asset-dir /path/to/assets. Cache files go under ~/.cache/alfa-updated/0.1.0; ALFA_CACHE_DIR can choose another directory.

Start the embedding service in a separate terminal:

.venv/bin/alfa-embedding

This launches the exact locked Ollama runtime and mxbai-embed-large model on localhost. It refuses occupied ports and does not replace other services. For another port, use alfa-embedding --port 11435 and set ALFA_EMBEDDING_URL=http://127.0.0.1:11435 when grading. This changes only the service address, not its model or scoring parameters.

Grade a model

Generate a command for each task using your model's own prompt/inference setup. Submit each command without repairing or rewriting it:

from icalfa import submit_command, close

try:
    score = submit_command(index=0, command="ls")
finally:
    close()

The compatible import remains icalfa; install this package in a separate environment from upstream icalfa. Task indices are zero-based, 0 through 299. Public task requests are in data/tasks.jsonl.

For a full batch, save exactly 300 unique rows in JSONL form:

{"index":0,"command":"ls"}
.venv/bin/alfa-grade --predictions commands.jsonl --output results/model-run

The output directory must be new. A run freezes commands and grader inputs, records per-task execution/checking evidence, and refuses changes to its definition during grading. Each fixture uses one owned container, resetting state between candidate/reference executions and subsequent tasks. The evaluator removes its containers and network on close.

Scoring and fairness

Output scoring is a documented hybrid: supported task-information contracts check correctness directly; remaining cases retain the locked embedding model, the original 1000-character prefix and strict similarity threshold > 0.75. Filesystem tasks check actual types, contents, permissions and links. No new generative LLM judging is added.

Repairs are applied equally to all models. Candidate outputs, requested scope, ordering and destructive effects are not relaxed to favor a model. Fixture and reference changes are explicit in data/fixture_errata.jsonl and data/reference_errata.jsonl. Positive and deliberately incorrect controls are included in tests/.

Report passes / 300 together with model failures and benchmark errors. Infrastructure errors have a null score and are disclosed; they do not justify shrinking the denominator. Live host/network observations can vary, and a passing embedding comparison is not proof of universal semantic correctness. Report the model, quantization, prompt, decoding settings and benchmark definition recorded in the run manifest.

Complete repair inventory

IDWhat changedWhy
R01Exact argv transport to BashPreserve nested quotes and the actual generated program.
R02Normal compound-CD interpretationAvoid the old interceptor rewriting valid commands.
R03Sequential reuse of one owned container per fixtureIdentical candidate/reference initial environment; no per-question container pairs.
R04Verified PAX snapshot reset, including mtimes and ignored mutable rootsPrevent deletion/checkout timing, /tmp loss and residual files from contaminating later questions.
R05Real type/content/mode/owner/link filesystem checksAvoid whitespace Git parsing, unchecked modifications, bad directory hashes and hardlink/symlink confusion.
R06Separate stdout/stderr/status and explicit benchmark errorsError strings are not requested information or successful checksums.
R07Symmetric ten-second deadlines and process cleanupStop timed-out children affecting later tasks.
R08Task-information validators with positive/negative controlsEquivalent counts/layouts may pass; wrong values, missing interfaces and wrong execution-local PID/UID must fail.
R09AES-256-CBC plaintext validation for task 104Random salted ciphertext differs legitimately; incorrect plaintext/password or extra plaintext remains wrong.
R10Explicit reference errata for 176/182/210Fix nonempty-directory removal, rounded-size threshold, and owner/group permission confusion.
R11Uniform standard-tool additions and Bash/GNU fixture 5Support legitimate tools and upstream Bash-specific references; disclosed environment extension.
R12Dependency/data/image/embedding locks and local archivesReproduction requires exact artifacts, not mutable latest names.
R13Retained public/internal query differences for 150/190/284Expose the actual requests models received rather than hide provenance differences.
R14Frozen per-run sources and inputs; no overwrite/resume patchingRepair only between runs and retain every earlier failure receipt.
R15Recorded relative-age fixtures; frozen calendar tasks onlyStable age predicates without falsifying live kernel boot/uptime facts.
R16Explicit task-local prerequisites for 6/11/52/60/62/63/64Missing targets and crontabs, or searching live /proc, did not provide valid independent tests. Sparse size-boundary files also reject incorrect thresholds.
R17Correct symlink references/destinations for 150/190/258Follow the public request's direction and link type. Hard links and reversed links must fail.
R18Task-local path-list and multiline sorting inputs for 248/284The original NUL-filled path list and single-line sorting source could not test the stated operation.
R19Explicit failure for unavailable live download resourcesHTTP error pages are not text-file downloads. Reference 198 uses curl --fail/--location; no synthetic service or hosts substitution. Unavailable resources remain unscored.
R20Correct references for 139/189/294; complete archive checksPrevent directory recursion from selecting out-of-scope files and prevent hashing a newly created output file itself. Verify complete archive contents, compression and requested metadata instead of incidental compressed bytes.
R21Diff/ping status handling and graceful deadline interruptdiff status 1 means differences; ping status 1 means no replies. Correct diagnostics can succeed without pretending the address responded.
R22Physical/usable RAM, rotating DNS, ASCII-space, timing/hash/path checksCheck actual information. Preserve spaces and complete sets; physical RAM and usable RAM are distinct observations.
R23Requested timestamp preservation and system-file content checksA copy that discards timestamps, or wrong system-file content, must fail even if filenames match.
R24Non-mutating inspection and access-time restorationThe checker changed find -amin results through its own reads. Restore access times and inspect with Linux O_NOATIME.
R25Reviewer-driven negative controls and task 264 reference correctionRetain archive-root metadata, bind embedding digest to its tag, check actual ping transmission, and reject incomplete/fabricated process reports. Docker's separate execution roots require a complete process-table tree for 264.
R26Crypto parameter verification, API freezing, and aborted-run summariesExplicit KDF, salt and base64 options are verified by actual decryption. Opaque settings remain unscored. Both APIs freeze definitions; failures retain evidence and never produce numeric model rewards.
R27Discriminating age fixtures, observed process facts, immutable bootstrapBoth old/recent files must be handled correctly. Process names, arguments, parent relationships and stable completeness are checked; unsupported formats remain explicit. Ordinary setup loads only locked assets.
R28Explicit bounded hybrid scoring and scope safeguardsCorrectness checks are documented changes to output equivalence, not an unchanged embedding-only judge. Headerless ps, numbered ip-link names and elapsed-only timing are valid layouts. DNS269 keeps the original comparator because RRset membership cannot establish first-answer/variable requirements.
R29Precise uptime duration parsing and disk-usage layoutsWeeks, scalar/one-column/full df reports preserve requested values. Uptime tolerance follows the displayed unit and observed interval, rather than accepting wrong exact-second values.
R30Complete ordinary IPv4 route records against captured kernel factsCheck destination/prefix/gateway/interface, printed metric and supported ip-route metadata. Independently observed resolver facts handle legitimate hostname labels; missing/extra/duplicate/wrong routes fail. Unsupported layouts remain explicit errors.
R31Complete difference and per-file MD5 recordsDiff status1 is normal. Compare complete difference multisets and complete file/hash associations or unlabelled per-file hash counts where format is unspecified. No embedding prefix or aggregate checksum can substitute for omitted records.
R32Supply only task243’s named destination directory; normalize resolved link targetsThe unchanged request requires links inside ~/newlinks. Original source list is unchanged; absolute and equivalent relative symlinks may pass, while wrong target/type/location and extra links fail.
R33Complete group/count/layout checks for 50/65/187/216/289Verify actual memberships, all requested counts, mountpoint and every disk-usage row. Optional filename labels and truthful report granularity do not change the answer.
R34Task129 DIRECTORY prerequisite and discriminating basenamesSupply the request's unset variable equally; reject case-sensitive-only or full-path vowel matches. The original aeiouy convention remains.
R35Correct content-only concatenation reference192Original -print contaminates the output with filenames. Verify complete source bytes in any whole-file order; reject added labels, omitted or repeated data.
R36Actual filesystem creation-time reference/check228Modification time is not birth time. Independently capture each restored inode's birth timestamp and permissions; unsupported platforms are explicit errors.
R37Preserve duplicate selected-path records and classify input mutationDuplicate membership/path rows must fail. Candidate checksum input mutation is a model failure; missing independent baseline facts remain an infrastructure error.

Verify the implementation

With prepared assets and the embedding service running:

.venv/bin/pytest -q -p no:cacheprovider tests

The lock pins task/errata/rule files, dependencies, prepared Docker images, embedding runtime and model. Docker build recipes are included for inspection; rebuilding with today's package repositories does not reproduce the locked image IDs. Use the accompanying verified archives for this release.

Credit and license

Based on westenfelder/InterCode-ALFA commit 2d3a69473a68569828ab0b4859073ef4a0ae482c, derived from Princeton NLP's InterCode. Canonical requests come from westenfelder/NL2SH-ALFA, revision a99cb5784cf5c2a42b1cc26c1903d9c3b35206ba.

The upstream MIT license is retained in LICENSE.txt. Embedding/runtime archives retain their own license notices. The repaired public API is submit_command or Evaluator; unsupported historical Gym and external judging interfaces are not part of this package.

dirac-run/ALFA-updated

Reproducible ALFA shell-command benchmark with documented environment and correctness repairs

Python

0

0 commits

updated Oct 5, 2026

See the code

README

ALFA-updated

ALFA-updated is a repaired execution environment and grader for the 300 public requests in InterCode-ALFA. It preserves task requirements and submitted model commands, while correcting documented transport, reset, fixture, reference and output-checking defects. Its scores must be identified as ALFA-updated, separately from original ALFA.

The complete change log lists every repair and its reason. REASONING.md explains the correctness boundaries, and known limitations states remaining weaknesses.

Install and prepare

The pinned release assets support Linux x86_64, Docker and Python 3.11+. Docker must be running and accessible to your user. The embedding service runs on CPU. It is an embedding comparator, not a new generative LLM judge.

git clone https://github.com/dirac-run/ALFA-updated.git
cd ALFA-updated
python3 -m venv .venv
.venv/bin/pip install -r requirements.lock
.venv/bin/pip install --no-build-isolation .
.venv/bin/alfa-prepare

From a GitHub checkout, preparation derives the asset download URL from its origin remote and the single release tag v0.1.0. It downloads only missing assets, verifies SHA256 checksums, loads the exact five Docker image IDs, and prepares the locked embedding runtime/model. It never rebuilds fixtures, updates the lock, or pulls a mutable model tag. The release downloads contain all seven archives and their checksum file.

For a source archive or installed wheel without Git metadata, pass the asset base URL with alfa-prepare --release-url URL, or download the release's seven .tar files into one directory and use alfa-prepare --asset-dir /path/to/assets. Cache files go under ~/.cache/alfa-updated/0.1.0; ALFA_CACHE_DIR can choose another directory.

Start the embedding service in a separate terminal:

.venv/bin/alfa-embedding

This launches the exact locked Ollama runtime and mxbai-embed-large model on localhost. It refuses occupied ports and does not replace other services. For another port, use alfa-embedding --port 11435 and set ALFA_EMBEDDING_URL=http://127.0.0.1:11435 when grading. This changes only the service address, not its model or scoring parameters.

Grade a model

Generate a command for each task using your model's own prompt/inference setup. Submit each command without repairing or rewriting it:

from icalfa import submit_command, close

try:
    score = submit_command(index=0, command="ls")
finally:
    close()

The compatible import remains icalfa; install this package in a separate environment from upstream icalfa. Task indices are zero-based, 0 through 299. Public task requests are in data/tasks.jsonl.

For a full batch, save exactly 300 unique rows in JSONL form:

{"index":0,"command":"ls"}
.venv/bin/alfa-grade --predictions commands.jsonl --output results/model-run

The output directory must be new. A run freezes commands and grader inputs, records per-task execution/checking evidence, and refuses changes to its definition during grading. Each fixture uses one owned container, resetting state between candidate/reference executions and subsequent tasks. The evaluator removes its containers and network on close.

Scoring and fairness

Output scoring is a documented hybrid: supported task-information contracts check correctness directly; remaining cases retain the locked embedding model, the original 1000-character prefix and strict similarity threshold > 0.75. Filesystem tasks check actual types, contents, permissions and links. No new generative LLM judging is added.

Repairs are applied equally to all models. Candidate outputs, requested scope, ordering and destructive effects are not relaxed to favor a model. Fixture and reference changes are explicit in data/fixture_errata.jsonl and data/reference_errata.jsonl. Positive and deliberately incorrect controls are included in tests/.

Report passes / 300 together with model failures and benchmark errors. Infrastructure errors have a null score and are disclosed; they do not justify shrinking the denominator. Live host/network observations can vary, and a passing embedding comparison is not proof of universal semantic correctness. Report the model, quantization, prompt, decoding settings and benchmark definition recorded in the run manifest.

Complete repair inventory

IDWhat changedWhy
R01Exact argv transport to BashPreserve nested quotes and the actual generated program.
R02Normal compound-CD interpretationAvoid the old interceptor rewriting valid commands.
R03Sequential reuse of one owned container per fixtureIdentical candidate/reference initial environment; no per-question container pairs.
R04Verified PAX snapshot reset, including mtimes and ignored mutable rootsPrevent deletion/checkout timing, /tmp loss and residual files from contaminating later questions.
R05Real type/content/mode/owner/link filesystem checksAvoid whitespace Git parsing, unchecked modifications, bad directory hashes and hardlink/symlink confusion.
R06Separate stdout/stderr/status and explicit benchmark errorsError strings are not requested information or successful checksums.
R07Symmetric ten-second deadlines and process cleanupStop timed-out children affecting later tasks.
R08Task-information validators with positive/negative controlsEquivalent counts/layouts may pass; wrong values, missing interfaces and wrong execution-local PID/UID must fail.
R09AES-256-CBC plaintext validation for task 104Random salted ciphertext differs legitimately; incorrect plaintext/password or extra plaintext remains wrong.
R10Explicit reference errata for 176/182/210Fix nonempty-directory removal, rounded-size threshold, and owner/group permission confusion.
R11Uniform standard-tool additions and Bash/GNU fixture 5Support legitimate tools and upstream Bash-specific references; disclosed environment extension.
R12Dependency/data/image/embedding locks and local archivesReproduction requires exact artifacts, not mutable latest names.
R13Retained public/internal query differences for 150/190/284Expose the actual requests models received rather than hide provenance differences.
R14Frozen per-run sources and inputs; no overwrite/resume patchingRepair only between runs and retain every earlier failure receipt.
R15Recorded relative-age fixtures; frozen calendar tasks onlyStable age predicates without falsifying live kernel boot/uptime facts.
R16Explicit task-local prerequisites for 6/11/52/60/62/63/64Missing targets and crontabs, or searching live /proc, did not provide valid independent tests. Sparse size-boundary files also reject incorrect thresholds.
R17Correct symlink references/destinations for 150/190/258Follow the public request's direction and link type. Hard links and reversed links must fail.
R18Task-local path-list and multiline sorting inputs for 248/284The original NUL-filled path list and single-line sorting source could not test the stated operation.
R19Explicit failure for unavailable live download resourcesHTTP error pages are not text-file downloads. Reference 198 uses curl --fail/--location; no synthetic service or hosts substitution. Unavailable resources remain unscored.
R20Correct references for 139/189/294; complete archive checksPrevent directory recursion from selecting out-of-scope files and prevent hashing a newly created output file itself. Verify complete archive contents, compression and requested metadata instead of incidental compressed bytes.
R21Diff/ping status handling and graceful deadline interruptdiff status 1 means differences; ping status 1 means no replies. Correct diagnostics can succeed without pretending the address responded.
R22Physical/usable RAM, rotating DNS, ASCII-space, timing/hash/path checksCheck actual information. Preserve spaces and complete sets; physical RAM and usable RAM are distinct observations.
R23Requested timestamp preservation and system-file content checksA copy that discards timestamps, or wrong system-file content, must fail even if filenames match.
R24Non-mutating inspection and access-time restorationThe checker changed find -amin results through its own reads. Restore access times and inspect with Linux O_NOATIME.
R25Reviewer-driven negative controls and task 264 reference correctionRetain archive-root metadata, bind embedding digest to its tag, check actual ping transmission, and reject incomplete/fabricated process reports. Docker's separate execution roots require a complete process-table tree for 264.
R26Crypto parameter verification, API freezing, and aborted-run summariesExplicit KDF, salt and base64 options are verified by actual decryption. Opaque settings remain unscored. Both APIs freeze definitions; failures retain evidence and never produce numeric model rewards.
R27Discriminating age fixtures, observed process facts, immutable bootstrapBoth old/recent files must be handled correctly. Process names, arguments, parent relationships and stable completeness are checked; unsupported formats remain explicit. Ordinary setup loads only locked assets.
R28Explicit bounded hybrid scoring and scope safeguardsCorrectness checks are documented changes to output equivalence, not an unchanged embedding-only judge. Headerless ps, numbered ip-link names and elapsed-only timing are valid layouts. DNS269 keeps the original comparator because RRset membership cannot establish first-answer/variable requirements.
R29Precise uptime duration parsing and disk-usage layoutsWeeks, scalar/one-column/full df reports preserve requested values. Uptime tolerance follows the displayed unit and observed interval, rather than accepting wrong exact-second values.
R30Complete ordinary IPv4 route records against captured kernel factsCheck destination/prefix/gateway/interface, printed metric and supported ip-route metadata. Independently observed resolver facts handle legitimate hostname labels; missing/extra/duplicate/wrong routes fail. Unsupported layouts remain explicit errors.
R31Complete difference and per-file MD5 recordsDiff status1 is normal. Compare complete difference multisets and complete file/hash associations or unlabelled per-file hash counts where format is unspecified. No embedding prefix or aggregate checksum can substitute for omitted records.
R32Supply only task243’s named destination directory; normalize resolved link targetsThe unchanged request requires links inside ~/newlinks. Original source list is unchanged; absolute and equivalent relative symlinks may pass, while wrong target/type/location and extra links fail.
R33Complete group/count/layout checks for 50/65/187/216/289Verify actual memberships, all requested counts, mountpoint and every disk-usage row. Optional filename labels and truthful report granularity do not change the answer.
R34Task129 DIRECTORY prerequisite and discriminating basenamesSupply the request's unset variable equally; reject case-sensitive-only or full-path vowel matches. The original aeiouy convention remains.
R35Correct content-only concatenation reference192Original -print contaminates the output with filenames. Verify complete source bytes in any whole-file order; reject added labels, omitted or repeated data.
R36Actual filesystem creation-time reference/check228Modification time is not birth time. Independently capture each restored inode's birth timestamp and permissions; unsupported platforms are explicit errors.
R37Preserve duplicate selected-path records and classify input mutationDuplicate membership/path rows must fail. Candidate checksum input mutation is a model failure; missing independent baseline facts remain an infrastructure error.

Verify the implementation

With prepared assets and the embedding service running:

.venv/bin/pytest -q -p no:cacheprovider tests

The lock pins task/errata/rule files, dependencies, prepared Docker images, embedding runtime and model. Docker build recipes are included for inspection; rebuilding with today's package repositories does not reproduce the locked image IDs. Use the accompanying verified archives for this release.

Credit and license

Based on westenfelder/InterCode-ALFA commit 2d3a69473a68569828ab0b4859073ef4a0ae482c, derived from Princeton NLP's InterCode. Canonical requests come from westenfelder/NL2SH-ALFA, revision a99cb5784cf5c2a42b1cc26c1903d9c3b35206ba.

The upstream MIT license is retained in LICENSE.txt. Embedding/runtime archives retain their own license notices. The repaired public API is submit_command or Evaluator; unsupported historical Gym and external judging interfaces are not part of this package.