AFAW: AI-assisted development framework for deterministic multi-agent orchestration. Enforces hard guardrails and AST validation on AI code to block hallucinations. Manages concurrent context via JSON state-machines and offloads testing to parallel CI/CD pipelines. Real systems engineering, zero probabilistic fluff.
See the codeA deterministic, multi-agent methodology for AI-assisted software development.
Several coding agents working on one repository fail in predictable ways: they collide on branches and on shared context files, claim results they never measured, and write tests that pass whatever the code does. AFAW is a repository-level method and boilerplate that lets agents write code in parallel while deterministic tools decide whether the result is acceptable and a human approves every merge.
AFAW is the method; the code here is one way to implement it. See what is core and what you can replace.
White paper (PDF): docs/paper/afaw.pdf · LaTeX source: docs/paper/afaw.tex · DOI 10.5281/zenodo.23050310
Live demo: agustindiazcano.github.io/afaw-anti-fragile-agentic-workflow · this repository's own dashboard: /live/
A dashboard built by scripts/build_state.py from mock data: a fictional project whose every task and number is invented (examples/mock_dashboard.py). It shows each traffic-light rule firing: CI failing, a blocker not done, no activity for three days, nothing measured.



Build it locally: python -m examples.mock_dashboard mock-output and open mock-output/index.html.
It costs less. Without it, an agent runs the whole test suite on your machine, the machine is slow, tests fail for reasons that have nothing to do with the code, and the agent retries for an hour. Every retry sends its whole context again: that is where the token bill goes. With AFAW the agent runs only the test it is writing; the cloud runs the rest in parallel and answers once. A few minutes of CI cost far less than an hour of an agent going in circles.
It works alone and in a team. Alone, you let the agents run and nobody checks them: they commit to the wrong branch, write tests that pass no matter what, and report results they never measured. The checks, the branch protection and the hooks do that reviewing for you.
With 2 or 10 people, everyone sees what the agents did. Each person runs their own agents, and without a shared record nobody knows what the others' agents changed, why, or where each task stands. Here every task, decision and status is in the same place and in the same format, for every person and every agent: the board, the generated LASTCONTEXT.md and the decision records. You set it up once, from the template.
The project documents itself. Every task leaves what was done, what is left and, for each decision, why it was taken. When agents do the work, that is usually lost: afterwards nobody knows what was done or why.
A new agent understands the project in minutes. Instead of reading every file from scratch, it reads one generated index (LASTCONTEXT.md) and follows links to the decisions it needs. Less reading means fewer tokens and less time, and the saving grows with the project.
Fewer status meetings. "What did you do, where does it stand, what's blocked, why was this decided": the board and the records already answer it, and an optional assistant can answer it in a chat. See Ask the project.
These are the expected gains, not measured ones yet. The dashboard already records timings and CI cost, and the paper's evaluation protocol says how to measure all four.
The method does not require these exact tools. Write your tests however you like, use any CI, skip mutation testing outside critical code. Some parts, though, are what make it work: remove them and the problem they solve comes back.
| Core: remove it and the problem returns | Replaceable: same idea, another tool | Optional |
|---|---|---|
| Each task writes its own state files; shared views are derived. (One shared context file breaks as soon as two agents work at once.) | The CI system, the test framework | The dashboard and its charts |
| Measured values, never written by the agent | The mutation engine | Mutation testing outside critical code |
| Heavy checks in the cloud, not on the agent's machine | Hooks, for another agent or tool | Timings and CI cost |
One branch per task, no push to main, a human approves | The format of decision records | Per-role context files |
| The red step checked by a tool, not reported by the agent | The language of the rules file | |
| The retrieval stack of the assistant | An assistant that answers from the records (ask the project) |
Because everything is recorded as the work happens, you can put an assistant with RAG on top: it searches the project's records and answers questions in a chat. It is not part of the method, just an optional use of what the method already records.
What it replaces: status meetings. "What did you do, where does it stand, what's blocked, why was this decided." The board, LASTCONTEXT.md and the decision records already answer that. The assistant answers the same questions any time, to anyone, without taking an hour from five people.
Why it works well here. Every document is about one thing: one decision per ADR, one trap per gotcha, one delta per task. They split into clean pieces. The front matter says whether a decision is still in force, so the assistant can skip the ones that were replaced. LASTCONTEXT.md is the entry point. And every fact has a path, so each answer can cite where it came from and anyone can check it.
What it doesn't replace.
Two rules for the assistant.
A model proposes and writes code; its own assessment is never trusted. Every number (tests, types, coverage, mutation score, task timings, the traffic light) comes from a deterministic tool, and a value an agent could have made up is rejected where it would be declared.
| Failure | Mechanism | What remains |
|---|---|---|
| Wrong-branch work | Step 0 (git status, git branch), one branch per task, one terminal per agent; a hook denies a push to main | Agents sharing one working tree can still overwrite uncommitted files |
| Shared-state conflicts | Each task writes only its own files; shared views are derived on demand and never committed | Agents do not see unmerged work on other branches |
| Unmeasured claims | Measured values are rejected in task files; files_touched must equal the diff; the red step is verified in CI | A person must still follow the CI link |
| Vacuous tests | Red-first check; mutation testing on the diff; equivalent mutants accepted only by the human | Mutation score is not correctness |
| Heavy local runs | Two tiers: single tests locally, everything else in CI, in parallel | CI queue latency and minutes |
| Unreviewed integration | Hooks, CI checks, branch protection, human approval | Hooks are best-effort |
Colors in all diagrams: blue = agent action, amber = human, green = CI or deterministic check, purple = state files, red = blocked.
An agent writes only two files: its task file state/tasks/task_NNN.json (declared values: title, role, type, status, owner, priority, difficulty, dependencies, branch) and its delta context/tasks/task_NNN_context.json (summary, next, files_touched). Their shape is defined once, in state/schemas/.
Everything shared is derived, never committed: python -m scripts.build_state --ref origin/main builds the pending view, the history of done tasks, the metrics, the dashboard and the LASTCONTEXT.md files into build/. Timestamps and CI results are measured from git and the CI API (python -m scripts.collect_facts), never declared; the traffic light is a rule:
stale_daysState says what; prose says why. Decisions are ADRs in docs/adr/, operational traps are gotchas in docs/gotchas/ (templates in docs/templates/). The generated LASTCONTEXT.md indexes them with one line each, next to the current state, what waits on the human and the next tasks. A record leaves the index only for an explicit reason (superseded, deprecated, resolved, or promoted to a mechanism), never for its age, and a budget tells the human when to prune. See ADR 0001.
Locally, the agent runs only the test it is writing. CI runs everything else, including the red-first check (python -m scripts.red_check): each test a pull request adds must fail on the code before the change.
Only new or modified files are mutated. In critical modules a file below the minimum blocks the pull request. The gate is K / (N − E_human): only the human records an equivalent mutant, in state/equivalent_mutants.json, bound to the file's hash (python -m scripts.mutation_gate). The mutation engine itself is project-specific.
No commits or pushes to main, no merge without an explicit human directive, and the human approves every pull request. Rules in AGENTS.md are advice, so they are also enforced by hooks, CI and branch protection.
A static HTML page (no scripts) with progress, tasks by status, tasks merged per day, the board with ages and lights, timing (agent time, review latency, lead time) and CI cost (runs, queue time, job minutes). .github/workflows/dashboard.yml rebuilds it on every push to main, every hour and on demand, and publishes it on GitHub Pages: the demo at the site root and this repository's own dashboard at /live/. A Pages site is public: do not publish the dashboard of a private project.
.
├── AGENTS.md # The rules: single source for every agent
├── CLAUDE.md # Imports AGENTS.md
├── PENDING.md # The human's roadmap (hook: never deleted or emptied)
├── .claude/
│ ├── settings.json # Hook wiring
│ ├── hooks/ # session_start (builds and injects LASTCONTEXT.md),
│ │ # safety_guard (deny/ask), lint_check
│ └── skills/ # commit, ship, tests, push-dev, trash
├── .github/
│ ├── pull_request_template.md
│ └── workflows/
│ ├── checks.yml # Every PR: task files, delta, red-first, docs, views
│ ├── scripts-ci.yml # Lint, types, tests (skipped on documents-only PRs)
│ └── dashboard.yml # Builds and publishes the dashboard
├── state/
│ ├── config.json # Code folders, mutation settings, dashboard settings
│ ├── schemas/ # JSON Schemas: task, delta, config, equivalent mutants
│ └── tasks/task_NNN.json # One file per task (never deleted)
├── context/tasks/ # One delta per task
├── docs/
│ ├── adr/ # Architecture decision records
│ ├── gotchas/ # Operational traps, closed as resolved or promoted
│ ├── templates/ # ADR and gotcha templates
│ ├── diagrams/ and img/ # Graphviz sources and rendered figures
│ ├── screenshots/ # Screenshots of the demo dashboard
│ └── paper/ # The white paper: afaw.pdf and its source afaw.tex
├── scripts/
│ ├── afaw_state/ # Model, facts, lights, durations, views, dashboard
│ ├── build_state.py # Derived views from a git ref
│ ├── collect_facts.py # Measured facts from git and the GitHub API
│ ├── check_task_files.py # Lifecycle, declared values only, unique ids
│ ├── check_task_delta.py # files_touched equals the diff (--write fills it)
│ ├── check_docs.py # ADRs, gotchas, links
│ ├── red_check.py # Added tests must fail on the old code
│ ├── mutation_targets.py # Mandatory and optional files to mutate
│ ├── mutation_gate.py # Score with human-accepted equivalences only
│ ├── rules_to_tex.py # The paper's appendix from AGENTS.md
│ └── setup_protection.sh # Branch protection and required checks
├── examples/mock_dashboard.py # The demo: dashboard and LASTCONTEXT.md from mock data
├── src/ # The project's code
└── tests/tooling/ # Tests of the scripts and hooks
build/ (the derived views) is ignored by git.
GitHub does not copy branch protection, required checks or Pages settings to a repository created from a template.
state/config.json: code folders, mutation_critical_paths, mutation_min_score, and the dashboard settings.main: bash scripts/setup_protection.sh (needs the gh CLI) requires lint, typecheck, tests and state-and-docs, and blocks direct pushes.ROLE: question.ROLE:. The session-start hook injects the generated LASTCONTEXT.md and PENDING.md.summary and next, fills files_touched with python -m scripts.check_task_delta --base origin/main --write task_NNN, records any decision as an ADR, and asks before opening the pull request.pytest, mypy --strict, Postgres, Terraform). Adapt sections 4–10 and 14–16 of AGENTS.md for another stack.for f in docs/diagrams/*.dot; do n=$(basename "$f" .dot); dot -Tsvg "$f" -o "docs/img/$n.svg"; dot -Tpng -Gdpi=200 "$f" -o "docs/img/$n.png"; done
The paper: see docs/paper/README.md. A pull request that changes afaw.tex, AGENTS.md or docs/img/ must also commit the rebuilt docs/paper/afaw.pdf; CI checks it.
Agustin Diaz-Cano, MS Candidate, Information Systems Engineering (UTN). ORCID 0009-0001-4336-490X.
MIT
Python
99.1%
AFAW: AI-assisted development framework for deterministic multi-agent orchestration. Enforces hard guardrails and AST validation on AI code to block hallucinations. Manages concurrent context via JSON state-machines and offloads testing to parallel CI/CD pipelines. Real systems engineering, zero probabilistic fluff.
See the codeA deterministic, multi-agent methodology for AI-assisted software development.
Several coding agents working on one repository fail in predictable ways: they collide on branches and on shared context files, claim results they never measured, and write tests that pass whatever the code does. AFAW is a repository-level method and boilerplate that lets agents write code in parallel while deterministic tools decide whether the result is acceptable and a human approves every merge.
AFAW is the method; the code here is one way to implement it. See what is core and what you can replace.
White paper (PDF): docs/paper/afaw.pdf · LaTeX source: docs/paper/afaw.tex · DOI 10.5281/zenodo.23050310
Live demo: agustindiazcano.github.io/afaw-anti-fragile-agentic-workflow · this repository's own dashboard: /live/
A dashboard built by scripts/build_state.py from mock data: a fictional project whose every task and number is invented (examples/mock_dashboard.py). It shows each traffic-light rule firing: CI failing, a blocker not done, no activity for three days, nothing measured.



Build it locally: python -m examples.mock_dashboard mock-output and open mock-output/index.html.
It costs less. Without it, an agent runs the whole test suite on your machine, the machine is slow, tests fail for reasons that have nothing to do with the code, and the agent retries for an hour. Every retry sends its whole context again: that is where the token bill goes. With AFAW the agent runs only the test it is writing; the cloud runs the rest in parallel and answers once. A few minutes of CI cost far less than an hour of an agent going in circles.
It works alone and in a team. Alone, you let the agents run and nobody checks them: they commit to the wrong branch, write tests that pass no matter what, and report results they never measured. The checks, the branch protection and the hooks do that reviewing for you.
With 2 or 10 people, everyone sees what the agents did. Each person runs their own agents, and without a shared record nobody knows what the others' agents changed, why, or where each task stands. Here every task, decision and status is in the same place and in the same format, for every person and every agent: the board, the generated LASTCONTEXT.md and the decision records. You set it up once, from the template.
The project documents itself. Every task leaves what was done, what is left and, for each decision, why it was taken. When agents do the work, that is usually lost: afterwards nobody knows what was done or why.
A new agent understands the project in minutes. Instead of reading every file from scratch, it reads one generated index (LASTCONTEXT.md) and follows links to the decisions it needs. Less reading means fewer tokens and less time, and the saving grows with the project.
Fewer status meetings. "What did you do, where does it stand, what's blocked, why was this decided": the board and the records already answer it, and an optional assistant can answer it in a chat. See Ask the project.
These are the expected gains, not measured ones yet. The dashboard already records timings and CI cost, and the paper's evaluation protocol says how to measure all four.
The method does not require these exact tools. Write your tests however you like, use any CI, skip mutation testing outside critical code. Some parts, though, are what make it work: remove them and the problem they solve comes back.
| Core: remove it and the problem returns | Replaceable: same idea, another tool | Optional |
|---|---|---|
| Each task writes its own state files; shared views are derived. (One shared context file breaks as soon as two agents work at once.) | The CI system, the test framework | The dashboard and its charts |
| Measured values, never written by the agent | The mutation engine | Mutation testing outside critical code |
| Heavy checks in the cloud, not on the agent's machine | Hooks, for another agent or tool | Timings and CI cost |
One branch per task, no push to main, a human approves | The format of decision records | Per-role context files |
| The red step checked by a tool, not reported by the agent | The language of the rules file | |
| The retrieval stack of the assistant | An assistant that answers from the records (ask the project) |
Because everything is recorded as the work happens, you can put an assistant with RAG on top: it searches the project's records and answers questions in a chat. It is not part of the method, just an optional use of what the method already records.
What it replaces: status meetings. "What did you do, where does it stand, what's blocked, why was this decided." The board, LASTCONTEXT.md and the decision records already answer that. The assistant answers the same questions any time, to anyone, without taking an hour from five people.
Why it works well here. Every document is about one thing: one decision per ADR, one trap per gotcha, one delta per task. They split into clean pieces. The front matter says whether a decision is still in force, so the assistant can skip the ones that were replaced. LASTCONTEXT.md is the entry point. And every fact has a path, so each answer can cite where it came from and anyone can check it.
What it doesn't replace.
Two rules for the assistant.
A model proposes and writes code; its own assessment is never trusted. Every number (tests, types, coverage, mutation score, task timings, the traffic light) comes from a deterministic tool, and a value an agent could have made up is rejected where it would be declared.
| Failure | Mechanism | What remains |
|---|---|---|
| Wrong-branch work | Step 0 (git status, git branch), one branch per task, one terminal per agent; a hook denies a push to main | Agents sharing one working tree can still overwrite uncommitted files |
| Shared-state conflicts | Each task writes only its own files; shared views are derived on demand and never committed | Agents do not see unmerged work on other branches |
| Unmeasured claims | Measured values are rejected in task files; files_touched must equal the diff; the red step is verified in CI | A person must still follow the CI link |
| Vacuous tests | Red-first check; mutation testing on the diff; equivalent mutants accepted only by the human | Mutation score is not correctness |
| Heavy local runs | Two tiers: single tests locally, everything else in CI, in parallel | CI queue latency and minutes |
| Unreviewed integration | Hooks, CI checks, branch protection, human approval | Hooks are best-effort |
Colors in all diagrams: blue = agent action, amber = human, green = CI or deterministic check, purple = state files, red = blocked.
An agent writes only two files: its task file state/tasks/task_NNN.json (declared values: title, role, type, status, owner, priority, difficulty, dependencies, branch) and its delta context/tasks/task_NNN_context.json (summary, next, files_touched). Their shape is defined once, in state/schemas/.
Everything shared is derived, never committed: python -m scripts.build_state --ref origin/main builds the pending view, the history of done tasks, the metrics, the dashboard and the LASTCONTEXT.md files into build/. Timestamps and CI results are measured from git and the CI API (python -m scripts.collect_facts), never declared; the traffic light is a rule:
stale_daysState says what; prose says why. Decisions are ADRs in docs/adr/, operational traps are gotchas in docs/gotchas/ (templates in docs/templates/). The generated LASTCONTEXT.md indexes them with one line each, next to the current state, what waits on the human and the next tasks. A record leaves the index only for an explicit reason (superseded, deprecated, resolved, or promoted to a mechanism), never for its age, and a budget tells the human when to prune. See ADR 0001.
Locally, the agent runs only the test it is writing. CI runs everything else, including the red-first check (python -m scripts.red_check): each test a pull request adds must fail on the code before the change.
Only new or modified files are mutated. In critical modules a file below the minimum blocks the pull request. The gate is K / (N − E_human): only the human records an equivalent mutant, in state/equivalent_mutants.json, bound to the file's hash (python -m scripts.mutation_gate). The mutation engine itself is project-specific.
No commits or pushes to main, no merge without an explicit human directive, and the human approves every pull request. Rules in AGENTS.md are advice, so they are also enforced by hooks, CI and branch protection.
A static HTML page (no scripts) with progress, tasks by status, tasks merged per day, the board with ages and lights, timing (agent time, review latency, lead time) and CI cost (runs, queue time, job minutes). .github/workflows/dashboard.yml rebuilds it on every push to main, every hour and on demand, and publishes it on GitHub Pages: the demo at the site root and this repository's own dashboard at /live/. A Pages site is public: do not publish the dashboard of a private project.
.
├── AGENTS.md # The rules: single source for every agent
├── CLAUDE.md # Imports AGENTS.md
├── PENDING.md # The human's roadmap (hook: never deleted or emptied)
├── .claude/
│ ├── settings.json # Hook wiring
│ ├── hooks/ # session_start (builds and injects LASTCONTEXT.md),
│ │ # safety_guard (deny/ask), lint_check
│ └── skills/ # commit, ship, tests, push-dev, trash
├── .github/
│ ├── pull_request_template.md
│ └── workflows/
│ ├── checks.yml # Every PR: task files, delta, red-first, docs, views
│ ├── scripts-ci.yml # Lint, types, tests (skipped on documents-only PRs)
│ └── dashboard.yml # Builds and publishes the dashboard
├── state/
│ ├── config.json # Code folders, mutation settings, dashboard settings
│ ├── schemas/ # JSON Schemas: task, delta, config, equivalent mutants
│ └── tasks/task_NNN.json # One file per task (never deleted)
├── context/tasks/ # One delta per task
├── docs/
│ ├── adr/ # Architecture decision records
│ ├── gotchas/ # Operational traps, closed as resolved or promoted
│ ├── templates/ # ADR and gotcha templates
│ ├── diagrams/ and img/ # Graphviz sources and rendered figures
│ ├── screenshots/ # Screenshots of the demo dashboard
│ └── paper/ # The white paper: afaw.pdf and its source afaw.tex
├── scripts/
│ ├── afaw_state/ # Model, facts, lights, durations, views, dashboard
│ ├── build_state.py # Derived views from a git ref
│ ├── collect_facts.py # Measured facts from git and the GitHub API
│ ├── check_task_files.py # Lifecycle, declared values only, unique ids
│ ├── check_task_delta.py # files_touched equals the diff (--write fills it)
│ ├── check_docs.py # ADRs, gotchas, links
│ ├── red_check.py # Added tests must fail on the old code
│ ├── mutation_targets.py # Mandatory and optional files to mutate
│ ├── mutation_gate.py # Score with human-accepted equivalences only
│ ├── rules_to_tex.py # The paper's appendix from AGENTS.md
│ └── setup_protection.sh # Branch protection and required checks
├── examples/mock_dashboard.py # The demo: dashboard and LASTCONTEXT.md from mock data
├── src/ # The project's code
└── tests/tooling/ # Tests of the scripts and hooks
build/ (the derived views) is ignored by git.
GitHub does not copy branch protection, required checks or Pages settings to a repository created from a template.
state/config.json: code folders, mutation_critical_paths, mutation_min_score, and the dashboard settings.main: bash scripts/setup_protection.sh (needs the gh CLI) requires lint, typecheck, tests and state-and-docs, and blocks direct pushes.ROLE: question.ROLE:. The session-start hook injects the generated LASTCONTEXT.md and PENDING.md.summary and next, fills files_touched with python -m scripts.check_task_delta --base origin/main --write task_NNN, records any decision as an ADR, and asks before opening the pull request.pytest, mypy --strict, Postgres, Terraform). Adapt sections 4–10 and 14–16 of AGENTS.md for another stack.for f in docs/diagrams/*.dot; do n=$(basename "$f" .dot); dot -Tsvg "$f" -o "docs/img/$n.svg"; dot -Tpng -Gdpi=200 "$f" -o "docs/img/$n.png"; done
The paper: see docs/paper/README.md. A pull request that changes afaw.tex, AGENTS.md or docs/img/ must also commit the rebuilt docs/paper/afaw.pdf; CI checks it.
Agustin Diaz-Cano, MS Candidate, Information Systems Engineering (UTN). ORCID 0009-0001-4336-490X.
MIT
Python
99.1%