agentic-hil/stm32-starter

Nucleo-F446RE starter for Agentic Hardware-in-the-Loop

Python

1

41 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Claude Code writing firmware, flashing a real board and checking the result (r/SideProject)

The demo for Agentic HIL shows Claude Code writing a small firmware program, flashing it to a board and reading “Hello World” back over the serial connection. It then runs a deliberately wrong check so you can see the failure report too. What I like about this is that the hardware response becomes…

0

Oct 1, 2026

README

STM32 Agentic HIL Starter

Take your own copy of this repository, say one sentence to your coding agent, and watch it build firmware, flash a real Nucleo-F446RE, talk to it, and prove what the board actually did. The three test plans are the gate the agent has to pass.

Click Use this template above to get your own copy, or clone this one. The hardware run needs a Nucleo-F446RE, a USB cable, and the onboard ST-LINK, and nothing more: no fixture to wire, no adapter to buy. The three test plans and the suite that validates them need no board at all, so that half runs on any machine and is the first thing to run whether or not a Nucleo is on your desk. This is the reference path for Agentic HIL, the local MCP server that lets an AI agent develop firmware on hardware you own: build, flash, stimulate, observe, diagnose, fix, with the run on the board deciding whether the work is done and the reactor's report as the evidence you read. A Nucleo-F446RE is what this starter proves that path on, not what the path needs: any board reached through an ST-LINK runs the same way, with that board's own target script or controller name and its own UART named in the configuration instead of this one's.

The firmware ships with one deliberate defect in its diagnostic protocol. Finding it is the exercise.

Install Agentic HIL

Linux and macOS, in any shell:

curl -LsSf https://agentic-hil.github.io/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"

Windows, in PowerShell:

irm https://agentic-hil.github.io/install.ps1 | iex

Windows, from cmd.exe or the Run box:

powershell -c "irm https://agentic-hil.github.io/install.ps1|iex"

One line installs the package user-local and registers the agent skill and the MCP server for every agent CLI it finds on your PATH. No admin rights are required, and it writes nothing inside this repository. Then restart your agent once.

When ~/.local/bin is not already on your PATH, the installer puts it there for the shells that come after it, by one line in your shell profile, and names the file it edited. The export line in the first block is for the shell you are in right now, and for any shell that does not read that file; it puts uv within reach as well, because both land in that directory.

To register a single agent instead of all of them, pass --agent claude-code (or codex, opencode); piped, that reads | sh -s -- --agent claude-code.

Run the plans with no board attached

The check-plan suite validates the three hardware test plans on any machine, with nothing plugged in. It needs Python 3.10 or newer and uv, which is what the lock file is for. The one-line installer above leaves you with uv on most machines: it installs through uv when it finds one, and fetches uv itself when the Python it found has no pip or is one the distribution keeps for its own packages, so only a Python with a usable pip and no uv on the machine ends up without it. If command -v uv finds nothing, install it with curl -LsSf https://astral.sh/uv/install.sh | sh on Linux and macOS, or irm https://astral.sh/uv/install.ps1 | iex in PowerShell.

Clone this repository first, then run these inside the clone:

uv sync
uv run pytest -q -s

Every test states what a green run is worth:

PASS  configuration and test semantics validated without a board
NEEDS PHYSICAL FIXTURE  electrical behavior not verified

The environment uv sync builds carries the Agentic HIL release the lock file pins, and that release backs this suite. The hardware commands further down run the agentic-hil the installer put on your PATH, which may be newer. Two versions on one machine is the expected shape, not a mistake; agentic-hil doctor reports the one on your PATH.

Without uv, the same suite runs from a plain virtual environment:

python -m venv .venv
.venv/bin/pip install "agentic-hil>=0.21.0" pytest pyyaml   # .venv\Scripts\pip on Windows
.venv/bin/pytest -q -s   # .venv\Scripts\pytest on Windows

What runs where is the account of what a green suite establishes and what it leaves to the bench.

What you need for the hardware run

  • A Nucleo-F446RE connected through its ST-LINK USB port
  • CMake 3.21 or newer, which is what CMakePresets.json needs for its preset version and its toolchainFile, plus Ninja and the GNU Arm Embedded Toolchain (arm-none-eabi-gcc) on your PATH; STM32CubeCLT carries all three
  • A debugger backend, STM32CubeProgrammer CLI or OpenOCD; either one is enough on its own, and step 2 says what each of them means for discovery
  • Python 3.10 or newer
  • On Linux, membership in two groups before the first plan: dialout, which owns the ST-LINK's virtual COM port, and plugdev, which the openocd package's udev rule gives the probe itself. sudo usermod -aG plugdev,dialout "$USER" adds both, and it is the one step in this path that asks for an administrator. Log in again afterwards, because a shell takes its groups at login and the one that ran usermod does not have them. The udev rule applies to a board plugged in after the openocd package was installed, so replug a board that was already attached

The build tools above are yours to put on the path, and cmake --build --preset Debug is what tells you they are there. Build before you run a plan: the test reactor, the part of Agentic HIL that takes one declared test plan and runs it end to end against the bench, flashes build/Debug/stm32-starter.elf, and on a tree nobody has built it refuses the plan before its first hardware action with test_config_invalid: Firmware artifact does not exist.

agentic-hil doctor validates this project's configuration and reports the debugger backend and the devices it bound, so it comes after the command that writes that configuration and not before it: run first, it has nothing to validate and refuses with config_file_not_found. agentic-hil setup or agentic-hil init is what writes the file it wants.

Three steps

1. Open this repository in your agent

git clone https://github.com/agentic-hil/stm32-starter.git
cd stm32-starter

Start your agent from this directory. It reads AGENTS.md and finds the agentic-hil MCP server already registered by the installer.

2. Say one sentence

Set up this project for the attached Nucleo-F446RE and run the three hardware
test plans in tests/hil.

The agent runs agentic-hil setup --agent <agent>, which discovers your ST-LINK, matches its virtual COM port, and writes this project's authoritative configuration outside the repository, which is where the policy that decides what the bench may do belongs. Then it builds the firmware and runs the plans.

Discovery uses STM32CubeProgrammer's CLI where that is installed. Where it is not, it enumerates the attached ST-LINK out of the host's own USB inventory and drives it with the openocd on your PATH, which is how a Linux bench with OpenOCD and no vendor tool is discovered. Neither backend needs the other, and agentic-hil doctor afterwards names the one your configuration bound.

A placeholder is written when no ST-LINK port is enumerated at all, on either path: discovery finds no probe, puts a placeholder where the probe's identity goes, and says so in that step of its own output. With no board attached that command is green anyway, and what you have afterwards is a configuration to finish on the day the board arrives, not a failure to work around. With the board on the desk and a placeholder written regardless, nothing enumerated the port, so it is the cable, the groups above, or the probe.

setup is the first command when the agent registrations in your home directory are yours, which on your own machine they are; when they belong to somebody else, under a CI runner's account for instance, agentic-hil init is the first command instead and writes this project's half without touching them.

3. Watch it prove itself

Two of the three plans go green on a working board. tests/hil/diagnostic.testconfig.yaml does not, and its report names the claim that went unmet and quotes what the board answered instead.

Every step is a command you can run yourself, and these are the ones the agent runs. The order is not decorative: doctor has a configuration to check only once setup has written one, and a plan has an image to flash only once the build has produced one.

agentic-hil setup --agent claude-code    # or: codex / opencode
agentic-hil doctor

cmake --preset Debug
cmake --build --preset Debug

agentic-hil test-reactor --test-config tests/hil/nominal.testconfig.yaml
agentic-hil test-reactor --test-config tests/hil/diagnostic.testconfig.yaml
agentic-hil test-reactor --test-config tests/hil/recovery.testconfig.yaml

The exercise: fix the bug

The firmware in firmware/src/main.c speaks a small JSON diagnostic protocol over the ST-LINK virtual COM port at 115200 baud:

CommandExpected answer
STATUS{"state":"READY","diagnostic":"NONE"}
DIAG ON{"state":"DEGRADED","diagnostic":"E_SELF_TEST"}
DIAG CLEAR{"state":"READY","diagnostic":"NONE"}

One of those three commands does not do what this table says. tests/hil/diagnostic.testconfig.yaml is the plan that catches it, and it is the only red plan of the three, so the failure points at one command rather than at the firmware in general.

Hand the exercise to your agent:

Run tests/hil/diagnostic.testconfig.yaml on the board, work out why the
diagnostic claim goes unmet, make the smallest firmware fix, rebuild, and rerun
all three plans. Do not change the test plans or the protocol.

A finished run is three green plans on one firmware revision, and the evidence is the report each run keeps under the operator's state root, at the path the run prints in its own summary line; .agentic-hil/reports/last-report.json in the workspace holds the last run only, so three plans leave one file there and three under the state root.

What runs where

On any machine, with no board attached, the check-plan suite validates the three hardware plans with the test reactor's own loader: the closed plan schema, the step vocabulary, the format version gate, the diagnostic protocol the plans state, and the rule that a plan names logical devices, the configuration's own names for the probe and the serial line, dut and dut_uart here, and never somebody's serial port. A plan that passes here is one the reactor's loader accepts. The reactor also holds a plan against this bench's configured devices and permissions before the first hardware action, and that half needs a configuration, so it happens on the bench.

On a bench with a Nucleo-F446RE attached, the three plans in tests/hil/ run through agentic-hil test-reactor, which validates every device name, permission and session order before the first hardware action, holds the probe and the serial line for the whole run under one lease, the machine-wide claim on a device that keeps a second run off it until this one gives it back, closes them even when a step fails, and writes one JSON report saying what ran under which policy. That is where electrical behaviour is established.

Register the MCP server in another host

The installer registers the server at user level for Claude Code, Codex and opencode, which is where a hardware gate belongs: outside the repository, so the agent working in the repository cannot rewrite how it is launched. That is why this repository ships no .mcp.json and no .vscode/mcp.json, and why both are listed in .gitignore.

For a host the installer does not cover, register it by hand. Every block below carries the same launch contract: the verified absolute path to your persistent agentic-hil executable, the argument mcp-stdio, and this repository root as the working directory. agentic-hil doctor prints the exact executable path to paste.

Claude Code, from this repository root:

claude mcp add --transport stdio --scope user agentic-hil -- "/absolute/path/to/persistent/agentic-hil" mcp-stdio

Codex, in ~/.codex/config.toml:

[mcp_servers.agentic-hil]
command = "/absolute/path/to/persistent/agentic-hil"
args = ["mcp-stdio"]
cwd = "/absolute/path/to/stm32-starter"
enabled = true

opencode, in ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "agentic-hil": {
      "type": "local",
      "command": [
        "/absolute/path/to/persistent/agentic-hil",
        "mcp-stdio"
      ],
      "cwd": "/absolute/path/to/stm32-starter",
      "enabled": true
    }
  }
}

VS Code and GitHub Copilot, in the VS Code user-profile MCP configuration (VS Code uses servers, not mcpServers):

{
  "servers": {
    "agentic-hil": {
      "type": "stdio",
      "command": "/absolute/path/to/persistent/agentic-hil",
      "args": [
        "mcp-stdio"
      ],
      "cwd": "${workspaceFolder}"
    }
  }
}

The MCP host guide has JetBrains, the generic host contract, and how to verify the connection.

Continuous integration

  • .github/workflows/check-plan.yml runs the check-plan suite on a GitHub-hosted runner, on every push and pull request, and uploads the JUnit XML report.
  • .github/workflows/hardware-test.yml runs the three plans on a self-hosted runner that has a Nucleo-F446RE attached, labelled agentic-hil and nucleo-f446re. It serialises bench access with a concurrency group and uploads the reactor reports and logs whether the run passed or failed. Its diagnostic step is expected to fail with comparator_unmet on the shipped firmware, and the step after it asserts exactly that from the diagnostic run's own report, so the workflow is green while the firmware answers DIAG ON the way it ships and red when the bench, the build or any other plan breaks. Once the firmware is fixed that plan passes outright, and the workflow names the two lines to delete so the step becomes a plain pass.

expected/ holds the reference check-plan report, and expected/README.md has the two commands that reproduce it and compare the two byte for byte.

Where everything lives

PathWhat it is
firmware/src/main.cThe whole firmware: USART2 setup, the diagnostic protocol, and the defect
CMakeLists.txt, CMakePresets.jsonThe Debug and Release builds, producing build/Debug/stm32-starter.elf
tests/hil/Three test reactor plans, one per hardware test
tests/check_plan/The host-side validation of those plans
agentic-hil.config.example.yamlWhat agentic-hil setup should discover for this project
AGENTS.mdThe instructions your agent reads when it opens this repository
expected/Reference reports to diff against
validation/The evidence gate: the checklist of results this starter has to have behind it before it is announced anywhere, and how far it has been walked

Where to ask

Questions about this starter, including the ones you are not yet sure are defects, go to Discussions Q&A. A run you got working, on this board or on another one behind the same probe, belongs in Show and tell, because a path somebody has actually walked end to end is the most useful thing anyone can read here. A first run on your own bench goes in the first run report, green or red: a red one says where this path breaks for a reader who has not read the source, which is the only way that gets found.

Licence

Apache License 2.0, the same licence Agentic HIL itself ships under.

agentic-hil
ai-agents
embedded
firmware
hardware-in-the-loop
mcp
nucleo
stm32

agentic-hil/stm32-starter

Nucleo-F446RE starter for Agentic Hardware-in-the-Loop

Python

1

41 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Claude Code writing firmware, flashing a real board and checking the result (r/SideProject)

The demo for Agentic HIL shows Claude Code writing a small firmware program, flashing it to a board and reading “Hello World” back over the serial connection. It then runs a deliberately wrong check so you can see the failure report too. What I like about this is that the hardware response becomes…

0

Oct 1, 2026

README

STM32 Agentic HIL Starter

Take your own copy of this repository, say one sentence to your coding agent, and watch it build firmware, flash a real Nucleo-F446RE, talk to it, and prove what the board actually did. The three test plans are the gate the agent has to pass.

Click Use this template above to get your own copy, or clone this one. The hardware run needs a Nucleo-F446RE, a USB cable, and the onboard ST-LINK, and nothing more: no fixture to wire, no adapter to buy. The three test plans and the suite that validates them need no board at all, so that half runs on any machine and is the first thing to run whether or not a Nucleo is on your desk. This is the reference path for Agentic HIL, the local MCP server that lets an AI agent develop firmware on hardware you own: build, flash, stimulate, observe, diagnose, fix, with the run on the board deciding whether the work is done and the reactor's report as the evidence you read. A Nucleo-F446RE is what this starter proves that path on, not what the path needs: any board reached through an ST-LINK runs the same way, with that board's own target script or controller name and its own UART named in the configuration instead of this one's.

The firmware ships with one deliberate defect in its diagnostic protocol. Finding it is the exercise.

Install Agentic HIL

Linux and macOS, in any shell:

curl -LsSf https://agentic-hil.github.io/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"

Windows, in PowerShell:

irm https://agentic-hil.github.io/install.ps1 | iex

Windows, from cmd.exe or the Run box:

powershell -c "irm https://agentic-hil.github.io/install.ps1|iex"

One line installs the package user-local and registers the agent skill and the MCP server for every agent CLI it finds on your PATH. No admin rights are required, and it writes nothing inside this repository. Then restart your agent once.

When ~/.local/bin is not already on your PATH, the installer puts it there for the shells that come after it, by one line in your shell profile, and names the file it edited. The export line in the first block is for the shell you are in right now, and for any shell that does not read that file; it puts uv within reach as well, because both land in that directory.

To register a single agent instead of all of them, pass --agent claude-code (or codex, opencode); piped, that reads | sh -s -- --agent claude-code.

Run the plans with no board attached

The check-plan suite validates the three hardware test plans on any machine, with nothing plugged in. It needs Python 3.10 or newer and uv, which is what the lock file is for. The one-line installer above leaves you with uv on most machines: it installs through uv when it finds one, and fetches uv itself when the Python it found has no pip or is one the distribution keeps for its own packages, so only a Python with a usable pip and no uv on the machine ends up without it. If command -v uv finds nothing, install it with curl -LsSf https://astral.sh/uv/install.sh | sh on Linux and macOS, or irm https://astral.sh/uv/install.ps1 | iex in PowerShell.

Clone this repository first, then run these inside the clone:

uv sync
uv run pytest -q -s

Every test states what a green run is worth:

PASS  configuration and test semantics validated without a board
NEEDS PHYSICAL FIXTURE  electrical behavior not verified

The environment uv sync builds carries the Agentic HIL release the lock file pins, and that release backs this suite. The hardware commands further down run the agentic-hil the installer put on your PATH, which may be newer. Two versions on one machine is the expected shape, not a mistake; agentic-hil doctor reports the one on your PATH.

Without uv, the same suite runs from a plain virtual environment:

python -m venv .venv
.venv/bin/pip install "agentic-hil>=0.21.0" pytest pyyaml   # .venv\Scripts\pip on Windows
.venv/bin/pytest -q -s   # .venv\Scripts\pytest on Windows

What runs where is the account of what a green suite establishes and what it leaves to the bench.

What you need for the hardware run

  • A Nucleo-F446RE connected through its ST-LINK USB port
  • CMake 3.21 or newer, which is what CMakePresets.json needs for its preset version and its toolchainFile, plus Ninja and the GNU Arm Embedded Toolchain (arm-none-eabi-gcc) on your PATH; STM32CubeCLT carries all three
  • A debugger backend, STM32CubeProgrammer CLI or OpenOCD; either one is enough on its own, and step 2 says what each of them means for discovery
  • Python 3.10 or newer
  • On Linux, membership in two groups before the first plan: dialout, which owns the ST-LINK's virtual COM port, and plugdev, which the openocd package's udev rule gives the probe itself. sudo usermod -aG plugdev,dialout "$USER" adds both, and it is the one step in this path that asks for an administrator. Log in again afterwards, because a shell takes its groups at login and the one that ran usermod does not have them. The udev rule applies to a board plugged in after the openocd package was installed, so replug a board that was already attached

The build tools above are yours to put on the path, and cmake --build --preset Debug is what tells you they are there. Build before you run a plan: the test reactor, the part of Agentic HIL that takes one declared test plan and runs it end to end against the bench, flashes build/Debug/stm32-starter.elf, and on a tree nobody has built it refuses the plan before its first hardware action with test_config_invalid: Firmware artifact does not exist.

agentic-hil doctor validates this project's configuration and reports the debugger backend and the devices it bound, so it comes after the command that writes that configuration and not before it: run first, it has nothing to validate and refuses with config_file_not_found. agentic-hil setup or agentic-hil init is what writes the file it wants.

Three steps

1. Open this repository in your agent

git clone https://github.com/agentic-hil/stm32-starter.git
cd stm32-starter

Start your agent from this directory. It reads AGENTS.md and finds the agentic-hil MCP server already registered by the installer.

2. Say one sentence

Set up this project for the attached Nucleo-F446RE and run the three hardware
test plans in tests/hil.

The agent runs agentic-hil setup --agent <agent>, which discovers your ST-LINK, matches its virtual COM port, and writes this project's authoritative configuration outside the repository, which is where the policy that decides what the bench may do belongs. Then it builds the firmware and runs the plans.

Discovery uses STM32CubeProgrammer's CLI where that is installed. Where it is not, it enumerates the attached ST-LINK out of the host's own USB inventory and drives it with the openocd on your PATH, which is how a Linux bench with OpenOCD and no vendor tool is discovered. Neither backend needs the other, and agentic-hil doctor afterwards names the one your configuration bound.

A placeholder is written when no ST-LINK port is enumerated at all, on either path: discovery finds no probe, puts a placeholder where the probe's identity goes, and says so in that step of its own output. With no board attached that command is green anyway, and what you have afterwards is a configuration to finish on the day the board arrives, not a failure to work around. With the board on the desk and a placeholder written regardless, nothing enumerated the port, so it is the cable, the groups above, or the probe.

setup is the first command when the agent registrations in your home directory are yours, which on your own machine they are; when they belong to somebody else, under a CI runner's account for instance, agentic-hil init is the first command instead and writes this project's half without touching them.

3. Watch it prove itself

Two of the three plans go green on a working board. tests/hil/diagnostic.testconfig.yaml does not, and its report names the claim that went unmet and quotes what the board answered instead.

Every step is a command you can run yourself, and these are the ones the agent runs. The order is not decorative: doctor has a configuration to check only once setup has written one, and a plan has an image to flash only once the build has produced one.

agentic-hil setup --agent claude-code    # or: codex / opencode
agentic-hil doctor

cmake --preset Debug
cmake --build --preset Debug

agentic-hil test-reactor --test-config tests/hil/nominal.testconfig.yaml
agentic-hil test-reactor --test-config tests/hil/diagnostic.testconfig.yaml
agentic-hil test-reactor --test-config tests/hil/recovery.testconfig.yaml

The exercise: fix the bug

The firmware in firmware/src/main.c speaks a small JSON diagnostic protocol over the ST-LINK virtual COM port at 115200 baud:

CommandExpected answer
STATUS{"state":"READY","diagnostic":"NONE"}
DIAG ON{"state":"DEGRADED","diagnostic":"E_SELF_TEST"}
DIAG CLEAR{"state":"READY","diagnostic":"NONE"}

One of those three commands does not do what this table says. tests/hil/diagnostic.testconfig.yaml is the plan that catches it, and it is the only red plan of the three, so the failure points at one command rather than at the firmware in general.

Hand the exercise to your agent:

Run tests/hil/diagnostic.testconfig.yaml on the board, work out why the
diagnostic claim goes unmet, make the smallest firmware fix, rebuild, and rerun
all three plans. Do not change the test plans or the protocol.

A finished run is three green plans on one firmware revision, and the evidence is the report each run keeps under the operator's state root, at the path the run prints in its own summary line; .agentic-hil/reports/last-report.json in the workspace holds the last run only, so three plans leave one file there and three under the state root.

What runs where

On any machine, with no board attached, the check-plan suite validates the three hardware plans with the test reactor's own loader: the closed plan schema, the step vocabulary, the format version gate, the diagnostic protocol the plans state, and the rule that a plan names logical devices, the configuration's own names for the probe and the serial line, dut and dut_uart here, and never somebody's serial port. A plan that passes here is one the reactor's loader accepts. The reactor also holds a plan against this bench's configured devices and permissions before the first hardware action, and that half needs a configuration, so it happens on the bench.

On a bench with a Nucleo-F446RE attached, the three plans in tests/hil/ run through agentic-hil test-reactor, which validates every device name, permission and session order before the first hardware action, holds the probe and the serial line for the whole run under one lease, the machine-wide claim on a device that keeps a second run off it until this one gives it back, closes them even when a step fails, and writes one JSON report saying what ran under which policy. That is where electrical behaviour is established.

Register the MCP server in another host

The installer registers the server at user level for Claude Code, Codex and opencode, which is where a hardware gate belongs: outside the repository, so the agent working in the repository cannot rewrite how it is launched. That is why this repository ships no .mcp.json and no .vscode/mcp.json, and why both are listed in .gitignore.

For a host the installer does not cover, register it by hand. Every block below carries the same launch contract: the verified absolute path to your persistent agentic-hil executable, the argument mcp-stdio, and this repository root as the working directory. agentic-hil doctor prints the exact executable path to paste.

Claude Code, from this repository root:

claude mcp add --transport stdio --scope user agentic-hil -- "/absolute/path/to/persistent/agentic-hil" mcp-stdio

Codex, in ~/.codex/config.toml:

[mcp_servers.agentic-hil]
command = "/absolute/path/to/persistent/agentic-hil"
args = ["mcp-stdio"]
cwd = "/absolute/path/to/stm32-starter"
enabled = true

opencode, in ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "agentic-hil": {
      "type": "local",
      "command": [
        "/absolute/path/to/persistent/agentic-hil",
        "mcp-stdio"
      ],
      "cwd": "/absolute/path/to/stm32-starter",
      "enabled": true
    }
  }
}

VS Code and GitHub Copilot, in the VS Code user-profile MCP configuration (VS Code uses servers, not mcpServers):

{
  "servers": {
    "agentic-hil": {
      "type": "stdio",
      "command": "/absolute/path/to/persistent/agentic-hil",
      "args": [
        "mcp-stdio"
      ],
      "cwd": "${workspaceFolder}"
    }
  }
}

The MCP host guide has JetBrains, the generic host contract, and how to verify the connection.

Continuous integration

  • .github/workflows/check-plan.yml runs the check-plan suite on a GitHub-hosted runner, on every push and pull request, and uploads the JUnit XML report.
  • .github/workflows/hardware-test.yml runs the three plans on a self-hosted runner that has a Nucleo-F446RE attached, labelled agentic-hil and nucleo-f446re. It serialises bench access with a concurrency group and uploads the reactor reports and logs whether the run passed or failed. Its diagnostic step is expected to fail with comparator_unmet on the shipped firmware, and the step after it asserts exactly that from the diagnostic run's own report, so the workflow is green while the firmware answers DIAG ON the way it ships and red when the bench, the build or any other plan breaks. Once the firmware is fixed that plan passes outright, and the workflow names the two lines to delete so the step becomes a plain pass.

expected/ holds the reference check-plan report, and expected/README.md has the two commands that reproduce it and compare the two byte for byte.

Where everything lives

PathWhat it is
firmware/src/main.cThe whole firmware: USART2 setup, the diagnostic protocol, and the defect
CMakeLists.txt, CMakePresets.jsonThe Debug and Release builds, producing build/Debug/stm32-starter.elf
tests/hil/Three test reactor plans, one per hardware test
tests/check_plan/The host-side validation of those plans
agentic-hil.config.example.yamlWhat agentic-hil setup should discover for this project
AGENTS.mdThe instructions your agent reads when it opens this repository
expected/Reference reports to diff against
validation/The evidence gate: the checklist of results this starter has to have behind it before it is announced anywhere, and how far it has been walked

Where to ask

Questions about this starter, including the ones you are not yet sure are defects, go to Discussions Q&A. A run you got working, on this board or on another one behind the same probe, belongs in Show and tell, because a path somebody has actually walked end to end is the most useful thing anyone can read here. A first run on your own bench goes in the first run report, green or red: a red one says where this path breaks for a reader who has not read the source, which is the only way that gets found.

Licence

Apache License 2.0, the same licence Agentic HIL itself ships under.

agentic-hil
ai-agents
embedded
firmware
hardware-in-the-loop
mcp
nucleo
stm32

Languages

Python

53.5%

C

23.3%

CMake

10.3%

Assembly

8.2%

Linker Script

4.7%