Skip to main content

Agentic HIL

PyPI version CI License

Your AI agent writes the firmware, flashes it to the board on your desk, drives UART and CAN against it, reads back what the hardware actually did, and fixes what it got wrong; the run on the real board is what decides whether the work is done, and you review the pull request with that run's evidence in it.

https://github.com/user-attachments/assets/8d39ba93-beeb-484e-b9e9-d9ce79538523

Nothing in that run is staged. One restart after the install line, in a freshly created firmware project, the first sentence makes the agent set the bench up itself and the second makes the board say Hello World and prove it said it: the configuration is created over MCP with the permissions reported out loud, the firmware is written on the spot, flash_firmware and com_read go through the gate, the twelve bytes come back off the wire, and the plan it pins is run once green and once against a wrong expectation, because a test that cannot fail proves nothing. What remains in the project afterwards is the plan as a reviewable file and the run's own report: lease released, safe state confirmed, nothing quarantined.

Install

Linux / macOS (any shell):

curl -LsSf https://agentic-hil.github.io/install.sh | sh

Windows, in PowerShell:

irm https://agentic-hil.github.io/install.ps1 | iex

The command lands in the user bin directory of the package manager that installed it, and the installer asks that manager where it went rather than assuming: uv tool dir --bin for a uv install, and the selected interpreter itself for a pip --user one. When that directory is not on your PATH already, the installer puts it there and says so: one line in the one shell profile your shell reads, or on Windows the directory in front of your own Path. Open a new shell and the command is there. Pass --no-path (-NoPath in PowerShell) to keep that edit for yourself, and the installer prints the exact line instead.

One line installs the package user-local and registers the agent skill and the MCP server for every agent CLI it finds on your PATH. No admin rights required, ever, and it touches nothing inside any repository. Finding no claude, codex or opencode CLI there, it says so and writes nothing of any agent's: install the agent CLI, then run agentic-hil agent-install --agent <claude-code|codex|opencode> yourself, which is the line the installer prints for that case. Then restart your agent once, and after that one restart your agent sets this project up itself, at the first hardware question you ask it.

The line above and the checksummed route in installation run the same installer, and on a machine with nothing here yet both install the same thing, the release from PyPI: the script resolves agentic-hil[can] against the package index and carries no branch or git reference at all, and --version <x.y.z> pins one exact release instead. On a rerun it repairs what is already installed rather than forcing the public release over it: an installation reporting a .devN version (an editable checkout of this repository) is kept untouched, and a uv-managed tool installed from a path, URL or git reference is refreshed from that same recorded source rather than switched to the index. What the checksummed route adds is the script itself, read before it runs: install.sh and its install.sh.sha256 come from the same release, one is checked against the other, and the file that runs is the file you checked.

The STM32 starter is the shortest way to watch that happen on hardware: three steps on a Nucleo-F446RE, with the firmware, the test plans and one planted defect already in place.

Claude Code can also take the skill from the plugin marketplace: /plugin marketplace add agentic-hil/agentic-hil, then /plugin install agentic-hil@agentic-hil. The plugin carries the skill and nothing else; the MCP server itself still comes from the install line above, which registers a verified absolute executable path outside your repository.

Installation has every other path: the cmd.exe spelling, the repair run (the same line again, which reinstalls in place when agentic-hil upgrade itself fails), that checksummed route in full, registering one agent instead of all of them, driving your own package manager, setup for a bench that is already attached, the optional extras, upgrading, and every platform and debugger backend. TROUBLESHOOTING.md covers what to do when something does not start.

A first command that needs no board, in a clone of this repository: agentic-hil check-plan examples/nucleo-f446re_demo/testconfig.yaml answers All 1 test plan(s) load through the reactor's loader. and exits 0, having loaded no configuration and touched no hardware.

Agentic Hardware-in-the-Loop (Agentic HIL) is a Python package that lets a coding agent develop firmware on the real board. It exposes bounded MCP tools for probing, flashing, resetting, artifact validation, serial and CAN stimulus/feedback, reports, and logs, all without giving an agent arbitrary host or debugger access. The run on the board is the gate: the work is done when the hardware behaved, and the report that run writes is the evidence a reviewer reads. What it supports is the debug probe with its backend rather than a board, ST-Link through OpenOCD or the STM32CubeProgrammer CLI and CMSIS-DAP probes through pyOCD, so any board behind such a probe runs the same software. Each project has exactly one authoritative configuration stored outside the repository, out of reach of the agent's own file tools.

Why

A green build is not enough in embedded development: firmware has to behave correctly on the real board, so the run on that board is the gate the work has to pass before it is done. Classic tools automate single steps (flash here, read a log there), but the moment real hardware has to respond, a human is back in the loop, which is what stops an agent from developing firmware through to that gate. Handing an agent a raw debugger shell or direct serial access instead is neither safe nor reproducible, and leaves a reviewer nothing to read.

The Agentic HIL loop: build, flash, stimulate, observe, then diagnose and fix, closing back onto build. Flash and stimulate write to the real board on your bench; observe reads back from it. Your agent runs the loop unattended; you review the pull request.

Agentic HIL closes the gap with a small, auditable gate:

An AI agent or CI reaches Agentic HIL over MCP stdio only. The authoritative configuration, owned by the operator outside the workspace, gates every action. Agentic HIL drives debug probes (OpenOCD, pyOCD, STM32CubeProgrammer), serial ports, and CAN buses shared through the broker, and answers with structured results, reports, and a SHA-256 audit chain.

Every hardware action is validated against the selected authoritative configuration, executed with timeouts, logged to .agentic-hil/logs/, and answered with a structured JSON result (ok, error_type, summary, likely_causes, report_path, log_path) that an agent can act on and a reviewer can read afterwards. What the agent may do at all is per device and per permission, and reaching for a debugger escape hatch is what takes flashing away.

What it drives

The unit it drives is the probe with its backend, not one board: three debugger backends (OpenOCD, pyOCD, and the STM32CubeProgrammer CLI), plus serial ports and CAN (PCAN, SocketCAN, or a custom bridge; several runs can share one bus), on Linux, macOS, and Windows, Python 3.10 or newer, all CI-tested. The worked example in examples/nucleo-f446re_demo/ runs the whole loop on an ST Nucleo-F446RE, the board this repository proves the path on rather than the boundary of what runs; installation has every backend and platform in detail.

The test reactor

A plan is how the gate is written down, and one YAML plan drives the whole bench: flash, reset, write, read with a comparator (exact text, a pattern, or a numeric range over a captured value), delays, and sessions that close themselves. Plans name logical devices; the bench configuration binds them to real hardware, so the same plan runs unchanged on every machine that has one. A failing step aborts the run, and the bench recovers itself: reap, reset into halt, probe, all attested in the run result. How plans work.

Security by construction

A device does not exist on this bench until your configuration declares it, and a call naming any other one is refused before a driver is opened: unknown_device where a run declares it, and com_port_not_configured or can_bus_not_configured where a port tool or a bus tool names it, each of those two naming the ones the configuration does declare. On a device it does declare, a generated configuration grants every permission that has a tool behind it and holds allow_raw_debugger_commands and allow_mass_erase false, the interlocked pair that refuses flashing while either is true, false by construction when MCP project_config_create generates with no loaded configuration, and the default agentic-hil init writes unless your agentic-hil.config.example.yaml opens one; a project_config_create regenerating an existing bench carries the loaded permissions back instead, either interlock included, keyed on the entry name and not on whether a probe was found before, so a pair an operator opened on a named dut placeholder survives the very call that first binds a probe to it and only a logical entry whose name the loaded configuration did not carry comes back false. agentic-hil revoke <key> takes any single grant back and agentic-hil grant <key> reopens it, from your own shell; over MCP an agent narrows its own authority and never widens it. Every hardware action is validated, leased machine-wide, and written to a SHA-256 audit chain. The authoritative configuration lives outside the workspace, where the agent cannot edit it. Enforcement sits in the tool rather than in the agent host on purpose: a host's permission system judges shell strings and differs per host, while the bench's permissions judge the hardware action itself and travel with the bench, so the CLI, pytest, CI and the test reactor all walk the same gate. A failed run still gives the bench back: it aborts with its verdict, the recovery action resets and re-reads the target, and the standing quarantine is kept for the one state no later contact can rebuild, a broken audit trail. The safety model is the short version, the security design the long one.

Quickstart: one real run

The worked example is a firmware project of its own. Plug the board in, build it, and point Agentic HIL at it from that directory:

cd examples/nucleo-f446re_demo
cmake --preset Debug && cmake --build --preset Debug   # → build/Debug/nucleo-f446re_demo.elf
agentic-hil setup --agent claude-code                  # or: codex / opencode
agentic-hil doctor

doctor checks the configuration against the attached bench and names what it finds (a missing toolchain, an unreachable probe, a target type this host cannot resolve) before anything is flashed. If the board arrived after setup ran, agentic-hil adopt-hardware fills in the probe serial, the backend executable and the COM device it left unset (--dry-run shows the plan first). On a host with OpenOCD and no STM32CubeProgrammer the probe is read from the host's own USB serial inventory, so a lone attached ST-Link binds without anybody retyping its serial; --probe-id <serial> names the board where a second probe that publishes no virtual COM port is attached beside it.

With the MCP host started from that directory, the agent drives four calls:

flash_firmware     {"image_path": "build/Debug/nucleo-f446re_demo.elf"}
com_session_start  {"port_id": "dut_uart"}
reset_target       {"mode": "run"}
com_read           {"port_id": "dut_uart", "wait_timeout_s": 5}
→ feedback contains "Hello World"

The same loop runs headless as a pytest regression: pytest tests/ in that directory flashes the ELF, resets the target and asserts the boot banner on the UART. examples/nucleo-f446re_demo/ walks through both, and docs/testing.md covers writing the run down as a reviewable YAML plan instead.

Where the depth lives

If you want Read
to install, upgrade, add CAN or pyOCD, or look up a command docs/installation.md
what the authoritative configuration declares and who may change it docs/configuration.md
the complete MCP tool surface and how a run is composed from it docs/mcp-tools.md
to register the server in a specific MCP host docs/mcp-hosts.md
to write hardware tests (YAML plans or pytest) docs/testing.md
why it is safe to leave an agent alone with the bench docs/safety-model.md and docs/security-design.md
a failure diagnosed TROUBLESHOOTING.md
to point your agent at this repository AI_AGENT_QUICKSTART.md and AGENTS.md

Names: the Python distribution/install target, CLI command, repository URL, and MCP server name use agentic-hil. Python imports, pytest plugin names, fixtures, and Python examples use agentic_hil.

Development

python -m pip install -e '.[dev]'
ruff check src tests evals tools
pytest
python -m build
twine check dist/*

The package is configured for PyPI publishing through GitHub trusted publishing in .github/workflows/workflow.yml. Contribution guidelines: CONTRIBUTING.md.

Security

Policy bypasses are treated as vulnerabilities; see SECURITY.md.

Support

Linux, macOS and Windows are supported equally, what is supported is the debug probe with the backend behind it (ST-Link through OpenOCD or the STM32CubeProgrammer CLI, CMSIS-DAP probes through pyOCD) rather than any list of boards, and issues are answered within 24 hours on workdays, security reports within seven days: docs/support.md is the whole promise, including what is not promised. Ask in Discussions Q&A, show a run in Show and tell, and put a first run on your own bench, green or red, in the first run report.

License

Apache-2.0. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentic_hil-0.21.5.tar.gz (2.2 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentic_hil-0.21.5-py3-none-any.whl (963.5 kB view details)

Uploaded Python 3

File details

Details for the file agentic_hil-0.21.5.tar.gz.

File metadata

  • Download URL: agentic_hil-0.21.5.tar.gz
  • Upload date:
  • Size: 2.2 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentic_hil-0.21.5.tar.gz
Algorithm Hash digest
SHA256 4a60748bfbf35923730738a69f9f26c2a97dddd62d7ca9562eb82d3f2d2972d5
MD5 974b115c5cec9a858fd4f233c0fb3f4f
BLAKE2b-256 abce63aaebcbd668a961c98d5ff28a70b5cc36b8a8c98e068188bb0d13a1c4d7

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentic_hil-0.21.5.tar.gz:

Publisher: workflow.yml on agentic-hil/agentic-hil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agentic_hil-0.21.5-py3-none-any.whl.

File metadata

  • Download URL: agentic_hil-0.21.5-py3-none-any.whl
  • Upload date:
  • Size: 963.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentic_hil-0.21.5-py3-none-any.whl
Algorithm Hash digest
SHA256 f8f04cd8606dab4f0abb6790ff4b28c9fce5c3482e66dd46afc9e79b0b94377e
MD5 441b68bd311947e05647b51fde8840c0
BLAKE2b-256 5c2c08eeb9cc273fdc9c894657db6b908474eb7cc6a5eaf1f4be26a22f817d37

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentic_hil-0.21.5-py3-none-any.whl:

Publisher: workflow.yml on agentic-hil/agentic-hil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.21.5 This release

2 files

0.21.4

2 files

0.21.3

2 files

0.21.2

2 files

0.21.1

2 files

0.21.0

2 files

0.20.0

2 files

0.19.0

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.1

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page