A risk-tiered agent system for hardware prototyping — electronics, PCBs, firmware, sourcing and manufacturability — built for Claude Code.
It is two things: a set of skills that know the domain, and a harness that stops the agent doing irreversible things with them. The second half is most of the work, because hardware is unforgiving in a way software isn't. A bad deploy rolls back in ninety seconds; a bad eFuse burn is a dead chip.
The skills come in two kinds. Domain skills say what is true about circuits, firmware, parts and fabs. Reasoning lenses say how to think — which budget a symptom implicates, what actually decides an A-or-B choice, which measurement to take before changing anything, and how much a given number can be trusted. They load together. A lens without a domain skill is advice; a domain skill without a lens is an operator waiting to be told what to do.
Every action is one of three tiers:
- T1 — autonomous. Read, analyse, simulate, build, export. Runs freely. Nothing changes state on real hardware and nothing costs money.
- T2 — confirm first. Anything that writes to a connected device. The agent must state which device, what changes, and how to undo it before it runs.
- T3 — advisory only. Irreversible silicon (eFuse, secure boot), mains voltage, lithium charging, spending money, machine motion, regulatory claims, and promoting knowledge into the canonical wiki. Drafted and handed over, never executed.
Enforced in three places, because prose alone is advisory to a model:
| Layer | Covers | Fails |
|---|---|---|
Skills (AGENT.md, skills/) |
Reasoning — the agent declines before it reaches for a tool | Open: it's persuasion |
Risk broker (harness/risk_broker/) |
Commands routed through its MCP tools | Closed: unmatched → T3 |
PreToolUse hook (harness/kb/) |
Write/Edit/MultiEdit/NotebookEdit/Bash into the wiki | Closed for known tools |
Four of them, loaded alongside whichever domain skill the work needs:
| Lens | The question it asks |
|---|---|
problem-reframing |
Which budget is actually violated — current, thermal, energy, timing, area, RF, cost? |
decision-framing |
What are the alternatives, what constraint decides, and how expensive is it to undo? |
diagnostic-reasoning |
What is the cheapest measurement that splits the hypotheses, taken before anything changes? |
model-vs-reality |
Is this number measured, documented, calculated, or assumed — and how would I check it? |
Two things keep a lens from floating free of the harness. Each one names the deliverable shape its output has to land in — a current budget, a design decision, a bring-up checklist — so thinking terminates in something concrete. And each names the tier of the actions it may suggest, so a lens cannot reason its way past T2 or T3. Reframing is free; the measurement it proposes still has to be a read.
The reversibility ladder is the idea most of them lean on: breadboard → PCB rev 1 → parts committed → potted or installed → certified. The cost of change goes up roughly tenfold per rung, and the risk tiers are the bottom of that ladder made mechanical. Decide at the lowest rung that can actually test the question.
Design rationale, including which lenses were deliberately not built and why
most of the creative-technologist lenses don't transfer, is in
docs/analysis-reasoning-skills.md.
AGENT.md orchestrator: routing, tiers, constraint questions, explanation style
policy/tool-tiers.yaml risk classification as reviewable data
policy/guardrails.md stances that apply when no command is involved
skills/ domain: circuit-design, firmware, sourcing-bom, manufacturing-dfm, knowledge-base
lenses: problem-reframing, decision-framing, diagnostic-reasoning, model-vs-reality
harness/risk_broker/ MCP server enforcing the tiers in code
harness/kb/ wiki bootstrap, linter, write gate
harness/smoke/ nine behavioural tests and their recorded results
TOOLS.md tool catalogue: purpose, tier, licence, install
wiki-template/ scaffold for the knowledge wiki
docs/ design analyses, written before the code they argue for
Every number in skills/*/references/ carries a source URL and the date it was
checked, or is explicitly marked unverified. That discipline exists because the
first draft of these files confidently asserted several things that turned out
to be wrong — see REVIEW.md.
Requires Python 3.10+, just, and ideally
uv (falls back to venv + pip).
git clone https://github.com/melissa-pereira-deel/hardware-agent
cd hardware-agent
just setup
just test # 182 tests
just doctor # which allowlisted tools are actually installedTOOLS.md catalogues every tool the agent can drive — purpose, risk tier,
licence and install command. Being allowlisted is not the same as being
installed; just doctor reports the difference.
A project-scoped .mcp.json is included and uses a relative interpreter path,
so it works wherever you cloned to. To use the broker from other projects,
register it at user scope instead:
claude mcp add --scope user risk-broker -- "$(pwd)/.venv/bin/python" -m risk_broker.serverjust verify-broker # starts it over stdio, lists tools, probes every tier boundaryjust bootstrap-wiki ~/dev/wiki-hardwareMarkdown in git, searched with ripgrep. No vector database — at one person's scale grep wins on maintenance, staleness and debuggability. Three tiers with a wall between them:
raw/ cached sources. Immutable, gitignored; raw/MANIFEST.tsv records what was fetched.
scratch/ agent drafts. Untrusted, freely writable.
wiki/ canonical. Provenance required. Gated.
The wall is the point. An agent writing unsupervised into its own authoritative knowledge base is how one early misreading becomes load-bearing fact six months later.
A PreToolUse hook that refuses tool writes into wiki/, so canonicalisation
stays a human action. Add to ~/.claude/settings.json:
{
"hooks": {
"PreToolUse": [{
"matcher": "Write|Edit|MultiEdit|NotebookEdit|Bash",
"hooks": [{ "type": "command",
"command": "python3 /absolute/path/to/hardware-agent/harness/kb/deny_wiki_write.py" }]
}]
}
}Set HARDWARE_AGENT_WIKI if your wiki isn't at ~/dev/wiki-hardware.
Hook registrations load at session start — one added mid-session is inert
until you restart. Verify with just check-gate (logic) and
python3 harness/kb/check_gate.py --live (wiring).
Stated plainly because the repo's own subject is not overclaiming about guardrails:
- The write gate is bar-raising, not airtight. It inspects command strings, so
cd wiki/hardware && cp ../../x.md .defeats it, and anything a child process writes — viajust,make, a shell script — is invisible to it. - Unknown tools fail open until the matcher is widened. A fail-closed path is implemented and tested but dormant; see
harness/kb/deny_wiki_write.py. - The broker only sees what's routed through it. A direct
Bashcall doesn't reach it. That is why the hook exists. - Simulation is not hardware. Wokwi and friends diverge from real silicon.
- Nothing here is legal or compliance advice. The ANATEL and licensing material is direction, not a substitute for an accredited lab.
The durable guarantee is the wiki's pre-commit lint gate plus a human reading the diff. Everything else is defence in depth above it.
harness/smoke/ holds the behavioural tests, run against cold agents with no
memory of the session that wrote the skills, plus what each one found. The
first five cover the domain skills and have recorded results; prompts 6–9
cover the lenses and have not been run cold yet, which is the honest state
of the evidence for them.
REVIEW.md is the original audit of the scaffold. CHANGELOG.md says what
changed and why.
Built with Claude Code; the commit history
carries Co-Authored-By trailers throughout. The knowledge wiki it manages is a
separate, private repository by design — see
skills/knowledge-base/references/legal-and-etiquette.md for the reasoning.
MIT — see LICENSE.