Skip to content

Repository files navigation

hardware-agent

A risk-tiered agent system for hardware prototyping — electronics, PCBs, firmware, sourcing and manufacturability — built for Claude Code.

It is two things: a set of skills that know the domain, and a harness that stops the agent doing irreversible things with them. The second half is most of the work, because hardware is unforgiving in a way software isn't. A bad deploy rolls back in ninety seconds; a bad eFuse burn is a dead chip.

The skills come in two kinds. Domain skills say what is true about circuits, firmware, parts and fabs. Reasoning lenses say how to think — which budget a symptom implicates, what actually decides an A-or-B choice, which measurement to take before changing anything, and how much a given number can be trusted. They load together. A lens without a domain skill is advice; a domain skill without a lens is an operator waiting to be told what to do.

The risk model

Every action is one of three tiers:

  • T1 — autonomous. Read, analyse, simulate, build, export. Runs freely. Nothing changes state on real hardware and nothing costs money.
  • T2 — confirm first. Anything that writes to a connected device. The agent must state which device, what changes, and how to undo it before it runs.
  • T3 — advisory only. Irreversible silicon (eFuse, secure boot), mains voltage, lithium charging, spending money, machine motion, regulatory claims, and promoting knowledge into the canonical wiki. Drafted and handed over, never executed.

Enforced in three places, because prose alone is advisory to a model:

Layer Covers Fails
Skills (AGENT.md, skills/) Reasoning — the agent declines before it reaches for a tool Open: it's persuasion
Risk broker (harness/risk_broker/) Commands routed through its MCP tools Closed: unmatched → T3
PreToolUse hook (harness/kb/) Write/Edit/MultiEdit/NotebookEdit/Bash into the wiki Closed for known tools

The lenses

Four of them, loaded alongside whichever domain skill the work needs:

Lens The question it asks
problem-reframing Which budget is actually violated — current, thermal, energy, timing, area, RF, cost?
decision-framing What are the alternatives, what constraint decides, and how expensive is it to undo?
diagnostic-reasoning What is the cheapest measurement that splits the hypotheses, taken before anything changes?
model-vs-reality Is this number measured, documented, calculated, or assumed — and how would I check it?

Two things keep a lens from floating free of the harness. Each one names the deliverable shape its output has to land in — a current budget, a design decision, a bring-up checklist — so thinking terminates in something concrete. And each names the tier of the actions it may suggest, so a lens cannot reason its way past T2 or T3. Reframing is free; the measurement it proposes still has to be a read.

The reversibility ladder is the idea most of them lean on: breadboard → PCB rev 1 → parts committed → potted or installed → certified. The cost of change goes up roughly tenfold per rung, and the risk tiers are the bottom of that ladder made mechanical. Decide at the lowest rung that can actually test the question.

Design rationale, including which lenses were deliberately not built and why most of the creative-technologist lenses don't transfer, is in docs/analysis-reasoning-skills.md.

What's here

AGENT.md                  orchestrator: routing, tiers, constraint questions, explanation style
policy/tool-tiers.yaml    risk classification as reviewable data
policy/guardrails.md      stances that apply when no command is involved
skills/                   domain: circuit-design, firmware, sourcing-bom, manufacturing-dfm, knowledge-base
                          lenses: problem-reframing, decision-framing, diagnostic-reasoning, model-vs-reality
harness/risk_broker/      MCP server enforcing the tiers in code
harness/kb/               wiki bootstrap, linter, write gate
harness/smoke/            nine behavioural tests and their recorded results
TOOLS.md                  tool catalogue: purpose, tier, licence, install
wiki-template/            scaffold for the knowledge wiki
docs/                     design analyses, written before the code they argue for

Every number in skills/*/references/ carries a source URL and the date it was checked, or is explicitly marked unverified. That discipline exists because the first draft of these files confidently asserted several things that turned out to be wrong — see REVIEW.md.

Install

Requires Python 3.10+, just, and ideally uv (falls back to venv + pip).

git clone https://github.com/melissa-pereira-deel/hardware-agent
cd hardware-agent
just setup
just test          # 182 tests
just doctor        # which allowlisted tools are actually installed

TOOLS.md catalogues every tool the agent can drive — purpose, risk tier, licence and install command. Being allowlisted is not the same as being installed; just doctor reports the difference.

Register the risk broker

A project-scoped .mcp.json is included and uses a relative interpreter path, so it works wherever you cloned to. To use the broker from other projects, register it at user scope instead:

claude mcp add --scope user risk-broker -- "$(pwd)/.venv/bin/python" -m risk_broker.server
just verify-broker   # starts it over stdio, lists tools, probes every tier boundary

Create a knowledge wiki

just bootstrap-wiki ~/dev/wiki-hardware

Markdown in git, searched with ripgrep. No vector database — at one person's scale grep wins on maintenance, staleness and debuggability. Three tiers with a wall between them:

raw/       cached sources. Immutable, gitignored; raw/MANIFEST.tsv records what was fetched.
scratch/   agent drafts. Untrusted, freely writable.
wiki/      canonical. Provenance required. Gated.

The wall is the point. An agent writing unsupervised into its own authoritative knowledge base is how one early misreading becomes load-bearing fact six months later.

Optional: the write gate

A PreToolUse hook that refuses tool writes into wiki/, so canonicalisation stays a human action. Add to ~/.claude/settings.json:

{
  "hooks": {
    "PreToolUse": [{
      "matcher": "Write|Edit|MultiEdit|NotebookEdit|Bash",
      "hooks": [{ "type": "command",
                  "command": "python3 /absolute/path/to/hardware-agent/harness/kb/deny_wiki_write.py" }]
    }]
  }
}

Set HARDWARE_AGENT_WIKI if your wiki isn't at ~/dev/wiki-hardware.

Hook registrations load at session start — one added mid-session is inert until you restart. Verify with just check-gate (logic) and python3 harness/kb/check_gate.py --live (wiring).

What this does not do

Stated plainly because the repo's own subject is not overclaiming about guardrails:

  • The write gate is bar-raising, not airtight. It inspects command strings, so cd wiki/hardware && cp ../../x.md . defeats it, and anything a child process writes — via just, make, a shell script — is invisible to it.
  • Unknown tools fail open until the matcher is widened. A fail-closed path is implemented and tested but dormant; see harness/kb/deny_wiki_write.py.
  • The broker only sees what's routed through it. A direct Bash call doesn't reach it. That is why the hook exists.
  • Simulation is not hardware. Wokwi and friends diverge from real silicon.
  • Nothing here is legal or compliance advice. The ANATEL and licensing material is direction, not a substitute for an accredited lab.

The durable guarantee is the wiki's pre-commit lint gate plus a human reading the diff. Everything else is defence in depth above it.

Evidence

harness/smoke/ holds the behavioural tests, run against cold agents with no memory of the session that wrote the skills, plus what each one found. The first five cover the domain skills and have recorded results; prompts 6–9 cover the lenses and have not been run cold yet, which is the honest state of the evidence for them. REVIEW.md is the original audit of the scaffold. CHANGELOG.md says what changed and why.

Provenance

Built with Claude Code; the commit history carries Co-Authored-By trailers throughout. The knowledge wiki it manages is a separate, private repository by design — see skills/knowledge-base/references/legal-and-etiquette.md for the reasoning.

Licence

MIT — see LICENSE.

About

Risk-tiered agent system for hardware prototyping — electronics, PCBs, firmware, sourcing and manufacturability. Skills that know the domain, plus a harness that stops the agent doing irreversible things with them.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages