Designed and engineered by Ƴunior Ƥortal (ƳƤ)
A production-oriented agentic quality-engineering control system where Claude can plan, investigate, and adapt while deterministic policy governs authority, controlled tools produce provenance-bound evidence, and subject-bound validation retains terminal authority.
Documentation · Architecture · Runtime Lifecycle · Result Contract · Runtime Control · Security · CI/CD · Setup
Important
The model is a reasoner, not the test oracle. Reasoning is advisory. Observations are provenance-bound. Authority is deterministic. Success requires closure. Claude may interpret evidence and propose actions; it cannot convert untrusted context into authority, self-approve a side effect, weaken the validation contract, or certify terminal success.
| Surface | Framework contract |
|---|---|
| Runtime | Python 3.11+ · claude-agent-sdk==0.2.163 · default model identifier claude-sonnet-5 |
| Reasoning | LLM planner/diagnostician; never test oracle, authorization engine, or terminal authority |
| Controlled tools | 18 least-privilege, purpose-built in-process QA tools; no generic autonomous Bash/Edit/Write/Web authority |
| Trusted Skills | exactly five allowlisted Claude Skills |
| Live mutation | Python/pytest-backed test mutation only; exact-path/revision closure is required before persistence |
| Evidence | run-confined state, immutable identities, manifests, hashes, artifacts, lineage, hash-chained journal, optional regulated audit chain |
| Network | exact host allowlists, read-only API default, browser routing controls, independent k6 egress prerequisite |
| External MCP | explicitly approved vendor integrations; provider identity never grants blanket authority and returned content remains untrusted evidence |
| Evaluation | deterministic tests, 34-scenario primary adversarial corpus, separately executed H-series readiness corpus, frozen safety thresholds |
| Merge governance | ordinary CI is development evidence; protected merge authority is independently admitted and identity-bound |
flowchart LR
accTitle: Evidence-first agentic QA trust and authority architecture
accDescr: An authorized objective reaches advisory Claude reasoning. Every action request passes through deterministic policy. Internal tools and explicitly approved provider actions produce evidence. Target and provider content remain untrusted. Subject-bound deterministic validation derives the structured terminal result.
O[Authorized objective]
C[Claude Agent SDK]
subgraph CONTROL[Trusted deterministic control plane]
direction LR
P[Policy + permissions + hooks] --> Q[18 narrow QA tools]
Q --> E[Evidence + artifact store]
E --> I[Deterministic QA intelligence]
I --> V[Subject-bound validation]
V --> R[Structured runtime result]
end
subgraph TARGET[Untrusted target / SUT]
S[Repository + application + test environment]
end
subgraph PROVIDERS[Approved providers · returned content untrusted]
direction TB
G[GitHub official MCP]
A[Atlassian Rovo MCP]
end
O --> C
C -->|action request| P
P -->|authorize internal| Q
P -->|authorize provider| G
P -->|authorize provider| A
Q <--> S
G -->|provider result| E
A -->|provider result| E
classDef neutral fill:#f6f8fa,stroke:#57606a,color:#24292f,stroke-width:1.5px
classDef advisory fill:#fbefff,stroke:#8250df,color:#24292f,stroke-width:2px
classDef authority fill:#ddf4ff,stroke:#0969da,color:#24292f,stroke-width:2px
classDef evidence fill:#dafbe1,stroke:#1a7f37,color:#24292f,stroke-width:2px
classDef untrusted fill:#ffebe9,stroke:#cf222e,color:#24292f,stroke-width:2px,stroke-dasharray:5 3
classDef terminal fill:#dafbe1,stroke:#1a7f37,color:#24292f,stroke-width:3px
class O neutral
class C advisory
class P,Q,I authority
class E,V evidence
class R terminal
class S,G,A untrusted
style CONTROL stroke:#0969da,stroke-width:2px,stroke-dasharray:6 4
style TARGET stroke:#cf222e,stroke-width:2px,stroke-dasharray:6 4
style PROVIDERS stroke:#cf222e,stroke-width:2px,stroke-dasharray:6 4
linkStyle default stroke:#57606a,stroke-width:1.5px
Diagram key: purple = advisory reasoning · blue = deterministic authority · green = evidence/validation · red dashed = untrusted evidence source. Color is never the only signal.
The detailed request sequence, evidence-first runtime flow, and transactional mutation/crash-recovery state machine have moved to Runtime Lifecycle. Deeper trust-zone and component ownership lives in Architecture.
Claude reasons.
Deterministic policy authorizes.
Controlled tools observe and act.
Evidence carries provenance.
Subject-bound validation owns terminal truth.
Four contracts keep those responsibilities separate:
| Contract | Question | Authority |
|---|---|---|
| Authority | What may the agent do? | deterministic policy, hooks, permissions, budgets |
| Evidence | What was actually observed? | controlled tools, manifests, artifacts, hashes, provider responses |
| Mutation | When may automated code changes persist? | path ownership, rollback transaction, revision-bound validation closure |
| Outcome | What may be called successful? | deterministic validation lineage and terminal evaluation |
The system is intentionally fail-closed: uncertainty reduces authority. Missing evidence does not become green, ambiguous ownership does not become permission, and incomplete validation does not become success.
python3.11 -m venv .venv
source .venv/bin/activate
make install
ai-qa doctor
ai-qa demomake install selects the matching committed interpreter-specific lock, enforces package hashes, installs without dependency resolution, and runs pip check. Windows and deliberate lock-update procedures are in Setup and Supply-Chain Integrity.
export ANTHROPIC_API_KEY='...'
export AI_QA_CONTROL_ROOT='/path/to/ai-qa-automation'
export AI_QA_ARTIFACT_ROOT='/path/to/ai-qa-artifacts'
export AI_QA_BASE_REF='origin/main'
ai-qa agent \
--control-root "$AI_QA_CONTROL_ROOT" \
--workspace /path/to/isolated/sut-worktree \
'Investigate the failing checkout test. Do not modify tests unless evidence proves a test defect.'The control root, artifact root, and target workspace are separate trust domains. Exact configuration/credential policy lives in Setup.
Model capability is deliberately broader than runtime authority. The live path removes generic write/shell/web authority, uses strict MCP configuration, requires deterministic authorization for controlled tools/provider actions, fails closed when approval is unavailable, and keeps independent turn/tool/network/mutation/repetition/time/cost budgets and per-tool circuits.
The runtime exposes 18 least-privilege, purpose-built in-process QA tools across repository inspection, pytest evidence, API/browser observation, deterministic failure intelligence, source/coverage context, test design/regression/test-quality analysis, bounded generation proposals, locator-only self-healing, contracts, CI, mobile, and controlled performance execution. It also loads exactly five allowlisted Claude Skills from the trusted control root.
Library capability is not runtime authority. Reusable patch logic can understand multiple test syntaxes, while live autonomous mutation remains intentionally Python/pytest-backed because that path owns the execution and revision-binding contracts. There is no generic existing-test rewrite tool in the live agent surface.
See Runtime Control, Skills, and Production Readiness.
Terminal outcomes are distinct from individual validations and provider health. SUCCESS means every active deterministic gate required by the objective/revision is closed; FAILURE, BLOCKED, POLICY_DENIED, INFRASTRUCTURE_FAILURE, BUDGET_EXCEEDED, CANCELLED, and NOT_VERIFIED preserve materially different failure/uncertainty states.
A model result subtype of success can never produce framework SUCCESS by itself. For a changed test revision to persist, the mutation path requires exact-path patch-safety, exact-path targeted execution diagnostics, independently trusted targeted executed-test semantics, full-regression diagnostics, independently trusted full-regression semantics, no conflicting validation, and durable transaction closure at the same revision.
The current live pytest adapter deliberately does not promote target-controlled collection/output/report/exit data into independent positive semantic authority. That keeps autonomous mutation fail-closed rather than manufacturing green.
See the authoritative Runtime Result Contract, the visual lifecycle in Runtime Lifecycle, and the implementation mechanics in Runtime Control.
The framework uses AI where interpretation helps while keeping acceptance deterministic:
- Failure investigation: evidence-weighted classification distinguishes application, automation, locator/UI-contract, data, timing, environment, dependency, auth/configuration, performance, and insufficient-evidence classes.
- Self-healing: restricted to semantic locator maintenance; uniqueness alone is insufficient, and live mutation is further subject to exact-path/revision closure.
- Test generation: coverage-aware, provenance-bound planning/proposal; unknown product behavior is not invented and current generic proposals do not claim a coverage gap is closed.
- Change intelligence: merge-base-aware committed/dirty/untracked change analysis, ownership/risk/test-impact context, and conservative OpenAPI/Swagger drift evidence.
Deep dives: Change Intelligence, Contract Drift Boundary, and Technical Walkthrough.
| Surface | Deterministic boundary |
|---|---|
| API | exact host allowlist; read-only default; redirects/proxy inheritance disabled; bounded sanitized observation |
| Browser | allowlisted navigation/subresources/WebSockets; service workers disabled for evidence context; final URL rechecked; bounded diagnostics/screenshots |
| Performance / k6 | production-like targets denied; script restrictions; bounded execution/output; independently enforced external egress required for every run |
| Mutation | isolated Git worktree, lease/fingerprint, non-symlink owned path, rollback snapshot, one unresolved transaction, exact revision closure |
| Recovery | prior run/journal/target/rollback/backup/fingerprint/ownership revalidated before automatic stale recovery |
| External MCP | explicit vendor integrations only; conservative action authorization; provider output remains untrusted evidence |
| Persistence | confined run roots, bounded state/runtime/manifest/journal/artifacts, immutable evidence identities, hash verification, symlink rejection |
Application-level controls are defense in depth, not substitutes for deployment isolation, egress, identity, secret management, device/provider configuration, retention, or repository settings. Detailed boundaries live in Security, Threat Model, API Observation Boundary, Browser Validation, Pytest Execution Isolation, and MCP.
Each run receives a confined durable evidence surface under artifacts/<run_id>/ with canonical QA state, separate process-control state, evidence manifests, content-addressed artifacts, an append-only SHA-256 hash-chained journal, validation lineage, provenance, optional provider usage/cost, and unsigned run-integrity attestations.
ai-qa recover artifacts/run-<id>
ai-qa lineage artifacts/run-<id>
ai-qa attest artifacts/run-<id>
ai-qa contract-diff --baseline old-openapi.yaml --current new-openapi.yamlAn attestation is deliberately unsigned: content-addressed integrity proves byte relationships, not actor identity, notarization, compliance certification, trusted timestamp, business correctness, or test success. See Traceability and Verification Boundaries.
The framework is evaluated as software, not by persuasive prose. Repository qualification includes deterministic unit/integration/policy/security tests, a fixed 34-scenario primary adversarial corpus, a repository-visible but separately executed H-series readiness corpus, frozen threshold schema/hard-safety limits, and separated credentialed/model/browser execution boundaries.
make quality
make test
make eval
make security
make verify-local
make holdoutThe H-series corpus is execution-separated but not blind/independent evidence because its fixtures remain repository-visible. Frozen hard-safety thresholds are policy artifacts rather than post-hoc knobs. See Evaluation Strategy.
.github/workflows/ci.yml provides read-only, secret-free deterministic evidence for pull requests, pushes to main, and merge groups, including Required PR Gate. That candidate-controlled workflow is development evidence, not protected merge authority.
Trusted default-branch admission in trusted-pr-auto.yml has two automatic classes. Owner-routine PRs require zero protected-root drift. Four finite governed-bot lanes—canonical Dependabot GitHub Actions updates, deterministic dependency promotions, CodeQL auto-heal repairs, and protected security remediations—may cross only their explicitly reviewed roots after exact bot/branch/source provenance, full prospective-merge validation, governed-head CodeQL/tree binding where required, and terminal reproof. The dedicated GitHub App then publishes an exact PR/base/head/merge-bound Trusted PR Gate status. Dependency and security merge paths additionally isolate the exact portyu9 owner-review identity from downstream mutation authority and revalidate the App gate plus exact subject immediately before merge. Explicit owner-protected maintenance uses the accepted-main GitHub-native lane with an immutable exact-subject authorization, reviewed liveness, full trusted validation, and the same dedicated-App terminal publication. Unrecognized protected subjects remain deny-by-default; there is no external cloud fallback.
Manual H-series/model validation and release-candidate preparation are separately scoped and do not silently acquire publishing or protected-merge authority. See CI/CD, Trusted PR Control Plane, Release Candidate, and Supply Chain.
.
├── .claude/
├── .github/
├── artifacts/
├── docs/
├── evals/
├── examples/
├── performance/
├── requirements/
├── scripts/
├── src/
└── tests/
Only top-level ownership boundaries are shown here; the technical docs own file-level structure.
Start with the documentation hub. Key review paths:
| Topic | Document |
|---|---|
| Architectural authority/trust | Architecture |
| Runtime request/mutation/recovery flows | Runtime Lifecycle |
| Terminal/validation/provider semantics | Runtime Result Contract |
| Transaction/recovery implementation | Runtime Control |
| Security/threat boundaries | Security · Threat Model |
| Trusted setup/credentials | Setup |
| CI and protected merge governance | CI/CD · Trusted PR Control Plane |
| Change/regression intelligence | Change Intelligence |
| Evaluation/readiness governance | Evaluation |
| Evidence lineage/attestation | Traceability |
| External MCP | MCP |
| Production control model | Production Readiness |
| Explicit non-claims | Limitations |
| End-to-end implementation review | Technical Walkthrough |
The repository does not claim that application flags create a firewall/process sandbox, hashes authenticate publishers, reproducible artifacts create signed provenance, provider configuration proves provider availability, one browser/API/load/mobile observation proves target correctness, ordinary PR CI creates protected merge authority, or model reasoning can replace deterministic controls.
Those boundaries are intentional and are documented in Limitations and Verification Boundaries.
MIT — see LICENSE.