Skip to content
howardc38Public

About

A development workflow for Claude Code and Codex: scope changes, verify behavior, combine parallel work, and track reviews and repairs with current evidence.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

English · 简体中文 · 繁體中文

vibeproof

Develop with a coding agent. Keep the work checkable.

vibeproof is a development workflow for Claude Code and Codex. Start with a scoped request, check the change, combine parallel work, and track reviews and repairs as your repo evolves.

You get a task record you can inspect: what was allowed to change, which checks ran, what evidence supports a repair, and which results still apply to the current code. Use those results to decide what to accept and what still needs work.

Start with one change in your repo → · Try the standalone demo · Explore the workflow

When to try it

  • Build a feature or fix a defect. Keep the requested outcome, allowed files and verification together while the agent works.
  • Combine parallel tasks. Work in separate Git worktrees, then run fresh checks on the combined code.
  • Maintain reviews and repairs over time. Track review findings through authorized repairs, and see when a review is incomplete or its evidence is out of date.

Best first fit: an existing Git repo with a real Python test suite, where you already use a coding agent and spend time checking its work.

From a request to work you can inspect

  1. Give the task a clear boundary. State the outcome and the files the agent may change. Confirm project facts used by the checks, and have the agent account for the requested work. That accounting helps you inspect delivery; it does not judge whether the request was met.
  2. Run the checks that apply. The framework identifies relevant checks for the scoped work and records their results. Supported editing and stop hooks help surface unresolved work. Your task policy distinguishes issues that block completion from issues that are reported for attention.
  3. Ask for evidence of the behavior. Run your real test suite. A review repair can be tied to a test that fails before the fix, passes after it and executes the target. Configured browser/UI tests and runtime queries can check other behavior. Useful assertions and project setup determine what is proven.
  4. Know when to check again. Earlier runs stay in the local history. Relevant code, configuration or check changes can make their results stale. Reviews are recorded against their inputs, so you can distinguish complete, partial and outdated reviews.
  5. Bring the work together and follow through. The coding-agent host coordinates tasks in worktrees, with revalidation after merging. Maintenance tracks authorized repair handoffs, can recheck a repair in its linked worktree, and reads back the original finding's state. Recurring reviews require an actual job in your host scheduler.
  6. Adapt the checks to your project. Add custom rules and validate them against fixtures. Use doctor to inspect wiring, and inspect declared risk coverage, trends, recorded cost and exported records. Monitor reviews examine changes to framework criteria and project facts; optional notices help route attention.

The host launches agents and schedules recurring work. The CLI supplies checks and recorded state; review judgment and business acceptance still need suitable tests and decisions. A delivered or acknowledged Telegram notice does not close a finding.

Full feature map · Maintenance operation · Technical reference

Use it in your repo

Follow the adoption guide →

Start with one feature or fix in a repo you already use with an agent. You provide the requested outcome, allowed files, real test command and decisions about unresolved risks. The guide walks through installation and the first task, with a pasteable agent prompt.

Full installation adds checkers, detectors, fixtures, hooks and prompts, verifies fixtures, and asks you to confirm facts about your repo. It takes longer than the short demo.

Choose your host: installation defaults to Claude Code. Use --hosts codex or --hosts both for Codex. --activate-hooks merges framework handlers while preserving unrelated settings; Codex hooks also need review and trust in the host. Task binding and permissions

Work to do Claude Code Codex
Make one scoped change /run $vibeproof-run
Coordinate tasks in separate worktrees /wave $vibeproof-wave
Review current code through applicable lenses /sweep $vibeproof-sweep
Inspect findings and coordinate maintenance /maintain $vibeproof-maintain

Demo: one change, three proof moments

This supporting example shows three parts of the workflow: missing test execution, a verified repair, and evidence that expires after another edit. It starts with a deliberately broken discount: 100 − 20 returns 120. The tests still pass.

Three moments: green tests with a wrong total; a repair verified before and after; another edit makes the old evidence STALE.

Watch the narrated demo · Run it yourself

1. Green tests. Wrong result.

The cart should total 80, but returns 120. An unrelated 2 + 2 test stays green. The ordinary Python checker reports:

the suite passed and executed none of the 1 changed file(s):
  checkout.py

That exposes missing execution evidence. It does not establish that every changed line was tested: importing a file can pass this ordinary check.

2. A fix that earns its proof.

A new regression test calls total(100, 20) and expects 80. It fails on the broken implementation. After the fix, the review checker verifies the same test fails before the fix, passes after it, and executes the target function. An unrelated green test is refused as repair evidence.

This is proof for the demonstrated repair. The test still needs a meaningful assertion; it does not prove every requirement or branch.

3. Another edit. The old proof expires.

The demo records a real checker attempt in a temporary ledger, then changes checkout.py again. The kernel reports:

ANSWERED → STALE
its subject moved: checkout.py

The old successful attempt stays in the history. It no longer counts as current evidence. STALE means revalidation is needed, not that a new bug has already been found. Restoring the exact checked source makes that evidence applicable again.

Run it yourself

You need Git and Python 3.12+, on macOS or Linux. Native Windows is not validated. No packages, API key or coding-agent subscription are needed for this demo.

git clone https://github.com/howardc38/vibeproof.git
cd vibeproof
python3 examples/first-proof/run.py

The script creates and removes a temporary repo. After cloning, it runs offline and does not install the framework into your project. Expected FAIL output is part of the demonstration; success ends with DEMO VERIFIED.

The images and videos replay measured output from the 2026-09-14 constructed example, including real checker execution and kernel evidence-state queries. Narration is synthetic and the presenter portrait is fictional. These are not a captured AI conversation or a complete installation/ship run. Source, transcript and controls

Supported checks and their limits

Claude Code and Codex adapters share the kernel; native host validation is on macOS. Host setup and limits

Structural and credential checks cover selected patterns according to language and repo facts. Test-change checks report count reductions and selected expectation/shape changes, with limited assertion-quality analysis.

Ordinary changed-file execution tracing is Python-only. Executable review repair paths include Python, Go and Node/V8, depending on the runner. Verified same-file Python function/method renames can preserve the original finding. Structural checks support Python, Go and TS/JS to different depths, with limited Rust support.

Full feature map · Technical reference · Facts format

Know what the result means

  • SHIP is the configured task decision. It does not deploy code or guarantee that every requirement was met.
  • Some findings initially report rather than block; test deletion is report-only by default. A finding on the standing repo-review task does not automatically block another task.
  • Hooks cover supported host payloads and can stand down when state is unavailable. Stop checks interrupt once per stop continuation; later independent turns can be checked again. Hooks are not a sandbox.
  • Local ledger records and hashes are not an immutable external trust service. Accepted-risk paths exist, including agent signatures.
  • Prompts, review lenses and recorded completion do not prove independent judgment. Business correctness and security still need suitable tests and human decisions.

Exact limits and exit codes · Full workflow

Try one real change

Tell us what happened: what it caught, what it got wrong, and whether you would keep it enabled for the next task. Sanitized logs are enough; private code and credentials are not needed.

Run the framework's own tests with python3 tests/run_without_silent_skips.py.

Maintaining vibeproof? Edit the canonical development repo; see contributing and release synchronization.

MIT licensed. License · Implementation specification

About

A development workflow for Claude Code and Codex: scope changes, verify behavior, combine parallel work, and track reviews and repairs with current evidence.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages