Skip to content

Compress feedback bench scenarios into typed intervals - #286

Merged
zhoubot merged 1 commit into
mainfrom
codex/feedback-scenarios
Oct 10, 2026
Merged

zhoubot merged 1 commit into
mainfrom
codex/feedback-scenarios

Conversation

@zhoubot

@zhoubot zhoubot commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

The feedback system bench repeated eight large conditional trees to select stimulus and expected outputs. Replace them with 1,051 typed, inclusive scenario intervals selected through the existing Table.first query. The source shrinks from 13,103 lines / 639,219 bytes to 1,116 lines / 90,403 bytes.

Preserve the full 64-bit phase domain and wrap, all 3,763 cycles, the constant tail, four guarded assertions, four unconditional logs and explicit DUT/Item construction. DUT sources, native/RTL/four-state oracles, storage, configuration and timeouts remain unchanged. No compiler changes or new API. Publish a fresh verified module receipt, update navigation, and correct stale language limitations following #284.

Validation:

  • Independent proof: 1,051 endpoints, 8,408 interval payload values, all 30,104 active field values, full known-u64 tail/wrap and exact structure; 19 malicious mutations rejected.
  • Standalone tests 3/3: module, full system and four-state. System completes 7,526 epochs and 30,104 observations across native workers 1/2 and Verilator. Original module oracles retain 7,527 known and 7,607 four-state Work rows, including genuine Icarus, reset/discard, token conservation and terminal failures.
  • Complete original and candidate native workers 1/2 and RTL observations agree. Only changed AST registration/site metadata is excluded from the cross-source comparison; events, values, instances, epochs and order remain checked.
  • Source-import and transformed artifacts reproduce both published source units byte-for-byte. Public execution and measured C++/RTL outputs consume identical final IR and complete generated-file inventories.
  • Related Python tests 113 passed; changed-file pre-commit and strict documentation passed.

Performance is a tradeoff, not a simulation speedup. Serial same-profile measurements reduce bench compilation 21.49→3.79 s, link 14.24→6.06 s, C++ emission 13.38→5.96 s and C++ build 4.11→3.31 s. Two alternating paired warm native runs have medians 30.56→31.59 s (+3.4%). Verilog build increases 2.35→3.64 s and runtime 0.69→0.89 s. Final IR falls 62.37→25.82 MB; generated C++ 10.82→7.89 MB and RTL 2.36→1.51 MB. These are limited measurements on one machine with LLVM 22 and -O0; README records the full table. An earlier baseline timing outlier was not reproduced and is not used to claim speedup.

Full catalog/nightly and platform matrices were not run. The known-phase proof does not claim injected X/Z phase equivalence or execution of 64-bit rollover. Remaining migration and scaling work stay in #272 and #265. Intermediate experiments, independent tools/reviews and raw candidate-bound evidence remain ignored under docs/gates/logs/feedback-bench-20261010/; tests do not depend on those files.

@zhoubot
zhoubot requested a review from xiekunpeng as a code owner October 10, 2026 03:07
@zhoubot

zhoubot commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

Final candidate 5bdbe00f0d1fe19b53ce1188984e8d2428ed6c01: independent code review APPROVE, architecture CLEAR, and performance-report watch resolved. Both required GitHub checks passed on this exact head.

Independent preservation covers the complete known-u64 phase domain, 30,104 active field values, 1,051 intervals and 19 rejected mutations. Standalone module/system/four-state tests pass 3/3; related Python tests 113/113, complete baseline/candidate traces, retained compiler stages and generated-output bindings pass. Final README uses matched serial phase measurements and discloses the measured 3.4% native runtime increase and small Verilog build/runtime increases. No simulation speedup, full nightly or executed-u64-rollover claim.

Author, independent test, architecture and code-review instances are separate. Code review used gpt-6.1-sol/high; architecture used the real architect preset gpt-6-astra/xhigh. Candidate-bound reviews and evidence remain in ignored local logs.

@zhoubot
zhoubot merged commit 716cfdf into main Oct 10, 2026
2 checks passed
@zhoubot
zhoubot deleted the codex/feedback-scenarios branch October 10, 2026 03:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant