Visual similarity alone cannot tell whether the program-specified action and end state actually occur.
A benchmark for programmable world models: 170 programmatically constructed episodes and 600 proxy videos, each logged as a replayable world record of entity states and timestamped events, so generated videos can be checked against the observable consequences of program execution.
Programmable world models separate executable dynamics from visual generation, but their visual adherence to explicit rules and interactions remains insufficiently evaluated. PROWBench renders synchronized views and proxy representations (e.g., coarse 3D, bounding boxes) from engine-recorded world records across first- and third-person perspectives, and evaluates entity control, long-horizon memory, and — with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate — whether timestamped events are visually realized on the prescribed timeline.
- [2026-10-02] Project page released.
- Project page
- Paper — PDF
- Benchmark data
- Evaluation code
Coming soon.
