Skip to content
AlayaLabPublic

About

PROWBench: evaluating whether video models faithfully render what a program specifies.

Resources

Stars

32 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Do Video Models Render
What the Program Specifies?

Alaya Lab

PROWBench overview

Visual similarity alone cannot tell whether the program-specified action and end state actually occur.

A benchmark for programmable world models: 170 programmatically constructed episodes and 600 proxy videos, each logged as a replayable world record of entity states and timestamped events, so generated videos can be checked against the observable consequences of program execution.

Programmable world models separate executable dynamics from visual generation, but their visual adherence to explicit rules and interactions remains insufficiently evaluated. PROWBench renders synchronized views and proxy representations (e.g., coarse 3D, bounding boxes) from engine-recorded world records across first- and third-person perspectives, and evaluates entity control, long-horizon memory, and — with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate — whether timestamped events are visually realized on the prescribed timeline.

📰 News

🚀 Release Roadmap

  • Project page
  • Paper — PDF
  • Benchmark data
  • Evaluation code

Citation

Coming soon.

About

PROWBench: evaluating whether video models faithfully render what a program specifies.

Resources

Stars

32 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors