runner: a run is resumed at any step of its queue trace, by replaying a firing journal - #41
Merged
Merged
Conversation
… a firing journal `wfpy run flow.py --resume-from wf-out/<run> --at-step N` carries a run on from the step the IDE's stepper shows, not only from where it failed. Every firing that completes now appends to `run.wf-journal.jsonl` what it took (by digest), what it put and the actor's state after; a resume at step N starts the run again from its inputs and replays each actor's firings up to the step instead of running them, then carries on live. Replaying per actor rather than restoring the state at step N is what makes it sound with parallel workers: a step is recorded while other firings are in flight, so the state "after step N" can hold half of one, but each actor sees the same sequence of tokens however firings interleave, so its k-th firing can be replayed whenever its inputs come. Which firings a step covers is chosen per actor too -- one journal cut-off races the same way. A replayed firing checks its inputs against the original's; when they differ, replay stops (`replayStopped` in the run record) and the run carries on live. The queue trace is written when a run fails too, and its steps carry `journalSeq` and `replayed`. Phase 3 of docs/proposals/resume.md, whose section is rewritten to the design built. Claude-Session: https://claude.ai/code/session_015VK7fH1c4aKbexnq2QcuKU
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 3 of
docs/proposals/resume.md, after #38–#40.wfpy run flow.py --resume-from wf-out/<run> --at-step N(orrun(..., resume_from=..., at_step=N)) carries a run on from the step the IDE's stepper shows, not only from where it failed.The journal. Every firing that completes appends one line to
run.wf-journal.jsonl: the actor's path (child/leafinside a nested workflow), the tokens it took per queue (count and digest), the tokens it put, and its state afterwards (task fields, FSM state, chat history, context if it changed). It is on whenever the queue trace is (the default).The resume. It starts again from the run's inputs (from the journal header) and replays each actor's firings up to step N instead of running them, then carries on live.
Why replay, not the state at step N. The proposal first suggested rebuilding the state after step N from per-step deltas. That is unsound with parallel workers: a step is recorded while other firings are in flight, and a consumer can even be recorded before the firing that produced its input. Replay per actor is sound under any scheduling, because in a dataflow network each actor sees the same token sequence however firings interleave. Which firings a step covers is also chosen per actor (
entries_at_step), since one journal cut-off races the same way. The first parallel test run showed exactly that race. The proposal section is rewritten to describe the design that was built.Divergence. A replayed firing checks its inputs' digest against the original's. When they differ (something upstream ran live and answered differently), replay stops,
replayStoppedgoes into the run record, and the run carries on live. Replay also stops onnotReplayablefirings and at quiescence with firings left to replay.Also:
journalSeqandreplayed(additive; the trace stays version 1 for the IDE's reader);--at-step, because it has no inputs to start from.Tests:
tests/test_resume_at_step.py(11). It resumes at every step of a run, with one worker and with parallel ones, and checks which firings ran against the trace. It also covers a failed run at its last step, a loop over a child workflow at every step, a resumed run resumed again, replayed answers kept, replay stopping when inputs differ, the refusals, and the CLI. Full suite: 690 passed, 9 skipped. The resume tests passed 15 repeated runs.https://claude.ai/code/session_015VK7fH1c4aKbexnq2QcuKU