Skip to content

runner: a firing that fails gives back the tokens it took - #38

Merged
endrix merged 1 commit into
mainfrom
feat/atomic-firings
Oct 1, 2026
Merged

endrix merged 1 commit into
mainfrom
feat/atomic-firings

Conversation

@endrix

@endrix endrix commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Phase 1 of docs/proposals/resume.md (#37).

Every actor kind took its inputs before it ran: an action after its guard, a tool and an agent on entry. A firing that raised had therefore lost them, and no checkpoint saved after the failure could bring them back.

How: every Queue.dequeue() / try_dequeue() made during a firing is logged per thread (_atomic_firing, around _step_actor). When the firing raises, each token goes back to the head of its queue in reverse order, so the queue is as it was. This is safe without holding the queues: one actor never fires twice at once, and each queue has one consumer. A producer appending to the tail is unaffected. The log is what phase 3 will record as a step's consumed.

Not wrapped: a nested workflow and an if / loop. One firing of theirs runs other actors' firings to completion, each atomic on its own. Giving back theirs too would hand a token to a firing that had already consumed it. A firing that succeeded keeps what it took, whatever fails after it.

Not restored: the actor's own state (self._done = True before a raise). That is the checkpoint's job, in phase 2.

Tests: tests/test_atomic_firings.py (8). It covers an action (single and multi-token), a tool, an agent, a failure inside a nested workflow (the token is held once, by the child), a failure inside a loop body after an earlier body firing succeeded, and the queue-level semantics. Without the give-back, 6 of them fail. Full suite: 634 passed, 9 skipped.

https://claude.ai/code/session_015VK7fH1c4aKbexnq2QcuKU

Every actor kind took its inputs before it ran -- an action after its
guard, a tool and an agent on entry -- so a firing that raised had lost
them, and no checkpoint saved after the failure could bring them back.
Now every dequeue made during a firing is logged per thread, and when
the firing raises each token goes back to the head of its queue, in
order, before the error propagates: the run stops where the failed
firing had not happened. This is phase 1 of docs/proposals/resume.md,
and the log is what a later phase records as a step's `consumed`.

A nested workflow and an if / loop are not wrapped. One firing of theirs
runs other actors' firings to completion, each atomic on its own, so
giving back theirs too would hand a token to a firing that had already
consumed it. A firing that succeeded keeps what it took, whatever fails
after it.

Claude-Session: https://claude.ai/code/session_015VK7fH1c4aKbexnq2QcuKU
@endrix
endrix merged commit 6774586 into main Oct 1, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant