You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
An orchestrator running live runners can get into a state where its single runner slot is permanently held by a stale session, yet it keeps accepting paid reservations for that slot indefinitely. Every reservation is paid through the remote signer at reserve time, then fails at POST .../app/stream with 409 runner already has an active session. Observed against a third-party orchestrator on Arbitrum mainnet during sustained load testing: 559 paid-but-undeliverable reservations in one hour, while three other orchestrators in the same run served 386/386 sessions cleanly.
Environment
Client: livepeer-python-gateway ja/live-runner @ 89ac0f7, load test looping reserve_session -> POST /stream -> drain -> stop_runner_session
Payments: remote signer (-remoteSigner, live per-second pricing), Arbitrum mainnet
Orchestrator: third party, recent build (advertises the new price schema from runner: Add fixed pricing #3999), exact version unknown
Timeline / two related failure modes
Run 1 (30 min): the first session's stop failed with HTTP 404; body='runner not found', stranding the session on the capacity-1 runner. Every subsequent attempt for the rest of the run did: reserve OK -> signer payment created -> POST /stream -> 409 runner already has an active session. 284 attempts, 0 successes, 165 payments signed to the orch's recipient.
Operator restarted the runner (new runner_id), discovery showed capacity_used: 0, capacity_available: 1.
Run 2 (60 min): zero stop failures from the client this time, yet the fresh runner 409'd from the very first stream attempt to the last: 559 attempts, 0 successes, every one preceded by a successful paid reservation. So the runner's internal "active session" state diverged from the orchestrator's capacity accounting without any client-side misbehavior, and the orchestrator sold that dead slot for a full hour.
Healthy control group in the same runs: three other orchestrators (including one running stock v0.9.0 images) at 100% success across 386 sessions, so client-side stop handling is not the trigger.
Client-side evidence
reserve/NoRunnerAvailableError: All runners failed (1 tried):
https://<orch>/apps/runner_dschhn5u/session/159bc96c/app/stream:
HTTP 409; body='runner already has an active session'
(x559, each after a successful /generate-live-payment on the signer)
Run 1 leak precursor:
stop/LivepeerGatewayError: HTTP empty POST error: HTTP 404; body='runner not found'
Expected behavior / suggestions
Reservation should verify the slot is actually deliverable before charging: if the runner reports an active session for its only slot, the reserve should fail (402/503), not collect a payment for a stream that can only 409.
Stale session TTL / reaping: a runner session with no traffic (no stream started, or stream gone) should be reaped after a timeout so a leaked or desynced session cannot hold a slot forever. The load-test docs already call this out ("no idle TTL").
On runner not found at stop time, the orchestrator should release its own capacity accounting for that session rather than leaving it to diverge.
Happy to provide full load-test logs (951-session run summary, per-session verbose log) on request.
Summary
An orchestrator running live runners can get into a state where its single runner slot is permanently held by a stale session, yet it keeps accepting paid reservations for that slot indefinitely. Every reservation is paid through the remote signer at reserve time, then fails at
POST .../app/streamwith409 runner already has an active session. Observed against a third-party orchestrator on Arbitrum mainnet during sustained load testing: 559 paid-but-undeliverable reservations in one hour, while three other orchestrators in the same run served 386/386 sessions cleanly.Environment
ja/live-runner@ 89ac0f7, load test loopingreserve_session -> POST /stream -> drain -> stop_runner_sessionlivepeer-example/flux-klein(persistent mode, capacity 1)-remoteSigner, live per-second pricing), Arbitrum mainnetTimeline / two related failure modes
Run 1 (30 min): the first session's stop failed with
HTTP 404; body='runner not found', stranding the session on the capacity-1 runner. Every subsequent attempt for the rest of the run did: reserve OK -> signer payment created ->POST /stream->409 runner already has an active session. 284 attempts, 0 successes, 165 payments signed to the orch's recipient.Operator restarted the runner (new runner_id), discovery showed
capacity_used: 0, capacity_available: 1.Run 2 (60 min): zero stop failures from the client this time, yet the fresh runner 409'd from the very first stream attempt to the last: 559 attempts, 0 successes, every one preceded by a successful paid reservation. So the runner's internal "active session" state diverged from the orchestrator's capacity accounting without any client-side misbehavior, and the orchestrator sold that dead slot for a full hour.
Healthy control group in the same runs: three other orchestrators (including one running stock
v0.9.0images) at 100% success across 386 sessions, so client-side stop handling is not the trigger.Client-side evidence
(x559, each after a successful
/generate-live-paymenton the signer)Run 1 leak precursor:
Expected behavior / suggestions
runner not foundat stop time, the orchestrator should release its own capacity accounting for that session rather than leaving it to diverge.Happy to provide full load-test logs (951-session run summary, per-session verbose log) on request.