Skip to content

live runners: orchestrator keeps selling paid reservations for a runner slot stuck with a stale session #4004

Description

@rickstaa

Summary

An orchestrator running live runners can get into a state where its single runner slot is permanently held by a stale session, yet it keeps accepting paid reservations for that slot indefinitely. Every reservation is paid through the remote signer at reserve time, then fails at POST .../app/stream with 409 runner already has an active session. Observed against a third-party orchestrator on Arbitrum mainnet during sustained load testing: 559 paid-but-undeliverable reservations in one hour, while three other orchestrators in the same run served 386/386 sessions cleanly.

Environment

  • Client: livepeer-python-gateway ja/live-runner @ 89ac0f7, load test looping reserve_session -> POST /stream -> drain -> stop_runner_session
  • App: livepeer-example/flux-klein (persistent mode, capacity 1)
  • Payments: remote signer (-remoteSigner, live per-second pricing), Arbitrum mainnet
  • Orchestrator: third party, recent build (advertises the new price schema from runner: Add fixed pricing #3999), exact version unknown

Timeline / two related failure modes

Run 1 (30 min): the first session's stop failed with HTTP 404; body='runner not found', stranding the session on the capacity-1 runner. Every subsequent attempt for the rest of the run did: reserve OK -> signer payment created -> POST /stream -> 409 runner already has an active session. 284 attempts, 0 successes, 165 payments signed to the orch's recipient.

Operator restarted the runner (new runner_id), discovery showed capacity_used: 0, capacity_available: 1.

Run 2 (60 min): zero stop failures from the client this time, yet the fresh runner 409'd from the very first stream attempt to the last: 559 attempts, 0 successes, every one preceded by a successful paid reservation. So the runner's internal "active session" state diverged from the orchestrator's capacity accounting without any client-side misbehavior, and the orchestrator sold that dead slot for a full hour.

Healthy control group in the same runs: three other orchestrators (including one running stock v0.9.0 images) at 100% success across 386 sessions, so client-side stop handling is not the trigger.

Client-side evidence

reserve/NoRunnerAvailableError: All runners failed (1 tried):
  https://<orch>/apps/runner_dschhn5u/session/159bc96c/app/stream:
  HTTP 409; body='runner already has an active session'

(x559, each after a successful /generate-live-payment on the signer)

Run 1 leak precursor:

stop/LivepeerGatewayError: HTTP empty POST error: HTTP 404; body='runner not found'

Expected behavior / suggestions

  1. Reservation should verify the slot is actually deliverable before charging: if the runner reports an active session for its only slot, the reserve should fail (402/503), not collect a payment for a stream that can only 409.
  2. Stale session TTL / reaping: a runner session with no traffic (no stream started, or stream gone) should be reaped after a timeout so a leaked or desynced session cannot hold a slot forever. The load-test docs already call this out ("no idle TTL").
  3. On runner not found at stop time, the orchestrator should release its own capacity accounting for that session rather than leaving it to diverge.

Happy to provide full load-test logs (951-session run summary, per-session verbose log) on request.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    status: triagethis issue has not been evaluated yet

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions