Skip to content

Panic at ai_orchestrator.go:612 when runner returns an unmapped HTTP status (502/503/504) #4026

Description

@NeuralStream-livepeer

Version: go-livepeer v0.9.0
Component: core/ai_orchestrator.go, ai/worker/worker.go, ai/worker/runner.gen.go
Impact: every affected request panics; the stream fails for the paying gateway. Observed on a production orchestrator at a ~5% failure rate on live-video-to-video session starts.


Summary

When an AI runner is unreachable or returns an unmapped HTTP status, Worker.LiveVideoToVideo returns (nil, nil) instead of an error. The caller checks only err == nil and then dereferences the response, causing a panic.

This is not an edge case tied to one deployment: there is currently no HTTP status a runner (or any proxy in front of it) can return to signal "busy" or "unavailable" without panicking the orchestrator. 502, 503, 504 and 429 all take this path. Only 500 is handled cleanly.

Any orchestrator running its runners behind a reverse proxy, a load balancer, or on a remote host is exposed.

Root cause

1. The generated response parser silently returns a nil payload.

ai/worker/runner.gen.go, ParseGenLiveVideoToVideoResponse:

switch {
case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200:
    // ... sets response.JSON200
case ... rsp.StatusCode == 400:
case ... rsp.StatusCode == 401:
case ... rsp.StatusCode == 422:
case ... rsp.StatusCode == 500:
}

return response, nil   // no default case: any other status falls through with all fields nil

There is no default branch. A 502 (or a 200 without a JSON Content-Type) matches nothing, so response.JSON200 stays nil and no error is produced.

2. The worker does not check the payload before returning it.

ai/worker/worker.go:

resp, err := c.Client.GenLiveVideoToVideoWithResponse(ctx, req)
if err != nil { return nil, err }
if resp.JSON400 != nil { ... }
if resp.JSON401 != nil { ... }
if resp.JSON422 != nil { ... }
if resp.JSON500 != nil { ... }

return resp.JSON200, nil   // JSON200 may be nil here

Result: (nil, nil).

3. The caller dereferences it.

core/ai_orchestrator.go L606-620:

func (orch *orchestrator) LiveVideoToVideo(ctx context.Context, requestID string, req worker.GenLiveVideoToVideoJSONRequestBody) (interface{}, error) {
    if orch.node.AIWorker != nil {
        workerResp, err := orch.node.LiveVideoToVideo(ctx, req)

        if err == nil {
            return orch.node.saveLocalAIWorkerResults(ctx, *workerResp, requestID, "application/json")  // L612 — panics when workerResp == nil
        } else {
            clog.Errorf(ctx, "Error processing with local ai worker err=%q", err)
            ...
        }
    }

err == nil is true, so the else branch that would have handled the failure is never reached.

Stack trace

goroutine 391868 [running]:
net/http.(*conn).serve.func1()
      .../net/http/server.go:1943 +0xd3
panic({0x43abe0?, 0x3ac63f0?})
      .../runtime/panic.go:783 +0x132
github.com/livepeer/go-livepeer/core.(*orchestrator).LiveVideoToVideo(...)
      /src/core/ai_orchestrator.go:612 +0x165
github.com/livepeer/go-livepeer/server.startAIServer.(*lphttp).StartLiveVideoToVideo.func12(...)
      /src/server/ai_http.go:1153 +0xf86
net/http.HandlerFunc.ServeHTTP(...)
github.com/oapi-codegen/nethttp-middleware.OapiRequestValidatorWithOptions.func1.1(...)
net/http.(*ServeMux).ServeHTTP(...)
github.com/livepeer/go-livepeer/server.(*lphttp).ServeHTTP(...)
      /src/server/rpc.go:229 +0xc6
net/http.serverHandler.ServeHTTP(...)
net/http.(*conn).serve(...)
      .../net/http/server.go:2109 +0x665
created by net/http.(*Server).Serve in goroutine 100

The panic occurs in the per-connection handler goroutine, so net/http's recover() keeps the process alive. Other concurrent streams are unaffected — but the request that triggered it dies with no response to the gateway.

Reproduction

  1. Point an orchestrator at a live-video-to-video runner through a reverse proxy (nginx or equivalent).
  2. Make the proxy return 502 — stop the runner, or saturate a connection limit so the proxy rejects before reaching any upstream.
  3. Send a POST /live-video-to-video request.
  4. The orchestrator panics at ai_orchestrator.go:612.

A 200 response without a JSON Content-Type header reproduces it identically.

Observed impact

On a production orchestrator over a 6-hour window:

  • 17 panics
  • 326 live-video-to-video session start attempts, 17 failed → 5.21% failure rate
  • Every failure was a stream that never started for a paying gateway

Correlating the orchestrator logs with the proxy's error.log showed two distinct upstream conditions producing the same panic:

  • connect() failed (111: Connection refused) — runner down (9 occurrences)
  • no live upstreams / upstream prematurely closed connection — proxy rejected before or during upstream contact (8 occurrences)

Both surface to go-livepeer as an HTTP 502.

Workaround

Making the proxy return 500 with Content-Type: application/json instead of 502 eliminates the panic entirely, since JSON500 is a handled case. Validated in production: panics dropped to zero, failures now surface as clean 500s the gateway can retry.

This is a deployment-side patch and shouldn't be the answer — it relies on rewriting a status code to lie about the failure type.

Suggested fixes

1. Add a default case to the generated parser so unmapped statuses produce an error rather than a nil payload:

default:
    return nil, fmt.Errorf("unexpected status %d from runner: %s", rsp.StatusCode, string(bodyBytes))

Since runner.gen.go is generated by oapi-codegen, this likely belongs in the OpenAPI spec or the generator configuration rather than the generated file itself.

2. Guard the return in worker.go:

if resp.JSON200 == nil {
    return nil, fmt.Errorf("unexpected empty response from runner (status %d)", resp.StatusCode())
}
return resp.JSON200, nil

3. Defensive nil check at the call site (ai_orchestrator.go:612), so a (nil, nil) contract violation anywhere downstream degrades to an error instead of a panic.

Fix 1 or 2 alone resolves the issue; 3 is cheap insurance.

Related

The same *workerResp dereference pattern appears in the other pipelines in this file (TextToImage and others). Those call synchronous endpoints where a nil payload may be less likely, but the same class of bug applies if a proxy or network layer returns an unmapped status.

Related but distinct: #3989 guards a nil in server; this is a separate nil path in core/ai_orchestrator.go.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

status: triagethis issue has not been evaluated yet

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions