Version: go-livepeer v0.9.0
Component: core/ai_orchestrator.go, ai/worker/worker.go, ai/worker/runner.gen.go
Impact: every affected request panics; the stream fails for the paying gateway. Observed on a production orchestrator at a ~5% failure rate on live-video-to-video session starts.
Summary
When an AI runner is unreachable or returns an unmapped HTTP status, Worker.LiveVideoToVideo returns (nil, nil) instead of an error. The caller checks only err == nil and then dereferences the response, causing a panic.
This is not an edge case tied to one deployment: there is currently no HTTP status a runner (or any proxy in front of it) can return to signal "busy" or "unavailable" without panicking the orchestrator. 502, 503, 504 and 429 all take this path. Only 500 is handled cleanly.
Any orchestrator running its runners behind a reverse proxy, a load balancer, or on a remote host is exposed.
Root cause
1. The generated response parser silently returns a nil payload.
ai/worker/runner.gen.go, ParseGenLiveVideoToVideoResponse:
switch {
case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200:
// ... sets response.JSON200
case ... rsp.StatusCode == 400:
case ... rsp.StatusCode == 401:
case ... rsp.StatusCode == 422:
case ... rsp.StatusCode == 500:
}
return response, nil // no default case: any other status falls through with all fields nil
There is no default branch. A 502 (or a 200 without a JSON Content-Type) matches nothing, so response.JSON200 stays nil and no error is produced.
2. The worker does not check the payload before returning it.
ai/worker/worker.go:
resp, err := c.Client.GenLiveVideoToVideoWithResponse(ctx, req)
if err != nil { return nil, err }
if resp.JSON400 != nil { ... }
if resp.JSON401 != nil { ... }
if resp.JSON422 != nil { ... }
if resp.JSON500 != nil { ... }
return resp.JSON200, nil // JSON200 may be nil here
Result: (nil, nil).
3. The caller dereferences it.
core/ai_orchestrator.go L606-620:
func (orch *orchestrator) LiveVideoToVideo(ctx context.Context, requestID string, req worker.GenLiveVideoToVideoJSONRequestBody) (interface{}, error) {
if orch.node.AIWorker != nil {
workerResp, err := orch.node.LiveVideoToVideo(ctx, req)
if err == nil {
return orch.node.saveLocalAIWorkerResults(ctx, *workerResp, requestID, "application/json") // L612 — panics when workerResp == nil
} else {
clog.Errorf(ctx, "Error processing with local ai worker err=%q", err)
...
}
}
err == nil is true, so the else branch that would have handled the failure is never reached.
Stack trace
goroutine 391868 [running]:
net/http.(*conn).serve.func1()
.../net/http/server.go:1943 +0xd3
panic({0x43abe0?, 0x3ac63f0?})
.../runtime/panic.go:783 +0x132
github.com/livepeer/go-livepeer/core.(*orchestrator).LiveVideoToVideo(...)
/src/core/ai_orchestrator.go:612 +0x165
github.com/livepeer/go-livepeer/server.startAIServer.(*lphttp).StartLiveVideoToVideo.func12(...)
/src/server/ai_http.go:1153 +0xf86
net/http.HandlerFunc.ServeHTTP(...)
github.com/oapi-codegen/nethttp-middleware.OapiRequestValidatorWithOptions.func1.1(...)
net/http.(*ServeMux).ServeHTTP(...)
github.com/livepeer/go-livepeer/server.(*lphttp).ServeHTTP(...)
/src/server/rpc.go:229 +0xc6
net/http.serverHandler.ServeHTTP(...)
net/http.(*conn).serve(...)
.../net/http/server.go:2109 +0x665
created by net/http.(*Server).Serve in goroutine 100
The panic occurs in the per-connection handler goroutine, so net/http's recover() keeps the process alive. Other concurrent streams are unaffected — but the request that triggered it dies with no response to the gateway.
Reproduction
- Point an orchestrator at a live-video-to-video runner through a reverse proxy (nginx or equivalent).
- Make the proxy return
502 — stop the runner, or saturate a connection limit so the proxy rejects before reaching any upstream.
- Send a
POST /live-video-to-video request.
- The orchestrator panics at
ai_orchestrator.go:612.
A 200 response without a JSON Content-Type header reproduces it identically.
Observed impact
On a production orchestrator over a 6-hour window:
- 17 panics
- 326 live-video-to-video session start attempts, 17 failed → 5.21% failure rate
- Every failure was a stream that never started for a paying gateway
Correlating the orchestrator logs with the proxy's error.log showed two distinct upstream conditions producing the same panic:
connect() failed (111: Connection refused) — runner down (9 occurrences)
no live upstreams / upstream prematurely closed connection — proxy rejected before or during upstream contact (8 occurrences)
Both surface to go-livepeer as an HTTP 502.
Workaround
Making the proxy return 500 with Content-Type: application/json instead of 502 eliminates the panic entirely, since JSON500 is a handled case. Validated in production: panics dropped to zero, failures now surface as clean 500s the gateway can retry.
This is a deployment-side patch and shouldn't be the answer — it relies on rewriting a status code to lie about the failure type.
Suggested fixes
1. Add a default case to the generated parser so unmapped statuses produce an error rather than a nil payload:
default:
return nil, fmt.Errorf("unexpected status %d from runner: %s", rsp.StatusCode, string(bodyBytes))
Since runner.gen.go is generated by oapi-codegen, this likely belongs in the OpenAPI spec or the generator configuration rather than the generated file itself.
2. Guard the return in worker.go:
if resp.JSON200 == nil {
return nil, fmt.Errorf("unexpected empty response from runner (status %d)", resp.StatusCode())
}
return resp.JSON200, nil
3. Defensive nil check at the call site (ai_orchestrator.go:612), so a (nil, nil) contract violation anywhere downstream degrades to an error instead of a panic.
Fix 1 or 2 alone resolves the issue; 3 is cheap insurance.
Related
The same *workerResp dereference pattern appears in the other pipelines in this file (TextToImage and others). Those call synchronous endpoints where a nil payload may be less likely, but the same class of bug applies if a proxy or network layer returns an unmapped status.
Related but distinct: #3989 guards a nil in server; this is a separate nil path in core/ai_orchestrator.go.
Version: go-livepeer v0.9.0
Component:
core/ai_orchestrator.go,ai/worker/worker.go,ai/worker/runner.gen.goImpact: every affected request panics; the stream fails for the paying gateway. Observed on a production orchestrator at a ~5% failure rate on live-video-to-video session starts.
Summary
When an AI runner is unreachable or returns an unmapped HTTP status,
Worker.LiveVideoToVideoreturns(nil, nil)instead of an error. The caller checks onlyerr == niland then dereferences the response, causing a panic.This is not an edge case tied to one deployment: there is currently no HTTP status a runner (or any proxy in front of it) can return to signal "busy" or "unavailable" without panicking the orchestrator.
502,503,504and429all take this path. Only500is handled cleanly.Any orchestrator running its runners behind a reverse proxy, a load balancer, or on a remote host is exposed.
Root cause
1. The generated response parser silently returns a nil payload.
ai/worker/runner.gen.go,ParseGenLiveVideoToVideoResponse:There is no
defaultbranch. A502(or a200without a JSONContent-Type) matches nothing, soresponse.JSON200staysniland no error is produced.2. The worker does not check the payload before returning it.
ai/worker/worker.go:Result:
(nil, nil).3. The caller dereferences it.
core/ai_orchestrator.goL606-620:err == nilis true, so theelsebranch that would have handled the failure is never reached.Stack trace
The panic occurs in the per-connection handler goroutine, so
net/http'srecover()keeps the process alive. Other concurrent streams are unaffected — but the request that triggered it dies with no response to the gateway.Reproduction
502— stop the runner, or saturate a connection limit so the proxy rejects before reaching any upstream.POST /live-video-to-videorequest.ai_orchestrator.go:612.A
200response without a JSONContent-Typeheader reproduces it identically.Observed impact
On a production orchestrator over a 6-hour window:
Correlating the orchestrator logs with the proxy's
error.logshowed two distinct upstream conditions producing the same panic:connect() failed (111: Connection refused)— runner down (9 occurrences)no live upstreams/upstream prematurely closed connection— proxy rejected before or during upstream contact (8 occurrences)Both surface to go-livepeer as an HTTP
502.Workaround
Making the proxy return
500withContent-Type: application/jsoninstead of502eliminates the panic entirely, sinceJSON500is a handled case. Validated in production: panics dropped to zero, failures now surface as clean500s the gateway can retry.This is a deployment-side patch and shouldn't be the answer — it relies on rewriting a status code to lie about the failure type.
Suggested fixes
1. Add a
defaultcase to the generated parser so unmapped statuses produce an error rather than a nil payload:default: return nil, fmt.Errorf("unexpected status %d from runner: %s", rsp.StatusCode, string(bodyBytes))Since
runner.gen.gois generated byoapi-codegen, this likely belongs in the OpenAPI spec or the generator configuration rather than the generated file itself.2. Guard the return in
worker.go:3. Defensive nil check at the call site (
ai_orchestrator.go:612), so a(nil, nil)contract violation anywhere downstream degrades to an error instead of a panic.Fix 1 or 2 alone resolves the issue; 3 is cheap insurance.
Related
The same
*workerRespdereference pattern appears in the other pipelines in this file (TextToImageand others). Those call synchronous endpoints where a nil payload may be less likely, but the same class of bug applies if a proxy or network layer returns an unmapped status.Related but distinct: #3989 guards a nil in
server; this is a separate nil path incore/ai_orchestrator.go.