Repository navigation
Audit records outcome=success for JSON-RPC error bodies over 512 bytes #6720
Description
Activity
- addedneeds-triageIssue needs initial triage by a maintainerIssue needs initial triage by a maintainer
on Sep 26, 2026 I'd like to work on this
Thanks for picking this up — please go ahead, and I am standing back rather than opening a competing PR.
I have no stake in which of the three shapes wins, so I will not argue for one. Two things from re-reading the report that may save you time, and one correction:
Correction on my own numbers. I wrote that the int-parameter case gave
1.21e-08in the pymc thread — that was a different repo, ignore it. Here the figures I can stand behind are the mirror program you ran, and they are enough: the cut lands at 512 and detection goes false, while 511 still works. That is the boundary.The
outcomequestion is the part worth getting right.pkg/audit/auditor.go:300-308has one more path than the three in the report: whateverdetectApplicationErrorreturns, an unparseable or truncated body is currently indistinguishable from a genuinely clean response. If you take shape 3, the marker has to be readable by whatever consumes audit records downstream, not just present in the record — otherwise the record still reads assuccessto the operator who is looking for the error.A cheap thing worth checking first: is
errorDetectionBufferSizeactually configurable at runtime, or only a package constant?pkg/audit/config.go:45-49describes the contract but the report readsauditor.go:110-114as a const. If an operator can already raise it, that is a documented workaround worth putting in the issue, and it bounds how urgent shape 1 is.The repo rules that will apply to your PR, in case they are not obvious from the template: commits need a DCO sign-off, the PR is capped at 1000 lines, and
taskis the only way to run build, test and lint — notgo test,go buildorgolangci-lintdirectly.pkg/audit/auditor.gosits in av1beta1operator API surface, so a new field on the recorded event is a breaking change; a marker inside the existing metadata map is not.If you land shape 2 or 3, I would like to review the truncation boundary specifically — whether a body of exactly 512 bytes is still detected, and whether a cut landing inside the
errorkey rather than inside the string is handled. Those are the two places a streaming detector usually diverges from the prefix one, and the v1beta1 constraint means getting it wrong is awkward to walk back.Thanks @feiiiiii5 for the context and @breken-ai for volunteering to work on this! 🙏 I'm assigning this to you, we can discuss the actual approach on the PR.
Reacted by Chen Yufeiyang- removedneeds-triageIssue needs initial triage by a maintainerIssue needs initial triage by a maintainer
on Oct 8, 2026
Bug description
The audit middleware's JSON-RPC error detection stops working once a response body is larger than its 512-byte detection buffer, and the event is recorded as
outcome: success.errorDetectionBufferSizeis 512 (pkg/audit/auditor.go:110-114) andresponseWriter.Writecopies only the bytes that fit, truncating the rest (pkg/audit/auditor.go:145-153).detectApplicationErrorthen hands that prefix tomcp.ParseMCPResponse(pkg/audit/auditor.go:456-469), which is deliberately lenient and returnsHasError=falsewhenjson.Unmarshalfails (pkg/mcp/response.go:39-46). A truncated JSON object is not valid JSON, sooutcomestaysOutcomeSuccess(pkg/audit/auditor.go:300-308) and nojsonrpc_error_code/jsonrpc_error_messagemetadata is attached.The config contract says a prefix is buffered "to detect JSON-RPC error fields, independent of the IncludeResponseData setting" (
pkg/audit/config.go:45-49, and the same wording indocs/operator/crd-api.mdand the vMCP CRD), withDetectApplicationErrorsdefaulting totrue. For error bodies over 512 bytes the operator gets neither the error nor any hint that detection was skipped, which is the failure class #4678 was filed about.The streamable-HTTP proxy writes a JSON-RPC response in a single
Write(pkg/transport/proxy/streamable/utils.go:134), so the cut lands mid-string.Steps to reproduce
Any tool call whose JSON-RPC error body exceeds 512 bytes. Mirroring the truncation and
ParseMCPResponsesemantics outside the repo:With an envelope of
{"jsonrpc":"2.0","id":1,"error":{"code":-32000,"message":"…"}}(63 bytes of overhead) the cut is at a ~449-character message. Error bodies that large come from servers that put a stack trace or the offending input inerror.data, and from the proxy's own wrapped upstream failures.Expected behavior
detectApplicationErrors: trueshould classify any JSON-RPC error response asoutcome: application_error, or the record should say why it could not.Actual behavior
outcome: success, with no error code or message, for every JSON-RPC error body over 512 bytes.Environment (if relevant)
main@ dfb0713task test)Additional context
I looked at three shapes for a fix and did not want to pick one unilaterally:
error.datawould still be missed.json.Decodertoken walk that stops at the top-levelerrorkey). Handles unbounded bodies without buffering them.outcome: unknown(or attach adetection_truncated: truemarker) when the body was cut off, so the record is never silently positive.I lean towards 2, with 3 as a cheap safety net, but 1 would also fix the common case. Happy to be assigned and put up a PR with tests either way — I did not find an existing test that feeds a response body over the buffer size through the middleware.