Skip to content

RuntimeError: There is no Stream(gpu, 3) in current thread at KVCache __del__ teardown #1888

Description

@dahai80

Summary

When running batched inference with paged KV cache, a RuntimeError: There is no Stream(gpu, 3) in current thread is raised at process teardown / GC time.

Traceback

  File ".../mlx_lm/generate.py", line 1562, in __del__
  File ".../mlx_lm/generate.py", line 1557, in close
RuntimeError: There is no Stream(gpu, 3) in current thread.

Repro

# fusion-mlx bench with paged cache (any small model, batch>=2)
fusion-mlx bench qwen3.5-4b-4bit --num-prompts 6 --max-tokens 48 --max-num-seqs 4 --use-paged-cache

The bench completes successfully (throughput is printed), then the error fires during AsyncEngineCore.__aexit__ / GC when the KV cache __del__ runs close() in a thread that does not own the stream.

Environment

  • mlx-lm: 0.31.3
  • macOS 15 (Darwin 25.6.0), Apple Silicon (M3, 8 GPU cores)
  • Python 3.12

Analysis

The KV cache creates/uses a stream during inference (in the async worker thread). At teardown, __del__ invokes close() which tries to access that stream, but __del__ runs in whichever thread the GC finalizer picks (often the main thread), which does not have the stream registered. This is a stream-ownership / finalizer race.

Suggest: make close() defensive (catch the RuntimeError, or free resources via the owning thread explicitly rather than in __del__).

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions