Skip to content

mlx_lm.server: generation hangs at 0% CPU right after prompt processing on ~22-26k-token streaming requests (gemma-4-26b-a4b, 0.31.3) #1493

Description

@TadeuszWolfGang

Summary

mlx_lm.server hangs (decode never starts, process at 0% CPU) immediately after prompt processing completes, on long (~22–26k token) streaming chat-completion requests from a real client (Obsidian Copilot). Reproduced 2/2 times with real client traffic; synthetic requests of the same size — including concurrent + streaming — complete fine, so the trigger seems to be a specific parameter/content combination rather than prompt length alone.

Environment

  • mlx-lm 0.31.3, mlx/mlx-metal 0.31.2 (installed via uv tool install mlx-lm, transformers==5.12.1)
  • macOS 26.5 (Darwin 25.5.0), Apple M5 Max, 128 GB
  • Python 3.12.13
  • Model: mlx-community/gemma-4-26b-a4b-it-8bit (model_type gemma4), pinned at startup
  • Server: mlx_lm.server --model /path/to/gemma-4-26b-a4b-it-8bit --host 127.0.0.1 --port 8082 --decode-concurrency 32 --prompt-concurrency 8
  • (Environment note: a local one-line models/gemma4_unified.py alias exists for a different model, mirroring PR Add gemma4_unified model alias for Gemma 4 12B support #1386; it is not imported by the affected instance, which serves a stock gemma4 checkpoint.)

Failing requests (2/2 hangs)

Client: Obsidian Copilot plugin ("3rd party (openai-format)" provider). Request shape: stream: true, temperature: 0.1, max_tokens: 16000, system prompt + long mixed Polish/English markdown context.

  • Hang 1: prompt 22,258 tokens — prefill completes in ~8 s, then nothing.
  • Hang 2 (after server restart): prompt 25,660 tokens — prefill completes in ~12 s, then nothing.

In both cases the log shows prompt processing reaching N/N and then goes silent; no traceback, no completion line. The server process stays alive but at 0.0% CPU (checked minutes later — 5–18 min). Client eventually gives up.

2026-07-05 08:19:39,913 - INFO - Prompt processing progress: 23032/25660
2026-07-05 08:19:40,840 - INFO - Prompt processing progress: 25080/25660
2026-07-05 08:19:41,170 - INFO - Prompt processing progress: 25659/25660
2026-07-05 08:19:41,208 - INFO - Prompt processing progress: 25660/25660
(no further log lines; process at 0% CPU)

Right before the big request the client issued two small chat completions (~60 and ~300 prompt tokens), which completed normally.

What does NOT reproduce it (all pass)

Same server instance, same model:

  1. Non-streaming 22,225-token prompt, max_tokens 60 → OK in 7.3 s.
  2. Non-streaming 10,525-token prompt → OK in 9.5 s.
  3. Concurrent mix mimicking the client: 2 small non-streaming + 1 streaming 22k-token request (max_tokens 200) fired simultaneously → all OK in ≤12 s.

So plain long prompts, streaming, and concurrency by themselves are fine. The failing requests differ mainly in max_tokens: 16000, a system prompt, temperature: 0.1, and real markdown content; also the small preliminary requests populate the prompt cache right before the big one (the log prints "Prompt Cache: N sequences" before each request), so a prompt-cache/trim interaction is a plausible suspect.

Expected behavior

Either generate tokens after prefill, or fail loudly. A silent hang with the worker at 0% CPU makes the failure invisible to KeepAlive-style supervision.

Happy to provide full logs or run patched builds to help narrow this down.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions