You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
mlx_lm.server: generation hangs at 0% CPU right after prompt processing on ~22-26k-token streaming requests (gemma-4-26b-a4b, 0.31.3) #1493
mlx_lm.server hangs (decode never starts, process at 0% CPU) immediately after prompt processing completes, on long (~22–26k token) streaming chat-completion requests from a real client (Obsidian Copilot). Reproduced 2/2 times with real client traffic; synthetic requests of the same size — including concurrent + streaming — complete fine, so the trigger seems to be a specific parameter/content combination rather than prompt length alone.
(Environment note: a local one-line models/gemma4_unified.py alias exists for a different model, mirroring PR Add gemma4_unified model alias for Gemma 4 12B support #1386; it is not imported by the affected instance, which serves a stock gemma4 checkpoint.)
Failing requests (2/2 hangs)
Client: Obsidian Copilot plugin ("3rd party (openai-format)" provider). Request shape: stream: true, temperature: 0.1, max_tokens: 16000, system prompt + long mixed Polish/English markdown context.
Hang 1: prompt 22,258 tokens — prefill completes in ~8 s, then nothing.
Hang 2 (after server restart): prompt 25,660 tokens — prefill completes in ~12 s, then nothing.
In both cases the log shows prompt processing reaching N/N and then goes silent; no traceback, no completion line. The server process stays alive but at 0.0% CPU (checked minutes later — 5–18 min). Client eventually gives up.
2026-07-05 08:19:39,913 - INFO - Prompt processing progress: 23032/25660
2026-07-05 08:19:40,840 - INFO - Prompt processing progress: 25080/25660
2026-07-05 08:19:41,170 - INFO - Prompt processing progress: 25659/25660
2026-07-05 08:19:41,208 - INFO - Prompt processing progress: 25660/25660
(no further log lines; process at 0% CPU)
Right before the big request the client issued two small chat completions (~60 and ~300 prompt tokens), which completed normally.
What does NOT reproduce it (all pass)
Same server instance, same model:
Non-streaming 22,225-token prompt, max_tokens 60 → OK in 7.3 s.
Non-streaming 10,525-token prompt → OK in 9.5 s.
Concurrent mix mimicking the client: 2 small non-streaming + 1 streaming 22k-token request (max_tokens 200) fired simultaneously → all OK in ≤12 s.
So plain long prompts, streaming, and concurrency by themselves are fine. The failing requests differ mainly in max_tokens: 16000, a system prompt, temperature: 0.1, and real markdown content; also the small preliminary requests populate the prompt cache right before the big one (the log prints "Prompt Cache: N sequences" before each request), so a prompt-cache/trim interaction is a plausible suspect.
Expected behavior
Either generate tokens after prefill, or fail loudly. A silent hang with the worker at 0% CPU makes the failure invisible to KeepAlive-style supervision.
Happy to provide full logs or run patched builds to help narrow this down.
Summary
mlx_lm.serverhangs (decode never starts, process at 0% CPU) immediately after prompt processing completes, on long (~22–26k token) streaming chat-completion requests from a real client (Obsidian Copilot). Reproduced 2/2 times with real client traffic; synthetic requests of the same size — including concurrent + streaming — complete fine, so the trigger seems to be a specific parameter/content combination rather than prompt length alone.Environment
uv tool install mlx-lm,transformers==5.12.1)mlx-community/gemma-4-26b-a4b-it-8bit(model_typegemma4), pinned at startupmlx_lm.server --model /path/to/gemma-4-26b-a4b-it-8bit --host 127.0.0.1 --port 8082 --decode-concurrency 32 --prompt-concurrency 8models/gemma4_unified.pyalias exists for a different model, mirroring PR Add gemma4_unified model alias for Gemma 4 12B support #1386; it is not imported by the affected instance, which serves a stockgemma4checkpoint.)Failing requests (2/2 hangs)
Client: Obsidian Copilot plugin ("3rd party (openai-format)" provider). Request shape:
stream: true,temperature: 0.1,max_tokens: 16000, system prompt + long mixed Polish/English markdown context.In both cases the log shows prompt processing reaching N/N and then goes silent; no traceback, no completion line. The server process stays alive but at 0.0% CPU (checked minutes later — 5–18 min). Client eventually gives up.
Right before the big request the client issued two small chat completions (~60 and ~300 prompt tokens), which completed normally.
What does NOT reproduce it (all pass)
Same server instance, same model:
max_tokens 60→ OK in 7.3 s.max_tokens 200) fired simultaneously → all OK in ≤12 s.So plain long prompts, streaming, and concurrency by themselves are fine. The failing requests differ mainly in
max_tokens: 16000, a system prompt,temperature: 0.1, and real markdown content; also the small preliminary requests populate the prompt cache right before the big one (the log prints "Prompt Cache: N sequences" before each request), so a prompt-cache/trim interaction is a plausible suspect.Expected behavior
Either generate tokens after prefill, or fail loudly. A silent hang with the worker at 0% CPU makes the failure invisible to
KeepAlive-style supervision.Happy to provide full logs or run patched builds to help narrow this down.