Environment: mlx-lm 0.31.3, mlx 0.32.0, macOS 26.5.2 (M3 Max, 36GB), Python 3.13.11, mlx_lm.server --prompt-concurrency 2, model unsloth/Qwen3.6-35B-A3B-UD-MLX-4bit.
Symptom (design issue): when Thread-1 (_generate) dies from any uncaught exception (we have observed several distinct ones in production under sustained batched load: TypeError: 'NoneType' object is not iterable from batch-merged logits_processors, RuntimeError: [metal::malloc] Resource limit (499000) exceeded per #1332, ValueError: Slice indices must be 32-bit integers, and a mask-vs-empty-KV broadcast_shapes error), mlx_lm.server keeps accepting connections:
GET /v1/models returns 200
- every
POST /v1/chat/completions hangs indefinitely
- nothing is logged beyond the single traceback; no 5xx, no exit
Health checks pass while the server is unrecoverable — a zombie. Before we added an external watchdog we measured real outages of multiple hours per occurrence; with a probing watchdog it is still ~6 minutes of silent hang per event.
Suggested fix (any of):
- Exit the process when the generation thread dies — supervisors (launchd/systemd) restart it. This is what we do today via a
threading.excepthook wrapper; recovery drops to ~40s.
- Restart the generation thread and fail all pending requests with 503.
- At minimum, make
/v1/models (or a health endpoint) fail once the generation thread is gone, so probes can detect the state.
This issue multiplies the severity of every generation-side bug: with it fixed, the crashes above would be brief, detectable blips instead of silent outages. Happy to provide full tracebacks for each of the observed death modes or test a patch.
Environment: mlx-lm 0.31.3, mlx 0.32.0, macOS 26.5.2 (M3 Max, 36GB), Python 3.13.11,
mlx_lm.server --prompt-concurrency 2, modelunsloth/Qwen3.6-35B-A3B-UD-MLX-4bit.Symptom (design issue): when
Thread-1 (_generate)dies from any uncaught exception (we have observed several distinct ones in production under sustained batched load:TypeError: 'NoneType' object is not iterablefrom batch-mergedlogits_processors,RuntimeError: [metal::malloc] Resource limit (499000) exceededper #1332,ValueError: Slice indices must be 32-bit integers, and a mask-vs-empty-KVbroadcast_shapeserror),mlx_lm.serverkeeps accepting connections:GET /v1/modelsreturns 200POST /v1/chat/completionshangs indefinitelyHealth checks pass while the server is unrecoverable — a zombie. Before we added an external watchdog we measured real outages of multiple hours per occurrence; with a probing watchdog it is still ~6 minutes of silent hang per event.
Suggested fix (any of):
threading.excepthookwrapper; recovery drops to ~40s./v1/models(or a health endpoint) fail once the generation thread is gone, so probes can detect the state.This issue multiplies the severity of every generation-side bug: with it fixed, the crashes above would be brief, detectable blips instead of silent outages. Happy to provide full tracebacks for each of the observed death modes or test a patch.