Перейти к содержимому

[2/2] Qwen 3.6 35B A3B 4bit on M5 Pro(Mac Mini / Macbook)48GB vs M5 Ultra (Mac Studio) 256GB

Elite Test Engineering

0:00 / 0:00

[2/2] Qwen 3.6 35B A3B 4bit on M5 Pro(Mac Mini / Macbook)48GB vs M5 Ultra (Mac Studio) 256GB

470 просмотров · 3 дня назад
Elite Test Engineering
458 подписчиков
470 просмотров · 3 дня назад
Part 1 here    • [1/2] Qwen 3.6 35B A3B 4bit on M5 Pro (Mac...   Speculative decoding's ~20-23% speed gain wasn't worth trading away reliability on my daily-use chat interface, so both machines are back on plain mlx_lm.server, which is stable and previously verified clean under concurrent load. MTP is unreliable under real chat usage — not a performance call, a correctness one. On an earlier day to day use we hit a real crashes on the MacBook: [broadcast_shapes] Shapes (1,2,2,256) and (1,2,0,256) cannot be broadcast. As you see on the previous video, on the very first real message through Open WebUI, using the MTP speculative drafter you will see a partial result then crash. We investigated the root cause rather than just patching around it: It's not caused by Open WebUI's concurrent title/tag-generation requests — those run strictly after the main reply finishes, never overlapping it. We tried three different repro attempts (sequential calls, the exact title/tag sequence, full streaming matching Open WebUI's real protocol) and couldn't reproduce it in isolation. There is a genuine timing-dependent race condition in the MTP drafter's KV-cache bookkeeping — the same bug class already documented from earlier work on a different model+drafter pairing on the Studio, described there as "intermittent... recurred after several successful requests." A race condition can't be worked around by usage discipline ("just use one chat at a time" doesn't help, since it can fire on any request regardless of concurrency).