3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
-
Updated
Aug 22, 2026 - Python
3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
Beautiful, zero-dependency realtime dashboard + live activity log for a local MTPLX inference server — reads the /metrics endpoint, no build step.
Run an isolated Claude Science app copy through local or OpenAI-compatible model backends.
Recover the native MTP predictor missing from the 8-bit MLX Qwen3.8-27B-Uncensored package, build a BF16 sidecar, and reproduce a 15.59 → 48.75 tok/s controlled M4 Max result with MTPLX.
Double Qwen 3.8 27B inference speed on Apple Silicon with a one-click local coding agent
Add a description, image, and links to the mtplx topic page so that developers can more easily learn about it.
To associate your repository with the mtplx topic, visit your repo's landing page and select "manage topics."