同一个 Qwen3.6,加上推测解码快了 3 倍
上周看到一篇 Medium 文章,说在 M5 Max 上用 DFlash 推测解码跑 Qwen3.6-27B,能到 65 tok/s。
我手边有台 M3 Ultra 256GB,顺手验证了一下。但我不光测了 DFlash,还把纯 MLX 也跑了一遍——同一个模型,开不开推测解码,到底差多少。
● ● ●
纯 MLX vs DFlash
都在 M3 Ultra 上,同一个 Qwen3.6-27B 4-bit 模型(15GB),同一个 prompt。唯一的变量是推理引擎。
| 场景 | 纯 MLX | DFlash | 加速比 |
|---|---|---|---|
| binary search | 30.6 tok/s | 94.0 tok/s | 3.1× |
| trie 数据结构 | 33.7 tok/s | 113.3 tok/s | 3.4× |
| LRU cache + TTL | 34.8 tok/s | 69.8 tok/s | 2.0× |
| quicksort 完整实现 | 32.7 tok/s | 99.3 tok/s | 3.0× |
| **平均** | **33.0 tok/s** | **94.1 tok/s** | **2.9×** |
差不多 3 倍。DFlash 额外占 3.2GB 显存放草稿模型,换来接近 100 tok/s 的体验。
● ● ●
DFlash 推测解码做了什么
正常的自回归推理是串行的:算一个 token → 用它预测下一个 → 再算 → 再预测。GPU 大部分时间在等。
DFlash 同时跑两个模型:
- 01草稿模型(3.2GB):并行预测好几个 token
- 02主力模型(15GB):一次验证草稿对不对
对就保留,错就扔掉重来。实测接受率 70-84%,平均每次验证拿到 3-6 个 token。比每次只算一个快得多。
● ● ●
Ollama 怎么样
顺便也测了 Ollama。但 Mac 版 Ollama 0.30 只支持 MLX 引擎,GGUF 格式的 4-bit 模型跑不了。能用的只有 mxfp8 版本——31GB,精度更高但也更慢:
| DFlash 4-bit | 纯 MLX 4-bit | Ollama mxfp8 | |
|---|---|---|---|
| 速度 | **94 tok/s** | 33 tok/s | 20 tok/s |
| 内存 | 18 GB | 15 GB | 31 GB |
DFlash 比 Ollama 快 4-5×,但这不完全是推测解码的功劳——MLX 引擎本身就比 llama.cpp 快。真正的对比是 DFlash vs 纯 MLX,3 倍差距是推测解码的净贡献。
● ● ●
怎么搭
mkdir -p ~/ws/apps/dflash-local && cd ~/ws/apps/dflash-local
uv venv && source .venv/bin/activate
uv pip install mlx mlx-lm
uv pip install "dflash-mlx @ git+https://github.com/bstnxbt/dflash-mlx.git"
HF_ENDPOINT=https://hf-mirror.com dflash serve \
--model mlx-community/Qwen3.6-27B-4bit \
--draft z-lab/Qwen3.6-27B-DFlash \
--port 8000 \
--chat-template-args '{"enable_thinking": false}'
模型约 18GB(主力 15 + 草稿 3.2),M3 Ultra 256GB 毫无压力。64GB 的 Mac 也够用。
连 Claude Code 的话加个 claude-code-router 做协议桥:
npm install -g @musistudio/claude-code-router ccr start # 启动在 :3456 export ANTHROPIC_BASE_URL=http://127.0.0.1:3456 export ANTHROPIC_AUTH_TOKEN=test claude
● ● ●
值不值得
如果你的 Mac 是 64GB+ 且主要做本地 coding:DFlash 值得装。多占 3.2GB,换来 3 倍速度。100 tok/s 写代码基本秒出。
如果 32GB:27B 可能勉强,试试 Qwen3.6-14B 的 DFlash 草稿模型。
核心结论:同一个模型,同一个 MLX 引擎,DFlash 推测解码净提速 3 倍。 公开数据里 MLX 比 Ollama 快 2-3 倍是引擎层面的差距,推测解码又在这个基础上再翻了一番。不算 Ollama 的话,DFlash 也比纯 MLX 快 3 倍——这是推测解码本身的贡献。
● ● ●
参考链接
- Medium 原文:https://medium.com/data-science-collective/run-claude-code-locally-on-a-mac-65-tok-s-with-a-4-bit-qwen3-6-27b-and-dflash-speculative-decoding-bdb346364cdc
- DFlash v2:https://lmsys.org/blog/2025-06-09-dflash-v2/
- Mac 推理对比:https://insiderllm.com/guides/lm-studio-vs-ollama-mac/
- Qwen3.5 基准:https://antekapetanovic.com/blog/qwen3.5-apple-silicon-benchmark/
- Ollama MLX 公告:https://ollama.com/blog/mlx