M5 Ultra

Mac Studio, 36-core CPU, 80-core GPU, 256 GB.
Quants: DeepSeek-V4.1-Flash own 3-bit LSQ · MiMo mxfp4 · Qwen3.8-Flash-Next oQ8e · Qwen3.8-27B oQ8e
Kernels and patches on GitHub · DeepSeek-V4.1-Flash split across Mac + RTX

Prefill at depth

Prompt tokens ÷ time to first token.

DeepSeek-V4.1-Flash

Measured in 8K chunks.

MiMo-V2.6-Flash

Measured at 130K and 522K.

Qwen3.8-Flash-Next

Measured through the server at 8K to 1M tokens.

Qwen3.8-27B

What made it faster

DeepSeek-V4.1-Flash, same day. Every change keeps greedy output bit-identical.