DeepSeek-V4.1-Flash on a Mac and two RTX PRO 6000s

DeepSeek-V4.1-Flash on its original FP4/FP8 weights, served as one model across a Mac Studio M5 Ultra (256 GB) and two RTX PRO 6000 Blackwell GPUs over a 10GbE cable. Measured through the production API, September 26, 2026.
Code on GitHub · All M5 Ultra results

128K prompt, first token
7.9s
was 51.1 s
1M prompt, first token
82s
was 521 s
Decode at 512K
94tok/s
was 66 tok/s
Two requests at once
107tok/s
was 78 tok/s

One model, two machines

Only four of the model's 40 layers produce KV, and the upper half reuses layer 20's. The prompt state is about 0.9 KB per token, so the model splits cleanly at layer 20.

Prefill

Prompt tokens per second of time to first token, fresh uncached prompts. Dashed: the previous Mac-only 3-bit build.

0 5K 10K 15K 20K 0 128K 256K 512K 768K 1M 3-bit, Mac only Mac + RTX

Decode at depth

Output tokens per second, mean of six 256-token samples with the full range. Speed follows speculative acceptance, not context length: the box step is 7.2–7.8 ms from 8K to 1M.

0 30 60 90 120 0 128K 256K 512K 768K 1M 3-bit, Mac only Mac + RTX

By prompt length

Mac + RTX first, the 3-bit build beneath.

PromptFirst tokenPrefill tok/sSpeedupDecode tok/s
8K8,238 tokens1.03 s3.29 s7,9752,4983.2×87.271.7
16K16,428 tokens1.50 s6.38 s10,9662,5734.3×—
32K32,814 tokens2.38 s12.6 s13,8122,6105.3×—
64K65,581 tokens4.19 s25.3 s15,6632,5916.0×—
128K131,116 tokens7.87 s51.1 s16,6552,5456.5×94.266.2
256K262,188 tokens15.5 s111 s16,9112,3607.2×86.463.5
512K524,335 tokens32.6 s235 s16,0812,2347.2×94.465.6
768K786,479 tokens54.1 s376 s14,5432,0907.0×—
1M1,040,046 tokens81.5 s521 s12,7561,9986.4×88.961.4

Concurrency

2 requests, 8K
107.5 tok/s62 each
2 requests, 128K
102.7 tok/s60 each
4 requests, 8K
97.2 tok/s2 at a time

Resuming a conversation

Turn 2 at 8K
0.79 s
Turn 2 at 128K
1.25 s3-bit 0.85 s
Turn 2 at 512K
2.93 s

Images

1 image, first token
0.46 s6/6 correct
2 images, first token
0.52 s5/6 correct

Quality, 27 cases

Code continuation NLL
0.0583-bit 0.062
Needles at 128K and 1M
recalled
Draft acceptance
66%3-bit 65%

How it runs

The box prefills layers 0–20 over the whole prompt and streams the state to the Mac in 8K chunks while it works. The Mac replays the last rows through layers 20–39, then each decode step crosses the link twice: draft tokens go to the box, which returns 41 KB of hidden state per row. Two requests pipeline, so each machine works on one while the other finishes the next.

A step on the box is bit-identical to prefilling the same rows, including rejected drafts, so a box engine restart mid-answer rebuilds the session and the text continues unchanged. If the box stays down, clients get a retryable error; no other model ever stands in.