DeepSeek-V4.1-Flash on a Mac and two RTX PRO 6000s
One model, two machines
Only four of the model's 40 layers produce KV, and the upper half reuses layer 20's. The prompt state is about 0.9 KB per token, so the model splits cleanly at layer 20.
- Embeddings, Engram, layers 0–19
- Layer-20 KV rows, vision tower
- Prefill the whole prompt
- 1–5 verify rows a step, ~7 ms
- Layers 20–39, output head
- DSpark drafter, up to 4 tokens
- Accept, stream, draft the next rows
- Prefix cache for resumed turns
Prefill
Prompt tokens per second of time to first token, fresh uncached prompts. Dashed: the previous Mac-only 3-bit build.
Decode at depth
Output tokens per second, mean of six 256-token samples with the full range. Speed follows speculative acceptance, not context length: the box step is 7.2–7.8 ms from 8K to 1M.
By prompt length
Mac + RTX first, the 3-bit build beneath.
| Prompt | First token | Prefill tok/s | Speedup | Decode tok/s |
|---|---|---|---|---|
| 8K8,238 tokens | 1.03 s3.29 s | 7,9752,498 | 3.2× | 87.271.7 |
| 16K16,428 tokens | 1.50 s6.38 s | 10,9662,573 | 4.3× | — |
| 32K32,814 tokens | 2.38 s12.6 s | 13,8122,610 | 5.3× | — |
| 64K65,581 tokens | 4.19 s25.3 s | 15,6632,591 | 6.0× | — |
| 128K131,116 tokens | 7.87 s51.1 s | 16,6552,545 | 6.5× | 94.266.2 |
| 256K262,188 tokens | 15.5 s111 s | 16,9112,360 | 7.2× | 86.463.5 |
| 512K524,335 tokens | 32.6 s235 s | 16,0812,234 | 7.2× | 94.465.6 |
| 768K786,479 tokens | 54.1 s376 s | 14,5432,090 | 7.0× | — |
| 1M1,040,046 tokens | 81.5 s521 s | 12,7561,998 | 6.4× | 88.961.4 |
Concurrency
- 2 requests, 8K
- 107.5 tok/s62 each
- 2 requests, 128K
- 102.7 tok/s60 each
- 4 requests, 8K
- 97.2 tok/s2 at a time
Resuming a conversation
- Turn 2 at 8K
- 0.79 s
- Turn 2 at 128K
- 1.25 s3-bit 0.85 s
- Turn 2 at 512K
- 2.93 s
Images
- 1 image, first token
- 0.46 s6/6 correct
- 2 images, first token
- 0.52 s5/6 correct
Quality, 27 cases
- Code continuation NLL
- 0.0583-bit 0.062
- Needles at 128K and 1M
- recalled
- Draft acceptance
- 66%3-bit 65%
How it runs
The box prefills layers 0–20 over the whole prompt and streams the state to the Mac in 8K chunks while it works. The Mac replays the last rows through layers 20–39, then each decode step crosses the link twice: draft tokens go to the box, which returns 41 KB of hidden state per row. Two requests pipeline, so each machine works on one while the other finishes the next.
A step on the box is bit-identical to prefilling the same rows, including rejected drafts, so a box engine restart mid-answer rebuilds the session and the text continues unchanged. If the box stays down, clients get a retryable error; no other model ever stands in.