On Tuesday, October 6, at 20:04:41 by the server's clock, a pod with two NVIDIA RTX PRO 6000 cards logged torch.OutOfMemoryError and died. At 20:08:11 a second build did the same. Both were trying to load DeepSeek-V4.1-Flash, a 763 billion parameter model, with every weight on the GPUs.
The next boot used the same pod, the same model and the same build as the second one. It loaded and started answering requests. The only change was a plan file: a list of which experts should live in the machine's RAM instead of on the cards.
That plan file is what our new paper is about. It looks like a hardware stunt. The finding that surprised us is about pricing: the server produces enough tokens in a month for 262 developers like the one we profiled, and it can give 19 of them the speed a coding agent needs.
This is the short version of a 47-page measurement study. Every number below comes from it, and the paper traces each one to a data file.
DeepSeek-V4.1-Flash is a mixture-of-experts model. Each of its 40 layers holds 384 experts and each token visits 6 of them. Most of the weights sit idle at any moment, which is why people keep trying to park them somewhere cheaper than VRAM.
The released checkpoint is 475.2 GiB. A community pack quantizes the routed experts to 3 bits (the EXL3 format) and brings the whole thing to 309.3 GiB. We read every tensor header of both checkpoints with HTTP range requests, without downloading a single weight, and added up what each GPU would have to hold. The answer was 102.1 GiB. The driver reports 95.6 GiB per card. That is before KV cache, before CUDA graphs, before the first token.
The two failed boots are that arithmetic, confirmed by the hardware.
The plan file moves 4,608 of the 15,360 routed experts, the two Engram lookup tables and the token embedding into pinned host memory. GPU kernels read them in place over PCIe when a token needs them. Nothing is copied into VRAM and the CPU does no math. The GPU side drops to 72.4 GiB per card.
You don't need to know how EXL3 works to follow the rest. What matters is the trade: about a third of the experts now sit across a PCIe link, and every question below is about what that costs.
Before running a benchmark we audited the stack. We fetched every component at its pinned commit, applied every patch, compared the vLLM patches line by line with their upstream pull requests and searched every tree for the offload features.
None of the clever parts are ours. The per-expert host placement, the split launch that runs RAM-resident experts on a side stream with a capped number of SMs, and the kernels that make it fast are Diffbot's. They are a patch on Victor Cruz's vllm-exl3 plugin, which brings EXL3 into vLLM on top of turboderp's ExLlamaV3 kernels. Upstream vLLM already ships zero-copy offload of whole parameter tensors. Our share is orchestration, configuration and measurement.
We started with one idea of what the paper would be and the evidence took it away. What was left is a narrower question: what does this third-party offload make possible, and what does it cost?
Part of the answer came early. We tried upstream vLLM's own UVA offload on the same experts, at the same host bytes. With the KV block size of 64 that the stack uses, it fails at initialization. With 128 it boots, and the engine dies on the first request because two attention kernels accept different block sizes. On these pins, Diffbot's hooks are what makes the model serve at all.
The coding-agent experiment is built on one usage profile: one developer's week of Claude Code, read from the provider's usage screen. 94.6 million tokens over 9 sessions on 5 active days.
Here is the question. On the third pod, four agents working at once got 168 output tokens per second out of the server in total. How many developers like that one can the pod serve?
The natural way to answer is by volume. A month of that output equals 262 of these developers at full utilization, or 131 if you assume the pod is busy half the time. Pick your number and keep it. We will come back to it.
To see why the answer is smaller, you need one distinction. Serving a language model is two jobs.
Prefill is reading. The model takes in the whole prompt in one pass. It is bound by compute, and its cost grows with the prompt.
Decode is writing. One token at a time, per request. It is bound by memory traffic, and in this stack part of that traffic crosses PCIe.
Start with one row of data. On the second pod, a single request decoded at 7.6 ms per token after a 4K-token prompt.
Now a second row. After a 256K-token prompt, the same pod decoded at 8.3 ms per token.
Sixty-four times the context, about 9% slower writing.
Reading is a different story. A fresh 128K-token prompt took 18.33 s to its first token. A 260,000-token prompt took 45.07 s. That is the real price of long context on this stack, and it is paid before the first word.
The way around it is the prefix cache. When the long context is shared and already cached, as with an agent working on one big codebase, four users on a 256K-token prefix saw a p95 time to first token of 1.66 s and a median of 16.6 ms per token.
Long context costs reading, not writing. Keep it cached.
A coding agent has a person waiting on it. We set the bar at a median of 40 output tokens per second per stream, which is 25 ms per token. Aggregate throughput doesn't count. Ten streams at 20 tokens per second are not five at 40.
On the third pod we ran six regimes: short independent sessions, streams branching from one cached 128K or 256K-token prefix, streams that each own a cached 128K or 256K context, and agent turns shaped by the usage profile (contexts of 88,000 tokens growing by 3,500 a turn, replies of 400 tokens).
The agent-turn rows:
| simultaneous streams | median tok/s per stream | p5 tok/s |
|---|---|---|
| 1 | 186 | 134 |
| 2 | 95 | 68 |
| 3 | 70 | 59 |
| 4 | 59 | 52 |
| 6 | 40 | 34 |
| 8 | 31 | 24 |
Every regime run that far kept the bar at four streams and lost it somewhere between five and eight. Independent 256K contexts were only run to three, all above it. Six agent streams cleared 40 by less than one token per second, which is noise on a single run, so we provision on four with a 10% margin.
Context length barely matters once it is cached. The 256K regimes were no slower than the short one at the same number of streams. The batch sets the pace, not the context. Five independent 128K contexts used at most 48% of the KV cache, against a vLLM estimate of 2.38 full-length requests. That estimate is conservative for this model.
Back to your guess.
Developers don't generate all day. If every output token of the profile came out at exactly 40 tokens per second, and every new input token were read at the prefill rate we measured, the developer would be generating for 2.8 hours a week. That is 1.7% of the week.
Spread over an 8-hour window on each of the 5 active days, the chance that the developer is generating at a given moment is 7.1%. With several developers acting independently, the number generating at once follows a binomial distribution. One pod with room for four streams can take 19 developers before the chance of a fifth simultaneous stream passes 1%. Accept 5% and it takes 28.
262 by volume. 19 by concurrency. That gap is the central result of the paper.
Think of a bar that can pour drinks for 262 people over a night and has four stools at the counter. Nobody runs out of beer. People wait for a stool. Going over four doesn't fail any request either: for that moment, every stream drops below 40 tokens per second.
Then the money. The pod costs 3.73 EUR an hour, 2,723 EUR a month around the clock. That buys 15.1 subscriptions at 180 EUR a seat. One pod with 19 seats lands just above break-even, and the 20th developer needs a second pod: at 20 seats, self-hosting costs 1,847 EUR a month more than subscriptions. Scale changes the picture, because the sum of independent demands is smoother than each one. 100 developers need 4 pods, 39% below the subscription bill. 250 need 7, 58% below.
Those are GPU rental numbers only. No engineering, no on-call, no storage. A subscription to a commercial coding agent buys a whole product, not just a model, and we are not claiming this model equals the one behind it.
The argument that doesn't fit in a cost table is control. With open weights you choose the model and the version, and the source code and secrets an agent reads stay out of a model provider's hands. This profile re-reads 234 tokens of context for every token it writes, so that is a lot of code going somewhere. The benefit is complete on your own hardware. On a rented pod, like ours, the data still sits on someone else's machine.
The third session's record notes the first pod's public IP. Same datacenter, same GPU model, same driver, same PCIe Gen5 x16 link. The CPU model and the GPU bus IDs are different. It is a different machine.
That matters more than it sounds. We ran the same subset on pods 1 and 2, with the same stack, plan file and workload file. Pod 2 was slower at every point, with a median time per token 1.17 to 1.51 times pod 1's. Pod 1 has both GPUs on one NUMA node. Pod 2 has one GPU on each socket, so with tensor parallelism and experts read from host memory, traffic crosses the link between sockets. That is a plausible cause. One run per pod can't prove it.
The second rule is cheap: before you benchmark or serve, run nvidia-smi topo -m. NODE between the two GPUs is what you want. SYS means the cloud gave you a different machine than the one on the label.
We compared the stack (thinking off) with DeepSeek's own API on fixed subsets of three benchmarks:
- GSM8K: 0.960 against 0.960, the same outcome on every one of the 150 items.
- IFEval: +0.013, with a 95% interval from -0.013 to +0.047.
- HumanEval: -0.018, with an interval from -0.049 to +0.012.
The intervals include zero, and 4.0% and 4.3% of the items changed outcome. With 150 to 164 items per task, differences of a few points can't be ruled out. The reference is a service whose precision and serving code we can't see, so this bounds the whole stack against the vendor, not the quantization alone.
The first HumanEval run against the API scored 0. On every item. The task ends with a partial assistant message that the model has to continue, and a plain chat completion opened a new turn instead. On the vLLM side, the DeepSeek tokenizer ignored the continue flag and closed the message, and that run scored 0.244. Both invalid runs are kept in the repository and marked as such. A benchmark harness can fail louder than a model.
- Each experiment comes from a single boot of a rented pod. No measurement was repeated across boots, and the one configuration measured on two hosts differed by far more than any interval inside a boot.
- There is no all-in-VRAM baseline. The reference run on two B300 GPUs (15.78 USD an hour) was designed and not run. We can say the offload makes the model fit. We can't say what it costs against a machine that fits it.
- We didn't separate the causes of the slowdown under load: PCIe reads of cold experts, the speculative decoding batch, prefill interference, or eager execution above the largest CUDA graph. The ablations for that are written and didn't run.
- Agent turns are synthetic, replies are fixed at 400 tokens, and the profile is one developer's week. The binomial model assumes independent developers. A team on one schedule with shared deadlines has a heavier tail, so 19 is optimistic for that team.
- Quality covers three small subsets with thinking off. The LiveBench scores that made this model worth serving belong to the unquantized model at maximum effort with thinking on, not to this stack.
Not every error was the model's. In the first session, one client-side disconnect among eight requests tripped our pre-registered stop rule (error rate above 5%) and ended a sweep. A rule evaluated on eight requests is too twitchy. Later sessions required at least two errors, and the first session's data stays as collected.
The whole controlled campaign used about 14 USD of pod time over three sessions (101, 43 and 57 minutes), plus 12 cents of API calls.
The paper asks two things: what the offload makes possible, and at what cost. We have the first answer. On two 96 GB cards it turns a model that doesn't load into one that serves four coding agents at 40 tokens per second each, with up to 256K tokens of cached context.
The second answer is still out. In the one comparison we ran, moving more experts to host memory lowered throughput at every load level, so the cost is real. How large it is next to a machine that holds everything in VRAM, we don't know yet.
The experiment that would tell us is written. Its workload is pinned and its runner sits in the repository next to the others. It is called E9, and it hasn't run.