DeepSeek just made reading your prompts cheaper than thinking

DeepSeek's new open model charges less compute to read a prompt than to answer it. The 8B/16B split signals where model pricing is heading.

DeepSeek just made reading your prompts cheaper than thinking

DeepSeek released V4.1-Flash on September 10 and the benchmark numbers are fine. Better than fine, depending on which leaderboard you trust. But the number I keep coming back to is 890 bytes.

That's the size of the model's global KV cache per token. For comparison, previous long-context models measured their caches in kilobytes per token. DeepSeek puts the figure at roughly a quarter of what V4-Flash needed. When you multiply 890 bytes across a million tokens of context, you get roughly 890 MB. A single 80GB GPU can hold the cache for a conversation the length of a novel series, with room left for activations.

At the same time, the model runs a strange experiment. It spends 8 billion active parameters to read your prompt and 16 billion to answer it. Reading costs half of thinking.

Why agents are the reason

An agent session is mostly input. A coding agent reads the repo, the logs, the test output, the browser state. Then it writes a few hundred tokens. Research agents are worse: they ingest half the web before producing a paragraph. The ratio of input to output in real agent sessions runs somewhere between 20:1 and 100:1.

The standard decoder-only architecture treats those two jobs identically. Every layer computes its own keys and values, every token gets processed by the same amount of active parameters whether it's being read or being written. That made sense when sessions were short. For a chat about dinner plans, prefill and decode are roughly balanced.

Agent sessions broke the balance. DeepSeek noticed and redesigned the model around the imbalance. V4.1-Flash splits into a 20-layer causal encoder and a decoder that takes its global attention state from the encoder's final layer, so the decoder never rebuilds what the encoder already computed. Combine that with shared attention across layers and an FP4-compressed cache, and reading gets structurally cheaper than generating. DeepSeek calls this Causal Encoder-Decoder, or CED. The term will be everywhere in six months.

The economics follow the architecture

The number that matters if you run agents on a budget: Featherless serves the model at $0.30 per million input tokens, $0.03 per million for cached input, and $1.20 per million output tokens. DeepSeek's own API charges the same on input and output, and drops input to $0.15 off-peak. Cached input is the quiet headline: as low as $0.003 per million tokens off-peak on DeepSeek's API, because a tiny KV cache makes prefix reuse cheap to keep around.

That's a 50x spread between cached and uncached reading. My working theory: the gap won't close by cheap models catching up on quality. It will close by labs splitting input from output pricing everywhere, because the 8B/16B split already decoupled the cost side.

Because here's what the 8B/16B split means. Input-heavy workloads, meaning agent sessions, get roughly half the per-token compute cost on the reading side compared to balanced workloads. Half the FLOPs per input token is a real margin difference when a single session reads a million tokens before writing a thousand. Prefill has been the hidden line item in every agent bill I've seen. V4.1-Flash turns prefill cost into an architectural feature, and once one lab prices input and output separately, the others have to answer.

What I like about the release

The MIT license matters less than it looks. Can you run a 763B-parameter model yourself? You need a multi-GPU node, and the full weights come to roughly 380GB at 4-bit. The license gives you the right, and the weights are on Hugging Face, but the practical answer for most teams is a hosted endpoint. For me the license's real value is the guarantee the model can't be revoked: the file is out there, MIT, forever.

The benchmarks I read the way I read every release-day benchmark: as directional, not decisive. They're self-reported by the lab, and the web is currently full of takes like "this changes everything" and "this changes nothing." Both camps are wrong. A model that reads at half the active compute of a balanced design will keep its cost advantage on agent workloads regardless of where benchmark scores land.

The 45 trillion training tokens are a sensible bet on what sessions will look like: mostly reading, occasionally writing.

What I don't trust yet

Independent numbers exist now. Artificial Analysis measured the model at 40 on its Intelligence Index v4.3, sixth of 113 models, and 212 output tokens per second. That is a real eval, but it is one team's read, and it says little about the two parts I care most about: whether CSA2's mode selection keeps quality stable at the far end of the context window, and how the 196B-parameter Engram memory component behaves outside the lab's own tests. Nobody has answered that yet.

The "890 bytes per token" figure deserves a specific caveat. The cache structure changed too, so the size of the cache and the quality of attention over the compressed representations are entangled. If CSA2's per-layer mode selection, choosing between full, reindex, and reuse modes, drops the right tokens, the compression is real. If it drops the wrong ones, you get a model that looks efficient and quietly misses things in the middle of its million-token context. Nobody outside the lab can tell you which one this is yet, and that's a normal situation two weeks after a release of this size. I'm comfortable waiting a few more weeks before drawing conclusions.

I don't fully trust the durability of any of these prices either. Featherless is a reseller, DeepSeek's own docs say prices can shift, and last time the price moved the move was abrupt. Treat the numbers above as today's price, not a law of nature, and re-check them alongside the benchmarks in a few weeks.

One more thing I don't trust: myself. After writing this post I realized I had written "CSA2" in my notes without ever expanding the acronym. I went back and checked: Compressed Sparse Attention 2. I had no idea what the 2 counted. Still don't, exactly. That's the honest state of information two weeks after a release of this size: you build your priors from a model card, an independent eval, and a price page, and update later.

The takeaway

The prices to know right now: $0.03 cached input at the low end, $1.20 output at the high end on Featherless, $0.15 off-peak input on DeepSeek's own API. The gap between reading and generating will keep widening, because the 8B/16B split turns reading into the cheap half of inference. If your workload is input-heavy, that's the release that matters this month. And if your agent sessions aren't input-heavy yet, wait. The architectures now being designed around input-heavy sessions guarantee they will be.