Analysis

DeepSeek just repriced the week. Nvidia boxed it in six days.

V4.1-Flash is an MIT-licensed, million-token, cache-hit bargain. The closed labs still sell the same tokens like a hotel minibar. Blackwell got a four-bit reprint before the memes died.

NVIDIA CEO Jensen Huang on stage at GTC Japan, 5 October 2016. Photo by Masaru Kamikura, CC BY 2.0, via Wikimedia Commons.
Photo: Masaru Kamikura / Wikimedia Commons (CC BY 2.0)

DeepSeek put V4.1-Flash on the public internet on September 10: a 552-billion-parameter mixture-of-experts, native vision, a million-token window, MIT license. The price list is the part that made procurement Slack light up. Off-peak cache-hit input is listed at $0.003 per million tokens; cache-miss $0.15; output $0.60. Peak hours cost double. DeepSeek’s own notes say the new Flash comprehensively beats last quarter’s V4 Pro and will absorb that traffic after September 14.

I am not asking you to trust a vendor slide. I am asking you to look at the other menus. OpenAI’s GPT-5.6 Sol list price, as cited in contemporary write-ups of the Flash launch, still lives in hotel-minibar country. Anthropic’s Opus 5 cache-hit is not a rounding error next to three-tenths of a cent. You do not need a perfect benchmark to understand a 100× gap. You need an invoice.

The architecture paper is for researchers. The cache-hit rate is for whoever pays the bill.

What the weights actually change

KDnuggets and DeepSeek’s own card emphasize the plumbing: causal encoder-decoder, compressed sparse attention, FP4 KV cache, tricks that make a huge MoE cheap to serve. Agent benches — DeepSWE, CyberGym, terminal use — are where Flash claims to stand next to Opus and Sol, not just next to last year’s open models. If those numbers hold outside the lab, the “open is for toys” talking point is now a luxury belief.

Nvidia published DeepSeek-V4.1-Flash-NVFP4 on September 16: a four-bit Blackwell build, four GB300 GPUs, no retraining, MIT, accuracy within a point and a half on the six benches they showed. DataNorth’s read is the right one. This is not a new model. It is a statement that a Chinese open-weight drop is now a first-class Nvidia SKU inside a week. GLM-5.3-Flash got the same treatment in August. The gap between “weights on a Tuesday” and “enterprise-shaped checkpoint on a Monday” is collapsing.

The closed answer is still a price cut or a story

Frontier labs have two honest replies. Ship something that is not a token, or cut the token. Astra’s Daybreak badge and Claude’s Mythos velvet rope are the first reply: sell access, not volume. Flash is the second reply happening to them. If your mid-tier API still assumes the customer cannot run a million tokens at home, you are pricing a habit that DeepSeek is training people to break.

Watch three meters: cache-hit rates in the wild, whether V4.1 Pro arrives as a tax on the people who need it, and whether the next Nvidia conversion is announced like a product or like a weather report. The weather is: open weights now show up already quantized for the only GPUs that matter.