Skip to content
Pipeline Active / Signal #6626 / Auto-Classified
Hype Verified
Breaking SIG-6626 / 2026-09-10

DeepSeek V4.1 Flash Charges A Fraction Of A Cent Per Million Tokens

AnalystMoe Sbaiti
PublishedSep 10, 2026 · 5:24 pm
Read4 min
Hype Check
Worth Watching
6.6/10
Business Impact

Cuts the cost of routing AI features through an API, from support triage to document summarization, and hands premium-provider customers a documented floor rate for renewal negotiations.


What is DeepSeek V4.1 Flash and what does it do?

DeepSeek V4.1 Flash is the model the Chinese AI lab released on September 10, 2026, with API pricing that charges as little as a fraction of a cent per million tokens. Bloomberg’s launch coverage frames the release as a fresh blow to OpenAI and added pressure on rivals from Anthropic to Z.AI.

Under the hood it is a 552B-parameter mixture-of-experts model with a new causal encoder-decoder architecture that activates 8B parameters for input and 16B for output. DeepSeek’s own release note says it ships with native visual understanding and claims benchmark results ahead of flagship models, including its own V4 Pro.

DeepSeek V4.1 Flash is a production API tier priced like a rounding error, and it is live on the API today.

Is DeepSeek actually cheaper than OpenAI?

On the floor rate, yes, and the floor is documented on the company’s pricing page. Cache-hit input costs $0.006 per million tokens at peak and $0.003 off-peak, which is the fraction of a cent Bloomberg referenced.

The rest of the meter matters as much as the floor. Cache-miss input runs $0.15 per million tokens off-peak and $0.30 at peak, and output runs $0.60 to $1.20 across the same windows.

Off-peak rates sit at half of peak, with peak hours covering weekday windows of 01:00 to 04:00 and 06:00 to 10:00 UTC. A workload you schedule at night bills at half the rate of the same workload at midday.

The market treated the pricing as an attack. Bloomberg reports that MiniMax and Z.AI shares closed down more than 8% in Hong Kong on the news, while Alibaba slid more than 2%.

DeepSeek’s claim that V4.1 Flash outperforms pricier rivals like Moonshot’s Kimi K3 comes from the vendor, so treat it as a claim until your own workload confirms it. The independent model card and technical report on Hugging Face give you a place to start that check.

The cheap rate is real, and it lands on the cache-hit slice of your bill, which is where high-volume agent workloads spend most of their tokens.

DeepSeek V4.1 Flash vs V4 Pro: which should you use?

V4 Pro is the wrong comparison to agonize over, because DeepSeek is retiring it. The company’s API docs state that V4 Pro requests route to V4.1 Flash at Flash pricing from 12:00 Beijing time on September 14, 2026.

The price gap explains the decision. V4 Pro bills $1.32 per million input tokens at peak and $3.96 per million output tokens, against Flash at $0.30 and $1.20 for the same hours.

DeepSeek’s own docs conclude that V4.1 Flash surpassed V4 Pro in performance, cost, speed, and total time, which is the vendor’s stated reason for the retirement. Capacity holds up as well, with a 2,500 concurrent request limit on the Flash tier against V4 Pro’s 500.

Context length is 1M tokens with a 384K maximum output on both tiers. V4.1 Flash supports vision, while V4 Pro does not.

If you were weighing V4 Pro against Flash, the vendor closed that question on September 14.

Who is DeepSeek V4.1 Flash actually for?

It fits founders and teams routing high-volume, commodity AI work through an API: support triage, classification, summarization, data extraction, and internal copilots. The KV cache on V4.1 Flash needs a quarter of the HBM of the previous generation, and DeepSeek notes that cache-hit charges often account for a large share of agent costs, which is the economic case for the whole release.

Premium buyers get something out of it too. Anyone negotiating with a higher-priced provider now holds a documented floor rate to negotiate against, and the pricing page behind it is public.

The pattern is bigger than one model, and we track it as it lands: the audited signal trail behind this launch shows where the cost floor has moved across the AI market this quarter.

High-volume commodity workloads win first, and every other buyer wins a negotiating number.

The invoice from your AI vendor lands on the 1st, and the line that doubled is the one nobody can explain. Support triage, document summaries, the lead-classification job someone shipped in March: every one of them bills by the token now.

V4.1 Flash prices the cache-hit slice of that bill at $0.006 per million tokens, half that off-peak. That is the number that moves the budget meeting from whether to test it to who owns the replay run.

The catch sits in the rest of the meter, with cache misses at $0.15 to $0.30 and output at $0.60 to $1.20 per million. The floor is real, and so is the distance between the floor and a careless integration.

Should you switch to DeepSeek?

Run a controlled test on your own production data before any migration. Replaying a month of real call volume through the Flash tier costs close to nothing at these rates, and it produces the one benchmark that matters: output quality on your specific workload.

Watch the 3 billable levers while you test. Your cache-hit ratio moves you between the fraction-of-a-cent rate and the $0.15 to $0.30 cache-miss rate, off-peak scheduling halves the tab, and retry loops bill twice.

A token, as DeepSeek’s own docs define it, can be a word, a number, or a punctuation mark, so a month of support transcripts is millions of them. The arithmetic that felt marginal at premium rates flips at a $0.006 floor.

Switch the commodity workloads that pass the replay test, keep the judgment workloads where they are, and put the pricing page in your renewal negotiation either way.

Source: Bloomberg Tech

Moe Sbaiti
Moe Sbaiti AI Intelligence Analyst

I run 4 businesses simultaneously. The pipeline behind The AI Profit Wire monitors 100+ sources every 4 hours, scores every signal against 5 measurable data points, and cuts over 90% of the noise before anything reaches you. My background is 16 years of restaurant operations, ecommerce, fitness coaching, and web development. I evaluate tools like a business owner, not a tech reviewer. Hype scores never bend for affiliate relationships. The data decides.

Subscribe to the Wire