Skip to content
Pipeline Active / Signal #6056 / Auto-Classified
Hype Verified
Industry SIG-6056 / 2026-07-23

Why Is DeepSeek so Cheap? Inside the Claude Distillation Fight

AnalystMoe Sbaiti
PublishedJul 23, 2026 · 11:34 am
Read4 min
Hype Check
Confirmed Signal
10.0/10
Business Impact

Explains why rival models like DeepSeek, Kimi, and Qwen undercut Claude and GPT on price, and flags the compliance risk of building workflows on distilled models with disputed data provenance while Congress considers making distillation a sanctionable offense.

What Did Anthropic Just Accuse Chinese AI Labs Of?

Anthropic says three Chinese AI labs ran an industrial-scale operation to strip-mine Claude’s intelligence. According to the company’s own report, DeepSeek, Moonshot, and MiniMax operated roughly 24,000 fraudulent accounts that together generated more than 16 million exchanges with Claude, with MiniMax alone responsible for about 13 million of them. Weeks later, Anthropic escalated the accusation, saying Alibaba’s Qwen team ran nearly 25,000 fake accounts to pull 28.8 million exchanges between April 22 and June 5, which Anthropic calls the largest known distillation attack against its models to date.

None of this is settled fact in a courtroom. It is Anthropic’s account of what its own detection systems found, corroborated by outside reporting from CNBC and Reuters. Congress is reportedly weighing whether to make this kind of distillation a sanctionable offense, which means the story is still moving.

What’s the Evidence Behind the Claim?

Anthropic does not sell commercial access to Claude inside China for national security reasons. Anthropic says the accused labs got around that by routing traffic through commercial proxy resellers, so the requests looked like ordinary retail users logging in from permitted regions. When one account hit a rate limit or got flagged, the operation allegedly shifted the workload to a fresh account and kept going.

What makes this different from someone just using Claude for research is the target. Anthropic says the accused labs specifically prompted Claude to walk through its chain-of-thought reasoning on hard problems, math, code, multi-step agent tasks, rather than just asking for final answers. That reasoning trail is the valuable part. It’s the data you need if you want to train a cheaper model that thinks the same way without redoing years of research to get there.

How Does Distillation Actually Work?

Distillation itself is not the scandal. Every major lab uses it. It’s how you take a massive, expensive “teacher” model and train a smaller, faster “student” model to behave almost as well, at a fraction of the compute cost. That’s how Claude Haiku, GPT’s mini-tier models, and Gemini Flash all exist.

The process has two ingredients most people never hear about. The teacher doesn’t just hand over its final answer, it hands over the full probability distribution behind that answer (soft labels), plus a written explanation of how it got there (chain-of-thought rationale). The student model is trained to match both, using a loss function that blends how close its probabilities are to the teacher’s with how often it gets the final answer right. Do that across millions of examples, and the student inherits a compressed version of the teacher’s reasoning.

Every lab does this to its own models. What Anthropic is alleging is different: doing it to a rival’s paid product, at scale, through accounts built specifically to dodge the access restrictions and rate limits that were supposed to prevent exactly this.

Picture a rival restaurant sending 25,000 different customers through your dining room over six weeks, each one ordering your signature dish and quietly writing down the plating, the seasoning ratios, and the exact cook time. No single ticket looks suspicious on its own.

The pattern only shows up once you lay all 25,000 receipts side by side. That is roughly what Anthropic says happened here. Nobody stole a file off a server. They allegedly ordered the same intelligence millions of times through disguised accounts, then reconstructed the recipe from the pattern of orders.

The lesson for a small business owner isn’t about geopolitics. It’s about knowing what’s actually inside the tool that’s priced 80% below the market leader. Cheaper isn’t free. Someone paid for that intelligence, it just might not have been the company charging you for it now.

What Does This Mean for Your Day-to-Day AI Stack?

If you’re choosing between a premium model and a much cheaper competitor, this story is a reason to ask a few direct questions before you commit a workflow to it. Ask the vendor where their training data came from and whether they’ll say so in writing. A vendor that gets specific is behaving differently than one that deflects.

Check whether the tool’s behavior is consistent under pressure, not just on the demo. A distilled model can inherit blind spots from its teacher, or skip the safety tuning the original lab spent months on, and you won’t see that gap until you hit an edge case with a real customer.

Watch the regulatory story even if you never touch a Chinese model. If lawmakers do move to sanction this kind of distillation, any vendor built on disputed training data becomes an availability risk, not just an ethics question. That’s a business continuity problem, not an abstract one.

Verdict

A rock-bottom price on an AI tool is not proof of a leaner cost structure, and treating it that way without checking the vendor’s data practices is a bet you’re placing on facts you can’t see. Ask where the training data came from, test the tool on your messiest real task before you trust it on your cleanest one, and keep an eye on the regulatory fight, because it could decide whether your cheap tool is still around next year.

For more signals like this, browse the Signals archive.

Source: Anthropic

Moe Sbaiti
Moe Sbaiti AI Intelligence Analyst

I run 4 businesses simultaneously. The pipeline behind The AI Profit Wire monitors 100+ sources every 4 hours, scores every signal against 5 measurable data points, and cuts 98.9% of the noise before anything reaches you. My background is 16 years of restaurant operations, ecommerce, fitness coaching, and web development. I evaluate tools like a business owner, not a tech reviewer. Hype scores never bend for affiliate relationships. The data decides.

Subscribe to the Wire