A data-driven look at China's most competitive open AI model — and the story everyone's getting wrong about it
You've probably heard some version of this: a Chinese lab built an AI model almost as good as America's best, on hardware that's supposedly cheaper and weaker. The real story — once you check the hardware, the benchmarks, and the copying allegations against the actual record — is messier, and more interesting, than that.
Moonshot AI's Kimi K3 — a 2.8-trillion-parameter open-weight model released in July 2026 — is one of the largest AI models ever built, and it's genuinely competitive with the best American systems on several hard benchmarks. But three separate claims about it keep getting bundled together into one popular narrative: that it runs on cheap hardware, that it's a fair fight with the American frontier, and that China simply copied its way there. Only one of those three survives contact with the actual sourcing. Here's the data, section by section.
Every serious AI lab in the world — American or Chinese — still builds its models on Nvidia GPUs, and U.S. export rules try to keep the newest, most powerful ones out of China. Multiple reports say Moonshot trained Kimi K3 on Nvidia's restricted Blackwell chips anyway, obtained through Chinese datacenter firms that sidestepped both U.S. export controls and China's own import rules by joining multiple eight-GPU Blackwell servers together across facilities.
For day-to-day inference, Moonshot uses the H20 — a legal, export-approved but deliberately weakened Nvidia chip, still needing at least 64 GPUs per server. What actually got cheaper is how efficiently Kimi K3 uses that hardware: it's a Mixture-of-Experts model that wakes only about 1.8% of its total parameters for any given task, which is real engineering efficiency — just not the "different, much cheaper hardware" story that gets repeated.
Standardized tests called benchmarks are how the industry compares models — but two things about them matter more than the headlines suggest. They wear out fast: as soon as a test is published, labs train toward it, so scores cluster near the ceiling within months. And being right more often isn't the same as being trustworthy — Kimi K3's rising hallucination rate alongside its rising rank is the clearest example of that gap in this dataset.
On raw scores, Kimi K3 leads outright on several hard tests — SWE Marathon, DeepSearchQA, and Frontend Code Arena — while trailing narrowly on BrowseComp, GPQA-Diamond, and Terminal-Bench, and trailing by a wide margin on Humanity's Last Exam. It's genuinely competitive. It is not uniformly ahead.
Two separate claims get conflated here, and the evidence treats them very differently. China's software-piracy history is real — but "never pays for software" overstates it; piracy has been shrinking, and China also has a large legitimate software market.
Distillation — training a new model by learning from an already-trained model's outputs, rather than building from scratch — is a documented pattern among Chinese AI labs, most convincingly in OpenAI's formal case against DeepSeek. Anthropic and a White House science advisor made a similar allegation against Kimi K3 specifically. But independent researchers pushed back hard on that one: K3 launched only about two weeks after the Claude model it supposedly copied, which most experts say is far too short a window for that to be technically possible at K3's scale.
Moonshot has not confirmed a training cost for K3 or publicly responded to the Blackwell-acquisition or distillation allegations cited here — several claims rest on U.S. government or competitor statements, not Moonshot's own confirmation. Benchmark figures come from third-party leaderboards using different methodologies; treat as directional, not exact to the decimal.
This is the abstract argument behind the K3 story above — why exponential AI progress like this normalizes into background noise before we've actually registered what changed.