ByteDance Is Pre-Training a 10 Trillion-Parameter Model — Bigger Than Anything China Has Shipped
The Financial Times reports ByteDance is training a model of up to 10 trillion parameters, more than three times the size of Moonshot's Kimi K3. It would put a Chinese lab in the same parameter class as Anthropic's Mythos 5.
ByteDance is pre-training an AI model with as many as 10 trillion parameters, the Financial Times reported Friday, citing people familiar with the effort. If it ships at that scale, it is the largest model any Chinese lab has attempted by a wide margin.
The comparisons are what make the number land. Moonshot AI’s Kimi K3, the current Chinese size leader, is 2.8 trillion parameters. DeepSeek’s V4-Pro and Meituan’s LongCat-2.0 both sit at 1.6 trillion. ByteDance is going more than 3x past the top of its domestic field in a single jump — and roughly 6x past the models that defined Chinese frontier work six months ago.
Against Western labs, industry estimates put Anthropic’s Mythos 5 near 8 trillion parameters and its guardrailed sibling Fable 5 around 5 trillion. A 10-trillion ByteDance model would not just close the parameter gap. On raw count it would clear it.
Two caveats before anyone rewrites the leaderboard. Reuters could not independently verify the FT’s reporting, and ByteDance declined to comment. And parameter count stopped being a clean proxy for capability a long time ago — sparsity, active-parameter ratios in MoE architectures, data quality, and post-training now decide where a model lands on benchmarks. A 10T sparse mixture-of-experts with 400B active parameters is a very different machine from a 10T dense model, and nobody outside ByteDance knows which one this is.
The timeline is the concrete detail. The model is in pre-training now, a phase that typically runs three to six months before fine-tuning and release. That puts a plausible launch window somewhere between late 2026 and early 2027 — assuming the run doesn’t diverge, which at this scale is not a small assumption.
The strategic read is more interesting than the number. ByteDance has the two things this requires and most Chinese labs don’t: recommendation-scale infrastructure already built and paid for, and revenue from TikTok and Douyin that doesn’t depend on the model succeeding. DeepSeek made its name on efficiency under export-control constraints — doing more with less because less was what it had. ByteDance is running the opposite play. It is spending, at frontier scale, on a bet that raw capacity still buys capability.
Both approaches are responses to the same compute ceiling. One optimizes around it. The other tries to buy enough of what’s available to make the ceiling irrelevant.
The number to watch isn’t 10 trillion. It’s whatever active-parameter count and benchmark suite ByteDance publishes when the model lands — and whether the run finishes at all.