ByteDance Trains a 10 Trillion Parameter AI Model
3 min readThe largest AI model anyone has publicly attempted is being trained in China, and it belongs to TikTok’s parent company. ByteDance is pre-training a system with up to 10 trillion parameters, roughly three times the size of Moonshot AI’s Kimi K3 and larger than industry estimates for Anthropic’s Mythos 5.
Scale Was Supposed to Be Over
Parameters are the internal weights a model learns during training, and for years more of them meant a more capable system. That assumption softened through 2025 and 2026 as labs found gains elsewhere: reasoning at inference time, better data, and distillation, the practice of training a small model to imitate a much larger one. The result was a wave of cheap, small, fast models. DeepSeek’s V4 Flash, released days ago, is the clearest example, delivering near-frontier coding performance at a fraction of the cost.
ByteDance appears to be betting the other way. Founder Zhang Yiming has reportedly instructed teams to pursue genuine capability rather than lean on distillation, a direct rejection of the shortcut that made many Chinese models so cheap to serve.
What the ByteDance AI Model Involves
According to Reuters, the model is still in early pre-training. That stage typically runs three to six months before fine-tuning begins, so a release is not imminent and the final parameter count could move. Industry estimates put Anthropic’s Mythos 5 at around 8 trillion parameters, which would make ByteDance’s run the largest disclosed anywhere.
The company has advantages suited to a project this size: enormous proprietary data from its consumer apps, heavy infrastructure spending, and revenue that does not depend on selling API tokens. It also faces the constraint every Chinese lab faces. Export controls limit access to the most advanced accelerators, so a 10 trillion parameter run implies either a very large cluster of permitted chips or considerable efficiency engineering.
Why It Matters
China’s labs have spent the past year competing on price and open weights. A frontier-scale training run is a different move. It targets capability leadership rather than cost leadership, and it would put a Chinese company at the top of the size table for the first time.
The open question is whether the weights get released. ByteDance has been less committed to open weights than DeepSeek, Alibaba or Moonshot. A 10 trillion parameter open release would reshape what is freely available worldwide. A closed one would simply make ByteDance a frontier lab. Watch which way it goes.
