ByteDance Is Pre-Training a 10-Trillion-Parameter Model
ByteDance is training a model with up to 10 trillion parameters — three times larger than the leading Chinese model and comparable to Anthropic Mythos 5.
ByteDance is building the largest language model in the company's history — up to 10 trillion parameters. According to the Financial Times, citing three sources familiar with the project, the model is currently in the pre-training phase, which typically takes three to six months. Development is led by the 2,000-person Seed team.
At that scale, the new model is three times larger than Moonshot's Kimi K3 — the current record holder among Chinese language systems. Analysts estimate it is comparable to Anthropic Mythos 5, which is thought to have roughly 8 trillion parameters. xAI is simultaneously training 6- and 10-trillion-parameter versions of Grok on the Colossus 2 cluster, as Elon Musk has reported.
ByteDance founder Zhang Yiming has personally tasked the team with achieving "world leadership in long-horizon model capabilities." The company deliberately avoided distillation — the technique of training smaller models on the outputs of larger ones — for more than a year, betting instead on original pre-training from scratch.
Parameters are not an absolute measure of quality: training data and methods matter as much as raw size. Even so, the scaling race signals the level of investment at stake: tech companies remain convinced that massive architectures have not yet unlocked their full potential — and that whoever cracks it first will gain a decisive strategic advantage.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →▌ Related guides

Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles