ByteDance Trains 10-Trillion Parameter AI Model
ByteDance has begun training a massive artificial intelligence model with up to 10 trillion parameters, signaling a major push by Chinese tech firms to surpass top American AI labs.

The parent company of TikTok is in the early phases of pre-training a massive large language model designed to scale up to 10 trillion parameters. This pre-training phase is expected to last between three and six months. If successful, the resulting model would be three times larger than Moonshot's Kimi K3, which currently stands as the largest Chinese model released to date. It would also surpass the estimated size of Anthropic's most advanced model, Mythos 5, which industry experts believe contains roughly 8 trillion parameters, as well as its 5-trillion-parameter Fable 5 model.
This ambitious project highlights the rapid progress of Chinese developers, whose models from Moonshot and Alibaba have recently shown competitive benchmark performance, trailing only Fable 5 in certain categories. While Mythos 5 remains restricted to approved organizations following a temporary ban last June, ByteDance is positioning itself to lead the domestic market. The company already operates the popular Doubao consumer model, which boasts 324 million monthly active users, alongside its advanced SeeDance video-generation system.
To support these scaling efforts, ByteDance has spent three years aggressively expanding its infrastructure, including its Volcano Engine cloud division. The model development is being handled by Seed, a 2,000-member team led by former Google DeepMind scientist Wu Yonghui. Notably, ByteDance has rejected model distillation—the practice of training smaller systems to mimic larger, existing models—for over a year. According to reports from Latepost and The Information, founder Zhang Yiming recently urged the Seed team to focus on achieving "world-leading model capabilities" rather than worrying about short-term delays.
For AI practitioners, ByteDance's commitment to independent, non-distilled training at this scale represents a significant shift. If a 10-trillion-parameter model succeeds without relying on Western teacher models, it could validate a highly independent development path for Chinese AI. This could ultimately provide global enterprises with a powerful, culturally distinct alternative to US-centric frontier models, intensifying competition in the enterprise cloud and API markets.
This is our own summary of reporting by Ars Technica AI



