ByteDance trains 10-trillion-parameter AI model to rival Anthropic
The TikTok parent company is in the pre-training phase of a massive new model, rejecting industry-standard distillation techniques in favour of a long-term push for world-leading capabilities.

ByteDance is training an artificial intelligence model with up to 10 trillion parameters, a scale designed to directly challenge Anthropic’s cutting-edge Mythos system. According to three people with knowledge of the matter, the project is currently in the pre-training stage, an initial phase that typically spans three to six months. This effort highlights the accelerating ambition of Chinese technology laboratories to not only close the gap with leading US firms but to outperform them in advanced AI capabilities.
The proposed model size significantly exceeds current domestic benchmarks, standing at approximately three times the size of Moonshot’s Kimi K3, which remains the largest Chinese model released to date. While ByteDance has not disclosed the final parameter count, industry estimates suggest Anthropic’s most advanced Mythos 5 system contains roughly 8 trillion parameters, with its Fable 5 model holding around 5 trillion. Although parameter count establishes the fundamental capacity for information storage, actual model performance remains dependent on data quality and training methodologies.
ByteDance’s approach diverges from many competitors by eschewing model distillation, a process that compresses large, complex models into smaller, faster versions by training them to mimic existing outputs. The company’s Seed team, led by former Google DeepMind scientist Wu Yonghui, has maintained this independent development strategy for over a year. The team comprises approximately 2,000 members across China and overseas, including core researchers, infrastructure engineers, and data labelling specialists.
Founder Zhang Yiming has reinforced this strategic direction, directing the Seed team to target world-leading model capabilities in the long term. During an internal meeting two weeks ago, Zhang reportedly advised the team not to be overly concerned with short-term competitive positioning, asserting that independent development is the only path to producing a model that outperforms rivals. This stance contrasts with the faster iteration cycles seen among some peers, though ByteDance’s management remains committed to its rigorous development standards.
The investment in large-scale AI infrastructure is part of a broader push by ByteDance to strengthen its enterprise offerings and technological moat. The company has aggressively expanded its cloud unit, Volcano Engine, which provides AI solutions to businesses, and is exploring the development of custom AI chips. Meanwhile, its consumer-facing model, Doubao, continues to dominate the local market with 324 million monthly active users, providing a substantial user base for data and feedback loops.
ByteDance has kept a low profile regarding its underlying AI research, keeping most models closed unlike some domestic peers. However, its recent advancements in video generation and its aggressive capital expenditure in data centres signal a serious commitment to the frontier of generative AI. The company did not respond to a request for comment regarding the specifics of the new model.
