ByteDance trains 10 trillion-parameter model, founder Zhang Yiming opposes distillation
The Financial Times reported on August 7 that ByteDance is pre-training an AI model with up to 10 trillion parameters, approaching Anthropic Mythos 5 in scale. Founder Zhang Yiming told the Seed team to reject distillation.
On August 7, the Financial Times reported, citing three people familiar with the matter, that ByteDance is pre-training an AI model with as many as 10 trillion parameters. The model is in early pre-training, which typically takes three to six months, and the final parameter count has not been locked in.
If the upper bound holds, the model would be more than three times the size of Moonshot AI Kimi K3, currently the largest open-weight model at 2.8 trillion parameters, and would approach Anthropic Mythos 5 in scale, placing it among the largest training runs ever publicly disclosed. The project is led by Xiang Liang, head of ByteDance Seed Foundation, alongside Shen Ke, who leads LLM pre-training data. Both are veterans of ByteDance search, ads, and recommendation systems.
Seed has roughly 2,000 members across China and overseas, covering research, infrastructure, and data labeling. Founder Zhang Yiming told a Seed all-hands two weeks ago that the company rejects model distillation, arguing that distillation replicates the capabilities of frontier models like Claude rather than creating new breakthroughs, and that ByteDance would rather fall behind domestic competitors temporarily than take that shortcut.
Context matters: more than half of Doubao token consumption comes from video generation model Seedance and image generation model Seedream, not from language models. Seedance growth slowed in Q2, dragging down Volcano Engine MaaS token volume from 120 trillion in March to 180 trillion in June, below the original 250-300 trillion target. ByteDance 2026 capex guidance is above $30 billion.