CourionAI
EN
Newsletter
← All news
ai-basics 2 min read

TikTok Owner ByteDance Is Reportedly Training China's Largest AI Model Yet

Three insiders told the Financial Times that ByteDance has a model with up to ten trillion parameters in pretraining, three times the size of Kimi K3.

An enormous lattice dome of connected nodes under construction, dwarfing the cranes and scaffolding around it

ByteDance, the company behind TikTok, is training an AI model with up to ten trillion parameters, according to a Financial Times report citing three people familiar with the work. If the number holds, it would be roughly three times the size of Moonshot’s Kimi K3, currently the largest Chinese model, and would put ByteDance in the same weight class as Anthropic’s Mythos 5, which industry estimates place around eight trillion. Anthropic has never published its own figure.

Parameters are the adjustable numbers inside a model, the settings it tunes during training to encode what it has learned. More parameters means more room to store patterns, so the count is a rough proxy for capacity. Rough is the operative word. Data quality, training method and how long the model trains matter at least as much, and the industry has repeatedly seen smaller, better-trained models beat larger sloppy ones. Treat parameter counts the way you would treat engine size in a car: informative, not decisive.

The model is reportedly in pretraining, the first and most compute-hungry phase, where it works through enormous amounts of text before any fine-tuning happens. That phase typically runs three to six months, so a release would be some way off. One source added a detail worth flagging: ByteDance has avoided distillation for over a year. Distillation means training your model on the outputs of somebody else’s model, a shortcut that gets you decent quality cheaply and has become a sore point between labs. Choosing not to use it is a claim about doing the expensive thing properly, though it is a claim from an anonymous source rather than a published fact.

The bigger picture is a return of raw scale as a competitive lever. Founder Zhang Yiming reportedly told ByteDance’s 2,000-person Seed team to aim for world-leading model capabilities over the long term. Elon Musk says xAI is training Grok variants at six and ten trillion parameters on its Colossus 2 cluster. After a couple of years in which efficiency and clever training were the fashionable story, several labs are once again betting that bigger, at enormous cost, still wins.

What this means for you: for now, nothing practical. This model does not exist as a product, may never be released outside China, and would face export and regulatory questions in the EU if it were. What it is useful for is calibration. When you next read that a model has some eye-watering parameter count, you now know that the number describes capacity, not quality, and that the interesting question is what was trained on and how. If you follow open-weight models, keep an eye on this one anyway: ByteDance has released open weights before, and a Chinese lab at ten trillion parameters would reset expectations for everyone.

Sources

Source: https://www.ft.com/content/9b8383b1-a28d-4940-8c4e-2f0cd21556ef

Next story

An OpenAI Developer Says Now Would Be a Good Time to Clean Up Your Exposed Passwords

The warning is blunt: anything sensitive sitting in public on GitHub or Pastebin is about to be found much faster, because models are cheap enough to look everywhere at once.

A single brass key lying in a pool of light on a doorstep, with a dark hedge behind it full of tiny watching eyes