ByteDance is reportedly training a ~10 trillion‑parameter model to chase the frontier

Multiple news outlets reported this week that ByteDance — the Chinese owner of TikTok — is training a very large AI model that industry sources told the Financial Times could reach about 10 trillion parameters. The project is reportedly in a pre‑training phase that typically lasts several months; ByteDance has not publicly confirmed the size or schedule.

Key takeaways

  • Multiple outlets cite Financial Times sources saying ByteDance is pre‑training a model that could reach about 10 trillion parameters.
  • The figure is an industry estimate and has not been confirmed by ByteDance; pre‑training is said to take several months before fine‑tuning.
  • Parameter count is an imperfect proxy for capability; architecture, data and training methods matter as much or more.
  • If true, a 10‑trillion model would outsize recent Chinese releases like Kimi K3 and approach some industry estimates of Anthropic’s Mythos class.

What the reports say

According to multiple summaries of the Financial Times story, people familiar with the effort told the FT that ByteDance has started pre‑training a model that could grow to around 10 trillion parameters. Those outlets note the figure is not final and that the model would move from pre‑training to fine‑tuning and testing before any public release. Several articles explicitly caution the reports could not be independently verified.

How big is "10 trillion" in context?

Parameter counts have become a common shorthand for scale in the industry, though experts routinely warn they are an imperfect proxy for performance. Industry reporting and benchmark trackers place some recent Chinese models and leading western systems in the trillions‑of‑parameters range:

  • Moonshot AI’s Kimi K3 is frequently reported at about 2.8 trillion parameters, making it one of the largest publicly discussed Chinese models so far.
  • Anthropic’s frontier Mythos class is widely estimated by industry observers to sit in the multiple‑trillion range; some reports cited in the press place Mythos 5 near 8 trillion parameters, though Anthropic does not publish parameter counts.
  • By comparison, a ~10 trillion‑parameter model would be more than three times Kimi K3’s reported size and near or above some industry estimates for Anthropic’s top systems.

What parameter count does — and doesn’t — tell us

Parameters are the learned numerical values inside a model and higher counts can enable representation of more complex patterns. But model capability depends on many other factors: architecture, training data quality, fine‑tuning techniques, compute efficiency, safety alignments and the suite of evaluation tests used.

Reporting about previous large Chinese models has repeatedly emphasized that efficiency techniques, activation sparsity and selective routing can reduce the operational cost of very large networks. Headlines that only report parameter totals risk overstating what the system can actually do in real‑world tasks.

Where ByteDance fits in China's AI push

ByteDance is already widely reported to run high‑profile multimodal and media generation models, and some of the reporting points to two advantages if it pursues a frontier model: a large user base through Doubao and TikTok for distribution and abundant in‑house experience with multimodal data. A German tech outlet quoted the FT’s reporting and noted ByteDance’s recent Seedance video models and a large internal team working on general models.

Timelines, verification and open questions

  • Stage: Outlets say the project is in pre‑training, a phase that independent reports estimate often takes three to six months before fine‑tuning.
  • Confirmation: None of the reports cite an official ByteDance statement confirming the parameter count or a release timeline. Several stories explicitly state they could not independently verify the FT’s source‑based claims.
  • Hardware and cost: Training at this scale requires substantial compute. Past reporting about other labs suggests tens of thousands of accelerator units can be involved when building multi‑trillion‑parameter systems, but no hardware figures for this ByteDance project were reported.

Implications for the global AI landscape

If ByteDance proceeds with a model at this scale, it would represent a clear signal of intent by a major Chinese internet company to compete at the frontier. Industry coverage places this development in a broader pattern of rapid Chinese model launches and improving benchmark performance across multiple domestic labs, which some analysts say is narrowing the gap with US labs on price and specific capabilities.

At the same time, the reporting underscores a continuing caveat: parameter counts alone do not settle which models lead in reasoning, coding, safety or other applied metrics used by enterprise customers and researchers.

Practical takeaways for US and international observers

  1. Treat the 10‑trillion figure as an unverified, industry‑sourced estimate rather than a confirmed specification.
  2. Assess capability by empirical benchmarks and task‑specific evaluations rather than parameter totals alone.
  3. Watch for official disclosures from ByteDance on architecture, evaluation results and access policy; those details matter for commercial, academic and regulatory decisions.
  • Recent months: Chinese labs have released or iterated large models (examples previously reported include Kimi K3 and several Qwen releases) that analysts say improved long‑horizon agent workflows and lowered per‑task costs on some benchmarks.
  • Late July–early August 2026: Financial Times reporting and subsequent summaries by multiple outlets identify a ByteDance project in pre‑training that could reach roughly 10 trillion parameters; outlets note the number is not yet final and that ByteDance has not confirmed it.

“The model is reportedly in the pre‑training phase and the final size has not been fixed,” — Financial Times, as reported by multiple outlets.

Reporting on this story relies on industry sources cited by the Financial Times and subsequent press summaries. We will update coverage when ByteDance or independent benchmarkers publish confirmatory technical details, evaluation results, or access plans.