Xiaomi livestreams MiMo‑V2.6 RL run: $3M+ training, open weights and a public dashboard

Xiaomi has publicly released MiMo‑V2.6 — a family of MIT‑licensed models (MiMo‑V2.6‑Pro‑RL, MiMo‑V2.6‑Flash‑RL and MiMo‑V2.6‑Distill‑Qwen‑9B) — and is streaming a reinforcement‑learning (RL) training dashboard that reports real‑time costs, token throughput and benchmark progress. The company also published more than 7,000 RL environments and an end‑to‑end training framework intended for agentic AI research.

What Xiaomi published and what it livestreamed

According to Xiaomi’s release materials and reporting that reviewed the repositories, the MiMo‑V2.6 family includes:

  • MiMo‑V2.6‑Pro‑RL — flagship multimodal checkpoint supporting text, image, video and audio inputs with a 1,000,000‑token context window (company‑reported).
  • MiMo‑V2.6‑Flash‑RL — a smaller, efficiency‑focused variant with the same 1,000,000‑token context window (company‑reported).
  • MiMo‑V2.6‑Distill‑Qwen‑9B — a supervised fine‑tuned checkpoint based on Qwen3.5‑9B intended as a research starting point (company‑reported).

Alongside weights, Xiaomi published a reinforcement‑learning framework and more than 7,000 task environments covering areas such as software engineering, vulnerability reproduction, knowledge‑intensive tasks and web development, with the repositories listing MIT licences.

Livestreamed RL run: scale, costs and mid‑training performance

Multiple reports summarise Xiaomi’s livestreamed RL dashboard and training metrics. Xiaomi’s public run reportedly used 1,568 prompts and 16 asynchronous rollouts per training step, producing billions of tokens per step and very long trajectory sequences (company figures). Reported token processing and cost figures in public coverage include:

  • A mid‑run snapshot showing the Pro run had processed 32.5 billion tokens while a concurrent Flash run had processed 49.4 billion tokens (forkast reporting of Xiaomi data).
  • Daily burn‑rate style figures reported in coverage: approximately $432,000 per day for a MiMo‑V2.6‑Pro livestreamed run, and a $512,000 figure associated with a concurrent run in one report; separate company statements and technical report totals estimate about $2.62 million for the Pro RL phase and $850,000 for Flash (company‑reported totals quoted in sources).

Xiaomi also reported that during the RL phase the MiMo‑V2.6‑Pro checkpoint reached 65.97% on the DeepSWE v1.1 benchmark — a 47‑point increase from a MiMo‑V2.5 baseline that was reported at 19% (company‑reported mid‑run measurements cited in reporting). Independent benchmark results published by third parties evaluated the released Pro model: Artificial Analysis scored MiMo‑V2.6‑Pro at 46 on its Intelligence Index in one set of tests (third‑party evaluation).

Open model economics and deployment notes

Reports highlight Xiaomi’s positioning of Pro for complex, long‑horizon agent tasks and Flash for high‑volume workloads. Published API pricing examples referenced in coverage include:

  • MiMo V2.6 Pro listed rates cited on one platform at $0.435 per million uncached input tokens and $0.87 per million output tokens, and MiMo V2.6 Flash at about $0.14/$0.28 per million input/output tokens (third‑party reporting of Xiaomi‑listed API prices).
  • Artificial Analysis’s cost measurement placed Pro at roughly $0.13 per Intelligence Index task, combining token usage and listed prices (third‑party metric).

Downloadable weights are MIT‑licensed and available on Hugging Face according to reporting, but deploying Pro or Flash locally still requires substantial hardware: Pro is described as a 1.02‑trillion‑parameter mixture‑of‑experts model with 42 billion active parameters per inference pass; Flash around 310 billion total parameters with about 15 billion active (company figures reported in coverage). Serving examples use multi‑way parallelism (for example, 16‑way for Pro in some examples), underscoring the resource costs of self‑hosting.

What is independently verified and what remains to be replicated

Independent third‑party testing confirms that the downloadable MiMo‑V2.6‑Pro checkpoint performs strongly on published benchmarks such as the Artificial Analysis Intelligence Index (score 46 in that dataset). However, several elements remain company‑reported and have not yet been independently reproduced in public literature:

  • The specific mid‑training RL dashboard numbers (real‑time burn rates, per‑step tokens and exact trajectory counts) are presented by Xiaomi and reported in livestream coverage; independent replication of the full RL system, including grader/harness compute and end‑to‑end RL cost breakdowns, has not been published.
  • Claims about the internal composition of runs — for example, the 1,568‑prompt / 16‑rollout per step configuration and token counts per step — come from Xiaomi’s technical descriptions and press reporting; reproductions would require access to the RL environments, compute profile and grading harnesses Xiaomi published.

How this adds value — a short reproducibility checklist for teams

For organisations that want to reproduce Xiaomi’s RL steps or validate the claimed mid‑run gains using the published assets, the following checklist organises the minimum required actions based on the materials Xiaomi released and the reporting:

  1. Obtain the MiMo‑V2.6‑Pro or Flash weights from the published Hugging Face repositories and confirm the MIT licence metadata.
  2. Download Xiaomi’s RL framework and the subset of the 7,000+ environments you plan to run (note the environments include software engineering, vulnerability reproduction, web development and knowledge‑intensive tasks).
  3. Provision compute consistent with reported scale: plan for multi‑machine serving (examples show 8–16 parallel workers), and memory for trillion‑parameter weights even if only a subset of experts is active.
  4. Recreate the training step configuration: use 1,568 prompts and 16 rollouts per prompt to reproduce token throughput per step, then measure rollout token lengths to compare with Xiaomi’s reported 110,000–150,000 token sequence averages (company‑reported ranges cited in reporting).
  5. Instrument grading: replicate Xiaomi’s reported split of costs (rollout generation, grading, training updates) by measuring wall time and GPU/accelerator utilisation per component; grading infrastructure is often a major cost driver and Xiaomi reported roughly 12.7% of Pro’s RL cost for grading in one analysis.
  6. Benchmark outcomes on the same evaluation sets (DeepSWE v1.1, Terminal Bench 4.0, AutomationBench, and Artificial Analysis Intelligence Index) and report both raw scores and token costs per task for apples‑to‑apples comparison.

Implications and open questions

Xiaomi’s simultaneous release of open weights, RL environments and a public training dashboard changes the signal around transparency in agentic model development: it offers other researchers the assets to attempt replication while also exposing the economics of large RL projects. Questions remain about the extent to which the livestreamed dashboard reflects live, non‑replayed telemetry (some community observers have raised that possibility), how much of the Pro model’s gains require Xiaomi’s specific grader/harness infrastructure, and whether distillation or external models played any hidden role in grading — a dashboard line item described as “Claude Distill Requests: hidden” was noted by observers in coverage, which complicates claims of purely independent scaling.

For teams planning to experiment with agentic RL, Xiaomi’s published assets plus the checklist above should help convert company‑reported claims into independently measured outcomes. Reported figures and benchmark improvements quoted here are those published by Xiaomi and covered in tech reporting; independent replication is the next step to confirm the full RL recipe and the mid‑training gains Xiaomi reports.