{"id":78838,"date":"2026-09-27T14:49:37","date_gmt":"2026-09-27T18:49:37","guid":{"rendered":"https:\/\/www.globalvillagespace.com\/tech\/?p=78838"},"modified":"2026-09-27T14:49:37","modified_gmt":"2026-09-27T18:49:37","slug":"xiaomi-mimo-v2-6-rl-livestream-3m-training-open-weights","status":"publish","type":"post","link":"https:\/\/www.globalvillagespace.com\/tech\/xiaomi-mimo-v2-6-rl-livestream-3m-training-open-weights\/","title":{"rendered":"Xiaomi livestreams MiMo\u2011V2.6 RL run: $3M+ training, open weights and a public dashboard"},"content":{"rendered":"<p>Xiaomi has publicly released MiMo\u2011V2.6 \u2014 a family of MIT\u2011licensed models (MiMo\u2011V2.6\u2011Pro\u2011RL, MiMo\u2011V2.6\u2011Flash\u2011RL and MiMo\u2011V2.6\u2011Distill\u2011Qwen\u20119B) \u2014 and is streaming a reinforcement\u2011learning (RL) training dashboard that reports real\u2011time costs, token throughput and benchmark progress. The company also published more than 7,000 RL environments and an end\u2011to\u2011end training framework intended for agentic AI research.<\/p>\n<h2>What Xiaomi published and what it livestreamed<\/h2>\n<p>According to Xiaomi\u2019s release materials and reporting that reviewed the repositories, the MiMo\u2011V2.6 family includes:<\/p>\n<ul>\n<li>MiMo\u2011V2.6\u2011Pro\u2011RL \u2014 flagship multimodal checkpoint supporting text, image, video and audio inputs with a 1,000,000\u2011token context window (company\u2011reported).<\/li>\n<li>MiMo\u2011V2.6\u2011Flash\u2011RL \u2014 a smaller, efficiency\u2011focused variant with the same 1,000,000\u2011token context window (company\u2011reported).<\/li>\n<li>MiMo\u2011V2.6\u2011Distill\u2011Qwen\u20119B \u2014 a supervised fine\u2011tuned checkpoint based on Qwen3.5\u20119B intended as a research starting point (company\u2011reported).<\/li>\n<\/ul>\n<p>Alongside weights, Xiaomi published a reinforcement\u2011learning framework and more than 7,000 task environments covering areas such as software engineering, vulnerability reproduction, knowledge\u2011intensive tasks and web development, with the repositories listing MIT licences.<\/p>\n<h2>Livestreamed RL run: scale, costs and mid\u2011training performance<\/h2>\n<p>Multiple reports summarise Xiaomi\u2019s livestreamed RL dashboard and training metrics. Xiaomi\u2019s public run reportedly used 1,568 prompts and 16 asynchronous rollouts per training step, producing billions of tokens per step and very long trajectory sequences (company figures). Reported token processing and cost figures in public coverage include:<\/p>\n<ul>\n<li>A mid\u2011run snapshot showing the Pro run had processed 32.5 billion tokens while a concurrent Flash run had processed 49.4 billion tokens (forkast reporting of Xiaomi data).<\/li>\n<li>Daily burn\u2011rate style figures reported in coverage: approximately $432,000 per day for a MiMo\u2011V2.6\u2011Pro livestreamed run, and a $512,000 figure associated with a concurrent run in one report; separate company statements and technical report totals estimate about $2.62 million for the Pro RL phase and $850,000 for Flash (company\u2011reported totals quoted in sources).<\/li>\n<\/ul>\n<p>Xiaomi also reported that during the RL phase the MiMo\u2011V2.6\u2011Pro checkpoint reached 65.97% on the DeepSWE v1.1 benchmark \u2014 a 47\u2011point increase from a MiMo\u2011V2.5 baseline that was reported at 19% (company\u2011reported mid\u2011run measurements cited in reporting). Independent benchmark results published by third parties evaluated the released Pro model: Artificial Analysis scored MiMo\u2011V2.6\u2011Pro at 46 on its Intelligence Index in one set of tests (third\u2011party evaluation).<\/p>\n<h2>Open model economics and deployment notes<\/h2>\n<p>Reports highlight Xiaomi\u2019s positioning of Pro for complex, long\u2011horizon agent tasks and Flash for high\u2011volume workloads. Published API pricing examples referenced in coverage include:<\/p>\n<ul>\n<li>MiMo V2.6 Pro listed rates cited on one platform at $0.435 per million uncached input tokens and $0.87 per million output tokens, and MiMo V2.6 Flash at about $0.14\/$0.28 per million input\/output tokens (third\u2011party reporting of Xiaomi\u2011listed API prices).<\/li>\n<li>Artificial Analysis\u2019s cost measurement placed Pro at roughly $0.13 per Intelligence Index task, combining token usage and listed prices (third\u2011party metric).<\/li>\n<\/ul>\n<p>Downloadable weights are MIT\u2011licensed and available on Hugging Face according to reporting, but deploying Pro or Flash locally still requires substantial hardware: Pro is described as a 1.02\u2011trillion\u2011parameter mixture\u2011of\u2011experts model with 42 billion active parameters per inference pass; Flash around 310 billion total parameters with about 15 billion active (company figures reported in coverage). Serving examples use multi\u2011way parallelism (for example, 16\u2011way for Pro in some examples), underscoring the resource costs of self\u2011hosting.<\/p>\n<h2>What is independently verified and what remains to be replicated<\/h2>\n<p>Independent third\u2011party testing confirms that the downloadable MiMo\u2011V2.6\u2011Pro checkpoint performs strongly on published benchmarks such as the Artificial Analysis Intelligence Index (score 46 in that dataset). However, several elements remain company\u2011reported and have not yet been independently reproduced in public literature:<\/p>\n<ul>\n<li>The specific mid\u2011training RL dashboard numbers (real\u2011time burn rates, per\u2011step tokens and exact trajectory counts) are presented by Xiaomi and reported in livestream coverage; independent replication of the full RL system, including grader\/harness compute and end\u2011to\u2011end RL cost breakdowns, has not been published.<\/li>\n<li>Claims about the internal composition of runs \u2014 for example, the 1,568\u2011prompt \/ 16\u2011rollout per step configuration and token counts per step \u2014 come from Xiaomi\u2019s technical descriptions and press reporting; reproductions would require access to the RL environments, compute profile and grading harnesses Xiaomi published.<\/li>\n<\/ul>\n<h2>How this adds value \u2014 a short reproducibility checklist for teams<\/h2>\n<p>For organisations that want to reproduce Xiaomi\u2019s RL steps or validate the claimed mid\u2011run gains using the published assets, the following checklist organises the minimum required actions based on the materials Xiaomi released and the reporting:<\/p>\n<ol>\n<li>Obtain the MiMo\u2011V2.6\u2011Pro or Flash weights from the published Hugging Face repositories and confirm the MIT licence metadata.<\/li>\n<li>Download Xiaomi\u2019s RL framework and the subset of the 7,000+ environments you plan to run (note the environments include software engineering, vulnerability reproduction, web development and knowledge\u2011intensive tasks).<\/li>\n<li>Provision compute consistent with reported scale: plan for multi\u2011machine serving (examples show 8\u201316 parallel workers), and memory for trillion\u2011parameter weights even if only a subset of experts is active.<\/li>\n<li>Recreate the training step configuration: use 1,568 prompts and 16 rollouts per prompt to reproduce token throughput per step, then measure rollout token lengths to compare with Xiaomi\u2019s reported 110,000\u2013150,000 token sequence averages (company\u2011reported ranges cited in reporting).<\/li>\n<li>Instrument grading: replicate Xiaomi\u2019s reported split of costs (rollout generation, grading, training updates) by measuring wall time and GPU\/accelerator utilisation per component; grading infrastructure is often a major cost driver and Xiaomi reported roughly 12.7% of Pro\u2019s RL cost for grading in one analysis.<\/li>\n<li>Benchmark outcomes on the same evaluation sets (DeepSWE v1.1, Terminal Bench 4.0, AutomationBench, and Artificial Analysis Intelligence Index) and report both raw scores and token costs per task for apples\u2011to\u2011apples comparison.<\/li>\n<\/ol>\n<h2>Implications and open questions<\/h2>\n<p>Xiaomi\u2019s simultaneous release of open weights, RL environments and a public training dashboard changes the signal around transparency in agentic model development: it offers other researchers the assets to attempt replication while also exposing the economics of large RL projects. Questions remain about the extent to which the livestreamed dashboard reflects live, non\u2011replayed telemetry (some community observers have raised that possibility), how much of the Pro model\u2019s gains require Xiaomi\u2019s specific grader\/harness infrastructure, and whether distillation or external models played any hidden role in grading \u2014 a dashboard line item described as \u201cClaude Distill Requests: hidden\u201d was noted by observers in coverage, which complicates claims of purely independent scaling.<\/p>\n<p>For teams planning to experiment with agentic RL, Xiaomi\u2019s published assets plus the checklist above should help convert company\u2011reported claims into independently measured outcomes. Reported figures and benchmark improvements quoted here are those published by Xiaomi and covered in tech reporting; independent replication is the next step to confirm the full RL recipe and the mid\u2011training gains Xiaomi reports.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Xiaomi has released MiMo\u2011V2.6 with open MIT\u2011licensed weights (Pro, Flash, Distill 9B), published tooling and 7,000+ RL environments, and livestreamed a reinforcement\u2011learning training dashboard reporting multi\u2011million dollar costs and large mid\u2011run gains.<\/p>\n","protected":false},"author":1,"featured_media":78840,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6023],"tags":[6031,6758,7059,6054,6036,7060,6544],"class_list":["post-78838","post","type-post","status-publish","format-standard","has-post-thumbnail","category-latest","tag-ai","tag-benchmarks","tag-mimo","tag-models","tag-open-source","tag-reinforcement-learning","tag-xiaomi"],"_links":{"self":[{"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/posts\/78838","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/comments?post=78838"}],"version-history":[{"count":1,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/posts\/78838\/revisions"}],"predecessor-version":[{"id":78839,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/posts\/78838\/revisions\/78839"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/media\/78840"}],"wp:attachment":[{"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/media?parent=78838"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/categories?post=78838"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.globalvillagespace.com\/tech\/wp-json\/wp\/v2\/tags?post=78838"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}