Xiaomi MiMo-V2.6-Pro Tops Open-Weight Index at $0.13
$0.13 per task 📉 Xiaomi's MiMo-V2.6-Pro tops the Artificial Analysis open-weight ranking at 46—20 points over its predecessor and cheaper than most proprietary systems per unit of work. The gains come from transparent RL training: ~750k trajectories, a new "agentic grader," plus a sparse MoE (42B active of 1.02T params) delivering ~125 tokens/sec. Multi-harness accuracy jumped from ~50% to 66%. Yet it's shadowed by Anthropic's allegation that Xiaomi extracted 400k+ conversations to train it a claim Xiaomi denies. If intelligence keeps dropping in price this fast, the gap isn't closing—it's redefining what it costs. Does $0.13/task intelligence change how your team buys models, or does the distillation cloud hold you back?
On September 22, 2026, a quiet revolution landed in open-weight AI. Xiaomi's MiMo-V2.6-Pro didn't just top the Artificial Analysis Intelligence Index's open-weight ranking with a score of 46—it did so at $0.13 per task. That figure, roughly 20 points ahead of its predecessor (MiMo-V2.5-Pro scored 26) and cheaper than most proprietary systems per unit of work, signals something larger than a single benchmark win.
The gains weren't a gift of exotic architecture. Underneath, MiMo-V2.6-Pro runs a sparse mixture-of-experts design: 1.02 trillion total parameters, but only 42 billion activated per token, supported by 70 transformer layers and 384 routed experts. Flash, the efficiency sibling, trims that to 309 billion total and 15 billion active. Both carry a 1-million-token context window and native multimodal input (text, image, video, audio). The main architectural twist is Grouped Query Attention with a 128-token Sliding Window—a recipe aimed at memory efficiency, not headline-grabbing novelty.
Where the real gains came from
The discontinuity lives in training methodology. Xiaomi published unusually transparent reinforcement-learning (RL) runs: six days, roughly 750,000 trajectories, 30 update steps, each using 1,568 prompts with 16 rollouts, consuming 3.5–3.7 billion tokens per step. The scaling happened along three axes: more data per step, a wider variety of task environments (coding, general agents, vision, cybersecurity), and more compute devoted to grading.
That last piece matters most. Xiaomi replaced a simplicity verifier with an "agentic grader" that examines execution traces—not just final answers—to build task-specific rubrics from contrasting rollouts (Groupwise Reward Synthesis) and rank passing trajectories online (Groupwise Advantage Redistribution). Multi-harness training lifted hold-out accuracy from roughly 50% to 66%. On DeepSWE v1.1, Pro climbed from 58.4 to 72.6; Flash from 48.8 to 65.7. Across complete agentic workloads, Pro pushed further (71.9 on DeepSWE, 53.1 on AutomationBench 16.0, 34.9 on Terminal Bench 4.0).
What $0.13 per task actually buys
Pricing confirms the efficiency thesis. Pro runs at $0.435 per million input and $0.87 per million output tokens; Flash drops to $0.14 input and $0.28 output. A speculative-decoding layer (five-layer decoder behind the MoE, predicting seven future tokens per forward pass) plus native FP8 weights deliver roughly 125 output tokens per second. UltraSpeed variants claim up to 20× faster output at ten times Flash's price.
The competitive picture shifted accordingly. MiMo-V2.6-Pro ranks ahead of Chinese rivals (Kimi K3, Qwen3.8 Max, GLM-5.3, DeepSeek V4.1 Flash) and outperforms proprietary systems like Claude Opus 5 on several agentic suites at a fraction of their cost. That positioning aligns with a broader shift in how open-weight models move across borders. On September 14, OpenRouter launched US in-region routing for business customers, keeping Chinese models—GLM-5.3, DeepSeek V4 Pro, and Kimi K3 among them—within US borders to address data residency concerns. Analysts describe the MiMo-V2.6 result as Pareto optimal: maximum measured intelligence per dollar spent.
The routing infrastructure reflects how far Chinese open-weight leadership has traveled. Chinese labs have now traded the open-weight crown repeatedly, and leadership no longer depends on US token processing. OpenRouter's move is a direct response to that reality, alongside data-sovereignty pressure: Nvidia's $12.9 billion acquisition of Hugging Face this same period and the launch of its open-weight Nemotron model signal the ecosystem's appetite for hosted, governable, region-local weight distribution. Enterprises benefit from reduced cross-border transfer risk and the ability to process sensitive data within jurisdiction. Meanwhile, Xiaomi keeps its own MIT-licensed weights self-hostable, and the full RL toolkit ships openly on Hugging Face itself.
The controversy shadow
The milestone carries a reputational cloud. Anthropic alleges Xiaomi extracted more than 400,000 user conversations from its own chatbot to train MiMo-V2.6, citing reports from March–April 2026. Xiaomi denies the claim. That dispute lands amid a broader pattern: on September 11, 2026, Anthropic published a threat intelligence report detailing coordinated illicit distillation by Chinese labs—Alibaba (>151 million exchanges with Claude between May and July, peaking near 3 million daily), Moonshot (>23 million), and DeepSeek (>12 million attacks over 14 days in July)—using fraudulent accounts routed through Singapore and Japan. Regardless of the outcome for Xiaomi specifically, the allegations invite regulatory scrutiny and may temper enterprise adoption in sensitive markets even as the technical record compels attention.
The trajectory
Xiaomi's own roadmap points to continued data- and recipe-driven scaling rather than new attention variants, and the six-day, ~$2.6 million training run demonstrates how fast a competitive checkpoint can emerge. If the cost per unit of intelligence keeps falling at this rate, the open-weight frontier isn't just closing the gap—it's redefining what that gap costs.
Comments ()