Skip to content
AI.info

The Pulse

Xiaomi Open-Sources MiMo-V2.6 Pro and Flash With Reinforcement-Learning Resources

Xiaomi has released the MiMo-V2.6-Pro and MiMo-V2.6-Flash models alongside their weights, technical report and reinforcement-learning resources. The source was checked for a named speaker and role; it contains no attributable quotation.

Xiaomi Open-Sources MiMo-V2.6 Pro and Flash With Reinforcement-Learning Resources

AI.info Team ·

Xiaomi has officially released and open-sourced the MiMo-V2.6 series, comprising two native fully multimodal models: MiMo-V2.6-Pro and MiMo-V2.6-Flash. The company describes the release as part of its work on recursive self-improvement, using reinforcement learning on complex, verifiable tasks to improve model performance through repeated exploration and feedback.

The announcement, updated September 22, 2026, says the models completed 30 reinforcement-learning steps each in less than six days. Together, the runs used approximately 750,000 trajectories and cost about $850,000 for Flash and $2.62 million for Pro. Xiaomi reports that the average pass rate on training tasks improved by 25% for Flash and 12% for Pro.

Open-Source Training Infrastructure

Xiaomi is publishing more than model weights. The release includes a complete technical report, training environments and reinforcement-learning code intended to help researchers reproduce and examine the company’s results.

The package contains more than 7,000 reinforcement-learning task environments covering software engineering, vulnerability reproduction, knowledge-intensive work, and web design and development. It also includes an end-to-end reinforcement-learning framework built on verl, uni-agent and mini-swe-agent. According to Xiaomi, the framework covers environment interaction, trajectory collection, reward evaluation and policy optimization.

The company has also released lightweight, composable harnesses that separate system prompts, tools and context management. These components are designed to let researchers assemble different training configurations and combine framework elements across tasks.

During training, each update used 1,568 samples, supported a 1-million-token context length and processed between 3.5 billion and 3.7 billion tokens per step. Xiaomi says the training system combined coding, general, visual and cybersecurity tasks through multiple harnesses. It also describes measures intended to stabilize training, including freezing the mixture-of-experts router and using reward design, adversarial evaluation, anomaly detection and validator cross-checking to reduce reward hacking.

Reported Reinforcement-Learning Results

On the DeepSWE v1.1 long-range software-engineering evaluation, Xiaomi reports Flash improving from 48.8 to 65.7 and Pro improving from 58.4 to 72.6. The company presents these figures as evidence of improved sample efficiency, continued learning and performance on tasks outside the training set.

Xiaomi says MiMo-V2.6-Pro reached 46 points on the Artificial Analysis Intelligence Index, surpassing Kimi K3 and Qwen3.8 Max on that index. The company also says Pro performed on most Agent Benchmarks at a level comparable to Claude Opus5 and GPT-5.6 Sol, while Flash outperformed MiMo-V2.5-Pro.

From Code Generation to Interactive Worlds

The MiMo-V2.6 series combines multimodal perception, 3D spatial reasoning and computer-use capabilities. Xiaomi says the models can turn text, images or video into runnable interactive worlds through a workflow involving multiple agents. The process can include scene construction, interactive-logic programming, visual verification and corrections based on rendering results.

The announcement also describes Blender modeling based on text descriptions or reference images. In an embodied-simulation environment, MiMo-V2.6 can use images from multiple camera views to control a Franka Panda robotic arm, grasp objects, match colors and place items precisely through visual feedback.

For computer-use tasks, Xiaomi says the model can understand graphical interfaces and operate common office and productivity tools for information retrieval, editing and data processing. It can also check results, troubleshoot problems and adjust later actions based on visual feedback.

Applications in Research and Digital Creation

Xiaomi presents MiMo-V2.6-Pro assisting with materials research by retrieving literature and patents, proposing metal-organic framework designs aimed at capturing persistent pollutants, and using open-source tools to simulate binding strength before selecting candidates for further testing.

The model has also been used to formalize the main theorem in Li and Yorke’s paper “Period Three Implies Chaos” in Lean 4. Xiaomi says the resulting project contains more than 6,000 lines of Lean source code and was verified by the Lean kernel without unproven placeholders.

Other examples cover front-end interfaces, Figma designs, presentations, SVG files, videos and music. Xiaomi says MiMo-V2.6-Pro can create an orchestral piece, convert the resulting score into MIDI format and apply knowledge of instrumentation and arrangement.

Hosted Access and Open Model Availability

The MiMo-V2.6 series is available through Xiaomi’s open platform and desktop client. API prices remain unchanged from the V2.5 series, and MiMo-V2.6-Pro has an UltraSpeed mode that Xiaomi says can provide up to 20 times the standard inference speed.

The company has also released MiMo-V2.6-Distill-Qwen-9B and reports improvements over its supervised fine-tuning baseline across 11 benchmarks after reinforcement-learning training. The API model names are mimo-v2.6-pro, mimo-v2.6-flash and mimo-v2.6-pro-ultraspeed.

Source

Xiaomi MiMo

Explore

More articles