MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
Xiaomi, LLM-Core, :, null, Xia, Bingquan, Shen, Bowen, Cici, null, Zhu, Dawei, Zhang, Di, Wang, Gang, Zhang, Hailin, Liu, Huaqiu, Xiao, Jiebao, Dong, Jinhao, Zhao, Liang, Li, Peidian, Wang, Peng, Yu, Shihua, Chen, Shimao, Wang, Weikun, Ma, Wenhan, Deng, Xiangwei, Huang, Yi, Song, Yifan, Jiang, Zihan, Ye, Bowen, Cai, Can, He, Chenhong, Zhang, Dong, Zhang, Duo, Wang, Guoan, Tian, Hao, Zhao, Haochen, Qu, Heng, Xu, Hongshen, Shi, Jun, Bao, Kainan, Fang, Kai, Zhou, Kang, Zhou, Kangyang, Li, Lei, Zhu, Menghang, Chen, Nuo, Wang, Qiantong, Liu, Shaohui, Li, Shicheng, Gu, Shuhao, Ren, Shuhuai, Liu, Shuo, Deng, Sirui, Zhuang, Weiji, Lv, Weiwei, Yang, Wenyu, Zhang, Xin, Yong, Xing, Zhang, Xing, Song, Xingchen, Xu, Xinzhe, Wang, Xu, Yan, Yihan, Tu, Yu, Tian, Yuanyuan, Wang, Yudong, Yu, Yue, Lin, Zhenru, Song, Zhichao, Yue, Zihao
–arXiv.org Artificial Intelligence
We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing strategy to strengthen the base model's reasoning potential. MiMo-7B-Base is pre-trained on 25 trillion tokens, with additional Multi-Token Prediction objective for enhanced performance and accelerated inference speed. During post-training, we curate a dataset of 130K verifiable mathematics and programming problems for reinforcement learning, integrating a test-difficulty-driven code-reward scheme to alleviate sparse-reward issues and employing strategic data resampling to stabilize training. Extensive evaluations show that MiMo-7B-Base possesses exceptional reasoning potential, outperforming even much larger 32B models. The final RL-tuned model, MiMo-7B-RL, achieves superior performance on mathematics, code and general reasoning tasks, surpassing the performance of OpenAI o1-mini. The model checkpoints are available at https://github.com/xiaomimimo/MiMo.
arXiv.org Artificial Intelligence
Jun-6-2025
- Country:
- Europe (1.00)
- North America
- United States > Minnesota (0.28)
- Canada > British Columbia
- Genre:
- Research Report (0.54)
- Technology: