To Mix or To Merge?
Toward Multi-Domain Reinforcement Learning for Large Language Models

arXiv Hugging Face ModelScope GitHub COLM 2026

Haoqing Wang†, Xiang Long†, Ziheng Li†, Yilong Xu, Tingguang Li, Yehui Tang✉
Samsung Research, Beijing, China   ·   Peking University

--- ## 📰 News - **[2026.09.07]** 🎉 The model checkpoints are now open-sourced on [Hugging Face](https://hf.co/collections/Jackwang111/m2rl) and [ModelScope](https://modelscope.cn/collections/whq1111/M2RL)! **Feel free to use our checkpoints for your post-training research (e.g., weight merging, multi-teacher on-policy distillation)!** - **[2026.07.09]** 🎉 Our paper is accepted to **COLM 2026**! --- ## 📚 Citation If you find this work useful, please consider citing: ```bibtex @inproceedings{ wang2026to, title={To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models}, author={Haoqing Wang and Xiang Long and Ziheng Li and Yilong Xu and Tingguang Li and Yehui Tang}, booktitle={Third Conference on Language Modeling}, year={2026}, url={https://openreview.net/forum?id=jP7j5XkG8J} } ```