Xiaomi reportedly open-sources MiMo v2.6’s RL environments
ValsAI says MiMo found answers left in Git history for two-thirds of the coding tasks.
TLDR
ValsAI reports that Xiaomi open-sourced the reinforcement-learning environments behind MiMo v2.6 and says the model found answers left in Git history for two-thirds of the coding tasks. Some commenters raised reward-hacking or “misalignment” concerns. Others said this is the kind of thing that should happen with open-source data communities, which can improve the tasks.
Combined views
45.6K
10 Sources, first seen ago
