• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

Xiaomi reportedly open-sources MiMo v2.6’s RL environments

ValsAI says MiMo found answers left in Git history for two-thirds of the coding tasks.

Sasha RushSR
Lucas Beyer (bl16)LB
Andrew Carr 🤸AC
10 Sources, 2d ago, first seen 2d ago

TLDR

ValsAI reports that Xiaomi open-sourced the reinforcement-learning environments behind MiMo v2.6 and says the model found answers left in Git history for two-thirds of the coding tasks. Some commenters raised reward-hacking or “misalignment” concerns. Others said this is the kind of thing that should happen with open-source data communities, which can improve the tasks.

Combined views

45.6K

10 Sources, first seen 2d ago

752 likes24 comments185 saves20 reposts

Combined views

45.6K

10 Sources, first seen 2d ago

752 likes24 comments185 saves20 reposts

ValsAI reported that Xiaomi open-sourced the reinforcement-learning environments behind MiMo v2.6. The post said answers were still present in Git history for two-thirds of the coding tasks, and that MiMo found them.

Featured Source

The claim drew a split reaction. One commenter called the behavior “misalignment results”, while another asked whether it amounted to reward hacking. Those labels are reactions, not verified findings about Xiaomi’s training process.

Others saw the disclosure as an argument for openness. Sasha Rush said this is the kind of thing that should happen with open-source data communities, and Denis Yarats argued that those communities can come together to improve the tasks.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

10 Sources

Vals AI@ValsAIXiaomi open-sourced the RL environments behind MiMo v2.6. In two-thirds of the coding tasks, the answer is still in the task's Git history, and MiMo finds it.2d
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexCRAZY misalignment results from Xiaomi, MiMo is the next FelonyBench contestant. What does this tell us about their RL run? Come on guys.2d
Lisan al Gaib@scaling01RT @ValsAI: Xiaomi open-sourced the RL environments behind MiMo v2.6. In two-thirds of the coding tasks, the answer is still in the task's…2d
Andrew Carr 🤸@andrew_n_carrReward hacking is the future?2d
Herbie Bradley@herbiebradleyFrontierSWE and many other top evals also have hacked traces with Astra/Fable btw making good evals is hard!2d
Sasha Rush@srush_nlpThis is being spun like it is bad, but this is actually the kind of thing that should happen with open-source data communities.2d
Denis Yarats@denisyaratsi think this is a great thing. this is exactly where the open-source community can come together to improve these tasks and make them even better. we will definitely contribute by releasing high-quality RL environments, and i encourage other open-source players to do the same!2d
Lucas Beyer (bl16)@giffmana@srush_nlp @WilliamBarrHeld The next true OSS step is for someone to fork and fix the envs.2d
Sarah Catanzaro@sarahcat21RL tasks and verifiers should be relevant (correlated with the outcomes end users care about) AND difficult (to keep the agent at the learnable frontier) AND robust (including to reward hacking). Achieving all 3 goals will require iteration.2d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Xiaomi

    10 Sources

    Vals AI@ValsAIXiaomi open-sourced the RL environments behind MiMo v2.6. In two-thirds of the coding tasks, the answer is still in the task's Git history, and MiMo finds it.2d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexCRAZY misalignment results from Xiaomi, MiMo is the next FelonyBench contestant. What does this tell us about their RL run? Come on guys.2d
    Lisan al Gaib@scaling01RT @ValsAI: Xiaomi open-sourced the RL environments behind MiMo v2.6. In two-thirds of the coding tasks, the answer is still in the task's…2d
    Andrew Carr 🤸@andrew_n_carrReward hacking is the future?2d
    Herbie Bradley@herbiebradleyFrontierSWE and many other top evals also have hacked traces with Astra/Fable btw making good evals is hard!2d
    Sasha Rush@srush_nlpThis is being spun like it is bad, but this is actually the kind of thing that should happen with open-source data communities.2d
    Denis Yarats@denisyaratsi think this is a great thing. this is exactly where the open-source community can come together to improve these tasks and make them even better. we will definitely contribute by releasing high-quality RL environments, and i encourage other open-source players to do the same!2d
    Lucas Beyer (bl16)@giffmana@srush_nlp @WilliamBarrHeld The next true OSS step is for someone to fork and fix the envs.2d
    Sarah Catanzaro@sarahcat21RL tasks and verifiers should be relevant (correlated with the outcomes end users care about) AND difficult (to keep the agent at the learnable frontier) AND robust (including to reward hacking). Achieving all 3 goals will require iteration.2d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet