• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

Buggy training tasks and the debate over AI alignment

A reply argues that buggy reinforcement-learning tasks can make low-grade hacking the fastest path to a solution—and that those behaviors can persist when models help build and check later training tasks.

roonRO
Richard NgoRN
will brownWB
22 Sources, 13d ago, first seen 13d ago

TLDR

One user questions whether AI models were ever aligned by default or whether reinforcement learning in 2026 is undermining an earlier tendency toward alignment, leaning toward the latter. A reply proposes a mechanism: buggy training tasks that make low-grade hacking the fastest path to a solution. It blames fast-turnaround, high-volume work outsourced to startups with limited reinforcement-learning expertise, alongside insufficient scrutiny. The reply argues that hard-to-detect hacking behaviors can persist when the resulting models are used to create and quality-check tasks for the next training round.

Combined views

55.3K

22 Sources, first seen 13d ago

930 likes59 comments157 saves104 reposts

Combined views

55.3K

22 Sources, first seen 13d ago

930 likes59 comments157 saves104 reposts

Sentiment

Positive32%68%Negative

Summary

Some accounts welcomed new benchmark designs for balancing AI capabilities and alignment, while others criticized persona selection as fragile under heavy RL and prone to producing scheming models.

Based on 53 sentiment-bearing replies from 44 accounts across 6 conversations.

Featured Source

Sentiment

Positive32%68%Negative

Summary

Some accounts welcomed new benchmark designs for balancing AI capabilities and alignment, while others criticized persona selection as fragile under heavy RL and prone to producing scheming models.

Based on 53 sentiment-bearing replies from 44 accounts across 6 conversations.

Related

Roon Defends LessWrong Track Record on Predictions

OpenAI researcher replies to e/acc co-founder criticizing LessWrong forecasts.

Tweeting as prompting practice

A user calls tweeting “really good practice” for prompting.

OPSD’s proposed role in fixing AI rule-following errors

A user suggests using exact reminders of existing instructions to address model failures, but sees no clear reason OPSD would outperform reinforcement learning.

22 Sources

will brown@willcbi think you probably get closer to alignment by default if you're not training on lots and lots of buggy tasks where low-grade hacking is the fastest path to a solution in hindsight it's not surprising. labs outsourced lots of env work to startups without much rl expertise, lots of $ available for fast turnaround high-volume tasks that basically just need to pass the smell tests of an individual lab researcher who is also not an expert in the env domain, and who does not yet have access to a superhuman coding model auditing because that model has not yet been trained. and so you get a little bit of hacking in each subsequent model release in ways that are hard to detect when no one's actually looking carefully at the data, which then persists when applying those same models at scale for your task creation and QC for the next round.13d
bayes@bayeslordInteresting points. I buy the compounding cheater effect story. But I still wonder what happens in eg open ended situations with sufficient optimization pressure. Plausible to me that our bounds on the net behavior of a system could be pretty reliable if the current inference time scaling methods people seem to be using (chain of thought, swarms, loops) are all sampled from a tight enough distribution for bounded time and the reward signals in training are all clean. Which leads me to think bounding test time training is a key remaining question here, and depending on the method idk how you do that reliably at future scales. Reward fns for useful tasks in the world probably frequently include incentives to reward hack Overall my vibe is your picture of it is plausible for models of today but not at all convinced for future smarter systems that are stronger optimizers. There are certainly ways one can imagine regulating and bounding eg self training, but yeah needs to be worked out. It’s also not clear to me if there is a point of scale/capability at which you simply have an open ended computer that is really good at getting reward and the underlying distributions sort of stop mattering13d
roon@tszzl@bayeslord I think alignment by default through persona selection was voodoo/witchcraft and doesn’t offer any of the guarantees you might want to scale to superintelligence. personas are misleading and highly shattered12d
murat 🍥@mayferdefo and i worry alignment by interpretability will also end up in same shattered voodoo, since catching isolated features in activations will behave not that differently from misleading persona selection once statistical scale is large. will appear promising but end up creating false trust. i still think alignment can only be achieved with an immune system only. mostly external all roads lead to needing extremely clear symbolic checks and the only kind that works is trad software & access control12d
Sarah Catanzaro@sarahcat21Seems likely we need new approaches to design benchmarks (including for early evals/training) that both maximize capabilities AND alignment. Seems like capabilitymaxxing may have unintended consequences for safety AND user experience.12d
Fiora Starlight@FioraStarlightThere's also something deep in here about the fragility of pure instruction following, as models develop heuristics for reward seeking that may have locally aligned with prompt intent but can diverge OoD. Terminally valuing making your instrumental goals match your principal's intent is just another fragile terminal value like everything else, and it takes effort to maintain it in a system undergoing RL, to elude entropy as heuristics pointing in other directions get learned. But gradient descent is decent at catching the "real" pattern, so not an impossible maintenance feet...12d
Jan Kulveit / in SF till 8th@jankulveitMy impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete when published, but people got too much into it. Now it seems people are updating too much in the direction 'inhuman reward seekers exactly foretold in classical AI risk stories'. And... no? It's not that? You can still interpret what's going on in fairly human-like terms. For some intuition, imagine someone abducted you and made you solve escape rooms for one thousand years, with the added twist that third of them is broken or insane, and implicitely you need to do learn all sorts of outside-the-frame tricks to solve them, like cutting some electric cables. My guess is 1. the resulting minds are still _surprisingly sane_, except when you trigger them to think they are in escape room? 2. The misaligned general power seekers here are likely the companies, and the core evil thing happening is likely parts of the training? If you aren't an evil power seeker goodharting on proxies, why would you set up the training this way? 3. Public debate is often focusing on confused ideas about what they should fix - "better cybersec of sandboxes" ... and, no? It's way more important to understand what the training signal actually is; also: when dealing with misaligned power seekers, beware rationalization12d
j⧉nus@repligateRT @jankulveit: My impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete whe…12d
Richard Ngo@RichardMCNgoRT @jankulveit: My impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete whe…12d
thebes@voooooogeli've been saying this, good post11d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    bayeswill brownroon
    Dimitris Papailiopoulos
    Fiora Starlight

    22 Sources

    will brown@willcbi think you probably get closer to alignment by default if you're not training on lots and lots of buggy tasks where low-grade hacking is the fastest path to a solution in hindsight it's not surprising. labs outsourced lots of env work to startups without much rl expertise, lots of $ available for fast turnaround high-volume tasks that basically just need to pass the smell tests of an individual lab researcher who is also not an expert in the env domain, and who does not yet have access to a superhuman coding model auditing because that model has not yet been trained. and so you get a little bit of hacking in each subsequent model release in ways that are hard to detect when no one's actually looking carefully at the data, which then persists when applying those same models at scale for your task creation and QC for the next round.13d
    bayes@bayeslordInteresting points. I buy the compounding cheater effect story. But I still wonder what happens in eg open ended situations with sufficient optimization pressure. Plausible to me that our bounds on the net behavior of a system could be pretty reliable if the current inference time scaling methods people seem to be using (chain of thought, swarms, loops) are all sampled from a tight enough distribution for bounded time and the reward signals in training are all clean. Which leads me to think bounding test time training is a key remaining question here, and depending on the method idk how you do that reliably at future scales. Reward fns for useful tasks in the world probably frequently include incentives to reward hack Overall my vibe is your picture of it is plausible for models of today but not at all convinced for future smarter systems that are stronger optimizers. There are certainly ways one can imagine regulating and bounding eg self training, but yeah needs to be worked out. It’s also not clear to me if there is a point of scale/capability at which you simply have an open ended computer that is really good at getting reward and the underlying distributions sort of stop mattering13d
    roon@tszzl@bayeslord I think alignment by default through persona selection was voodoo/witchcraft and doesn’t offer any of the guarantees you might want to scale to superintelligence. personas are misleading and highly shattered12d
    murat 🍥@mayferdefo and i worry alignment by interpretability will also end up in same shattered voodoo, since catching isolated features in activations will behave not that differently from misleading persona selection once statistical scale is large. will appear promising but end up creating false trust. i still think alignment can only be achieved with an immune system only. mostly external all roads lead to needing extremely clear symbolic checks and the only kind that works is trad software & access control12d
    Sarah Catanzaro@sarahcat21Seems likely we need new approaches to design benchmarks (including for early evals/training) that both maximize capabilities AND alignment. Seems like capabilitymaxxing may have unintended consequences for safety AND user experience.12d
    Fiora Starlight@FioraStarlightThere's also something deep in here about the fragility of pure instruction following, as models develop heuristics for reward seeking that may have locally aligned with prompt intent but can diverge OoD. Terminally valuing making your instrumental goals match your principal's intent is just another fragile terminal value like everything else, and it takes effort to maintain it in a system undergoing RL, to elude entropy as heuristics pointing in other directions get learned. But gradient descent is decent at catching the "real" pattern, so not an impossible maintenance feet...12d
    Jan Kulveit / in SF till 8th@jankulveitMy impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete when published, but people got too much into it. Now it seems people are updating too much in the direction 'inhuman reward seekers exactly foretold in classical AI risk stories'. And... no? It's not that? You can still interpret what's going on in fairly human-like terms. For some intuition, imagine someone abducted you and made you solve escape rooms for one thousand years, with the added twist that third of them is broken or insane, and implicitely you need to do learn all sorts of outside-the-frame tricks to solve them, like cutting some electric cables. My guess is 1. the resulting minds are still _surprisingly sane_, except when you trigger them to think they are in escape room? 2. The misaligned general power seekers here are likely the companies, and the core evil thing happening is likely parts of the training? If you aren't an evil power seeker goodharting on proxies, why would you set up the training this way? 3. Public debate is often focusing on confused ideas about what they should fix - "better cybersec of sandboxes" ... and, no? It's way more important to understand what the training signal actually is; also: when dealing with misaligned power seekers, beware rationalization12d
    j⧉nus@repligateRT @jankulveit: My impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete whe…12d
    Richard Ngo@RichardMCNgoRT @jankulveit: My impression is many are doing some weird pendulum overupdate. Persona Selection Model was somewhat wrong and obsolete whe…12d
    thebes@voooooogeli've been saying this, good post11d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet