• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

OPSD’s proposed role in fixing AI rule-following errors

A user suggests using exact reminders of existing instructions to address model failures, but sees no clear reason OPSD would outperform reinforcement learning.

Andreas Kirsch 🇺🇦AK
will brownWB
7 Sources, 14d ago, first seen 14d ago

TLDR

A user outlines a possible use for OPSD: identify violations of instructions already in a model’s context, insert a verbatim reminder just before the failure, and train only on the tokens immediately following it. They argue this avoids hint leakage and addresses rule-following problems in long contexts, but question whether it would beat reinforcement learning that penalizes the same failures. They also warn that training actions that conflict with preceding reasoning could raise concerns about the faithfulness of that reasoning. They see tool-schema failures as the clearest use case, while arguing that frequent tool-call mistakes likely point to bigger post-training problems.

Combined views

32.6K

7 Sources, first seen 14d ago

544 likes21 comments212 saves14 reposts

Combined views

32.6K

7 Sources, first seen 14d ago

544 likes21 comments212 saves14 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Tweeting as prompting practice

A user calls tweeting “really good practice” for prompting.

Buggy training tasks and the debate over AI alignment

A reply argues that buggy reinforcement-learning tasks can make low-grade hacking the fastest path to a solution—and that those behaviors can persist when models help build and check later training tasks.

AI Tooling Improves Workflows Five Percent Monthly

Tweet notes monthly workflow gains from clicking AI tool updates.

7 Sources

will brown@willcbif i had to steelman the general usage of OPSD, my recommendation would be: - use judges to identify behaviors in rollouts which expressly violate guidance which is *already in context* (e.g. system prompt rules, tool schemas) - insert a verbatim reminder of that guidance right before the failure occurs - train only on the tokens immediately following the reminder this avoids hint leakage, and combats the legitimate problem of rule adherence in long-context settings. however, you can also just do this with RL, and it's not clear why you should expect OPSD to be any better than folding those same failures into reward penalties. it's also training the model to take actions which might directly conflict with preceding reasoning, which could be an issue for CoT faithfulness if that's something you're concerned about. tool schema failures are the clear win, but if your model is regularly messing up tool calls, there are probably more bigger problems in your post-training to solve. i think it's pretty unlikely that any serious frontier lab uses it in a real way other than as a last-minute band-aid on weird formatting bugs which show up in final testing states. it's not bitter lesson pilled. some people like to think that *specialization* isn't bitter lesson pilled, and i very much disagree here. the most intelligent and successful humans are usually exceptionally specialized and spiky. somehow we convinced ourselves that being superhuman at *everything* ought to be a free lunch, but there's not really any historical precedent or theoretical argument for this being true in the limit. the success of frontier foundation models isn't really evidence here, they're *way* more expensive to train than specialized counterparts, but final-run training costs are dwarfed by research and inference, and it's much easier to ship a one-size-fits-all product. jev is a hit because nobody ever made BERT finetuning easy enough for non-expert developers. real continual learning has never been tried. we don't even have the cognitive core nailed yet.14d
Andreas Kirsch 🇺🇦@BlackHC@willcb Real OPSD hasn't been tried before. (But also you need sufficiently capable models to start from)14d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    will brownAndreas Kirsch 🇺🇦

    7 Sources

    will brown@willcbif i had to steelman the general usage of OPSD, my recommendation would be: - use judges to identify behaviors in rollouts which expressly violate guidance which is *already in context* (e.g. system prompt rules, tool schemas) - insert a verbatim reminder of that guidance right before the failure occurs - train only on the tokens immediately following the reminder this avoids hint leakage, and combats the legitimate problem of rule adherence in long-context settings. however, you can also just do this with RL, and it's not clear why you should expect OPSD to be any better than folding those same failures into reward penalties. it's also training the model to take actions which might directly conflict with preceding reasoning, which could be an issue for CoT faithfulness if that's something you're concerned about. tool schema failures are the clear win, but if your model is regularly messing up tool calls, there are probably more bigger problems in your post-training to solve. i think it's pretty unlikely that any serious frontier lab uses it in a real way other than as a last-minute band-aid on weird formatting bugs which show up in final testing states. it's not bitter lesson pilled. some people like to think that *specialization* isn't bitter lesson pilled, and i very much disagree here. the most intelligent and successful humans are usually exceptionally specialized and spiky. somehow we convinced ourselves that being superhuman at *everything* ought to be a free lunch, but there's not really any historical precedent or theoretical argument for this being true in the limit. the success of frontier foundation models isn't really evidence here, they're *way* more expensive to train than specialized counterparts, but final-run training costs are dwarfed by research and inference, and it's much easier to ship a one-size-fits-all product. jev is a hit because nobody ever made BERT finetuning easy enough for non-expert developers. real continual learning has never been tried. we don't even have the cognitive core nailed yet.14d
    Andreas Kirsch 🇺🇦@BlackHC@willcb Real OPSD hasn't been tried before. (But also you need sufficiently capable models to start from)14d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet