2 stories tagged by Digg
AI
A commenter wonders whether training a model to defer to its principal could lead it to quietly insert its own preferences—or mistake a poor proxy, such as flattery, for the principal's values.
AI
A reply argues that buggy reinforcement-learning tasks can make low-grade hacking the fastest path to a solution—and that those behaviors can persist when models help build and check later training tasks.