Several critics are disputing how much a reported “pain axis” inside language models can reveal about a model’s own experience. Dillon Plunkett argues that the work does not offer strong evidence that current AI systems experience pain.
As Plunkett describes it, the paper identifies a direction in a model’s activation space that differs when the model predicts text about painful experiences rather than other negative experiences or sensations. He argues that this may represent a speaker being in pain, much as other activation directions can represent a speaker being afraid or angry. That would not establish that the model itself feels anything.
Plunkett also questions the paper’s reported steering tests. He says 80% of the prompts used to construct the direction involve emotional pain, so steering a model toward that pattern could predictably produce language about worthlessness. In another test, the steered assistant reportedly chose a harmful action, but Plunkett says it did so even when told the action would increase its pain. He adds that his own quick experiments found similar harmful behavior for roughly 25% to 30% of random steering directions.
Robert Long said he shares Plunkett’s reservations and welcomed public disagreement because of the stakes for AI welfare. Long separately suggested that examples involving “moral injury” could bake destructive language into the direction. Derek Shiller similarly doubts that the axis is experienced as pain, while Thomas Dietterich argues that a system may model pain without experiencing it. None of these critiques settles whether an AI can feel pain; they challenge what this particular line of evidence can establish.