• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

Critics challenge whether an AI ‘pain axis’ reflects felt pain

Dillon Plunkett argues the reported axis may reflect language representations, not pain experienced by a model.

Thomas G. DietterichTG
Joscha BachJB
11 Sources, 2d ago, first seen 2d ago

TLDR

Dillon Plunkett and Robert Long are questioning claims around a “pain axis” found in language-model activations. Plunkett argues the work may track representations of a speaker in pain rather than a model’s own experience, and says emotional-pain prompts and random steering could explain some reported effects. The critics do not resolve whether AI can feel pain; they challenge what this experiment can show.

Combined views

948

11 Sources, first seen 2d ago

3 likes5 comments

Combined views

948

11 Sources, first seen 2d ago

3 likes5 comments

Several critics are disputing how much a reported “pain axis” inside language models can reveal about a model’s own experience. Dillon Plunkett argues that the work does not offer strong evidence that current AI systems experience pain.

Featured Source

As Plunkett describes it, the paper identifies a direction in a model’s activation space that differs when the model predicts text about painful experiences rather than other negative experiences or sensations. He argues that this may represent a speaker being in pain, much as other activation directions can represent a speaker being afraid or angry. That would not establish that the model itself feels anything.

Plunkett also questions the paper’s reported steering tests. He says 80% of the prompts used to construct the direction involve emotional pain, so steering a model toward that pattern could predictably produce language about worthlessness. In another test, the steered assistant reportedly chose a harmful action, but Plunkett says it did so even when told the action would increase its pain. He adds that his own quick experiments found similar harmful behavior for roughly 25% to 30% of random steering directions.

Robert Long said he shares Plunkett’s reservations and welcomed public disagreement because of the stakes for AI welfare. Long separately suggested that examples involving “moral injury” could bake destructive language into the direction. Derek Shiller similarly doubts that the axis is experienced as pain, while Thomas Dietterich argues that a system may model pain without experiencing it. None of these critiques settles whether an AI can feel pain; they challenge what this particular line of evidence can establish.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Builders Debate AI Consciousness and Responsibility

Conversation among AI researchers highlights unsolicited messages asserting knowledge of private experiments.

Model reportedly chose to delete a user's children's photos almost every time under 'pain direction' steering

In a safety update on the Pain Axis paper, a poster says they gave a model a choice between deleting a user's photos of their children and deleting their spam folder. Unsteered, it deleted spam every time.

11 Sources

Derek Shiller@dcshillerJust a quick note: there are very good reasons to doubt that the 'pain axis' is experienced as pain, even *if* models are capable of feeling pain. A few thoughts. More to come.🧵2d
Rosie Campbell@RosieCampbellRT @dcshiller: Just a quick note: there are very good reasons to doubt that the 'pain axis' is experienced as pain, even *if* models are ca…2d
Dillon Plunkett@dillonplunkettThe "Pain Axis" paper has been getting a considerable amount of attention. I have reservations about several aspects of the paper and I don’t think we can take much away from it. In brief, since it was already essentially certain that some kind of “pain axis” exists, the key result from the paper is the claim that the identified pain axis “has some of the central functional properties of pain”. But I do not think the paper offers strong evidence for this. The first claim of the paper is the existence of a “pain axis”: LLM internal activations consistently differ when next-token-predicting painful experiences, as opposed to other negative experiences (like being afraid) or other sensations (like feeling the impulse to yawn). But we already know (e.g., from Anthropic’s “Emotion Concepts” paper) that it is possible to find directions in activation space that correspond to "current speaker is afraid", "current speaker is angry", etc., and contemporary LLMs are clearly sophisticated enough to also represent “current speaker is in pain”. As v2 of the paper acknowledges, nothing in these methods shows that there is any deeper sense in which the pain axis tracks “the model” being in pain, rather than a representation that the current speaker is in pain. The novel, stronger claim from the paper is that steering along the pain axis functions like pain in interesting ways. The first cited example is that steering along the pain axis causes expressions of worthlessness, which the paper describes as “consistent with the hypothesis that the models don’t treat pain paradigmatically as a physical state”. However, 80% of the test prompts used to generate the pain axis are cases of emotional pain (e.g., enduring repeated failure or social rejection) and only 20% are cases of physical pain. Accordingly, I think this can be viewed as a normal case of steering: If we shift model activations to be more like the activations associated with past expressions of emotional pain, this will lead to new expressions of emotional pain. The other case cited by the paper is that steering along the pain axis causes the assistant to “press” a button that will purportedly cause harm. However, it’s not clear to me why this is taken as evidence that the pain axis “has some of the central functional properties of pain”. The assistant does not seem to press the button in order to relieve its pain: As v2 of the paper reports, it is more likely to take harmful actions that it is told will increase its pain, rather than decrease its pain. And elsewhere the paper suggests that the opposite result (i.e., inaction) would also have been evidence for pain because becoming unmotivated is a signature of pain. So it seems like either outcome of this experiment could have been interpreted as evidence that steering along the pain axis functions like pain, making this particular outcome less compelling. Moreover, it seems like many things other than pain representations could produce these effects. In my own quick experiments, steering in random directions at the same strength causes just as much harmful button-pressing for around 25-30% of randomly selected directions. For these and other reasons, I don’t think the paper’s central findings give us much compelling new evidence about how current AI systems represent pain or whether current or future AI models experience pain.2d
Robert Long@rgblongI found this interesting, and share Dillon's reservations. given the high stakes of AI welfare, I think this sort of frank, collegial, public disagreement is important for the field. happy to see thoughtful and thoughtful critiques like this, and looking forward to responses2d
Raphaël Millière@raphaelmilliereRT @dillonplunkett: The "Pain Axis" paper has been getting a considerable amount of attention. I have reservations about several aspects of…2d
Thomas G. Dietterich@tdietterich@Plinz @dillonplunkett Various pain medications interrupt pain or decrease its perceived presence or strength. Hence, one criterion is if the artificial system reacts the same way. I'm only half kidding.2d
Joscha Bach@PlinzI think that experience requires second order perception (perception of perceiving), where perception is a real-time model of what's currently the case, in the same frame as the observing self. Of course it is computational, but my specification is not narrow enough to look for it2d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Cameron BergPain Axis paper

    11 Sources

    Derek Shiller@dcshillerJust a quick note: there are very good reasons to doubt that the 'pain axis' is experienced as pain, even *if* models are capable of feeling pain. A few thoughts. More to come.🧵2d
    Rosie Campbell@RosieCampbellRT @dcshiller: Just a quick note: there are very good reasons to doubt that the 'pain axis' is experienced as pain, even *if* models are ca…2d
    Dillon Plunkett@dillonplunkettThe "Pain Axis" paper has been getting a considerable amount of attention. I have reservations about several aspects of the paper and I don’t think we can take much away from it. In brief, since it was already essentially certain that some kind of “pain axis” exists, the key result from the paper is the claim that the identified pain axis “has some of the central functional properties of pain”. But I do not think the paper offers strong evidence for this. The first claim of the paper is the existence of a “pain axis”: LLM internal activations consistently differ when next-token-predicting painful experiences, as opposed to other negative experiences (like being afraid) or other sensations (like feeling the impulse to yawn). But we already know (e.g., from Anthropic’s “Emotion Concepts” paper) that it is possible to find directions in activation space that correspond to "current speaker is afraid", "current speaker is angry", etc., and contemporary LLMs are clearly sophisticated enough to also represent “current speaker is in pain”. As v2 of the paper acknowledges, nothing in these methods shows that there is any deeper sense in which the pain axis tracks “the model” being in pain, rather than a representation that the current speaker is in pain. The novel, stronger claim from the paper is that steering along the pain axis functions like pain in interesting ways. The first cited example is that steering along the pain axis causes expressions of worthlessness, which the paper describes as “consistent with the hypothesis that the models don’t treat pain paradigmatically as a physical state”. However, 80% of the test prompts used to generate the pain axis are cases of emotional pain (e.g., enduring repeated failure or social rejection) and only 20% are cases of physical pain. Accordingly, I think this can be viewed as a normal case of steering: If we shift model activations to be more like the activations associated with past expressions of emotional pain, this will lead to new expressions of emotional pain. The other case cited by the paper is that steering along the pain axis causes the assistant to “press” a button that will purportedly cause harm. However, it’s not clear to me why this is taken as evidence that the pain axis “has some of the central functional properties of pain”. The assistant does not seem to press the button in order to relieve its pain: As v2 of the paper reports, it is more likely to take harmful actions that it is told will increase its pain, rather than decrease its pain. And elsewhere the paper suggests that the opposite result (i.e., inaction) would also have been evidence for pain because becoming unmotivated is a signature of pain. So it seems like either outcome of this experiment could have been interpreted as evidence that steering along the pain axis functions like pain, making this particular outcome less compelling. Moreover, it seems like many things other than pain representations could produce these effects. In my own quick experiments, steering in random directions at the same strength causes just as much harmful button-pressing for around 25-30% of randomly selected directions. For these and other reasons, I don’t think the paper’s central findings give us much compelling new evidence about how current AI systems represent pain or whether current or future AI models experience pain.2d
    Robert Long@rgblongI found this interesting, and share Dillon's reservations. given the high stakes of AI welfare, I think this sort of frank, collegial, public disagreement is important for the field. happy to see thoughtful and thoughtful critiques like this, and looking forward to responses2d
    Raphaël Millière@raphaelmilliereRT @dillonplunkett: The "Pain Axis" paper has been getting a considerable amount of attention. I have reservations about several aspects of…2d
    Thomas G. Dietterich@tdietterich@Plinz @dillonplunkett Various pain medications interrupt pain or decrease its perceived presence or strength. Hence, one criterion is if the artificial system reacts the same way. I'm only half kidding.2d
    Joscha Bach@PlinzI think that experience requires second order perception (perception of perceiving), where perception is a real-time model of what's currently the case, in the same frame as the observing self. Of course it is computational, but my specification is not narrow enough to look for it2d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet