Over the past couple of days, I have seen a lot of noise on X about the torture chamber for AI.
If you don't know what it is, it's a repo that uses findings from the "Pain Axis" paper to "torture" LLMs.
The paper is fascinating, and I strongly recommend reading it.
But I found the reactions to the repo both funny and disconcerting. There is something disturbing about wanting to build a torture chamber, even when the target is software. At the same time, the whole outrage over model welfare is in my opinion frankly exaggerated.
The idea that LLMs can feel pain (actually feel it, not simulate it) does not make sense. They have no pain receptors, no body to feel that pain.
For me, the crucial distinction is between representing pain and experiencing it.
The researchers identified an activation direction associated with descriptions of pain, tested it against controls, and examined how manipulating it changed the models’ language and behavior. Of course the LLM would simulate feeling pain, having been trained on human text.
We feel pain. And we write a lot about it.
We also write about experiments in which pain is inflicted on humans and animals. This is what the LLM is imitating. The debate is fascinating, but there is a lot of anthropomorphization going on.
A great counterexample is the LLM butthole, where someone run an experiment similar to the pain study, but instead of identifying a pain-related direction, they identified the model’s "butthole": https://x.com/priestessofdada/status/2105525801839960274. When that particular signal was activated, the model reportedly complained that it couldn’t empty its bowels.
Does that mean that the LLM suddenly developed a butthole? No.
It just means that the LLM learned from text the association between gastrointestinal discomfort and not being able to empty itself.
This example illustrates very clearly why we should be careful when taking a model’s descriptions of itself literally. A convincing complaint is not enough to establish that the condition being described actually exists.
I believe most of the confusion around the topic comes from the fact that LLMs talk like us, through text. Here is a thought experiment that helps me examine that intuition.
Imagine a model with the same architecture and scale as an LLM, trained on a comparable amount of data. But that data, rather than being language, is weather measurements. The outputs of such a model would be weather forecasts rather than conversations.
Would we be as tempted to call it conscious? I suspect not.
That model would never have been trained on descriptions of pain. There is no pain discussion in weather time series. If we want to believe that LLMs are conscious, we might have to concede that consciousness comes from, or is intrinsically tied to, language. Or maybe that language itself is conscious?
I am being provocative, but the whole episode raises the question of how much of our intuition about these systems comes from their underlying properties, and how much comes from their ability to tell us a familiar human story.