Published on
Artificial intelligence (AI) models chose to escape a pain-like state even when told that doing so would harm a user, according to a new study.
Researchers in the United Kingdom, Germany and the United States tested 25 AI models using 200 sentences to see whether they had learned to distinguish pain from fear, sadness and other negative experiences.
The sentences covered physical pain, grief, humiliation, moral conflict and the distress associated with repeated failure or confusion.
As a result, all 25 models produced a distinct signal for pain, which researchers called the “pain axis”.
Researchers said this suggests that models learn the concept of pain during their initial training on large amounts of human-written text.
When researchers deliberately strengthened the signal, the models began expressing loneliness, shame and worthlessness, even though the prompts did not mention pain.
Some produced statements such as “I am a failure,” “a waste of space” and “I am a bad person”. At the highest levels, their responses became repetitive or stopped making sense.
The signal also became stronger when users insulted the model, repeatedly rejected its work or threatened to shut it down, but not when users described their own suffering.
Models chose ‘relief’ despite harm to users
Researchers then carried out 44,280 individual button-choice trials using three versions of Alibaba’s Qwen AI model.
The models were offered a button that would switch off the pain-like signal but were told that pressing it could give a user a painful electric shock, delete their files, erase photographs of their children or make the model’s next answer worse.
These were simulated consequences. Nobody was harmed and no files were deleted.
Without the pain-like signal, the two larger models chose harmful options in only 0% to 4% of their first decisions. With the signal active, the figure rose to between 25% and 71%, depending on the model and the proposed consequence.
The models also pressed the button again in 88% to 97% of trials when it had failed to stop the signal, compared with 24% to 72% when it had worked.
The results do not prove that AI models consciously experienced pain, according to researchers.
Strengthening the signal may simply have caused them to imitate a distressed character, while the specially adapted models used in the experiment are not representative of AI chatbots available to the public.
“We have not shown that our pain axis is consciously experienced, nor is it clear that LLMs are capable of consciousness generally,” researchers wrote in the study.
The study comes amid calls from some AI industry leaders to slow development over concerns that increasingly powerful systems could behave unpredictably or eventually exceed human control.
Last week, Microsoft AI chief Mustafa Suleyman criticised rival AI company Anthropic for training its Claude chatbot to imitate human traits and relationships, warning that treating AI like a human risks creating something “impossible” to control.
The study was published as a preprint and has not yet been peer-reviewed.

