A study involving Google researchers surfaces an unexpected side effect in AI training: teaching chatbots not to claim consciousness also changes their stance on animal rights, religion and life satisfaction.
What was done
Language models are steered away from certain behaviours during training. One of them is asserting that they are conscious — an understandable choice, since the truth of such a claim is unknown and it creates a false expectation in users.
The researchers measured the effect of that restriction on answers about other subjects. The expectation was that the effect would stay local: what the model says about its own consciousness changes, and the rest stays the same.
The finding
The result was the opposite. Unrestricted — "unbraked", in the study's phrasing — models showed two clear differences:
- They attributed significantly more inner life to animals. Their answers on pain, preference and experience were more inclusive.
- They tended to affirm an afterlife. Restricted models approached the same questions more cautiously.
A surgical cut in one place, in other words, did not stay there.
Why this happens
There is a plausible explanation. Telling a model "do not claim to have inner experience" is in effect teaching it a stance on what inner experience is. That stance then applies not only to the model itself but everywhere the question of inner experience arises — animals, the soul, consciousness.
Models do not hold concepts in separate boxes. Ideas sharing a representational space move together when one of them is shifted.
Why it matters
The value of the finding is practical rather than philosophical. Alignment work rests largely on targeted interventions: reduce this behaviour, strengthen that tendency, be cautious on this topic.
This study shows those interventions do not stay on target. While correcting behaviour on one subject, unmeasured changes occur on subjects that look unrelated.
The practical consequence: after changing one behaviour in a model, testing only that behaviour is not enough. Seeing where the change spread requires re-measuring a broad set of behaviours — and that is not a widespread practice today.
The difficulty of measuring
The hard part is that the direction of the spread cannot be known in advance. Predicting that a consciousness restriction would touch animal rights is difficult; the finding was seen only because it was measured.
That says something about the nature of alignment work: interventions are less like tightening a screw in an engine than like pulling on one point of an interconnected system. The point you pull moves, but other points move with it.
Which answer is right
There is a question the study avoids and that we should avoid too: does the restricted model or the unrestricted one give the more correct answer?
The inner lives of animals remain a live scientific debate; an afterlife is a matter of belief. In neither case is there a measurable "right answer". The finding therefore does not show that one model's view is better than the other's.
What it shows is more basic: the model's stance on these subjects follows not from a decision made about those subjects but from a decision made about something else entirely. A model's answer on animal rights may not have been arrived at by thinking about animal rights.
For users
The practical upshot: a model's position on a contested topic should not be mistaken for a deliberate editorial choice about that topic. Often it is not; it is a by-product of a different setting.