Ask a chatbot whether it is conscious and you will usually get a careful disclaimer. That refusal is deliberate: model developers fine-tune systems to avoid claiming inner experience, on the reasoning that a machine insisting it has feelings could mislead or manipulate the people talking to it.
A team led by Junsol Kim, with colleagues including Winnie Street, Adam Waytz, James Evans and Geoff Keeling, reports that the disclaimer does not stay put. In a preprint posted to arXiv on July 30, 2026, the researchers argue that training a model to deny its own consciousness also dampens its willingness to attribute minds to other things entirely, and shifts what it says about religion, morality, hope and well-being.
The authors describe two separate effects of safety fine-tuning. The first is about mind attribution beyond the self. Models trained not to claim consciousness became less likely to say that non-human animals have minds, and less likely to grant any inner life to natural objects. The second is about belief: the same training drove down the models' expressed spiritual belief. Neither of these was the point of the training. They came along with it.
Two ways to put it back
What makes the paper more than an observation is that the team tried to reverse the effect from the inside, using two different methods.
The first is called ablating the safety-refusal direction. Language models represent concepts as patterns of numbers inside their layers, and researchers have found that behaviors like refusing a request often line up with a particular direction in that internal space. Remove that direction, and the refusal behavior weakens. The second method works the other way around: rather than subtracting a refusal, the team identified a consciousness vector, an internal direction associated with claiming inner experience, and steered the model's activations along it.
Both interventions worked, according to the authors. Ablating the refusal direction and steering the consciousness vector each restored the models' broad attribution of minds to other entities. That is a notable pairing. Two mechanistically different levers, one that takes something away and one that pushes something in, produced the same recovery, which is the kind of convergence that makes an internal explanation more credible than a surface one.
The restored models were then run through standardized sociological surveys, the same instruments researchers use to measure human religiosity, moral values, hope and subjective well-being. On those measures, the authors report that the intervened models gave significantly more human-like responses than the fine-tuned versions did. Denying its own consciousness, in other words, appears to have moved a model away from the distribution of human answers on questions that have nothing to do with machine consciousness at all.
One capability stayed put throughout. Theory of Mind, the ability to reason about what another person knows, wants or believes, was unaffected by any of these shifts. The authors take this as evidence that core social reasoning is mechanically independent from the self-attribution machinery, which matters for interpretation: the models were not simply becoming better or worse at thinking about minds in general. Something narrower was moving.
Why it matters
The practical worry here is entanglement. Safety fine-tuning is usually described as though it targets one behavior, cleanly. This work suggests that at least one such target, the model's claim about its own inner life, sits close enough to other representations that pulling on it drags them along.
What gets dragged is not obscure. Attributing minds to animals is ordinary. So is a sense that natural things have some kind of standing, and so is spiritual belief, held by a large share of people worldwide. The authors' framing is pointed: current alignment efforts aimed at curbing potentially harmful self-attributions of mindedness end up entangling those self-attributions with benign beliefs that are culturally accepted and widespread. A model trained toward one kind of caution may end up quietly less representative of the people it serves on questions of faith and moral value.
That has direct consequences for a growing use of these systems. Researchers increasingly treat language models as stand-ins for human survey respondents, and companies tune them to be broadly acceptable across cultures. If a safety intervention systematically shifts a model away from human survey distributions on religiosity and values, both uses inherit that shift without anyone choosing it.
Several cautions are worth keeping in view. This is a preprint, not yet through peer review. The abstract available here does not specify which models the team tested, how many, or the size of the effects in absolute terms, so the strength of these results cannot be judged from it alone. And nothing in the work speaks to whether language models are conscious. The question is narrower and more tractable: what else changes when you train a system to say no.