A digital twin of a jet engine earns its keep by being right about the engine. Feed it sensor readings, and it tells you where the metal is fatiguing. Nobody worries that the model's advice will change the engine's mind about how to behave, or that predicting the past well might not translate into choosing well.

Medicine borrowed the idea anyway. Health digital twins, computational stand-ins for individual patients, are being built to compare treatments, project how someone's illness will unfold, and help clinicians decide what to do next. In a preprint posted to arXiv on July 29, Nikki L. B. Freeman and five colleagues argue that copying the engineering blueprint into the clinic invites two distinct failures. They name them the fidelity trap and the feedback trap.

Two ways to be wrong

The fidelity trap is the assumption that a model accurate enough to reproduce what happened is thereby accurate enough to say what would have happened instead. The authors treat these as different jobs. Prediction asks what comes next given how the world has been running. Counterfactual reasoning asks what would follow from an intervention, a treatment given to a patient who, in the observed data, did not receive it. A twin can fit past patient trajectories closely and still, in the authors' words, miss the mark when ranking treatments.

The reason is that observational medical data carries the fingerprints of the decisions that produced it. Sicker patients get the aggressive drug. A model that learns the pattern faithfully learns, among other things, that the aggressive drug is associated with worse outcomes. It reproduces the world it observed with high fidelity. It also, if asked which treatment to give, gives the wrong answer. Accuracy is not the safeguard people assume it is, and the authors' point is that no amount of additional accuracy fixes it, because the problem is not in how well the model fits but in what question it was fit to answer.

The feedback trap is the stranger of the two, and it only exists because health twins are meant to be used. A twin recommends a treatment. Clinicians act on the recommendation. The resulting patient outcomes become new data. The twin is refit on that data, including the part its own advice helped generate. Freeman and colleagues describe what can follow: the model recovers a biased relationship and grows more confident in it even as the data grows thinner.

That last phrase is worth sitting with. Confidence normally rises because evidence accumulates. Here it can rise because the twin has narrowed the range of situations it ever sees. If it stops recommending a treatment, it stops observing what that treatment does, and the remaining data offers less and less to contradict it. The twin is not so much learning as listening to an echo of itself. Standard statistical machinery reads the shrinking variation as certainty.

What they propose instead

Rather than better fidelity, the authors argue for a different architecture: health twins conceived as causally valid, modular, and evolving systems. Modularity means separating out the specific data and models needed to support an interventional claim, so the piece of the twin that recommends a treatment is not tangled up with the piece that merely describes a patient. Causal validity is what licenses the recommendation in the first place, grounding it in reasoning about interventions rather than correlations. Governed evolution means updating the twin while explicitly accounting for the fact that its own recommendations reshaped the data it is now learning from.

The closing line of the abstract is the paper's thesis in miniature. The standard for a health twin, the authors write, should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.

Why it matters

This is a position paper, filed under Other Statistics, not a study with patients or a simulation with results. What is available publicly is the abstract and metadata; the argument's supporting detail lives in the full manuscript. Readers should weigh it as a conceptual case from six statisticians, not as evidence that any particular deployed system has failed.

Still, the timing matters. Digital twins are moving into health care on the strength of an analogy to engineering, and analogies tend to carry their assumptions in quietly. The engine does not respond to what the model says about it. A patient population does, through the clinicians who read the model's output and change what they prescribe. That single asymmetry is what generates the feedback trap, and it has no counterpart on the factory floor.

The practical upshot for anyone evaluating one of these systems is a change in the question. Not "how closely does it match the record?" but "what evidence supports its treatment claims, and what is happening to the data as it operates?" A twin that scores beautifully on the first question can be quietly failing the second. The authors are arguing that the field should decide this now, while the architectures are still being drawn, rather than after a generation of twins has been trained on data it helped write.