Sixteen of the sixty patients the software flagged as mysteriously losing weight were, in fact, trying to lose weight. They had said so to their doctor. The doctor had written it down. The problem was where: in the free-text clinical notes, the running prose of a consultation, which the software could not read.
That gap sits at the centre of a mixed methods pilot published in JMIR Cancer by Javiera Martinez-Gutierrez and colleagues at the University of Melbourne, working with collaborators in Oxford, Singapore and Chile. The team switched on a new module in Future Health Today, a clinical decision support tool already installed in Australian general practices, and pointed it at one of the hardest signals in primary care: unintended weight loss.
Weight loss nobody asked for is a genuine warning sign. It shows up ahead of several cancers, and also ahead of depression, thyroid disease and gut disorders, which is exactly what makes it hard to act on. A general practitioner seeing a patient who has dropped a few kilograms has to decide, on thin evidence, whether to order blood tests, book a scan, or wait. The module was built to help: it scans records for weight-loss terms, applies a cancer risk algorithm developed in the UK, and returns one of three recommendations depending on whether the estimated risk falls below or above two percent and whether blood results are normal.
Five practices took part, four metropolitan and one rural, drawn from forty that already had the platform installed. Each averaged roughly 17,600 active patients and ten doctors. The algorithm surfaced 62 patients; the researchers audited 60. Thirty-six of those, sixty percent, turned out to have genuine unintended weight loss when a clinician read the full record. The rest were mostly people on a diet, post-surgery, or with no real weight change at all.
What the notes hid
The misclassification has a mundane cause. Australian GPs write in prose. Intention to lose weight appeared in the free-text notes for 98 percent of patients, and in a structured field the algorithm could actually parse for only 20 percent. One practice nurse told the interviewers she could not work out why some patients had appeared at all, since several had only a single weight ever recorded.
The algorithm also failed in the other direction. It was supposed to exclude anyone diagnosed with cancer in the previous five years, and it let three through: cases of leukaemia, myeloma and breast cancer that predated the weight-loss visit. The authors call this an unexpected error and trace it to the same coding problem.
The human findings are quieter but arguably more useful. The team interviewed seven staff across the five practices, and a clean split emerged. Practices with a dedicated quality improvement person, usually a nurse or pharmacist with protected hours, found the tool easy to fold into their week. Practices without one did not. The single GP interviewed did the review after hours, because Australian GPs have no allocated time for this kind of work. In three of the five practices, the clinicians doing the reviewing never touched the software at all; someone else printed them a list. Two practice managers had switched off the point-of-care pop-ups entirely, worried about alert fatigue, even though weight loss turns up in only about 1.5 percent of patients and the alerts would have been rare.
Then there is the result that complicates the whole exercise. Only five of the 36 confirmed patients were formally recalled. The other 31 were deferred, most because their doctor judged them already under appropriate care. Six months later, 34 of the 36, or 94 percent, had received follow-up regardless: a median of eight visits, blood tests for 86 percent, imaging or endoscopy for 61 percent. Four had cancer, but only one, a renal cancer, was diagnosed after the weight-loss consultation. Mental health conditions were the most common diagnosis linked to the weight loss.
Why it matters
A tool that flags patients who were going to be followed up anyway is not obviously worth the workflow disruption, and the authors say so directly. They ask whether these high follow-up rates reflect Australian general practice broadly, or whether the five practices that volunteered for a research study are simply better than average. Nobody knows, because the study did not include a comparison group. They also could not count false negatives: patients with real unintended weight loss the algorithm never flagged, since reviewing unflagged records was beyond the budget.
With 36 patients across five clinics, this is a pilot and the authors treat it as one. What it demonstrates cleanly is a constraint that applies well beyond weight loss. Decision support systems are only as good as the data they can read, and in primary care most of the clinically important detail is prose. The authors suggest AI scribes might eventually make that prose machine-readable, while noting that evaluations of current scribes find frequent omissions and occasional serious errors. Until something closes that gap, a tool reading only the structured fields is reading the smaller half of the chart.