Twenty-seven thousand times over thirteen years, an older adult in a large research study filled out a ten-question checklist and came out the other side labelled as having probable depression. In 13,356 of those cases, close to half, the same person had just said they were happy and were not feeling depressed.
That contradiction is the whole finding of a new preprint from Malcolm Forbes, John McNeil and Michael Berk, posted to medRxiv on 31 July 2026. It has not yet been through peer review. The authors did not run an experiment or test a treatment. They went back through data that already existed and asked a question that sounds almost too simple to be worth asking: when a screening questionnaire says someone is probably depressed, do they actually report feeling depressed?
The tool under scrutiny
The questionnaire is the CES-D-10, a ten-item short form of the Centre for Epidemiological Studies Depression Scale. It is one of the workhorses of population research on ageing. A participant rates how often, in the past week, various things were true of them, and the answers are added into a single sum score. Researchers then draw a line: a score of 8 or above, or in some studies 10 or above, gets counted as clinically significant symptoms or probable depression. That single number goes on to appear in headlines, in policy documents and in studies estimating how common depression is among older people.
Forbes and colleagues drew their data from ASPREE-XT, the long-running follow-up to the ASPirin in Reducing Events in the Elderly trial, which enrolled 19,114 older participants in Australia and the United States. Because participants were assessed repeatedly over more than a decade, the researchers were not counting people but observations: individual moments when someone filled out the form. Across thirteen years, 27,417 of those observations crossed the threshold of 8 or more.
Then came the check. Two of the ten items ask directly about mood. Item 3 is "I felt depressed." Item 8 is "I was happy." If the sum score is doing its job as a proxy for a major depressive episode, you would expect most people above the line to be endorsing sadness and denying happiness, since low mood and loss of pleasure are the two symptoms clinicians treat as the core of the diagnosis. Instead, 13,356 threshold-positive observations, 48.7 percent of them, showed minimally depressed mood together with preserved happiness. Roughly one in two.
Where the score comes from instead
The paper does not claim these people were feeling fine. A score of 8 has to be built from something, and the other eight items on the CES-D-10 cover territory like restless sleep, trouble concentrating, feeling that everything is an effort, and loneliness. Those are real experiences worth knowing about. They are also, in later life, the sorts of things that arthritis, poor sleep, bereavement, a recent hospital stay or simply the physical work of getting through a day can produce in someone whose mood is intact.
That is the mismeasure in the title. The scale adds up symptoms without weighting the ones that define the illness, so a person can accumulate enough points to be counted as probably depressed while consistently reporting that they feel neither sad nor unhappy. The authors' conclusion is measured: CES-D-10 sum scores, they write, should be used with caution as a proxy for major depressive disorder in older adults.
Some limits are worth keeping in view. This is a descriptive analysis, not a validation study against clinical interviews, so nobody in this work sat down with a psychiatrist to establish who did and did not have depression. The ASPREE cohort was healthy and living in the community at enrolment, which may not resemble a clinic population. And self-report has its own problems: older adults in particular may under-report low mood, meaning some of those 13,356 observations could belong to people who were genuinely struggling and did not say so on two specific lines of a form.
Why it matters
Estimates of how common depression is in later life rest heavily on screening scales like this one. If roughly half the observations above the threshold lack the affective core of the diagnosis, then prevalence figures built on those thresholds are measuring a broader and blurrier thing than the word depression suggests to most readers.
The consequences run further than statistics. Studies that use threshold-crossing as an outcome, asking whether a drug, a diet or a social programme reduces depression in older adults, are working with a signal that is partly mood and partly the ordinary friction of ageing bodies. A treatment that improves sleep or energy could shift scores without touching depression at all, and a real antidepressant effect could be diluted by everyone in the sample who was never depressed to begin with.
For a reader who has ever been handed a mood questionnaire in a waiting room, the practical takeaway is modest and useful. These instruments are screens, not diagnoses. Crossing a line on one is a reason for a conversation with a clinician, not a verdict. The authors are not arguing that the CES-D-10 should be thrown out. They are arguing that a sum score is a summary, and that summaries lose things, including, in this case, whether the person filling out the form felt sad at all.