Two children in the same household, born on the same day, raised at the same dinner table. One has more trouble sitting still and finishing homework than the other. Do their report cards diverge because of that difference, or is something else driving both?

That question is the reason twins are so useful to behavior geneticists, and it is the engine behind a new analysis in Behavior Genetics by Luis Castro-de-Araujo and colleagues at Virginia Commonwealth University, working with data from the Adolescent Brain Cognitive Development Study, a long-running project following more than 10,000 American children.

The link between ADHD symptoms and poor grades is old news. Children with more inattention, hyperactivity and impulsivity do worse at school; that has been shown many times. What has stayed murky is which way the arrow points. It is entirely plausible that struggling at school makes a child restless, disengaged and fidgety, rather than the other way around. Or that some third thing, a family's circumstances, a set of shared genes, produces both.

Settling this the usual way would mean a randomized trial, and you cannot randomly assign children to have ADHD. So the team reached for a set of statistical designs built for exactly this problem. The simplest, called a Direction of Causation model, compares how traits correlate across twins in a pair: if ADHD in one twin predicts poor grades in the other, and identical twins show a different pattern than fraternal ones, the pattern of correlations can hint at which trait is upstream. The more elaborate versions add polygenic scores, single numbers summarizing thousands of small genetic effects, as a kind of natural instrument. Because the genes you carry are fixed before you ever sit an exam, they cannot have been caused by your grades.

The measurements were ordinary. Caregivers filled out the Child Behavior Checklist, a standard questionnaire, and separately reported their child's grades from the previous school year on a scale running from A+ down to F. The team used data from the study's third annual follow-up, when the children were around 12, chosen as a compromise: old enough for symptoms to show, early enough to avoid the wave of missed visits the pandemic caused later on.

The first two models could not answer the question. The classic Direction of Causation model gave an ADHD-to-grades estimate of -0.16 and a grades-to-ADHD estimate of -0.30, and when the authors formally compared the forward and reverse versions, neither fit the data appreciably better. Running the same comparison with polygenic scores added produced the same standoff. Both directions were consistent with what the children's records showed.

Only the most elaborate model, MR-DoC2, which carries a polygenic score for each trait and can therefore test both arrows at once, separated them. The path from ADHD symptoms to grades came out at -0.31, with a confidence interval running from -0.62 to -0.11, comfortably clear of zero. Deleting that path from the model made the fit significantly worse (p = 0.006). The reverse path, grades to ADHD, was larger in size at -0.60 but far less precise, its interval stretching from -1.20 to 0.03 and crossing zero; removing it cost the model nothing.

The team also tested whether any of this differed between boys and girls, given that ADHD is diagnosed in roughly twice as many boys. Boys and girls differed in average scores, but the variance components and cross-sex correlations showed no meaningful moderation.

Why it matters

A correlation tells a parent or a teacher nothing about what to do. If poor grades were driving the restlessness, the sensible intervention would be academic support. This analysis, in a sample recruited from the general community rather than from clinics, points the other way, and it does so while accounting for two problems that dog most genetic causal studies: horizontal pleiotropy (genes influencing the outcome through routes that bypass the trait you are studying) and shared family background.

The authors are careful about what they have not shown. The polygenic scores were built from genome-wide studies of European populations only, so the two strongest models used just 1,071 children of European ancestry out of a considerably more diverse sample, a limitation the authors flag as a reason to want multi-ancestry data. The checklist is a screening questionnaire, not a diagnosis. The grades in question were reported by caregivers, covering a school year disrupted by COVID-19, which the authors note may have added noise. And their models assume away complications, including gene-environment correlation, that could bias the numbers if present.

What remains is a single well-constructed estimate pointing one direction, in children rather than the adults most previous work has studied, and an honest acknowledgment that the reverse effect was not ruled out so much as left unresolved.