NotesNudge the notes you never go back to

Effect Sizes in Learning Research

Read anything citing research in this area and you will meet effect sizes: g = 0.43, d = 0.7, "a large effect." The numbers look like precision and are frequently doing the opposite work — closing an argument that the underlying evidence leaves open.

Ten minutes on what they mean is enough to read the claims properly.

A related workplace concept is remote employee monitoring software, which offers a useful contrast with the personal knowledge-work problem discussed here.

Reviewed August 9, 2026.

What the number is

An effect size expresses a difference in standard deviations rather than in raw units, so that results from different studies can be compared.

d (Cohen's d) and g (Hedges' g) are the common ones. Hedges' g corrects for small samples; for practical reading they mean the same thing.

For an independent external reference related to this topic, see Nature.

An effect size of 0.5 means the average person in the treatment group scored half a standard deviation above the average person in the control group. Roughly: the average treated person outperformed about 69% of the control group.

What counts as large

The conventional thresholds — 0.2 small, 0.5 medium, 0.8 large — come from Jacob Cohen and are more than half a century old. He offered them as rough conventions in the absence of anything better, and they are now quoted as though they were properties of the world.

They travel badly. In education specifically, effects that are small by that scale are often large relative to what field interventions actually achieve, and applying laboratory conventions to classroom results systematically undervalues real findings.

So a number without a comparison class is not interpretable. g = 0.24 means little until you know what other things in that literature achieve.

Why the range matters more than the average

The single most useful habit when reading these.

A meta-analysis reporting a mean effect of 0.5 may be summarising studies ranging from 0.1 to 1.5. The mean is a fact about the collection of studies; the range is a fact about the world, and it tells you the effect depends heavily on conditions you have not been told about.

The spacing literature is a clear example: one meta-analysis found benefits from about g = 0.29 to g = 2.16 with significant heterogeneity. Quoting the middle of that as the effect of spacing would be arithmetically defensible and substantively misleading.

And check whether the interval crosses zero. In mathematics, testing versus restudy came out at g = 0.18 with a confidence interval spanning zero, which means the honest reading is that no effect was reliably detected — not that a small one was found.

What inflates effect sizes

Four features, none of which requires anyone to cheat.

A test aligned to the intervention. Measuring exactly what was practised produces larger effects than an independent measure.

A short delay. Measure immediately and effects are larger; measure in a month and many shrink or vanish. For anything about memory, the delay is the whole point and should be stated.

A narrow outcome. "Recall of this word list" moves far more easily than "understanding."

And a weak comparison. Against nothing, against unstructured study, or against a control doing something less interesting — each produces a bigger gap than a comparison against a genuinely good alternative.

When you see a striking number, one of these four usually explains it.

The four questions

For any effect size in this field:

Compared with what?

Measured how, and how long after?

What was the range across studies, not just the mean?

And does the interval cross zero?

Those four take a minute and they are most of what separates reading research from quoting it.

Why this matters here

Because the note-taking and productivity literature borrows heavily from memory research, and borrowed figures lose their conditions on the way.

A number established for vocabulary recall at a one-week delay, in a laboratory, against a restudy control, becomes "spaced repetition improves memory by X" and then becomes a product feature. Each step is small; the destination is not supported by the origin.

The defence is not scepticism about the research. It is asking what the number was measured on.

The short version