← All insights

When a scale stops discriminating

Three bands, one of them used STRONG MODERATE EMERGING 204 0 0 BEFORE THE PARSER WAS FIXED, EIGHT CASES SAT IN THE BOTTOM BAND ALL EIGHT WERE THE ONES THE PARSER COULD NOT READ the variation was the bug, not the evidence

A library of case studies carried an evidence tier on every entry: strong, moderate, or emerging. The rule was simple and looked reasonable. Four or more references counted as strong; fewer dropped the case into a lower band.

Eight of two hundred and four cases read emerging. That is the sort of distribution that invites no questions. A small tail of thinner cases in a mostly well-sourced library is exactly what you would expect, and it is what a reader would conclude from the field without ever asking how it was computed.

The eight were not thinner. They were the eight the reference counter could not read.

The library used two templates. The older one headed its citations References; a later, longer one headed them Further Reading. The counting function matched References, Sources and Bibliography, and not the fourth. Fifteen cases written to the later template therefore returned zero references each, while carrying five full academic citations apiece. The library reported 867 references where it held 942.

Fix the parser and every one of the two hundred and four cases lands in the top band. The tier had been reporting a bug in the code that computed it, and nothing else. The eight cases a reader would have treated as weakest evidence were, if anything, the ones written to the more thorough template.

There are two lessons here and the second is the one worth keeping.

The first is procedural. When a derived variable has a small number of cases in an unusual band, look at those cases before you look at anything else. Not because they are interesting, though they may be, but because a small anomalous group in a derived variable is more often a property of the deriving code than of the world. The question to ask is whether the cases in the odd band have anything in common other than the score, and here they had everything in common: they were the fifteen most recent, written to one template, by one process, in one region.

The second is about what happens after the fix, and it is less comfortable.

Once the parser was corrected, the tier returned strong for all 204 cases and the other two branches became unreachable. The function still ran. It still wrote a field into every record. Every case still displayed a tier. And the field now carried no information at all: a variable with zero variance is a constant with a label on it.

That state is worse than the bug was, because the bug was at least visible as variation somebody could investigate. A field that reads strong on everything reads as an assessment having been made. It invites a reader to believe that some cases could have been graded lower and were not. Nothing on the page says the grading has one possible outcome.

This is a common shape and it usually arrives the same way. A threshold is set against the data as it looked when the rule was written. The library grows. Everything added afterwards clears the threshold, because the threshold encoded what was normal at the time. The rule has not changed and the data has not degraded. The rule has simply stopped distinguishing.

The diagnostic is one line and almost nobody runs it: tabulate the derived variable. Not the inputs, which are usually checked, and not the rule, which is usually documented. The output. If the mode accounts for more than about nine tenths of the cases, the field is doing no work and should be defended or dropped.

There are three honest responses when it fails, and the reason to name them is that only one of them is a code change.

Raise the threshold, so the bands separate the library you actually have rather than the one you had.

Weight the signals the rule is short-circuiting. Here, whether a case rested on a randomised trial or a meta-analysis was recorded on every entry and never reached the calculation, because the reference count settled the tier before those were consulted.

Or drop the field, and stop implying a distinction the data does not carry.

Which of the three is right is an editorial judgement about what the library is claiming, not a bug to be patched, and it belongs to whoever owns that claim. What is not defensible is leaving a constant on the page dressed as an assessment. A reader cannot tell the difference, and the field is read by people who will never see the function that produced it.