The 2.8-consonant step, and its falsification three papers later
The 2.8-consonant step, and its falsification
This is the story of a finding that lasted three papers and that the series itself brought down. I tell it whole because the demolition is the result, not an accident along the way.
Seventh family: a step appears
When the series reached seven families measured with the same code, a regularity appeared that six had not shown.
A catalogue’s null — how much the procedure would group by shuffling concepts within each language, that is, how much chance alone groups — did not decline smoothly with word length. It jumped.
The two families with fewer than three consonants per form — Nuclear Trans New Guinea at 2.41 and Austronesian at 2.79 — had nulls of 37 % and 44 %. The five above 3.03 sat between 4.7 % and 15 %.
An eightfold difference on either side of a very narrow step.
It made mechanical sense: with short words the skeleton space saturates, and anything resembles anything. In Trans-New Guinea, 35 % of the resulting codes are single-consonant, and so are fourteen of the fifteen with the widest reach. n groups a hundred languages for eat, and as many for you, I, mother and sleep.
It was the family where the instrument performs worst in the whole series — 13 times chance, against 216× for Turkic and 104× for Indo-European — and it was chosen precisely for that. A described floor is worth more than an intuited one.
Eighth family: the prediction is registered, and confirmed
A finding made by looking at seven cases is a description, not a law. To be anything more, it had to be risked.
So before computing anything about Sino-Tibetan, the prediction was written down: this family has 2.22 consonants per form, therefore its null must be high.
It was confirmed. The Sino-Tibetan null is 42.0 %, against 4.7 % for Afro-Asiatic. Collision is 71.9 %, the highest in the series.
A secondary part of the prediction failed, and the previous paper stood corrected by what was measured there. Out of sample, which is the only test that counts.
So far, so good. Too good.
Ninth family: the step collapses
Austroasiatic entered the series on a new criterion — distributed branches × long word — and brought with it the case that was missing.
| family | consonants per form | null | languages |
|---|---|---|---|
| Austroasiatic | 2.75 | 14.6 % | 94 |
| Austronesian | 2.78 | 39.0 % | 582 |
Same word length. Twenty-four points apart on the null.
The step does not exist where it was said to be. What separates these two families is not the word: it is the languages. Ninety-four against five hundred and eighty-two.
What was actually going on
With all nine families measured, the correlations are clear and pull in opposite directions:
- r(arity, null) = −0.84 — the longer the word, the lower the null
- r(log languages, null) = +0.70 — the more languages, the higher the null
And the two variables are largely independent of one another.
The null is a function of word length and of sample size. The two had been confounded because, until now, every short-word family was also a large one.
Seven families were not enough to separate them. Eight were not either. It took the ninth — chosen on a criterion that was not even looking for this — for the case that breaks the confound to appear: a family with short words and a small sample.
Why this is published this way
It could have been rewritten. All nine families were known before any was published; nothing forced us to leave on record a finding the series itself was about to knock down.
But then the one thing that makes a series of this kind defensible would be invisible: that each paper registers its prediction before measuring, and that the next one can kill it.
A catalogue of consonantal codes is not defended by saying it gets things right. It is defended by showing where it stops getting them right, and who found out.
The papers:
- Consonant codes in 196 Trans-New Guinea languages: the floor of the method — where the step appears
- Consonant codes in 267 Sino-Tibetan languages — where it is confirmed out of sample
- Consonant codes in 94 Austroasiatic languages — where it collapses