The family where the root was already a theory: the consonantal skeleton against the triliteral root
The family where the root was already a theory
There is a problem with instruments you define yourself: they can only be checked against themselves.
In Austronesian, Turkic, Uralic, Pama-Nyungan and Indo-European, the consonantal skeleton was our object. We defined it, we measured it, we compared it against its own null. Five families measuring the ruler with the ruler.
Afro-Asiatic is different, for a reason no other family in the series offers.
The Semitic triliteral root — k-t-b, “write” — is exactly our object, described by another route and more than a thousand years ago.
Arabic and Hebrew philology reached the consonantal skeleton by a path that has nothing to do with ours, and left it described with a precision that needs nothing from us. That makes Afro-Asiatic the only place in the corpus where the instrument can converge with or diverge from an external description.
It is the most informative test in the series. It is also the riskiest: if the skeleton fails to rediscover the root where the root exists and is described, the problem is the skeleton’s.
The contrast was registered before anything was run
The hypothesis was written down — with its null, its declared direction and its falsifiers — before looking at the data. Without that, the test is worthless: one always finds what one goes looking for afterwards.
The main hypothesis is confirmed.
Semitic concentrates 50.6 % of its forms in skeletons of exactly three consonants, against 31.3 % for Chadic. A gap of 19.3 points, at p = 0.0001 under a null that permutes branch over languages.
And it is not that Semitic words are shorter. Mean consonants per form across the four branches are 2.99 · 3.13 · 3.02 · 3.02 — practically identical.
It is concentration, not length. A procedure blind to grammar rediscovers the shape of the triliteral root without having used it.
And two hypotheses from the same preregistration fail
This is the part that matters more, and it is published in the same detail as the part that worked.
Two of the registered hypotheses are not confirmed. What they diagnose is not the family: it is the instrument. A preregistration from which only the successful part is published is not a preregistration — it is a selection made after the fact.
What the catalogue recovered on its own
2,121 codes grouping 11,082 forms from 111 languages. The procedure groups 17.4 % of the available forms and its precision exceeds chance by a factor of 44: 5.26 % against 0.119 %.
Over grouping without the three-language threshold, the advantage is 25.0 points — 29.7 % observed against 4.7 % null, 6.3 times chance.
And the highest-reach codes reproduce the vocabulary Afro-Asiatic comparison had already established, without anything having been reconstructed:
m·t|dieacross 38 languages in two branchesd·m|bloodacross 29s·m|nameacross 25
None of them was searched for. All came out of requiring that a skeleton recur for the same concept in three distinct languages.
What this catalogue is not
Worth saying before anyone reads too much into it.
It does not reconstruct Proto-Afro-Asiatic and does not consult it. It does not use the cognacy judgements that exist for this family. It proposes no etymologies. And it does not claim that the highest-reach codes are cognates, though many are: the catalogue cannot separate inheritance from contact, and will not pretend it can.
There is also a coverage limit, declared before the results rather than after. Of the 111 languages, 72 % are Chadic, and two of the four branches hold two languages each. Omotic and Egyptian are missing; the corpus does not carry them.
This catalogue does not cover Afro-Asiatic. It covers the four branches the corpus contains, and it should be cited that way.
The paper: Consonant codes in 111 Afro-Asiatic languages: the family where the consonantal root was already a theory (PDF) · Versión en español
The method: The catalogue of consonantal codes