What the catalogue recovers without looking for it

A procedure that does not know families have branches, that has never seen a protoform, and that demands one thing only — that a consonantal skeleton recur for the same concept in three distinct languages — produces, family after family, a list specialists recognise.

It did not look for it. It came out.

This post collects those recoveries across the nine families, and then says how much must be discounted from them. The second part matters more than the first.

Austroasiatic: the mat and the muh

This is the first family in the series where the widest-reaching codes are not single-consonant.

  • m·t|eye appears in 64 languages and in eleven of the twelve branches
  • m·h|nose, in 58 and in ten

These are the Austroasiatic mat and muh, the two comparanda this family’s comparison has been citing for a century.

Sino-Tibetan: the miŋ, the luŋ

Crossing six to ten branches, without anything reconstructed:

  • m·n|name in 53 languages — the miŋ
  • l·n|stone in 44 — the luŋ
  • m·k|smoke in 46

The comparanda Sino-Tibetan comparison established a century ago.

Austronesian: the pan-Austronesian vocabulary

  • l·m|five in 252 languages
  • m·t|eye in 246
  • t·l·n|ear in 54

Uralic: what any Uralicist recognises instantly

n·m|name across 20 languages at 378× enrichment, v·r|blood, k·l|fish, m·n|I, t·n|THOU.

Pama-Nyungan, Afro-Asiatic, Indo-European

m·r|hand in 69 Pama-Nyungan languages, from one end of the continent to the other. m·t|die in 38 Afro-Asiatic languages across two branches, d·m|blood in 29, s·m|name in 25. And n·m|name in 103 Indo-European languages, the widest-reaching code in the entire corpus worked so far.

In Turkic the catalogue does something different and perhaps more telling: it separates branches without knowing they exist. Of its 297 sibling-code pairs, the largest group Chuvash with Siberian, and Oghuz with Western Kipchak.

Now, the discount

This is where an honest post parts company with an announcement.

The Sino-Tibetan paper records 1,819 codes crossing three branches or more. It is an impressive figure and it is not published on its own, because on close inspection it breaks:

code arity branch-crossing rate
unary (one consonant) 28.0 %
binary 14.4 %
ternary 3.3 %

58 % of those 1,819 codes are single-consonant. And the crossing rate falls with arity.

The shorter the code, the more it crosses — which is exactly what one expects if crossing comes from availability and not from inheritance.

The defensible figure is not 1,819. It is 762.

The same holds at the other end of the performance range. In Trans-New Guinea, n groups a hundred languages for eat, and as many for you, I, mother and sleep. That is not a finding: it is a saturated skeleton space.

What can be said, then

One thing only, and that is why the method is cheap: that certain form–meaning correspondences recur across languages more than chance would produce, with chance measured explicitly.

That m·t|eye appears in 64 Austroasiatic languages across eleven branches does not say those words are cognates. It may be common inheritance, an old loan, a recent loan, convergence or coincidence. The catalogue does not choose, because without protoforms it cannot, and it will not pretend otherwise.

What it does say is not nothing: that an instrument built from two observable facts — form and concept — and no genealogical hypothesis rediscovers, on its own, part of what two centuries of the comparative method established.

Where the comparative method has been applied, that is a calibration. Where it has not — 224 macrosystems out of 225 — it is a list of testable hypotheses that did not exist before.


The series papers are in the Consonantal Code Catalogues section of Publications.

The method: The catalogue of consonantal codes (PDF)