Indo-European blind

Indo-European is the extreme end of an inequality. It is the family with the densest reconstruction, the most thoroughly reviewed expert cognacy and the deepest bibliography in existence. Two centuries of excellent work, available, indexed, one click away.

This catalogue was built without consulting any of it.

The commitment

This is not a rhetorical flourish. It is a list of concrete renunciations:

  • No protoform is used. Neither to build the catalogue nor to read it.
  • No cognate list is used. The IE-CoR source contributes 25,731 forms to this corpus and carries expert cognacy judgements: they have not been read. Forms and glosses are taken from it as from any other source.
  • No sound law is cited, neither by name nor by content.
  • No result is compared with the established classification beyond the branch labels the corpus already carries as metadata, used exactly as in the four preceding families.

The four previous families — Turkic, Uralic, Austronesian, Pama-Nyungan — were chosen partly because their comparative literature is thinner, so giving it up cost little.

Here it costs everything. And that is why it is done.

Because if the instrument only works where nobody can check it, it does not work.

The figures

10,116 codes grouping 94,152 forms from 207 languages.

Precision of 8.73 % against a chance rate of 0.084 % — that is, 104 times chance, the second highest in the series after Turkic. Collision is 52.5 % and mean arity 3.35, the highest, tied with Pama-Nyungan. The catalogue covers 1,946 concepts, more than double any earlier family.

The widest-reaching code in the whole corpus

concept code languages enrichment branches
name n·m 103 379× 6
three t·r 81 130× 10
two d 79 128× 9
nose n·s 79 150× 6

n·m|name, at 103 languages, is the widest-reaching code in the entire corpus worked so far — ahead of Pama-Nyungan m·r|hand, which reaches 69, and of any Austronesian code.

And here one must be exact, because this is where the reader runs ahead. We write n·m|name and never n·m = name. The code does not mean “name”. It is the shape “name” repeatedly takes in this family, under these information conditions. The branch column counts how many branches a code appears in; it does not claim that this distribution has a historical cause, which is a different question and not this method’s.

Anyone who knows the family will recognise what lies behind these. The point is that the procedure did not.

Sixty-one branch pairs

Indo-European is the first family in the series with enough material to compute substitutions over 61 branch pairs, and it yields the sharpest alternation concentrations measured so far.

Geography comes out monotonic across the six cubes — and this despite being the family with languages transplanted to other continents, the case where one would most expect the geographical signal to break.

And a fourth consecutive negative result

Automatic affix discovery once again worsens the advantage over chance. Fourth family in a row.

It is not an accident that it repeats: it is a result. Morphological stripping pays off when the rule is declared, and damages the catalogue when it is discovered by resemblance. It had already been seen in Turkic, Uralic and Pama-Nyungan; Indo-European confirms it in the family where most is known about morphology.

What remains to be said

Unary codes head the reach list: four of the fifteen widest-reaching codes have a single consonant. They stay outside the figures this paper defends, but their presence says something worth not hiding — in a large family, the instrument produces coincidence easily.

That is the price of working blind, and it is published alongside the result.


The paper: Consonant codes in 207 Indo-European languages: a catalogue built blind (PDF) · Versión en español

The method: The catalogue of consonantal codes