A catalogue without genealogy

Historical linguistics has a powerful and expensive instrument. It reconstructs protoforms, establishes correspondences, produces trees. It works, and it has been working for two centuries.

But it has been applied with enormous inequality. The corpus this series works on holds 225 macrosystems, and expert-reviewed cognacy judgements exist for one of them. One out of 225.

An uncomfortable consequence follows, and it is worth stating before anything else: any comparative study resting on the historical layer of the lexicon will measure, without meaning to, the density of the comparative literature rather than the past of the languages.

The question this series asks is whether anything useful can be done by giving up reconstruction. Not improving it, not replacing it: doing without it, and seeing what remains.

What remains

What remains is a more modest and far cheaper object: a catalogue of consonantal codes.

It is built from three things and three only — the consonantal skeleton of a form, the concept it expresses, and the requirement that the match occur in at least three distinct languages. Nothing else enters. No branch classification, no protoform, no cognate list, no hypothesis about what descends from what.

The procedure does not know that families have branches.

Genealogy enters at the end, and only to read what came out. When codes are crossed with branch labels, it is done over an already-closed catalogue, and the question is not “which codes does the classification predict?” but the reverse: does the catalogue, on its own, rediscover something the classification recognises?

That order is not a stylistic preference. It is what makes the instrument useful where there is no reconstruction, which is exactly what it was designed for. A method that needed the classification in order to build its catalogue could not be applied precisely where it is needed.

What is not attempted

The object is easy to over-read, so it is worth bounding precisely.

No protoforms are reconstructed. No ancestral form, not even implicitly.

No relatedness is claimed. That two words share a code does not say they descend from a common one.

Inheritance is not distinguished from borrowing. Without protoforms it is impossible, and there will be no pretending otherwise. A group of words sharing a code may be common inheritance, an old loan, a recent loan, convergence or coincidence. The catalogue does not choose.

No meaning is attributed to a code. We write n·m|name and never n·m = name. The code does not mean “name”: it is the shape “name” repeatedly takes, under those information conditions.

Only one thing can be asserted, and that is why the method is cheap: that certain form–meaning correspondences recur across languages more than chance would produce, with chance measured explicitly. Everything else stays outside.

It is less than the comparative method gives. And it can be obtained where the comparative method has not been applied.

The test, in the one place where the answer is known

An instrument that cannot be checked is worth nothing. Indo-European has gold cognacy reviewed by experts, so that is where the measurement happens.

Against that cognacy, concept plus identical skeleton yields 93.3 % precision against a 26.0 % baseline, with 13.6 % coverage.

High precision, low coverage. Which is exactly what a catalogue should be: it finds a reliable minority of the relations that exist, and does not pretend to find them all.

Three corrections that cost five papers

The method paper does not publish only the protocol. It also publishes the three things that had to be corrected, because five successive applications forced them:

The null is computed exactly, not estimated. Estimating it with eighty thousand shuffles published, across three papers, a confidence interval that measured not the uncertainty of the data but the uncertainty of our own draw. With the closed form, the published figures changed: 225× → 218×, 77× → 78×, 39× → 37×.

Beating a shuffle null licenses the assertion of a correspondence and nothing more. Reading it as migration requires a protoform the method does not have.

The threshold is fixed between two bounds — the reachable ceiling and the false-positive rate — because the 2× used across five papers turned out to be exceeded by chance in 64 % of cases for one of the statistics employed.

An apparatus described without the failures that produced it is a list of good intentions.

Where this sits, and where it does not

Reducing a word to its consonants is not new: it is Dolgopolsky’s consonant class matching, and Turchin, Peiros and Gell-Mann’s. The variant used here differs in two ways.

It does not truncate: classic consonant class matching compares the first two consonants, whereas here the whole sequence is kept, because arity is an object of study and not a fixed parameter.

And the output is not a distance matrix but a catalogue queryable by concept. The difference is not one of format. A matrix answers “how similar are these two languages?”; a catalogue answers “what shape does this concept repeatedly take in this macrosystem?”. The second question does not average over the lexicon, and so each code can be inspected, disputed and refuted one at a time.

Which is, in the end, the only reason to publish it.


The paper: The catalogue of consonantal codes: a comparative method without genealogy, cognacy or reconstruction (PDF) · Versión en español

The infrastructure: Integrative Corpus — 3.79 million forms, 4,735 doculects, 225 macrosystems.