Four consecutive negative results

There is an obvious temptation in a method that reduces a word to its consonants: strip the affixes first. If Turkish gelmek is “come” plus an infinitive, the infinitive is in the way. Remove it and the catalogue will improve.

It seems obvious. It was tried on four families. It fails on all four.

Turkic

Automatic affix discovery gains 6.3 points of grouping and 13.9 points of null.

That is: it groups more, and chance groups a great deal more. The advantage over chance, which is the only thing that matters, gets worse.

And there is a diagnostic that explains it better than any figure. The discoverer found 365 prefixes in a family without prefixes. Turkic is agglutinating and suffixing; the prefixes the procedure “discovered” do not exist. They are initial resemblances between words, treated as morphology because the procedure has no way of telling one from the other.

Uralic

Grouping up ten points, null up eighteen. Second agglutinating family, same result.

Pama-Nyungan

Grouping up twelve points, null up twenty-six. Third consecutive negative result, declared as such in the paper.

Indo-European

The advantage over chance worsens once again. Fourth family in a row — and this is the family where more is known about morphology than anywhere in linguistics.

The rule that follows

Morphological stripping pays off when the rule is declared, and worsens the result when it is discovered by resemblance.

A declared affix — “in this language the infinitive is -mek / -mak” — is external information the catalogue did not have. An affix discovered by substring frequency is not information: it is the catalogue looking at itself in the mirror. Removing what recurs makes what remains resemble itself more, and chance rises for exactly the same reason the signal does.

It is the same trap the series meets again and again, wearing different faces.

Two more times aggregation lied

The substitution matrix. In Turkic, aggregated over the whole family, it loses the correspondence that defines the family’s primary division — its most firmly established feature, invisible. Only when restricted to the branch pair does it reappear, and then at 6.1 times chance.

Averaging over a whole family is not an overview: it is an erasure.

Maximum enrichment. The code with the greatest advantage over chance turned out, in Turkic, to be a loan again — Russian, this time. And the mechanism can be named: it is produced by the rare concept, not the rare code. A concept almost no language has, borrowed by a few from the same source, yields a spectacular enrichment that says nothing about the family.

And one time it did work

So that the list does not read as a mere collection of failures: the closure of the substitutions had gone two families producing nothing, and in Pama-Nyungan it yielded three clean classes that turned out to be places of articulation.

Nobody asked for them. They came out.

Why publish this

A negative result does not sell. Four in a row, less so.

But the argument of this series is not that the instrument works. It is that the instrument is described precisely enough that one can see where it stops working, and that this description is published in the same paper as the results rather than in an appendix.

A catalogue publishing only its successes would be checkable by nobody. And an instrument that cannot be checked is not an instrument: it is an opinion with tables.


The papers: