Five axes of the code

In 2025 the endolinguistic code was a simple object to state: a bounded, ordered combination of two or three OAS classes, drawn from the normalised root of a word by way of its pronunciation.

That remains true. And it remains insufficient — because practice uncovered that the same definition admits readings which do not yield the same object.

Three real examples, all three from the corpus:

  • Lat. rīvus is Λ·Φ·Σ read whole, and Λ in its descendant Sp. río. Which is “the code”?
  • Eng. ignite is Χ·Ξ·Θ, but the Latin ignis it comes from is Ξ·Ξ·Σ: the velar had nasalised ([ˈɪŋ.nɪs]) and English restored it by reading the letter. Which is “the code” of that network?
  • Eng. tooth is Θ·Θ, but its coderivative network is predominantly Θ·Ξ·Θ. Which one is taught?

None of the three has an answer until one declares which axis is being spoken on. Hence this document, which proposes no new theses: it orders.

Axis I · Level of abstraction — the ladder

The same material reads at three heights:

level what it is example specificity coverage
skeleton the actual canonical consonants d·n·t maximum minimum
code the ordered classes Θ·Ξ·Θ medium medium
nucleus the classes without order {Θ,Ξ} minimum maximum

Measured over 19,623 networks with three languages or more, mean coverage of the modal value is 56.4 % at the skeleton, 65.1 % at the code and 73.7 % at the nucleus.

There is no correct height: there is one per question. And every claim must say which one it was made at.

Below the skeleton lies the cualo — the originating consonantal sensation, continuous and pre-symbolic; above the code lies the molecule, the linked system forming a compound word. Neither is a rung of this ladder: they are its floors.

Axis II · Extent read — word or root

  ROOT code HEARD-WORD code
contains the root without affixes what reaches the ear, affixes included
serves comparative and descent claims claims about psychic legibility
the affix is language-specific noise part of the datum: the hearer does not segment before resonating
they cross never within a single test  

And the affix is not simply noise, for two reasons — the second is the strong one. The hearer receives the whole word; and this is how root stems arise: an extension sticks and stops being an affix, so the root/affix boundary is not ontological but diachronic.

Of the corpus forms, 539,456 have a computed root skeleton, and in 100 % of the networks where it is computed the root code differs from the word code. The choice is never innocuous.

There is a third state: the bare root, the word that is its own root. Recognising these took the usable population of one test from 3,232 to 70,260 networks.

Axis III · The bearer — the form or the network

A word does not have a code. It belongs to a coderivative network, and the code is what is preserved across it.

Teil is Θ·Λ because it sits in the network Teil · deel · del · dails. A consonant pattern without a network is not a code: it is a consonant pattern.

The measurable object is not a value but a network profile. Worked example, TOOTH, collapsed network, 138 languages:

level modal coverage distinct values
skeleton d·n·t 21.0 % 39
code Θ·Ξ·Θ 34.8 % 24
nucleus {Θ,Ξ} 58.7 % 12

English tooth (Θ·Θ) is second, at 17.4 %.

And networks must be collapsed before anything is measured: 157,327 cognate sets reduce to 69,416 networks. Without that collapse, a test can come out significant by counting the same etymon several times — it already happened (p = 0.0155 → 0.133).

Axis IV · Provenance — where the network’s code came from

Four states, a first-class attribute:

provenance what it means witness
inherited the code descends from the root Sl. *bělъ ← *bʰelH-
supervened the code arose during the language’s history Lat. bellus ← *duenelos, where the B-L was not there
borrowed the code arrived with the word Sp. bello ← Occitan bel
restituted a lost class returns by reading the spelling Eng. ignite Χ·Ξ·Θ against Lat. ignis Ξ·Ξ·Σ

The fourth is new, and its signature had to be corrected. The first version said “the loan has more classes than its etymon” — and it fails on its own witness: ignite Χ·Ξ·Θ against ignis Ξ·Ξ·Σ has the same three classes. Restitution is not a lengthening: it is a substitution.

Correct signature, and it requires THREE generations: the child recovers a class the GRANDPARENT had and the PARENT had lost.

Computed over 183,250 three-generation chains: 3,369 restituted forms. And the signature finds the phenomenon on its own — Eng. magnes: grandparent Ϻ·Χ·Ξ·Σ → Lat. magnes [ˈmaŋ.neːs] Ϻ·Ξ·Ξ·Σ → Eng. Ϻ·Χ·Ξ·Σ. An independent instance of ignis, found without looking for it.

Its recall is limited by the data, and this is declared: ignite itself does not fall under the signature, because its own chain has no grandparent with a code.

Axis V · The state of each position

A code can shorten for three different reasons that the model used to write the same way. Two are events in the language; one is an event in us.

notation state in the word? in the code? where it happened
Φ present yes yes
[Φ] hidden — the instrument does not show it yes yes, invisibly in us
(Φ) destructured — passed to the qualic layer yes no in the language
absent — no longer in the form no no in the language

Sp. río = Λ·(Φ): the w of rīvus, facing the vowel, dissolved into it. Ger. rein ← *hrainiz = ∅·Λ·Ξ: the h was lost with no transit.

And every /j/ in the corpus is [hidden], because cell J has no OAS class assigned: 320,845 segments invisible to the code — which is our blindness, not the object’s silence.

The full grid

A claim about a code is determined by five coordinates:

level × extent × bearer × provenance × state

And most disagreements about codes are disagreements of coordinate, not of fact. Tooth is Θ·Θ and its network is Θ·Ξ·Θ: both are true, on different bearers.

And the negations, which in practice help more

A code is not a consonant string · it is not read off the spelling · it is not a morpheme or an etymology · it is not a meaning · it is not the sum of its classes’ profiles · it is not a set, because order counts · and it does not belong to a word: it belongs to a network.

A code means nothing. It orients a field of semantic possibility, observable only through the coderivative networks that instantiate it.


The paper: The Forms of the Endolinguistic Code (PDF) · Versión en español