Digital One Foundation The Inherited Default
Paper 002 · Construction

A convention is a count, not a judgement.

A practice several teams arrived at separately is the organisation's, and something has to decide when that has happened. These notes derive the decision from the shape of its output — one sentence, published to every project in the organisation, carrying a count of distinct scopes as the whole of its warrant — which forces two independent scopes and no model in the loop. Those properties rest on a comparison of two strings, and one byte of punctuation broke it.

§1A convention is published before anyone asks for it

Start from the output. The promotion stage produces no score and no ranking. It produces one sentence, in the organisation's own words, addressed to every project inside it and asserted as a thing the organisation does — carrying, as the whole of its warrant, the count of distinct scopes the practice was found in.[1]

Nothing about it is negotiated afterwards. A project that did not exist when any of the material was written is handed the organisation scope and nothing else, so this promotion is the only route by which one repository's practice can reach it.[2] The sentence arrives already believed.

One requirement falls straight out of that. A scope marked sensitive contributes to no convention, because a convention's destination is the whole organisation.[1] A promotion that could be talked into reading a sensitive scope would not leak a candidate — the persuasion would itself be the publication.

The requirement, precisely

A practice is promoted only when statements of it are found in at least two distinct scopes. The unit of the count is the scope — a repository, a bounded place where one team writes things down — not the sentence, the file or the author: counting sentences would measure how thoroughly a team documents itself, counting scopes how many separate places arrived at the practice. A scope contributes once, however often it repeats itself.

It does not over-fire at the size measured. Run against the local store on 18 September 2026, the pass had promoted exactly one convention, cited to three scopes, among 326 distilled facts.[3]

§2A counting argument can be re-run; a judgement cannot

There is an obvious alternative: hand the candidates to a model and ask whether they express one organisational convention. It is declined, not because the answer would be poor, but because a sentence published as something an organisation holds has to survive the question why does this say that. A count answers in public — these scopes, this overlap, this threshold, over this corpus — every term of it re-derivable by whoever holds the corpus. A judgement answers with a provider, a model version, a prompt and a date, none of which is in the corpus.

Three costs follow. The organisation's institutional layer would depend on a vendor's prompt, so a revision nobody inside the organisation saw could change what it is recorded as believing. It would cost a call for every candidate statement. And it would drift: two runs over an unchanged corpus can disagree, so what an organisation believes could move on a day when nothing in the organisation moved.[1] The overlap measure standing in its place is cruder and has the opposite properties — auditable, free, identical each run. That does not make it right; it makes it printable, and the next section prints one of its failures.

What must never be traded away

The sentence and its warrant travel together. A convention's warrant is the count and the corpus it was counted over; once that warrant moves into a model's weights the sentence keeps its authority and loses its evidence, and nothing on the page distinguishes those two states from outside.

Turn the construction over and the exposure is plain. Independence is a count of scopes — but only of the scopes whose statements the comparison matched — and declining the model is worth something only if that comparison is right about what two statements say. It reduces each statement to a set of content tokens and scores the intersection over the union: set resemblance, in the form Broder reduced to "set intersection problems".[4] A pair is one practice at 0.55 and above, a candidate runs 25 to 400 characters, and at most 1,200 candidates are considered — each number configured rather than derived.[1]

§3The defect: a full stop made one word into two

The token pattern is /[a-z0-9][a-z0-9._/-]{1,}/g, and the dot sits inside the character class on purpose: crunch.md, package.json, feature/x-y and rate-limit are single identifiers here.[1] The cost was charged at the other end of the sentence. A word ending a sentence kept its stop, so production. and production were two different tokens, and the last content word of every sentence that ended in one matched nothing.

Defect · found by re-measuring the token comparison, 18 September 2026

A trailing full stop made the last word of a sentence a different word

Nothing exotic reproduces it: two statements of one practice, one per repository, identical to the byte except that one of them ends the way sentences end.

the failing input · one practice, written down in two scopes
scope 1   Every service ships with a health endpoint.
scope 2   Every service ships with a health endpoint
                                                    ^ one character

token-set overlap on four such pairs · the bar a pair must clear is 0.55
                                                     before   after
Every service ships with a health endpoint.           0.600   1.000
Database access goes through the repository layer.    0.714   1.000
We deploy on Thursday.                                0.250   0.667
Secrets are read from the environment.                0.400   0.750

two of the four under the bar: held in two scopes, counted in one
strip a TRAILING run of [._/-] only, and crunch.md stays one token

For those two the organisation held a practice in two places and the count saw one; the other two survive with their margin eaten. The fix is one line — strip a trailing run of [._/-] and nothing else — pinned by a test named for the failure rather than the fix: that the comparison "does not let a full stop make the last word of a sentence a different word".[5]

The arithmetic underneath is worth setting out. Of Every service ships with a health endpoint. the comparison sees four words — service, ship, health, endpoint — short tokens and a stop list never reaching it, a trailing plural folding to the singular. Two sets of four differing in one member share three of the five words in their union, and three fifths is the 0.600 that was measured. The stop takes one word out of the intersection and puts a second, unmatchable word into the union, so an overlap of i over u arrives as i−1 over u+1. Two of the four do not reach 1.000 even after the strip, so those statements differ in more than their punctuation.

ONE PRACTICE, WRITTEN TWICE, ONE BYTE APART token-set overlap before the trailing stop is stripped, and after before after 0.0 bar 0.55 1.0 health endpoint repository layer deploy on Thursday secrets from the environment two of the four sat under the bar: a practice held in two scopes, counted in one stripping a trailing run of [._/-] lifts the four, and crunch.md stays one token
Figure 1 The bar is where two statements become one practice, so a trailing byte decides which side of it a pair lands on — and a pair landing short is a practice the organisation holds and the count does not see.

It never surfaced as an error, only as a smaller count, which is indistinguishable from an organisation that has not converged yet. A convention that is not minted leaves no trace, because the pass reports what it promoted, what it already knew and what it retired, never what it should have promoted and did not.

A second instance has the same shape and a different byte: three repositories held one convention and two converged, because one had written services ship with a health endpoint where the others wrote service ships. Folding trailing plurals recovered the third, and both fixes are pinned by tests named for the sentence they protect.[5]

§4What a count cannot establish, stated before anyone asks

The threshold is a configured default, never tuned against labelled data. Measured over a planted corpus before any grouping, unweighted overlap within one convention runs 0.667 to 1.000 and across different conventions 0.429 to 0.667.[3] The ranges meet, so no single unweighted threshold separates them: 0.55 is not what makes the grouping work — the weighting above it is, which is the previous paper's subject.[6]

The plural fold is blunt on purpose: trailing plurals, and neither tense nor derivation, with carve-outs so that class, status and analysis survive, and one collision it knowingly accepts — https folds to http.[1] Neither the fold nor the comparison sees word order.

The largest limit is in the count itself. Two distinct scopes are two places, not two judgements: two repositories scaffolded by one team in one week are a single decision wearing two names, and nothing inside a count of scopes can tell. The same gap leaves the sharing of conventions across an organisational boundary an open question rather than a feature: two sub-organisations sharing a platform team are not two independent observations either, and whatever settles that is about provenance rather than similarity. Not built, not measured.

None of it is re-runnable by a reader today: the engine's repository is not public, and the store figure in §1 was taken against a private store.[3] What is printed here is enough to re-derive the arithmetic, not enough to reproduce the measurement. A count answers how many places wrote a practice down. It does not answer how many times that practice was decided.

On the provenance of this material

The stage described here was built inside SkilledMind, a Digital One property; it is the provenance of this method, not its subject. Its repository is not public, so what is quoted above is an account of that code rather than something a reader can open today — which is why the failing input and the four measured pairs are printed here in full. Both those and the store figure come from one workstation on 18 September 2026, and the store is private. Deliberately absent: prices, plans and tiers, and deployment internals.

References

  1. The promotion stage — packages/consolidation/src/converge.ts, file header: the rules quoted in §1 and §2, and the defaults beside them — threshold 0.55, length 25–400 characters, maxCandidates 1200, the token pattern, the plural fold. Not public.
  2. The birth-brief stage — apps/skilledmind-api/src/genesis.ts (ORG_SCOPE) and its test file, holding the counterfactual leak test read in §2. Same standing as [1].
  3. The measurements of 18 September 2026, on one workstation: the four pairs in §3 before and after the fix, and the pairwise ranges in §4 over a planted corpus with plain unweighted overlap. The store figure in §1 is private and cannot be re-run by a reader.
  4. A. Z. Broder, On the resemblance and containment of documents, SEQUENCES '97, IEEE, abstract: "to reduce these issues to set intersection problems". His resemblance is unweighted, and the measure is older — Jaccard's coefficient of community — but neither Jaccard paper could be opened for these notes, so nothing here quotes them. cs.princeton.edu/courses/archive/spring13/cos598C/broder97resemblance.pdf, read 18 September 2026.
  5. The regression tests — packages/consolidation/test/converge.test.ts: the one named does not let a full stop make the last word of a sentence a different word, the assertion that crunch.md and package.json survive tokenisation whole, and the plural fold. Not public; the failing input is printed in §3.
  6. The Inherited Default, Paper 001 — Rarity is a property of the corpus: the planted corpus whose generator it prints, and the weighting that answers the ranges quoted in §4.