Digital One Foundation Back to the record
The Text Layer

The document is not the text.

A PDF is a set of drawing instructions. The text a reader sees is one layer of it — sometimes present, sometimes not, never the whole. When a document is handed to a language model, something has to decide what the model actually receives: a rendered picture of every page, the text pulled out from underneath it, or a plain-text rendering built from the document's structure. That decision is paid for on every page, on every pass, and it decides whether a number the model quotes is still attached to the row it came from.

These are working notes on that decision: what conversion to plain text keeps and what it loses, how a saving should be measured before it is claimed, and what still lies beyond the text layer — the pictures, the drawings and the plans that no text layer describes, and the question of where converted text should live.

Why publish thisThe first number this series had to correct was its own

Claims about what a document costs a model are easy to publish and awkward to check. A percentage is quoted, the corpus behind it is private, the counter that produced it is unnamed, and the reader is left with a choice between trust and silence. The number that opens this series was published that way, at 64%, and was wrong: it had been counted with one vendor's tokenizer to support a claim about another's, on documents nobody outside could re-run. Measured with the vendor's own counter on a public corpus it is 53%, and the honest figure moved in the unflattering direction.

That is the argument for publishing the method rather than the result. A measurement described precisely enough to be re-run — the documents named by URL and hash, the counter named by endpoint and model, the arithmetic printed — can be attacked by people who owe the authors nothing, which is the only review that counts. These notes are written for that reader: someone deciding what to demand of a document pipeline before any model sees a page, not someone deciding whether to buy one.

They are contributed to the Foundation's public record under an open licence. The Foundation is being established under Dutch law as an independent steward of exactly this kind of material — open evidence, open standards, and neutral infrastructure meant to outlive any single company. Where a method was developed inside a commercial system, that provenance is stated plainly; where a direction is research rather than result, it is marked as open rather than dressed up.

The standard these are held to

Every figure in these papers is either reproducible from a corpus named by URL and hash, counted with the counter the claim is about, or a named vendor's published figure cited to its source with the date it was read. Every limitation is stated in the paper that would otherwise be read as claiming otherwise. A hypothesis is called one.

The papersTwo notes to begin with: the method, and what lies beyond it

ScopeWhat these notes deliberately leave out

Three categories are absent by design, and it is better to say so than to let a reader hunt for them.

Everything else — the measurements, the corpus, the arithmetic, the failure modes and the reasoning — is here in full, because a measurement you cannot re-run is a measurement you cannot check.

Not legal analysis

Where a paper mentions European instruments or a jurisdiction, it does so to explain why a technical property matters — where a document is processed, what a record can later show. Those references are contextual. They are not legal analysis, they are not advice, and where a legal question is genuinely unsettled the papers say so rather than resolving it in passing.