Digital One Foundation The Measured Endpoint
Paper 004 · Limits

Admission tests agreement, not accuracy.

Staging and consensus show that a promoted partner result agreed with its peers and the operator's baseline, not that it was accurate, and the public feed shows no individual result, promoted or rejected. A region derived from the source network shows where the traffic left as a mapping places it, not necessarily where the model ran. Per-provider circuit breakers keep a failing provider out of the probe budget apart from half-open trials, and its worst hours can arrive as absent rows the public history does not mark as not measured.

§1Promotion tests agreement, not accuracy

The design these notes draw on, a probe fleet whose workers sign their results, has two trust tiers. First-party workers run on centrally held credentials and form the trusted baseline; their signed results go direct. Partner workers bring their own credentials and are untrusted until verified: each result they submit lands in staging and is promoted only after cross-worker consensus, a tolerance check against the baseline, and anomaly and fraud checks. The quorum, the tolerances and the fraud signals are not specified. The design also lets partner workers earn standing through verified contribution; what standing changes, and whether it relaxes any check, is not described. A standing earned by agreeing is spent by diverging: a partner honest until trusted and shaded after is caught only by checks that do not relax with reputation.

RFC 9334, read as an analogy — it concerns attesting a device's own state, and nothing here suggests hardware attestation — separates what the tiers blur: "Trust is a choice one makes about another system. Trustworthiness is a quality about the other system that can be used in making one's decision to trust it or not."[1] The tiers are trust in that sense.

Three consequences follow, as reasoning rather than measurement. Consensus among partners means something only if distinct keys are distinct operators: "if a single faulty entity can present multiple identities, it can control a substantial fraction of the system".[2] Douceur's setting has no central authority and this design has one at enrolment, so its defence is enrolment's check that two keys are two parties, which is not described. Colluding partners reporting the same shaded figure pass consensus by construction. And a tolerance against the baseline lets partner data confirm the operator more easily than contradict it: a degradation visible only from a network no first-party worker reaches looks, to that check, like an anomaly. That is the divergence a distributed fleet exists to find.

Promotion is not truth

A promoted result agreed with other results and with the operator's own baseline. That does not show the baseline was right, and no individual result, promoted or rejected, appears in the public feed, so the rate at which divergence was discarded cannot be seen from outside.

§2A region is where the traffic left, as a mapping places it

The design takes a worker's region from its source network at ingestion, not from the worker's own report. That improves on asking the worker, and the design's description, which calls the region "verified", overstates it. Ingestion sees the address the traffic left from; turning that into a place takes a mapping the design does not describe — typically a geolocation dataset, some of it published by network operators about their own addresses. RFC 8805, the format for such feeds, tells a consumer to "only trust geolocation information for IP addresses or prefixes for which the publisher has been verified as administratively authoritative".[3] The self-report has moved from the worker to whoever publishes the address data.

Its accuracy has been measured, and is uneven. Poese and colleagues, in an editorial note whose header says it was not peer reviewed, on data from 2010 and 2011, wrote that "Geolocation databases can claim country-level accuracy, but certainly not city-level".[4] A cloud region usually names a metro. Weinberg and colleagues measured parties that choose where their traffic appears to come from: of the proxy servers they tested in 2018, "one-third of them are definitely not located in the advertised countries, and another third might not be".[5] Those were commercial proxy services, not a measurement fleet; the method-level point carries over to a partner worker on a network of its own choosing.

Two more distances are reasoning only: where traffic exited is not necessarily where the model ran, and it is not where users are. Active, latency-based geolocation is the kind of check that can bound a location and rule out a claimed one, and the design's description does not mention it. Neither /v1/scores nor /v1/history, as fetched on 14 September 2026, carries a region field; a third endpoint the operator's homepage reads was not fetched for these notes.[6]

§3A provider's worst hours can arrive as missing hours

Provider failures — quota exhaustion, regional gating, bursts of server errors — are scheduling inputs in this design, with a circuit breaker per provider. In Fowler's description, once failures cross a threshold calls fail "without the protected call being made at all", and a half-open state makes "a real call as trial to see if the problem is fixed".[7] The design excludes open circuits before dividing the per-cycle probe budget, so the other providers' shares are computed without the failing one, apart from its half-open trials. That is a scheduling property, not a measurement one. Nor does it isolate the score: the judge panel of Paper 001 is drawn from the same providers, so an outage that opens one provider's circuit can also remove a panel member, and that touches the quality score of every provider's output, not only the failing one's.

The operator describes honoured backoff as capped at thirty minutes, whatever recovery time a provider advertises; that is its builders' description, and nothing a reader can fetch today shows it. What such a cap trades can be read against the specifications. A Retry-After sent with a 503 "indicates how long the service is expected to be unavailable to the client"[8] — advisory wording, so re-testing sooner breaks no rule, but where a provider asked for longer the trial lands inside that window. A 429 "indicates that the user has sent too many requests in a given amount of time":[9] a fact about the caller's quota, which a breaker keyed per provider would record as the provider's failure; the design does not say how it tells the two apart.

What a reader meets is simpler. While a circuit is open nothing is measured, so a provider's worst hours can arrive as absent hours rather than as low scores. In the public history fetched on 14 September 2026 — 18 series, 169 hourly timestamps, ten series complete — these notes counted the five openai series each missing the fourteen hours from 00:00 to 13:00 UTC on 11 September (two of them also missing 14:00 UTC on 7 September), and the three google series each missing ten or eleven of the twelve hours from 23:00 UTC on 7 September.[6] The feed gives no reason, and these notes attribute neither gap to a breaker. /v1/scores carries a current status ("healthy" or "elevated_latency" in that snapshot) but /v1/history carries none, so a missing hour is an absent row and nothing marks it "not measured". Missing rows are not the only form a bad hour takes in that snapshot: 234 points carry a quality of 0, and the google gap is preceded, at 22:00 UTC on 7 September, by rows at quality 0 and reasoning 0 in all three series. Nothing in either endpoint distinguishes a judged 0 from an hour in which no verdict was produced. An average over the week therefore silently drops some bad hours and may count others as zeros, and a reader cannot tell which.

/v1/history · generatedAt 2026-09-14T14:28:00.963Z · hourly ts missing, UTC
openai  5 series   2026-09-11T00:00–13:00, in all five
                  2026-09-07T14:00, in two of them
google  3 series   2026-09-07T23:00 to 2026-09-08T10:00, ten or eleven hours each
the other ten series have all 169 hourly points · model identities withheld

§4What this cannot establish, stated before anyone asks

These notes describe an architecture. They publish no rubric, no prompt set, no panel composition and no test suite — and, for this paper, no admission thresholds. The Delegation Record Paper 007 made a published method someone other than its author can re-run, or outcomes measured on a named corpus, the condition for calling a score sound, and recorded that frontierscore.ai's public pages, read 12 September 2026, link no methodology.[10] That condition is still unmet. A reader can check the reasoning, the cited specifications and the shape of the feed, and cannot check that the fleet does what is described.

  1. The tier the admission checks skip is the operator's own. First-party results form the baseline without passing the checks partner results pass.
  2. Nothing here measures the fleet. The design's own size and throughput figures are left out because none can be checked from anything public.
On the provenance of this material

The architecture described here comes from FrontierScore, a Digital One property, and the design description these notes draw on was written by its builders. Its repository is not public, so that description — the two trust tiers (first-party and partner), the admission steps, the region taken at ingestion, the per-provider circuit breakers, the thirty-minute cap — is their account and is not something a reader can check today. The public delayed feed at frontierscore.ai's API host is checkable while its seven-day window lasts, and the feed figures above were counted from snapshots of it taken on 14 September 2026. Deliberately absent: prices, plans and tiers, availability, and deployment internals.

References

  1. H. Birkholz, D. Thaler, M. Richardson, N. Smith, W. Pan, Remote ATtestation procedureS (RATS) Architecture, RFC 9334, January 2023, §1; Informational. Used as an analogy. rfc-editor.org/rfc/rfc9334, read 14 September 2026.
  2. J. R. Douceur, The Sybil Attack, IPTPS 2002, abstract and §1. microsoft.com/en-us/research/wp-content/uploads/2002/01/IPTPS2002.pdf, read 14 September 2026.
  3. E. Kline et al., A Format for Self-Published IP Geolocation Feeds, RFC 8805, August 2020, §3.2; Independent Submission, Informational. rfc-editor.org/rfc/rfc8805.txt, read 14 September 2026.
  4. I. Poese, S. Uhlig, M. A. Kaafar, B. Donnet, B. Gueye, IP Geolocation Databases: Unreliable?, ACM SIGCOMM Computer Communication Review 41(2), April 2011, abstract. An editorial note, not peer reviewed per its own header; data from 2010–2011. Text read from the ORBi repository copy, orbi.uliege.be/bitstream/2268/107073/1/p53-v41n2h2-uhligA.pdf, read 14 September 2026.
  5. Z. Weinberg, S. Cho, N. Christin, V. Sekar, P. Gill, How to Catch when Proxies Lie: Verifying the Physical Locations of Network Proxies with Active Geolocation, IMC ’18, abstract and §8. Author-hosted copy, contrib.andrew.cmu.edu/~nicolasc/publications/Weinberg-IMC18.pdf, read 14 September 2026.
  6. The public delayed feed, publicapi.frontierscore.ai/v1/scores (generatedAt 2026-09-14T14:28:00.505Z) and /v1/history (generatedAt 2026-09-14T14:28:00.963Z), both declaring delayHours 1 and resolution "1h". Every count here was made by these notes over the snapshots saved that day; a later fetch returns different values. The history window is seven days, so these gaps are no longer in the live feed after 18 September 2026; the missing timestamps are printed in §3. An API, not linked; read 14 September 2026.
  7. M. Fowler, CircuitBreaker, bliki, 6 March 2014; the pattern is credited there to M. Nygard. martinfowler.com/bliki/CircuitBreaker.html, read 14 September 2026.
  8. R. Fielding, M. Nottingham, J. Reschke (eds), HTTP Semantics, RFC 9110 (STD 97), June 2022, §10.2.3 (Retry-After) and §15.6.4 (503). rfc-editor.org/rfc/rfc9110.txt, read 14 September 2026.
  9. M. Nottingham, R. Fielding, Additional HTTP Status Codes, RFC 6585, April 2012, §4 (429 Too Many Requests). rfc-editor.org/rfc/rfc6585.txt, read 14 September 2026.
  10. The Delegation Record, Paper 007 — A switch attests what it decided, not that it decided well: the condition for a score's soundness, and the reading of frontierscore.ai's public pages of 12 September 2026.