Digital One Foundation The Measured Endpoint
Paper 001 · Method

A judge is a party to the score.

Free-form model output is scored by another model, and models used as judges have shown measurable preferences, including for text like their own: where a judge's vendor also makes models it scores, the judge is a party to the score. A panel drawn from several providers dilutes that preference without removing it, and a rule that rejects the outlier among few judges — no panel size is published — protects at most against one member failing alone, not against anything two members share. These notes publish no rubric, prompt set or panel composition, so no score can be re-run from them.

§1Free-form output needs a judge, and the judge has a stake

A model asked a question with one correct answer can be scored by comparison. Asked to explain, plan or write, it produces text that no answer key matches, and scoring that text hour after hour, across model series from several providers, is not work people can do at the rate an endpoint changes. So a model scores it. That moves the question rather than answering it: the measurement is now partly a property of the model doing the measuring.

The literature is specific about how. Zheng and colleagues, examining models used as judges, name position bias, self-enhancement bias and verbosity bias — an LLM judge that "favors longer, verbose responses, even if they are not as clear, high-quality, or accurate as shorter alternatives" — alongside limited reasoning ability.[1] On self-enhancement their own result is inconclusive: "our study cannot determine whether the models exhibit a self-enhancement bias."[1] The firmer evidence came later. Panickssery, Bowman and Feng describe self-preference, "where an LLM evaluator scores its own outputs higher than others’ while human annotators consider them of equal quality", and find its strength tracks a model's ability to recognise its own text.[2] Their setting is summarisation with three models, and they call their causal case evidence rather than proof.[2] Neither measures a current judge, so neither sizes the effect in one.

The premise, precisely

Where the vendor of a judge also makes models among those the judge scores, the judge is not outside the measurement. It is a party to the score, and a method that uses it has to say how that is contained, and where the containment stops.

What follows is the part of that method which answers the stake — a panel and a rule for rejecting its outliers — each with the limit that belongs to it. The design is an operator's, named at the end. Nothing on this page meets the condition this record has already named for a score's soundness to be answerable, and the last section says so before it is asked.

§2A panel from several providers dilutes a preference

The design describes a panel: judge models drawn from different providers, each given the same transcript and the same rubric, each returning integer quality and reasoning scores from 0 to 100, which an aggregator combines. As designed, no single vendor's judge decides an aggregate alone. There is published support for the shape. Verga and colleagues had "each evaluator model independently" score an output, pooled the scores, and found that in their settings a panel of smaller models from disjoint families outperformed a single large judge in agreement with human judgements and "exhibits less intra-model bias due to its composition of disjoint model families".[3]

Their bounds travel with the result: three evaluator settings, a three-member panel, pooling by max voting or by average, and no outlier rejection, with the choice of which models belong on a panel left "to future work", and the authors work for the maker of one of the three panel members.[3] The same paper reports what the premise predicts: "the highest positive delta for each individual model being scored occurs when it is judged by itself."[3]

So a panel does not remove a judge's preference. It dilutes it, and unevenly. If members come from providers whose models are also scored, a provider with a judge on the panel has its models scored by a panel containing at least one member of its own family, and a provider with none does not. Averaging shrinks that family's weight; it does not make the aggregate neutral, because neutrality is agreement with a reference, and nothing these notes could read names one. The design does not say whether a member scores its own vendor's models, or how such a verdict is marked.

Independence is assumed, not shown

Judges from different vendors are not thereby judges with independent errors. Models trained on overlapping public text can share a blind spot, and a panel that agrees has shown agreement, not correctness. These notes hold no measurement of how correlated the panel's errors are.

As designed, provider diversity also buys fault tolerance: when one provider is down, the panel is meant to lose a member rather than fall silent, and nothing in the public delayed feed (read in Paper 002) shows which happened in any given hour. That is where the instrument changes. A panel short of a member is a different estimator, so a score series running through an outage mixes two, and nothing a reader receives marks the window.

§3Rejecting the outlier among few judges is a policy

The design rejects outliers per verdict: a member whose scores disagree with the rest is excluded before aggregation, and the exclusion is recorded. The statistic, the threshold, the minimum number of members that must survive, and whether quality and reasoning are tested separately or together are not described.

The size of the panel decides what such a rule can mean. Take three, the smallest panel in which one member can be outvoted — an illustration, since no panel size is published. Among three, "outlier" is defined only by the other two, so rejecting the most distant member hands the aggregate to the two closest. That protects against one member failing alone: a truncated reply, a refusal read as a zero, a judge misreading the rubric. It protects against nothing two members share, and where two are wrong in the same direction it can discard the one that was right.

Statistics said as much long before models judged anything. The NIST/SEMATECH handbook warns that Z-scores "can be misleading (particularly for small sample sizes)", that "we typically do not want to simply delete the outlying observation", and that deletion is warranted only "if it can be determined that an outlying point is in fact erroneous".[4] The authors it cites recommend that modified Z-scores with an absolute value above 3.5 "be labeled as potential outliers" — labelled, not removed.[4] The section assumes roughly normal univariate data, which integer scores from a handful of judges are not, so it supports a caution and endorses no rule. In that illustration, with two members left — one provider down — there is no outlier to define at all, and the safeguard is off when the panel is weakest.

What makes the rule defensible is not the rejection but the flag. An excluded member stays in the record with rejected set, so whoever holds the record can ask afterwards, for as long as rows are retained (a period the design does not state), how often each member was overruled, and on whose models. What that record holds, and what it cannot show a reader, is the subject of Paper 002.

§4What a panel cannot show a reader

The Delegation Record's Paper 007 names what would make a score's soundness answerable, "a published method someone other than its author can re-run, or … outcomes measured on a named corpus",[5] and that condition is unmet here.

Neutrality is not measured. No agreement between the panel and a named human-labelled or keyed reference is described, so whether the aggregate is free of vendor preference is a measurement these notes do not have.

The panel is unnamed. Its composition is not published, so that its members come from different providers is the operator's statement and not something a reader can check. There is a stated reason to withhold it — a named judge is a target an endpoint could be tuned to please — and a shape already in this record that keeps the reason while allowing the check: a commitment by hash,[6] made to each composition when it takes effect and revealed after a stated delay. The design describes nothing of the kind.

On the provenance of this material

The architecture described here comes from FrontierScore, a Digital One property, at frontierscore.ai; it is the provenance of this method, not its subject. Its repository is not public. The design description these notes draw on is its builders' own, and no part of it is checkable today, which is why its scale figures are left out. Deliberately absent: prices, plans and tiers, availability, and deployment internals.

References

  1. L. Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023 Datasets and Benchmarks Track, arXiv:2306.05685v4, abstract and §3.3; judges studied are 2023 models. arxiv.org/abs/2306.05685, read 14 September 2026.
  2. A. Panickssery, S. R. Bowman, S. Feng, LLM Evaluators Recognize and Favor Their Own Generations, arXiv:2404.13076v1, abstract and limitations; summarisation tasks only. arxiv.org/abs/2404.13076, read 14 September 2026.
  3. P. Verga et al., Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models, arXiv:2404.18796v2, abstract, §2.2, §3.1 and §4.4; the panel had three members and used max voting or average pooling, with no outlier rejection. arxiv.org/abs/2404.18796, read 14 September 2026.
  4. NIST/SEMATECH, e-Handbook of Statistical Methods, §1.3.5.17 "Detection of Outliers"; the modified Z-score recommendation is Iglewicz and Hoaglin's, as summarised there. itl.nist.gov/div898/handbook/eda/section3/eda35h.htm, read 14 September 2026.
  5. The Delegation Record, Paper 007, A switch attests what it decided, not that it decided well; its provenance note records that frontierscore.ai's public pages linked no methodology, read 12 September 2026.
  6. The Project Handshake, Paper 002, Two trees.