All field notes
Research 7 minute read

How to blind-judge AI model outputs

Reduce brand bias in model comparisons with randomized output identity, a fixed rubric, and disagreement review.

In brief

  • Randomize labels and presentation order.
  • Define the rubric before seeing outputs.
  • Resolve disagreements with evidence, not majority alone.
01

Why identity changes judgment

Reviewers carry expectations about brands, model families, writing styles, and price. Those expectations can influence ratings even when outputs are otherwise identical in usefulness.

02

Create the blind packet

Normalize formatting that reveals identity without altering substantive content. Randomize output order per task, assign opaque labels, and preserve raw outputs for audit.

  • Same prompt and context
  • Opaque candidate labels
  • Randomized order
  • Fixed rubric
  • Independent first-pass review
03

Score dimensions separately

Ask reviewers to score correctness, completeness, clarity, constraint adherence, and risk. A single preference vote makes disagreements difficult to diagnose.

04

Treat disagreement as data

High disagreement may reveal an ambiguous task, weak rubric, or a real tradeoff between approaches. Review evidence together and record why the final judgment changed or remained split.

Frequently asked

Questions, answered plainly.

Can writing style reveal the model?+

Sometimes. Normalize obvious wrappers and brand references, but do not rewrite the substance. Perfect blinding is less important than reducing avoidable cues.

How many judges are needed?+

Use enough independent review to expose disagreement for the stakes involved. Two reviewers plus adjudication can be useful for an internal comparison.

Should price be hidden too?+

Judge output quality first, then combine quality with cost and latency in a separate decision stage.

Sources and next paths

Check the living surfaces.

Put it to work

One interface. Your choice of model.

Run the same task through live GPT, Claude, Gemini, and Xpersona models without rebuilding your client.

Try Xpersona chat
How to blind-judge AI model outputs | Xpersona Blog