In brief
- Randomize labels and presentation order.
- Define the rubric before seeing outputs.
- Resolve disagreements with evidence, not majority alone.
Why identity changes judgment
Reviewers carry expectations about brands, model families, writing styles, and price. Those expectations can influence ratings even when outputs are otherwise identical in usefulness.
Create the blind packet
Normalize formatting that reveals identity without altering substantive content. Randomize output order per task, assign opaque labels, and preserve raw outputs for audit.
- Same prompt and context
- Opaque candidate labels
- Randomized order
- Fixed rubric
- Independent first-pass review
Score dimensions separately
Ask reviewers to score correctness, completeness, clarity, constraint adherence, and risk. A single preference vote makes disagreements difficult to diagnose.
Treat disagreement as data
High disagreement may reveal an ambiguous task, weak rubric, or a real tradeoff between approaches. Review evidence together and record why the final judgment changed or remained split.
Frequently asked
Questions, answered plainly.
Can writing style reveal the model?+
Sometimes. Normalize obvious wrappers and brand references, but do not rewrite the substance. Perfect blinding is less important than reducing avoidable cues.
How many judges are needed?+
Use enough independent review to expose disagreement for the stakes involved. Two reviewers plus adjudication can be useful for an internal comparison.
Should price be hidden too?+
Judge output quality first, then combine quality with cost and latency in a separate decision stage.
Sources and next paths
