
Task Testing
Part of AI output quality testing
Comparing outputs with a blinded review
Prepare anonymised AI outputs, vary their order, score against a fixed rubric and resolve reviewer disagreements before revealing providers.
A blinded review lets reviewers judge competing outputs without knowing which provider produced each. Fix the task and scoring rule first, conceal identifying cues where practical, vary display order, and give reviewers the material needed to check facts. Blinding may reduce one influence on judgement; it cannot make different workflows identical or guarantee anonymity.
Prepare the comparison set
Use outputs produced for the same task and approved inputs. Keep substantive instructions and acceptance rules equivalent where each interface allows, and record differences in format or setup. Decide whether reviewers will see first attempts or results after the same allowed number of revisions. Do not compare one candidate's repaired response with another's first attempt without saying so.
Save an unchanged copy of each output. Assign neutral labels such as A and B, and keep the provider key outside the review sheet. Remove interface decorations and incidental signatures.
Preserve content needed to judge the work, including qualifications, refusals and formats, even if it hints at the source. Record any case where the provider remains recognisable.
Give reviewers a complete pack
Each reviewer needs the task brief, approved source or answer key, fixed rubric and labelled outputs. Vary output order across reviewers or cases. Ask reviewers to decide each critical condition before recording an overall preference, so pleasing style does not conceal a factual failure.
Review field / What to record
- Critical checks
- Pass, fail or cannot judge, with the relevant passage
- Preference
- Which acceptable output better serves the task, or a tie
- Repair
- Changes needed before approval
- Confidence
- Missing evidence or ambiguity limiting the verdict
Reviewers may find that neither output passes. Forcing a winner would turn two unacceptable results into a misleading recommendation.
Reviewer Assessment Criteria
- Critical Checks
- Pass, fail or cannot judge — with relevant passage cited
- Preference
- Which acceptable output better serves the task, or a tie
- Repair
- Changes needed before approval
- Confidence
- Missing evidence or ambiguity limiting the verdict
Resolve and then reveal
Compare reviewers' reasons while labels remain blind. If one finds a wrong figure and another prefers that output's clarity, check the figure against the source first. Use the audience brief and rubric anchors for a genuine style disagreement. Keep the initial reviews and the resolution.
Reveal providers after recording quality verdicts. Then add operational context concealed during review: setup, plan, manual preparation and correction effort. Report ties, failures and imperfect blinding, and limit the conclusion to the tested task and conditions.
Key Outcomes from Blinded Reviews
- Prevents bias from provider identity
- Reduces influence of brand or interface familiarity
- Does not guarantee anonymity
- May still reveal source through content style or formatting
- Supports fair comparison
- Ensures evaluations based on merit, not origin
- Highlights need for transparency
- Report ties, failures, and imperfect blinding



