
Task Testing
Part of AI output quality testing
Building a rubric before reading competing outputs
Turn a defined task into observable quality criteria, pass and fail anchors, and a decision rule before reviewing AI outputs.
Write the scoring rubric before opening the competing outputs. Decide what a usable result must contain, which errors make it unacceptable and where reviewers may exercise judgement. This keeps an attractive response from changing the standard after the fact.
Start with the task
Take one defined output, such as a draft reply to a fictional service enquiry using an approved information sheet. Write down the reader's need and the next person's job. “Professional”, “helpful” and “high quality” are too loose alone. Use observable checks: answers the question, keeps an eligibility condition, identifies an unanswered point and makes no invented promise.
Separate requirements from preferences. A missing condition may make the draft unusable. A formal greeting may be a style choice unless the organisation requires it. If a reviewer cannot explain how a criterion serves the task, rewrite or remove it.
Give each criterion an anchor
State what a pass and fail look like. Use the approved source to settle factual points; a generated response is not its own answer key.
Pass anchor
- Required condition
- The condition appears beside the offer it limits
- Unanswered detail
- The draft asks for the detail or leaves the point open
- Reader action
- The next step is clear and authorised
- Clarity
- The reader can follow the response without substantial repair
Fail anchor
- Required condition
- The offer appears without the condition
- Unanswered detail
- It supplies a plausible but unsupported detail
- Reader action
- It promises an unapproved action
- Clarity
- Wording obscures the decision or next step
These are sample anchors, not a universal scorecard. Replace them with criteria for the actual task. Keep the rubric short enough to apply every item to every output.
Set and check the decision rule
Mark critical criteria as pass or fail. For qualities such as clarity, use an anchored scale only if it distinguishes useful differences. State whether one failed critical criterion rejects the output regardless of its other scores. Include “cannot judge” when the approved source is missing or ambiguous; do not force a factual verdict from an incomplete answer key.
Give reviewers a field for the passage behind each judgement and the repair needed. If several people will review, try the rubric on separate calibration examples before the comparison set is opened. Discuss disagreements and clarify ambiguous wording. Then fix the rubric version for the comparison round; record later changes as a new version so scores under different rules are not mixed.
The finished rubric should show another reviewer both the rule and why an output met or missed it. It still requires someone with the relevant subject knowledge to settle consequential facts.



