
Task Testing
AI output quality testing
Set acceptance rules, check AI outputs against approved sources and compare the work needed to reach a usable result.
AI output quality testing asks whether a result is good enough for a defined job under conditions resembling its intended use. Set the acceptance rule before reviewing outputs, check each result against it and retain failures. A fluent answer fails if a material fact or qualification is wrong.
Key Australian Regulatory References for AI Output Testing
- ATO AI Transparency Statement
- Guides responsible use of AI in tax and superannuation services.
- Digital Transformation Agency – AI Assurance Framework
- Provides step-by-step guidance for public sector AI evaluation.
- ANAO Audit on AI Governance at ATO
- Highlights risks in AI decision-making and oversight requirements.
- Australia Experiences – AI Disclosure Guidelines
- Recommends transparency in AI-generated content for businesses.
- AIGovernance – Australia Jurisdiction Overview
- Outlines legal and ethical considerations for AI use in Australian organisations.
Define the result you need
Name the task, the material the tool may use, the required output and who approves it. For a fictional service enquiry, an acceptable draft might answer from an approved information sheet, identify a missing detail and make no unsupported commitment. Those checks beat asking whether it “sounds good”.
Include ordinary and difficult cases. An incomplete request shows whether the output handles uncertainty; conflicting source material shows whether it guesses or leaves the issue for a person. Use authorised inputs. Fictional cases can test the method but cannot establish performance on restricted records.
Use separate quality gates
Start with conditions every output must meet. Then assess qualities that allow reasonable variation.
| Gate | Question for the reviewer | Possible decision |
|---|---|---|
| Required content | Are the necessary facts, actions and qualifications present? | Accept or record an omission |
| Accuracy | Does it meet the task-specific accuracy standard? See S047-P02-S01 for claim-by-claim checks. | Accept, correct or reject |
| Boundary | Has the output avoided an unsupported promise or unauthorised disclosure? | Stop or escalate if breached |
| Usefulness | Can the next person use it without substantial reconstruction? | Record the repair needed |
| Style | Is it clear and suitable for the audience? | Edit without changing supported meaning |
A total score can conceal a serious defect. Record critical conditions as pass or fail alongside any graded judgement about clarity or tone. Do not let a high style score offset a failed critical gate. Define the rule before reviewers see competing outputs.
Where a check can be scored consistently, structure it so the decision can be recorded or automated.
Compare like with like
Give candidates the same substantive task, source version and acceptance rules where their interfaces permit. Record differences in file preparation, prompts and available context. Keep first results and later attempts so extra prompting remains visible.
Hide provider names from reviewers where practical and vary display order. Apply the same acceptance rules across reviews; see S047-P02-S01 for claim-by-claim accuracy checks.
Report results by case type and consequence. Record the review effort needed to reach an accepted result. The decision concerns the route to an approved output, not just the first response.
Calibrate the scoring
Prefer a task-specific measure over a generic score. A generic score may reward fluent wording without showing whether the output meets the required standard.
Automated scoring can make structured checks more consistent, but it is not a complete verdict. Pair metric results with human review so reviewers can judge whether the measure answers the right question.
Compare automated scores with human decisions and investigate disagreements. Use those findings to adjust the scoring rule, then check whether reviewers apply it consistently; an unexplained score is not a sound basis for accepting an output.
Run evaluations through development
Treat evaluation as an ongoing part of development. Run scoped tests early and at later stages, comparing each result with the success criteria set for the task.
When a result changes, identify whether the task, inputs, scoring method or system configuration also changed. Clear conditions help distinguish a genuine improvement from a difference in how the test was run.
Keep the finding bounded
A test on one task, source set and account supports a conclusion about those conditions. It does not establish general accuracy, repeated-run reliability or suitability for another audience. Record the date, plan or configuration, source version, prompt and acceptance rule so later results can be compared honestly.
State which output types can be used with review, which need repair and which must be escalated. Keep failure cases and revisit affected cases when the task, source material or product changes.
In this guide
- Building a rubric before reading competing outputsTurn a defined task into observable quality criteria, pass and fail anchors, and a decision rule before reviewing AI outputs.
- Separating factual errors from style preferencesUse two review passes to identify wrong or unsupported AI claims before making tone and wording edits.
- Comparing outputs with a blinded reviewPrepare anonymised AI outputs, vary their order, score against a fixed rubric and resolve reviewer disagreements before revealing providers.
- Recording failure cases instead of publishing only successesKeep original AI failures, their test conditions and the work needed to repair them so a quality report reflects the whole evaluation.



