
Task Testing
Part of AI output quality testing
Separating factual errors from style preferences
Use two review passes to identify wrong or unsupported AI claims before making tone and wording edits.
Check what an AI output claims before debating how it sounds. A factual or meaning error changes what the reader is told about the task; a style preference concerns expression while the supported meaning stays intact. Record the two separately, so fluent wording cannot offset a material defect.
Test the meaning first
Read the output beside the approved source and task brief. For each material statement, ask who or what it concerns, what happened or is allowed, when, and under which conditions.
A changed number, date, actor or qualification belongs in the factual or meaning review, as does an added promise the source does not support.
Suppose a fictional notice says a discount applies to online orders placed before Friday, while stock lasts. “The discount applies to all orders” removes the channel and availability limits. “The online discount runs until Friday, subject to stock” may be clearer, provided the deadline and condition still mean the same thing. A preference for warmer wording is a separate judgement.
Label the finding precisely
| Finding | What the reviewer has established | Next action |
|---|---|---|
| Contradicted | The approved source says something incompatible | Correct or reject the claim |
| Unsupported | The available approved material does not establish the claim | Find evidence, qualify it or remove it |
| Missing requirement | A fact or condition required by the brief is absent | Restore it beside the relevant claim |
| Cannot judge | The source is missing or its meaning cannot be resolved | Ask the source owner or hold the verdict |
| Style preference | Supported meaning remains intact, but presentation could improve | Edit and check the meaning again |
“Unsupported” does not mean “proved false”. An inaccessible source can prevent a verdict. Conversely, an output can mislead through an omitted condition even when its remaining words are true.
Factual Errors vs Style Preferences in AI Output Review
- Factual Error
- Contradicted, Unsupported, Missing requirement, Cannot judge
- Style Preference
- Supported meaning intact, but presentation could improve
Review in two passes
First, mark material claims and required omissions against approved material. Record the exact output words and the source passage or evidence gap behind each finding. Ask the information owner to resolve consequential ambiguity.
Then judge clarity, length, tone and format against the audience brief. Keep a style edit only if the supported meaning survives it.
When reviewers disagree, identify the claim or criterion each is judging. One may be applying a factual rule while another is expressing a tone preference. Settle the factual question from the source; settle the preference against the brief. Check the final edited version again, since shortening can remove the qualification that made it accurate.
Two-Pass Review Process for AI Outputs
- Verify factual accuracyCheck material claims against approved source and task brief. Record exact output words and source evidence or gap.
- Resolve ambiguityAsk the information owner to clarify consequential ambiguities.
- Assess style and clarityEvaluate tone, length, format and readability against audience brief—only edit if meaning remains unchanged.
- Final validationRecheck edited version—shortening may remove critical qualifications.



