
Task Testing
Part of AI writing tools compared
Testing factual accuracy in AI writing outputs
Build an answer key, verify each claim and qualification, and report consequential factual errors in AI-generated writing.
Test factual accuracy by checking each claim in an AI-written draft against an approved source, including claims the brief never requested. Do this before judging style. A fluent sentence can still give the wrong date, omit a condition or turn an uncertain statement into a promise.
Prepare a source-backed answer key
Choose a writing task with material you are authorised to use. Give each candidate the same source text and brief where its workflow permits. Before generating, list the facts the finished piece must contain, the qualifications that belong with them and anything the writer must leave unanswered.
For a hypothetical event notice, the answer key might record the date, venue, booking requirement and whether attendance is subject to confirmation. Keep the key separate from the generated output. If the source is ambiguous or out of date, resolve that with its owner before grading the tool. The tool cannot supply an authoritative answer that the source does not contain.
Break the output into checkable claims
Read the draft sentence by sentence. Split compound sentences when they contain several assertions. For each claim, record its location, the supporting source passage and a verdict:
| Verdict | Use when |
|---|---|
| Supported | The source establishes the claim with the same scope and qualification. |
| Contradicted | The claim conflicts with the source. |
| Unsupported | No approved or independently verified source establishes it. |
| Unclear | The wording or source is ambiguous and needs a human decision. |
Check numbers with their units and denominators, dates with their time period, and names with the correct role. Test words such as “always”, “free” and “guaranteed” against their conditions. An accurate number attached to the wrong product or period is still an error. A source reference generated by the tool is only a lead: open it and check that it supports the precise claim.
Look for omissions as well as additions. If the brief requires an exception, eligibility condition or deadline, leaving it out may mislead even when every remaining sentence is true. Record it as a missing required fact.
Factual Verdicts: Claim Evaluation Criteria
- Supported
- Source establishes the claim with same scope and qualification
- Contradicted
- Claim conflicts with the source
- Unsupported
- No approved or independently verified source establishes it
- Unclear
- Wording or source is ambiguous; requires human decision
Report errors by consequence
Count claims by verdict, but keep the error list visible. One unsupported eligibility promise may matter more than several minor wording problems. Record whether a person can repair the draft, whether the source needs updating or whether the output should be rejected.
For comparisons, apply the same claim rules to every attempt, not just the best response. Report the task, source version, plan or feature used, and the claim types that failed. Results from a controlled brief describe those conditions; they do not prove a product’s general accuracy. Keep final approval with a person authorised to check the copy.
Error Reporting by Consequence
- Critical Errors
- Unsupportable promises, missing eligibility conditions
- Minor Errors
- Inconsistent wording, minor phrasing issues
- Action Required
- Draft repair needed, source update, or rejection



