Verify AI facts before approval: Check every claim against approved sources, not just requested ones; Split sentences to test each fact individually with source proof; Record errors by impact: missing rules or false promises matter most
Image: AI Tool Review Desk

Task Testing

Part of AI writing tools compared

Testing factual accuracy in AI writing outputs

Build an answer key, verify each claim and qualification, and report consequential factual errors in AI-generated writing.

Test factual accuracy by checking each claim in an AI-written draft against an approved source, including claims the brief never requested. Do this before judging style. A fluent sentence can still give the wrong date, omit a condition or turn an uncertain statement into a promise.

Prepare a source-backed answer key

Choose a writing task with material you are authorised to use. Give each candidate the same source text and brief where its workflow permits. Before generating, list the facts the finished piece must contain, the qualifications that belong with them and anything the writer must leave unanswered.

For a hypothetical event notice, the answer key might record the date, venue, booking requirement and whether attendance is subject to confirmation. Keep the key separate from the generated output. If the source is ambiguous or out of date, resolve that with its owner before grading the tool. The tool cannot supply an authoritative answer that the source does not contain.

Break the output into checkable claims

Read the draft sentence by sentence. Split compound sentences when they contain several assertions. For each claim, record its location, the supporting source passage and a verdict:

VerdictUse when
SupportedThe source establishes the claim with the same scope and qualification.
ContradictedThe claim conflicts with the source.
UnsupportedNo approved or independently verified source establishes it.
UnclearThe wording or source is ambiguous and needs a human decision.

Check numbers with their units and denominators, dates with their time period, and names with the correct role. Test words such as “always”, “free” and “guaranteed” against their conditions. An accurate number attached to the wrong product or period is still an error. A source reference generated by the tool is only a lead: open it and check that it supports the precise claim.

Look for omissions as well as additions. If the brief requires an exception, eligibility condition or deadline, leaving it out may mislead even when every remaining sentence is true. Record it as a missing required fact.

Factual Verdicts: Claim Evaluation Criteria

Supported
Source establishes the claim with same scope and qualification
Contradicted
Claim conflicts with the source
Unsupported
No approved or independently verified source establishes it
Unclear
Wording or source is ambiguous; requires human decision

Report errors by consequence

Count claims by verdict, but keep the error list visible. One unsupported eligibility promise may matter more than several minor wording problems. Record whether a person can repair the draft, whether the source needs updating or whether the output should be rejected.

For comparisons, apply the same claim rules to every attempt, not just the best response. Report the task, source version, plan or feature used, and the claim types that failed. Results from a controlled brief describe those conditions; they do not prove a product’s general accuracy. Keep final approval with a person authorised to check the copy.

Error Reporting by Consequence

Critical Errors
Unsupportable promises, missing eligibility conditions
Minor Errors
Inconsistent wording, minor phrasing issues
Action Required
Draft repair needed, source update, or rejection

More from Task Testing

Task Testing

AI image and design tools

Choose an AI image or design workflow by commercial use, text accuracy, editability and the file your team needs to deliver.

Task Testing

AI output quality testing

Set acceptance rules, check AI outputs against approved sources and compare the work needed to reach a usable result.