How AI reviews must be tested: Test method must state inputs, settings, and acceptance rules used.; Each attempt must be tracked, including errors and repairs made.; Conclusions need to match evidence type: observed, documented or judged.
Image: AI Tool Review Desk

Task Testing

Part of Publishing evidence-based AI reviews

Explaining exactly what was tested in an AI review

Show readers the cases, conditions, attempts, results and limits behind an AI review’s test-based claims.

A test-based review should state which task, inputs, account, settings, attempts and acceptance rule produced its conclusion. It should also state what those conditions cannot establish. “We tested the tool” leaves readers guessing.

State the test boundary

Put a compact method description near the verdict. Identify the product and feature, plan or supplied access, test period, relevant settings and connected sources.

Include a model or version identifier only if the tested interface exposed one reliably. If it was unavailable, say so. Record any human preparation that changed the input.

A hypothetical review might ask a tool to draft replies from an approved, fictional information sheet. The method description would identify the supplied questions, the rule for an acceptable reply and who judged the outputs. This describes a possible test, not a result.

Show all the attempts

State how many cases and first attempts were examined, how cases were selected and whether any were excluded. Include errors, refusals and outputs that needed repair. A screenshot of the best response cannot represent the whole set. If routine, missing-information and exception cases were included, report their results separately because they ask different things of the tool.

Reader question / Detail to provide

What entered the tool?
Prompt, source version, file type and preparation
What counted as acceptable?
Required facts, prohibited claims and permitted edits
What happened first?
Original output, error or refusal for each case
What made it usable?
Follow-up prompts, corrections and reviewer effort
What remains unknown?
Untested inputs, plans, languages or workflows

State whether examples were fictional, cleared for use or drawn from live work. Results on fictional material do not establish performance on restricted customer records.

Match each verb to its evidence

Use “observed” for a saved run, “the provider documents” for a feature checked only in product material and “may suit” for a reasoned fit judgement. If a supplier configured the account or supplied the demonstration, explain that involvement. Without access to the underlying output or a repeatable run, describe a demonstration rather than an independent test.

A conclusion from a small or selected set can describe those cases and the human checks they required. It cannot establish a general accuracy rate. Identify a consequential failure even when other examples looked good.

More from Task Testing

Task Testing

AI image and design tools

Choose an AI image or design workflow by commercial use, text accuracy, editability and the file your team needs to deliver.