
Task Testing
Part of Publishing evidence-based AI reviews
Explaining exactly what was tested in an AI review
Show readers the cases, conditions, attempts, results and limits behind an AI review’s test-based claims.
A test-based review should state which task, inputs, account, settings, attempts and acceptance rule produced its conclusion. It should also state what those conditions cannot establish. “We tested the tool” leaves readers guessing.
State the test boundary
Put a compact method description near the verdict. Identify the product and feature, plan or supplied access, test period, relevant settings and connected sources.
Include a model or version identifier only if the tested interface exposed one reliably. If it was unavailable, say so. Record any human preparation that changed the input.
A hypothetical review might ask a tool to draft replies from an approved, fictional information sheet. The method description would identify the supplied questions, the rule for an acceptable reply and who judged the outputs. This describes a possible test, not a result.
Show all the attempts
State how many cases and first attempts were examined, how cases were selected and whether any were excluded. Include errors, refusals and outputs that needed repair. A screenshot of the best response cannot represent the whole set. If routine, missing-information and exception cases were included, report their results separately because they ask different things of the tool.
Reader question / Detail to provide
- What entered the tool?
- Prompt, source version, file type and preparation
- What counted as acceptable?
- Required facts, prohibited claims and permitted edits
- What happened first?
- Original output, error or refusal for each case
- What made it usable?
- Follow-up prompts, corrections and reviewer effort
- What remains unknown?
- Untested inputs, plans, languages or workflows
State whether examples were fictional, cleared for use or drawn from live work. Results on fictional material do not establish performance on restricted customer records.
Match each verb to its evidence
Use “observed” for a saved run, “the provider documents” for a feature checked only in product material and “may suit” for a reasoned fit judgement. If a supplier configured the account or supplied the demonstration, explain that involvement. Without access to the underlying output or a repeatable run, describe a demonstration rather than an independent test.
A conclusion from a small or selected set can describe those cases and the human checks they required. It cannot establish a general accuracy rate. Identify a consequential failure even when other examples looked good.



