
Task Testing
Part of AI research tools
Recording unsupported claims in a research tool evaluation
Keep a reproducible record of unsupported AI research claims, source gaps, consequences and corrections.
Record each unsupported claim as a specific finding. Note the output’s exact wording, the evidence offered, what the reviewer inspected and the consequence of using the claim.
Define the finding
In an evaluation, an unsupported claim is one the appropriate evidence available for that run does not establish. It may lack a citation, point to an unrelated passage or add a condition absent from the source. The verdict does not prove the claim false everywhere.
Keep separate verdicts for a contradicted claim, where evidence says something incompatible; an uncheckable claim, where the relevant material cannot be accessed or resolved; and a partly supported claim, where only part is established. Do not count a tool’s honest statement that a document set lacks an answer as an unsupported assertion.
Keep a reproducible register
Use one row per material claim:
Field / What to record
- Run context
- Product, plan, date, source mode, question and prompt version
- Claim
- Exact wording and location in the output
- Offered evidence
- Citation, document name or absence of evidence
- Check
- Passage inspected, source version and the gap or conflict
- Verdict
- Unsupported, partly supported, contradicted or uncheckable
- Impact
- Consequence if the claim were used uncorrected
- Action
- Remove, qualify, replace evidence, seek an owner or rerun
Suppose a hypothetical answer says an internal policy “guarantees a reply within one day”, while the policy says the team “aims to reply within one business day”. Record the changed strength and timing, not just “citation wrong”. If the tool used an old policy, log that source-version problem separately.
Retain the original response even when a later prompt produces a better one. Log manual edits and follow-up prompts needed to reach an acceptable answer, so the evaluation reflects the work required.
Use the register to decide
Group findings by cause and consequence. A made-up date, a missing eligibility condition and an inaccessible citation need different repairs. Count affected claims and outputs, but describe consequential cases individually; one unsupported promise may matter more than several background errors.
Comparing candidates on the same question, sources and acceptance rule belongs to research-tool selection (S047-P04). This register stays with the evidence findings for each configuration.
At the end of a trial, state which claim types review can repair, which need a different source or instruction, and which make the proposed workflow unsuitable. Assign unresolved findings an owner and revisit affected cases after a source or product change.



