
Task Testing
Part of AI tool reliability
Measuring performance with longer input documents
Check whether an AI tool accepts longer documents and finds the right evidence, using a controlled length comparison and separate format checks.
To assess an AI tool on longer documents, check whether it accepts the input and retains the evidence needed as length increases. Fix the question, answer key, file format and question wording for a controlled length comparison; otherwise a layout change may look like a length effect. A stated context or page limit describes capacity only, not whether the tool finds every required detail.
Build a controlled length ladder
Choose an authorised document family, such as fictional supplier procedures. Create short, medium and long versions that preserve the same facts needed for each question, adding realistic surrounding material such as definitions, exceptions and appendices.
Prepare a human answer key with the exact passages needed. Ask about a detail near the beginning, one in the middle and one near the end. Include a question that requires two separated passages and one the document cannot answer; the latter checks whether the tool leaves an unsupported point open.
Check format and layout separately.
Record acceptance and answer coverage
For each length, record the upload outcome, any displayed limit or warning, time to a usable response, and answer-key points supported by the source. If the interface gives no reliable indication of what it read, an accepted upload does not prove the whole file reached the answer process.
Result / What to record
- Input accepted
- Format, size, pages and any rejection or warning
- Answer supported
- The passage behind each material answer point
- Answer incomplete
- Missing condition or missed passage
- Unsupported addition
- A claim the document does not establish
- Workaround and effort
- Splitting, conversion, re-uploading, follow-up and review time
Inspect the original passages beside the answer. A plausible sentence may attach a condition to the wrong supplier or period. Keep that error separate from file rejection.
Key Metrics for Long-Document AI Performance
- Document length tested
- Short, medium, long (exact page counts or word counts)
- Upload success rate
- Percentage of files accepted without error
- Answer accuracy by section
- Start: X%, Middle: X%, End: X%
- Unsupported claims detected
- Number of incorrect or unsupported assertions
- Average response time
- Time to usable response (seconds)
Compare the route fairly
Holding the conditions above constant, record any conversion, splitting or manually prepared summary, because each changes the task. Splitting may separate a rule from its exception; check the joined answer against the complete original and label the result.
Do not infer a universal length where accuracy falls. Report the document structure, questions and lengths examined, where coverage weakened, and whether a person could repair the result before the deadline. Approve only the file types, lengths and question types that the proposed check supports.



