Testing invoice extraction accuracy: Build answer key from invoices using verified values for each field.; Score field-level accuracy and exact-match rate separately for reliable results.; Mark incorrect values, missing fields, or ambiguous data—no partial credit allowed.
Image: AI Tool Review Desk

Task Testing

Testing extraction from invoices with known answers

An invoice-extraction test needs a verified answer key.

Build the answer key from each invoice, then run the same file through extraction and compare every returned field with its verified value. Score whether the value is correct and attached to the correct field or line-item row, not whether the output merely looks plausible.

Report field-level accuracy and exact-match rate separately. The first shows which fields are often wrong; the second shows how many invoices have every answerable field correct.

End-to-End Invoice Extraction Testing Workflow

  1. Build a verified answer key from each invoice
  2. Run the same file through extraction system
  3. Compare every field and line item against the answer key
  4. Score each fieldcorrect, missing, incorrect, ambiguous
  5. Calculate field-level accuracy and exact-match rate
  6. Report results by invoice type and overall performance

Construct a representative sample

Include invoices from different suppliers, digital PDFs, scans, multiple pages, credits, discounts and varied line-item structures. This reflects the range of invoice formats and quality, including scanned documents and digital PDFs.

Choose fields the workflow needs and record the expected value for each invoice: invoice number, invoice date, vendor name and address, due date, amount due, purchase-order reference and payment terms. For line items, record quantity, description and unit price; include subtotal, tax amount and grand total where they appear.

Review each invoice against its source and record one expected value per field, rather than copying the extraction output. Keep an invoice identifier beside each field name and value; for line items, keep each row’s quantity, description and unit price together.

Mark values absent from the invoice as absent, and mark genuinely ambiguous values as ambiguous rather than guessing. Use documents you are authorised to process and handle sensitive data under your organisation’s rules.

Invoice Extraction QA Checklist

  • Include diverse invoice typesdigital PDFs, scans, multi-page, credits, discounts
  • Record expected values for key fieldsinvoice number, date, vendor, due date, amount due
  • For line itemscapture quantity, description, unit price; subtotal, tax, grand total if present
  • Mark absent or ambiguous values—do not guess
  • Use authorised documents and comply with ATO data handling rules

Score business errors

Run the same authorised files and fixed field list through each candidate, then compare the output with the answer key field by field. Apply the same scoring rules to every run so the results can be compared fairly.

Mark each field correct, missing, incorrect or ambiguous. A value in the wrong field or line-item row is incorrect, even if the value itself appears elsewhere; do not award partial credit.

Calculate field-level accuracy as correct field checks divided by all answerable field checks, multiplied by 100. Exclude genuinely ambiguous source values from that calculation and report them separately; count an unexpected value for a field marked absent as incorrect.

Calculate exact-match rate as invoices with every answerable field correct divided by invoices with answerable keys, multiplied by 100. One wrong or missing value makes an invoice fail this measure, even when its other fields are correct.

For example, if the answer key’s invoice date appears in the output’s invoice-number field, count the invoice-number field as incorrect and the missing date as missing. If line-item values match but the grand total does not, score those fields separately and record a consistency-check failure.

Check whether line items, subtotal, tax and grand total reconcile, but do not assume a matching sum proves every line is correct. Count errors that would block payment or create a duplicate separately from cosmetic formatting differences.

Choose which fields are critical and what score is acceptable for your workflow before testing, then apply that rule consistently. Report results by invoice type as well as overall, so a single average does not conceal field-specific errors.

Scoring Rules for Field-Level Errors

Correct
Value matches verified source and is in the right field/row
Missing
Expected value not returned by extraction system
Incorrect
Value is wrong, or placed in wrong field/row—even if it appears elsewhere
Ambiguous
Source document does not allow clear interpretation; do not infer

More from Task Testing

Task Testing

Comparing table extraction from different document layouts

A table can be transcribed word for word yet lose the relationship between a row label and its amount.

Task Testing

Measuring correction effort after automated extraction

The cost of document extraction includes work after the model responds.

Task Testing

AI image and design tools

Choose an AI image or design workflow by commercial use, text accuracy, editability and the file your team needs to deliver.

Task Testing

AI output quality testing

Set acceptance rules, check AI outputs against approved sources and compare the work needed to reach a usable result.