The extraction endpoint is healthy. Yet finance staff spend longer fixing invoices than last week. The dashboard measures availability, while the business problem lives in document quality and downstream acceptance.
Start with an operator’s question: which invoices failed, where did they fail, and what should we do next?
Follow one invoice through the system
Use a stable workflow ID from ingestion to accounting-system confirmation, and a separate attempt ID for each retry.
| Stage | Record | Failure it explains |
|---|---|---|
| Ingestion | Document hash, pages, source | Duplicate or incomplete upload |
| Extraction | Model and prompt revision, field outputs | Regression after release |
| Validation | Failed rules and source locations | Totals that do not reconcile |
| Review | Corrected fields and review time | Automation creating more work |
| Export | Action reference, response, retries | Timeout or duplicate write |
| Acceptance | External record ID and final state | Output never used |
Store document references rather than copying every document into telemetry. Give investigators access through application permission checks. Decide which fields to redact and how long to retain diagnostic data before enabling verbose logs.
OpenTelemetry’s generative AI conventions offer a starting vocabulary. Keep business outcome fields alongside model telemetry; a model span cannot tell you whether finance accepted an invoice.
Separate three failures
Extraction failure: the system reads the wrong amount. Inspect the document, OCR result, prompt revision, and field-level evaluation.
Business-rule failure: extracted fields are correct, but the total does not match the line items. Route to review; asking the model to invent a matching value hides the discrepancy.
Integration failure: the approved invoice cannot be written to the accounting system. Inspect permissions, schema mapping, external status, and retries. Changing a prompt will not fix an expired credential.
Each category needs an owner and recovery procedure. Otherwise categorization produces another chart without changing what happens next.
Choose useful denominators
Track cost per accepted invoice, including OCR, inference, retries, evaluation, review labor, and allocated operating costs.
Show review rate beside correction rate and rejected-export rate. A falling review rate can look positive while incorrect records reach the destination. Separate routine invoices from unfamiliar layouts so easy traffic does not hide weak performance on difficult documents.
Show active processing time and total elapsed time including human queues. A fast model does not help a document waiting two days for approval.
Build alerts around decisions
Start with a growing unresolved-export queue, sustained increases in correction rate, and releases failing a known regression set. Define a comparison window and minimum sample size so one unusual document does not create constant alarms.
Choose thresholds with the operator. An illustrative pilot might escalate unresolved external writes older than one business day; that is a proposed service rule, not a universal standard. Time-sensitive work needs its own limit.
Make releases explainable
Before changing a prompt, model, or retrieval configuration, run the same versioned evaluation set. Include historical failures that forced corrections. Compare field accuracy, exception routing, cost, and downstream acceptance.
Link configuration versions to results. When quality drops, decide whether to roll back, narrow supported document types, or change review policy. One aggregate score cannot make that decision.
The handover should include a sample trace, dashboard definitions, alert routing, retention settings, and a worked incident investigation. Discuss an operational review if your system produces outputs but cannot explain what happens afterward.