ModelShifts
← Blog & insightsAI Operations

AI Observability: Diagnose Failed Workflows Before Adding More Dashboards

Trace document AI from ingestion to acceptance, separate model errors from integration failures, and build alerts operators can act on.

ModelShifts3 min readUpdated 11/09/2026
Editorial diagram: Request to Retrieval to Decision to Outcome

The extraction endpoint is healthy. Yet finance staff spend longer fixing invoices than last week. The dashboard measures availability, while the business problem lives in document quality and downstream acceptance.

Start with an operator’s question: which invoices failed, where did they fail, and what should we do next?

Follow one invoice through the system

Use a stable workflow ID from ingestion to accounting-system confirmation, and a separate attempt ID for each retry.

StageRecordFailure it explains
IngestionDocument hash, pages, sourceDuplicate or incomplete upload
ExtractionModel and prompt revision, field outputsRegression after release
ValidationFailed rules and source locationsTotals that do not reconcile
ReviewCorrected fields and review timeAutomation creating more work
ExportAction reference, response, retriesTimeout or duplicate write
AcceptanceExternal record ID and final stateOutput never used

Store document references rather than copying every document into telemetry. Give investigators access through application permission checks. Decide which fields to redact and how long to retain diagnostic data before enabling verbose logs.

OpenTelemetry’s generative AI conventions offer a starting vocabulary. Keep business outcome fields alongside model telemetry; a model span cannot tell you whether finance accepted an invoice.

Separate three failures

Extraction failure: the system reads the wrong amount. Inspect the document, OCR result, prompt revision, and field-level evaluation.

Business-rule failure: extracted fields are correct, but the total does not match the line items. Route to review; asking the model to invent a matching value hides the discrepancy.

Integration failure: the approved invoice cannot be written to the accounting system. Inspect permissions, schema mapping, external status, and retries. Changing a prompt will not fix an expired credential.

Each category needs an owner and recovery procedure. Otherwise categorization produces another chart without changing what happens next.

Choose useful denominators

Track cost per accepted invoice, including OCR, inference, retries, evaluation, review labor, and allocated operating costs.

Show review rate beside correction rate and rejected-export rate. A falling review rate can look positive while incorrect records reach the destination. Separate routine invoices from unfamiliar layouts so easy traffic does not hide weak performance on difficult documents.

Show active processing time and total elapsed time including human queues. A fast model does not help a document waiting two days for approval.

Build alerts around decisions

Start with a growing unresolved-export queue, sustained increases in correction rate, and releases failing a known regression set. Define a comparison window and minimum sample size so one unusual document does not create constant alarms.

Choose thresholds with the operator. An illustrative pilot might escalate unresolved external writes older than one business day; that is a proposed service rule, not a universal standard. Time-sensitive work needs its own limit.

Make releases explainable

Before changing a prompt, model, or retrieval configuration, run the same versioned evaluation set. Include historical failures that forced corrections. Compare field accuracy, exception routing, cost, and downstream acceptance.

Link configuration versions to results. When quality drops, decide whether to roll back, narrow supported document types, or change review policy. One aggregate score cannot make that decision.

The handover should include a sample trace, dashboard definitions, alert routing, retention settings, and a worked incident investigation. Discuss an operational review if your system produces outputs but cannot explain what happens afterward.

Keep exploring.

All articles ↗