ModelShifts
← Blog & insightsAI Agents

AI Workflow Reliability: Retries, Duplicate Actions, and Human Handover

Design AI workflows that survive timeouts, duplicate events, and uncertain results without losing customer requests.

ModelShifts3 min readUpdated 11/09/2026
Editorial diagram: Receive to Validate to Commit to Reconcile

A resident sends an enquiry, your system creates a CRM record, and the CRM times out before returning a confirmation. Should the agent retry? If it does, it may create a duplicate. If it does not, the enquiry may disappear from the team’s queue.

The model can understand a request perfectly while the surrounding workflow loses the work. Start by designing what happens when execution is uncertain.

Design the state before the prompt

For an illustrative community enquiry workflow, use an explicit sequence:

received → validated → ready_for_review → approved → submitted → confirmed

Keep needs_information, rejected, and submission_unknown separate. A timeout belongs in submission_unknown; it is not proof that the write failed. Store state in a durable database so a worker restart does not erase progress.

Each transition should record a request ID, previous state, new state, actor, timestamp, and evidence. A model may propose a transition, but application code should check whether it is allowed.

Treat reads and writes differently

OperationRecoveryCheck first
Read appointmentsBounded retryIs the response still fresh?
Extract document fieldsReprocess the same versionHas a reviewer already corrected it?
Create CRM enquiryReconcile before retryDoes the external record exist?
Send customer messageCheck delivery stateWas it accepted by the provider?

An idempotency key identifies the same intended action across attempts. Combine the enquiry ID, action type, and approved revision. Do not generate a new key on every retry: that defeats duplicate protection. A changed action needs a new revision and, where appropriate, new approval.

Some APIs support this directly; Stripe documents one implementation. Other systems need an operation ledger and reconciliation against external records. An internal ledger alone cannot guarantee exactly-once behavior in a third-party service.

Handle the uncertain middle

Suppose the CRM accepts a request but your worker crashes before recording the external ID. On restart, look up the CRM record using the stable reference you submitted. If it exists, save the external ID and confirm the operation. If the service cannot support reliable lookup, send the item to a reconciliation queue instead of retrying blindly.

The operator should see what was attempted, which system was contacted, the last confirmed state, and the available recovery actions. A generic error message is not an operational handover.

Build a failure test set

Before adding more autonomous steps, test these scenarios against a sandbox:

  1. Deliver the same incoming event twice. Expect one intended business action.
  2. Crash immediately after the remote write succeeds. Expect reconciliation on restart.
  3. Remove document access. Expect an access failure, not an invented answer.
  4. Change the enquiry after approval. Expect renewed review before execution.
  5. Disable a downstream service. Expect bounded retries and a visible queue.

These are proposed acceptance checks, not reported production results. Set latency and retry budgets from the business process: a resident enquiry can tolerate a different delay from an interactive booking screen.

Measure the queue, not only the model

Track duplicate actions, unresolved submissions, oldest queue item, manual recovery time, and completed enquiries. Record model quality separately. A high extraction score does not compensate for requests stranded between systems.

For each unresolved state, agree who responds and within what period. Add an operator action to reconcile or safely cancel the work, with the reason recorded. A queue without an owner becomes another inbox nobody checks.

The delivery package should include the state diagram, integration contracts, failure tests, an operator runbook, and a named queue owner. That lets the client operate the workflow after handover.

Explore the community workflow concept or discuss an integration.

Keep exploring.

All articles ↗