One click, then silence

The entry has already passed review. Its supplier reference, amount and supporting document are visible beside a green approval mark. The operator presses Submit and waits. A spinner turns. Then the interface says the request failed.

What failed is the conversation between two services. The business action may have succeeded. That distinction is missing from the screen, and the operator is being asked to make a decision without it.

A second click is understandable. If the application gives the retry a new identity, the receiving service may understand it as a second intended entry. The duplicate begins as a design problem before it becomes an accounting problem.

Keep the identity of the intent

The first repair is to give the intended business operation a stable identity, preserved across transport attempts. Amazon’s account of idempotent APIs explains how a client-supplied identifier can let a service recognise a retry. The receiving API must actually provide that contract; adding a random field to a request creates no guarantee.

In our imagined integration, the operator’s second click should resume the original operation. A changed invoice is a different matter. The system needs a correction policy that distinguishes retrying the same intent from authorising a new one.

Unknown is a state someone can work with

Now change the screen. It says that the submission was attempted and confirmation is missing. It shows the business reference and offers the next safe action: check the destination under that reference. It does not invite the operator to create a fresh entry.

Pat Helland’s position paper on life beyond distributed transactions discusses application patterns for coordinating work across separately managed entities. The practical connection here is responsibility for partial progress. If the integration has no shared transaction across both services, its workflow has to account for the interval in which they disagree or cannot communicate.

The application can remember that interval. Prepared, submitted, confirmed and uncertain are different states. They deserve different recovery actions.

Known stateUseful next action
PreparedSubmit under the stable operation identity
Submitted, no confirmationReconcile with the destination
ConfirmedShow the destination reference
Rejected by validationCorrect the cause before another attempt

The operator opens the record

The record contains an operation identifier, attempt time, destination and last confirmed result. It points to the reviewed input without copying unnecessary private contents into every log. The operator can see enough to establish the business state.

Google’s SRE guidance on monitoring separates useful operational signals from a flood of undirected data. For this workflow, we would make unresolved operations visible as a queue with an owner. A service can respond successfully to health checks while the queue of uncertain business work keeps growing.

If the destination confirms acceptance, update the local record. If it confirms rejection, correct the error. If the outcome is unknown and there is no safe way to retry, pass the case to a person who can investigate.

Rehearse the silence

Before releasing the integration, interrupt it after the destination accepts a request and before the caller receives confirmation. Then ask someone who did not build it to recover the operation. Watch the business state at the destination, not just the caller’s logs.

Repeat with a duplicate event and with a corrected invoice. The same identity should preserve the same intent under the API’s documented behaviour; a correction should follow the agreed business rule. The idempotency contract needs these details to be useful.

The incident ends well when the operator can say what happened, show the evidence and resume work without creating a second effect. That is the feature the integration was missing.

If confirmation is missing, check whether the operation completed before repeating it.

Sources & further reading

  1. Amazon Builders’ Library — Making retries safe with idempotent APIs
  2. Helland — Life beyond Distributed Transactions: an Apostate’s Opinion (CIDR 2007)
  3. Google SRE — Monitoring Distributed Systems

Primary sources inform the technical background. The examples and proposed working methods are MGLO’s own; Business scenarios are identified in the text.