Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk10 min

The API Worked. The Architecture Didn’t.

A 200 OK proves one interaction succeeded. It does not prove the order was charged, reserved, and recorded everywhere it needs to be. Here is how lost responses, dual writes, and half-finished workflows cause that gap, and how to design around them.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A successful API response tells you that one interaction went one way. It does not tell you that the order was charged, reserved, and recorded in every system that needs to know. When the endpoint reports success and the business process still stalls, the gap is usually in how the architecture handles lost responses, retries, partial completion, and event publication, not in the endpoint’s code.

What a successful response actually promises

Teams often use “success” to mean several different things. Before designing recovery, decide which of the following levels a given endpoint guarantees, and document it in the API contract.

Level What it establishes What it does not establish
Received The server got the request and responded. The request was valid, and nothing has necessarily been stored.
Accepted The request passed validation and the server took responsibility for it (a 202 Accepted response is typical). The work has run. Accepted work can still fail later.
Queued The work was placed on a queue or topic. A consumer has picked it up, let alone finished it.
Processed The handling service finished its own local work. Other services have acted on the change.
Durably committed The state change is saved and survives a restart. Downstream systems, notifications, and dependent workflow steps are complete.

A 200 response on a synchronous call usually means “processed and committed locally.” A 202 means “accepted,” and the client must check later. Neither means the business workflow is complete. Status codes differ between APIs, so check the actual contract for each endpoint instead of assuming a convention.

How business state diverges while the API looks healthy

Three failure shapes account for most of the divergence engineers encounter. Each one can occur without any error being logged at the endpoint.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The response is lost after the commit

The server commits the change, then the connection drops, a load balancer times out, or the client’s read deadline expires before the response arrives. The client sees a failure and retries. If the endpoint is not idempotent, the second call creates a second payment, a second shipment, or a second reservation. The server did its job both times; the business effect happened twice.

The database write and the event publication are separate steps

A service that updates its database and then publishes an event to a broker performs two operations with no shared transaction boundary. If the process crashes between them, the state change is durable but no event is sent, so downstream systems never learn about it. Reversing the order creates the opposite problem: an event announces a change that was never committed. AWS’s guidance on the transactional outbox pattern describes this dual-write motivation directly.

A multi-step workflow stops in the middle

An order flow might reserve inventory, authorize payment, and then create a shipment. If the payment is declined after the inventory reservation succeeded, the inventory service reports success for its step, and the payment service reports a failure for its own. Each service’s logs look correct in isolation. Nobody owns the question of whether the order as a whole is in a consistent state, and the customer sees an order stuck in “processing.”

Make retries safe before making them frequent

Retries are necessary, because transient failures such as timeouts, throttling, and brief unavailability are common in distributed systems. The danger is that retries multiply the effects of an operation that was already partly or fully applied. AWS’s retry-with-backoff guidance warns that retries without idempotency can corrupt state, and that excessive retries can make a degraded service worse. Treat retries as a contract with three parts:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Classify errors. Retry transient failures (timeouts, 503 responses, throttling signals). Do not retry permanent rejections such as validation errors or business declines. Retrying a declined payment only generates noise and can trip fraud controls.
  2. Back off and bound. Use exponential backoff with added jitter so that many clients do not retry in lockstep, and cap the number of attempts. Set an overall deadline for the whole operation so that nested retries across services do not multiply.
  3. Make state-changing operations idempotent. The client sends a stable operation identifier, often called an idempotency key. The server stores that key together with the result of the operation, in the same local transaction as the business change. A repeated request with the same key returns the stored result rather than repeating the effect.
  4. Return the original outcome. If a retry arrives while the first attempt is still in progress, the server should return a “still processing” status rather than starting a second attempt.

For example, a checkout service might send a charge request with the key order-8841-charge-1. If the response is lost, the retry carries the same key. The payment service finds the stored record and returns the original authorization, so the customer is charged once. The key must be scoped to the business operation, not generated freshly for each attempt, or the protection disappears.

Publishing state changes without a dual write

The fix for the database-plus-event divergence is to stop treating the two writes as independent. The transactional outbox pattern, as described in AWS Prescriptive Guidance, records the event in an outbox table inside the same local database transaction as the business change.

Approach Behavior if a crash occurs between the state change and the publication Duplicates and ordering Main cost
Write to database, then publish directly State is saved, event is lost unless extra retry logic exists. Retries can produce duplicates; ordering depends on the broker and retry logic. Simple, but the failure window is structural.
Publish first, then write Event announces a change that may never be committed. Consumers can act on phantom state. Simple, but unsafe for most business events.
Transactional outbox Business row and outbox row commit or roll back together, so the event is never lost once the transaction commits. Delivery is typically at least once, so consumers must handle duplicates. Ordering needs explicit design, such as keying events by aggregate. An extra table, a relay process, and cleanup of sent records.

In practice, the outbox works like this:

  1. Within one local transaction, insert or update the business row and insert an outbox row describing the event.
  2. A relay process reads unsent outbox rows in order, publishes them, and marks them as sent only after the broker acknowledges.
  3. If the relay crashes after publishing but before marking the row, the event is sent again. Consumers must therefore record processed message identifiers and skip duplicates.
  4. Monitor the age of the oldest unsent row. A growing age is the first visible sign that the relay has stopped.

The outbox guarantees that a committed change will eventually be published from its own database. It does not coordinate a business transaction across several services, which is the job of the next section.

Coordinating multi-service workflows with sagas

A saga breaks a business operation into a sequence of local transactions, each in one service or data store. Every step has a defined follow-up: either the workflow continues to the next step, or, if a step fails, compensating actions undo the earlier steps. Microsoft’s Saga design pattern guidance and AWS Prescriptive Guidance describe the same core model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sagas bring real trade-offs:

  • Eventual consistency. The system passes through intermediate states that other readers can see. A saga does not give the isolation of a single database transaction.
  • Compensation is business logic. Undoing a shipment or refunding a payment is not a database rollback. Compensations can themselves fail, and they may not perfectly reverse the effect, for example when a customer has already been notified.
  • Idempotency is required. Each step and each compensation may be retried, so each must tolerate repeated delivery.
  • Testing is hard. Microsoft’s guidance notes that integration testing across services is difficult. Tests should deliberately fail each step and restart services between steps.

Choreography or orchestration

Dimension Choreography Orchestration
Control Each service reacts to events published by others. No central controller. A coordinator sends commands to participants and tracks the workflow state.
Coupling Services are loosely coupled to each other but tied to the shared event contracts. Participants are coupled to the coordinator’s commands.
Visibility Harder to track as participants grow, because the flow is spread across event handlers. Easier to see overall state in one place.
Failure risk Gaps in event flow can leave a workflow silently incomplete. The coordinator becomes a dependency that must be made highly available and recoverable.

AWS’s saga orchestration guidance treats the central coordinator as a deliberate design choice with its own operational costs. For a workflow with only a few participants and clear ordering, orchestration is often easier to debug. For loosely related consumers that react to facts, choreography may fit better.

Decide whether to retry forward or compensate

When a step fails, the recovery action depends on the cause:

  • Transient failure in an idempotent step: retry the step forward, within a bounded budget.
  • Permanent business rejection: stop moving forward and run compensations for the completed steps in reverse order.
  • Compensation fails: retry the compensation with the same idempotency protection. After the attempt limit, route the workflow to a human queue with the full state history. Do not mark it complete.
  • Step outcome unknown (timeout): query the participant by operation identifier before choosing either action. Resending or compensating blindly can make the state worse.

Recovering from partial completion

Every workflow should have an explicit list of the partial states it can enter. The table below is a starting template for a multi-step order flow; the states, signals, and thresholds will differ for each system.

Partial state Detection signal Recovery action
Remote operation committed, response lost Client timeout with no response, followed by a retry Look up the operation by idempotency key and return the stored result. Do not resend under a new key.
Local change committed, event not yet published Unsent outbox row older than the expected relay interval Let the relay retry. Alert if the age keeps growing, which usually indicates a broker or relay failure.
Duplicate event delivered Message identifier already recorded by the consumer Skip the message and acknowledge it.
Step N complete, step N+1 permanently rejected Workflow stuck in a waiting or failed step for longer than its deadline Run compensations for earlier steps, or move to manual review if compensation is not defined.
Compensation failed Compensation attempts exceed the retry budget Escalate with full state history. Keep the workflow open until a person or reconciliation job resolves it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Observability that follows the business workflow

Endpoint uptime and error rates are necessary, but they cannot reveal a workflow that is stuck between services. Instrument the business process itself:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Workflow identifier everywhere. Put the same business or correlation identifier on every log line and trace span that touches the operation, across all participating services.
  • Log state transitions. Record the from-state and to-state, the step, and the reason, so that an engineer can reconstruct the path without replaying traffic.
  • Separate request metrics from workflow metrics. Track “requests succeeded” alongside “workflows completed within deadline.”
  • Watch for stuck work. Examples include the age of the oldest in-progress workflow, the count of outbox rows not yet sent, and the number of records that exist in one system but have no match in another after a reconciliation window.
  • Make recovery actionable. Each alert should name the workflow, the step, and the runbook action from the partial-state table.

The thresholds in that list are examples to tune against your own process timings. None of them is a standard metric set.

A diagnostic sequence for an integration that “worked” but failed

  1. State exactly what the response guaranteed: received, accepted, queued, processed, or durably committed.
  2. Take one business operation and trace it across every participant using its workflow identifier. Compare the request outcome with the final business state, and write down where they differ.
  3. Ask what happens if the remote side committed but the response was lost. Check whether a retry can recognize the completed operation.
  4. Find any place where a state change and an event publication are separate writes. If one exists, confirm that an outbox or another explicit delivery contract covers it.
  5. List the partial states the workflow can enter, and assign a recovery action to each, choosing between forward retry and compensation.
  6. Confirm that stuck or unmatched business work is monitored, not only endpoint availability.

What the evidence does and does not establish

The failure modes above are well documented as design concerns. The available sources do not measure how often they occur. AWS Prescriptive Guidance and Microsoft’s architecture guidance describe patterns and their trade-offs, not incidence rates. Two accounts describe the symptom of a successful call followed by an unfinished business process: a vendor article from Rigg Technologies dated August 15, 2026, which discusses lost responses and mismatched transaction records, and an individual technical essay by Prem Chandak on Medium dated April 7, 2026, which walks through an illustrative end-to-end scenario. Both are useful for framing the problem. Neither presents independently verified prevalence data, and the counts they use should not be read as general findings. This article does not describe any particular system or incident.

Several of these patterns are complementary rather than competing. An outbox can reliably publish the event that starts a saga, and idempotent steps make both forward retries and compensations safe.

Frequently Asked Questions

Do I need idempotency on every endpoint?

Every endpoint that changes business state and can be retried by a client, a gateway, or a message consumer should be idempotent. Read-only endpoints generally do not need idempotency keys, because repeating them has no business effect.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When is a saga worth the complexity?

A saga is worth considering when one business operation spans services that each commit their own changes, and a failure needs a defined continuation or undo. A single downstream call protected by idempotency and bounded retries may be enough on its own.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.