An integration is easy to demonstrate when both systems respond quickly and every request succeeds. The harder design work starts when a supplier accepts an order but the response disappears, a rate limit interrupts a large import, or an event arrives twice. These are ordinary conditions in connected systems, not unusual exceptions that can be left until later.
A reliable integration gives each business operation an understandable state, bounds the work it performs and provides a route to resolve uncertainty. The objective is not to hide every failure. It is to make sure temporary problems do not silently become lost orders, duplicated records or permanent confusion for the people using the application.
Define the business contract before the transport
Start with the meaning of the exchange. What event causes a request? Which system owns the resulting record? What does acceptance mean, and how will completion be confirmed? An HTTP success response might mean that work has finished, or only that another system has placed it in a queue.
Write examples of normal and exceptional outcomes with the business owner. Include an invalid customer, an unavailable product and an operation that takes longer than expected. Distinguish failures that the user can correct from failures requiring the systems to recover.
Record the external contract's limits: authentication, allowed request rate, payload size and any versioning expectations. Those constraints influence the user experience. A supplier that processes requests slowly may require a visible pending state rather than an interface promising an immediate result.
Give every important operation a stable identity
When a client times out, it cannot always tell whether the server performed the action. Blindly repeating a request can therefore produce duplicate effects. Use a stable operation identifier and an agreed duplicate-handling mechanism where the external service supports it.
For a hypothetical order export, retain a local record connecting the internal order, its revision and the external submission identifier. Repeating the same submission should either retrieve the known result or be recognised as the same business operation. A changed order may require a new revision rather than reusing the original identity.
An idempotency key is not magic by itself. Define how long it remains valid, which request details it represents and how conflicts are handled. If a supplier has no suitable duplicate protection, the integration may need a lookup and reconciliation process before retrying an uncertain write.
Retry only when another attempt can help
Some failures are temporary, such as brief connectivity problems. Others, such as invalid input or missing permission, will not improve through repetition. Classify responses before choosing whether to retry. Microsoft's Retry pattern discusses transient failures, delays between attempts and the importance of considering idempotency.
Set a maximum number of attempts and an overall deadline. Increase delays where appropriate and avoid having every client retry at exactly the same moment. Honour the provider's retry guidance and rate limits instead of creating extra demand while it is already struggling.
Check for retries at multiple layers. A browser, gateway and service client can each repeat an operation, multiplying attempts beyond what any one developer expected. Decide which layer owns the retry policy and expose enough information to understand the actual number of external calls.
Stop applying pressure to an unhealthy dependency
A persistent outage needs a different response from a brief interruption. Continuing to send requests can consume application resources while providing no useful result. A circuit-breaker mechanism can temporarily stop calls after a pattern of failures and allow controlled attempts to determine whether the dependency has recovered.
The user experience still needs a defined outcome. A pending request, a manual route or an explicit temporary-unavailability message may be appropriate. Returning an apparently successful result while discarding the work usually is not.
Keep the scope of failure contained. An unavailable document-conversion provider should not prevent staff from viewing existing customer records unless that dependency is truly required. Review which operations need the external service and which can continue with clearly described limitations.
Keep a durable record of unfinished work
If an operation must survive a process restart, store its state in a durable place. Record when it was accepted, which attempts were made and what evidence confirms completion. An in-memory list is not a reliable record of business commitments after the application stops.
Consider the gap between saving a business change and announcing it to another system. If those actions happen separately, a failure between them can leave a record with no corresponding message. An outbox approach can record the intended message alongside the business change, then publish it asynchronously; delivery and duplicate handling still require their own design.
Make exceptions visible to an operator. A queue of failed items should include a useful reason, the related business record and a safe next action. Avoid a single generic “retry all” control that can repeat expensive or irreversible effects without review.
Design for duplicates and out-of-order events
Webhook and message consumers should not assume every event arrives once or in chronological order. A delayed update can arrive after a newer one, and a sender may repeat a notification when it did not receive acknowledgement. Use the provider's event identifiers and ordering information where available.
Check whether the integration should apply an event directly or retrieve the current authoritative record. The right choice depends on the external contract and whether intermediate states matter. Document how deleted records, corrections and delayed notifications are handled.
Validate the sender and payload before processing. Keep authentication and business validation separate: a request from the expected provider can still contain an unsupported event or an invalid relationship. Record enough evidence for investigation without retaining unnecessary credentials or sensitive payloads.
Reconcile outcomes, not just successful requests
Compare important records across the boundary periodically. For an order integration, confirm that accepted internal orders have the expected external identifiers and states. Investigate unmatched records and disagreements, including operations that appeared successful at the transport layer.
Test ambiguous failure conditions in a controlled environment: a response lost after acceptance, a rate limit, a repeated webhook and a restart during processing. These exercises tell you whether the state model remains understandable when the happy path breaks.
Worktechlabs builds application integrations with explicit ownership, recoverable processing and verifiable outcomes. Pair the integration with production monitoring so a temporary supplier problem produces an actionable state instead of becoming a hidden business error.

