Six integration failure modes that only appear in production
Every call can arrive twice and every partner will go down. Design idempotency, failure paths and daily reconciliation before go-live.

Every integration programme begins with optimism and a partner API document that is at least two versions out of date. The engineering question is not whether the partner’s system will behave unexpectedly, but whether your architecture notices when it does.
Assume every call can arrive twice
Networks retry. Webhooks fire again after a timeout. Queues redeliver after a consumer restart. If an inbound payment instruction or order creation is not idempotent, duplicate work is not an edge case — it is a scheduled event. Every mutating handler needs a natural or synthetic key, a check before action, and a record of what it already did.
Design the failure path before the happy path
For each integration, answer four questions in writing:
- What happens when the partner is unavailable for four hours?
- Where do messages go when they cannot be processed, and who looks at them?
- How is a stuck item replayed safely, and how do we know replay is safe?
- How does finance reconcile totals between the two systems each day?
If any answer is “we would find out”, the integration is not ready for production traffic.
Version contracts, including the inbound ones
Outbound APIs usually get versioning discipline. Inbound partner payloads frequently do not — a partner changes a field type and something downstream quietly stores a string where a decimal belongs. Validate inbound payloads against an explicit schema, reject rather than coerce, and record the raw message alongside the parsed result so a disputed transaction can always be reconstructed.
Make reconciliation a first-class feature
Reconciliation dashboards are usually treated as reporting. They are better understood as the integration’s test suite running in production. Publish daily totals on both sides, flag differences, and give operations a workflow to resolve them. An integration with a visible reconciliation figure is one that operations trusts; one without it is a permanent argument waiting to happen.
Observability that names the business event
“HTTP 500 from service-b” is not useful at 3am. “Order 88213 failed to reach the warehouse system” is. Structure logs and traces around business events, include the correlation identifier you already use for idempotency, and alert on business-level symptoms rather than raw error rates.