Most ERP integrations on Shopify are not broken. They are rotting: quietly, in production, in ways that never page anyone. The drift shows up during a reconciliation, an audit or a failure window, and by then it has been running for months.
The integrations we have inherited from other agencies have rotted in roughly the same five ways. If you run a Plus store with a NetSuite, Brightpearl or Microsoft Dynamics integration, these are the five things to check before one of them costs you a weekend.
1. Webhooks without idempotency
Almost every integration we inherit has webhook handlers that assume each event arrives exactly once. Shopify's webhooks do not guarantee that. Under network retry conditions the same event can arrive anywhere from two to ten times, and if your handler does not dedupe by event ID you double-create orders, double-decrement inventory and double-bill the customer.
What to check: Does every webhook handler look up a deduplication key (idempotency key, event ID or hash) before processing? Is the dedupe store durable across deploys?
The fix: A 30-line Redis-backed dedupe layer at the top of every handler. On inherited integrations it is the first thing we retrofit.
2. Retry queues without dead-letter monitoring
The default response to "the upstream API failed" is to retry, and most integrations do. Almost none of them show the queue depth, the retry count or the dead-letter depth to anyone who is watching.
We have found dead-letter queues holding weeks of events that nobody knew about until a reconciliation turned them up.
What to check: Where is retry queue depth visible? Where is dead-letter depth visible? Who is paged when either crosses a threshold?
The fix: Every queue we deploy has depth metrics in Grafana, dead-letter alerts to PagerDuty above thresholds, and a daily Slack post of yesterday's queue health.
3. Schema drift without versioning
Your ERP changed its API last quarter. The integration did not. It still mostly works, except for the new tax field on shipping line items, which is now silently dropped, or the new fulfillment status enum value, which fails the integration's switch statement and falls through to the catch-all default.
This is the most expensive class of rot because it never fails loudly. It produces wrong data quietly.
What to check: Is there a versioned schema for the data flowing in and out? Is there a contract test that runs against the upstream's real API? When the upstream changes, what alerts you?
The fix: A typed schema (Zod, Pydantic, or codegen from an OpenAPI spec) on every boundary, and a contract test that runs daily against the upstream's real responses and fails loudly when fields appear or disappear.
4. Time-bomb credentials
OAuth tokens, API keys and webhook secrets all expire eventually, and they do not always tell you when. We have seen tokens silently fail to refresh for weeks because the refresh was wired to a deprecated OAuth scope, with the only error a 401 in a log nobody read.
What to check: Where are your credentials? When do they rotate? What is the alert path if a refresh fails?
The fix: Credentials in a vault (1Password, AWS Secrets Manager, Doppler) with a rotation schedule. A health check on every external integration that authenticates and reports green or red. A failed health check pages on-call.
5. Missing reconciliation jobs
The hardest rot to see is the kind where every individual transaction processed correctly and the aggregate is still wrong. Say one order in 50,000 drops a line item to a race condition. The order looks fine and the customer was charged correctly, but the line item never reached the ERP, so inventory is right in Shopify and wrong in NetSuite.
Without a reconciliation job you find out at month-end close, or at audit, or never.
What to check: Is there a reconciliation job that compares order counts, revenue totals and inventory positions across systems? How often does it run? Where are its results visible?
The fix: A nightly reconciliation that compares Shopify, ERP and 3PL totals across the previous seven days, alerts on drift above 0.1%, and writes the results to a dashboard that ops reviews weekly.
What an integration audit covers
When we inherit an integration we run a 30-item checklist before touching it. The five above are the ones that turn up in nearly every audit. Others on the list:
- Webhook signature verification on every endpoint
- Rate limit handling that backs off instead of failing
- Bulk operations isolated from sync paths
- Replay capability for any event
- Versioned migrations for schema changes
- Documented failure modes
If a marketplace freelancer built your integration at $80 an hour two years ago, none of these will be in place. That is no criticism of the freelancer. They were paid to ship the happy path, and the infrastructure work does not fit inside a marketplace ticket.
What to do about it
If you are looking at an inherited integration:
- Audit before you refactor. A two-week audit tells you what is load-bearing and what is rotting. Then decide between refactor and rebuild.
- Stand up monitoring before you change anything. You need the metrics in place to know whether a change helped.
- Fix one thing at a time. Take the highest-impact item, watch the metrics, then take the next.
- Document. Every integration we ship has a runbook for its failure modes. Every integration we inherit gets one written before we close out.
Most integrations rot because they get handed to whoever is available, then handed off, then handed off again. Nobody owns them long enough to notice the drift.