Skip to content
Leivadev
Back to Blog

A slow provider hung our Workers

An upstream call with no deadline kept Worker invocations alive for ~31 seconds and collided with the database's own timeout. The fix is two ideas: bound every call, and retry only what is safe to repeat.

· 5 min read · ReliabilityTimeoutsEdge

The day after fixing a payment race, I went looking at production logs for anything else that smelled wrong, and found Worker invocations that were living for around thirty-one seconds. On an edge runtime where requests are supposed to be short, that is a long time to be doing nothing but waiting.

The cause was almost embarrassingly simple: our outbound calls to payment providers had no timeout. When a provider stalled, fetch() just waited. The invocation stayed alive, and worse, it started colliding with the edge database’s own storage timeout, so a slow provider could turn into a database error somewhere else entirely. One missing deadline was quietly coupling two unrelated subsystems.

This post is about the two changes that fixed it, because they are really two separate reliability ideas that people tend to reach for together and then apply too broadly.

Idea 1: every network call needs a deadline

A call to something you do not control can take forever, so you have to decide how long “forever” is allowed to be. I bounded every upstream request with an abort signal, and, just as importantly, surfaced a timeout as its own kind of failure rather than a generic server error.

async request(url: string, opts: RequestOptions = {}) {
  // Caller can pass a shared signal; otherwise each call gets its own deadline.
  const signal = opts.signal ?? AbortSignal.timeout(15_000);
  try {
    return await fetch(url, { ...opts, signal });
  } catch (err) {
    // Key off signal.aborted, not the error's name, so it survives runtime quirks.
    if (signal.aborted) throw new ApiTimeout('Upstream timed out', 504);
    throw err;
  }
}

Two details earned their place here. First, detecting the timeout with signal.aborted instead of matching on an error name or message, because runtimes are inconsistent about what they throw when a fetch is aborted, and I did not want the classification to break on a runtime upgrade. Second, mapping the timeout to a 504 instead of letting it fall through as a 500. “The upstream did not answer in time” and “we threw an exception” are different facts, and a caller, a dashboard, or an on-call engineer should be able to tell them apart at a glance.

There was one more subtlety. One provider flow made two sequential calls: validate, then create. Give each its own 15-second deadline and the worst case is thirty seconds, which is exactly the hang I was trying to kill. So the two calls share a single timeout budget: one signal, passed into both, so the whole operation is bounded, not each leg of it.

// Share one deadline across a multi-step operation.
const signal = AbortSignal.timeout(15_000);
await api.validateMerchant({ signal });
await api.createOrder({ signal }); // worst case ~15s total, not ~30s

Idea 2: retries are only safe when the operation is

The edge database is backed by a single-threaded store that can occasionally be reset out from under an in-flight query. You see it as a transient “object was reset” or “network connection lost,” and the platform’s own guidance is to retry. So I added a small retry helper with exponential backoff and jitter, capped at three attempts so a struggling store does not get hammered harder.

async function withRetry<T>(op: () => Promise<T>): Promise<T> {
  for (let attempt = 1; ; attempt++) {
    try {
      return await op();
    } catch (err) {
      if (!isTransient(err) || attempt >= 3) throw err;
      await sleep(backoff(attempt) + jitter()); // spread the retries out
    }
  }
}

The backoff matters because retries that all fire at the same instant just recreate the stampede that caused the problem. The jitter spreads them out. The cap matters because a retry loop with no ceiling is a denial-of-service attack you wrote against yourself.

But the part I want to underline is what I did not wrap. It is tempting to put every database call behind the retry helper and move on. That is a bug. A transient failure can arrive after the write has already committed. You get an error, you retry, and now you have written the row twice.

So the rule I settled on: retry reads, retry updates, retry idempotent creates (the INSERT ... ON CONFLICT from the previous incident), because repeating those is harmless. Never retry a plain insert into a table without a uniqueness constraint, because a post-commit transient error plus a retry equals a duplicate row, and nobody is there to catch it.

// Safe to retry. Repeating them changes nothing.
withRetry(() => repo.findByReference(ref));
withRetry(() => repo.markPaid(id));
withRetry(() => repo.createIfNotExists(tx)); // idempotent by construction

// NOT retried. A retry after a post-commit failure would duplicate the row.
repo.create(rawEvent);

The whole strategy collapses into one question asked of every write:

flowchart TD
  E[Transient database error] --> D{Is this write safe to repeat?}
  D -->|Yes: reads, updates, ON CONFLICT inserts| R[Retry with capped backoff and jitter]
  D -->|No: plain inserts| S[Do not retry. A post-commit error would duplicate the row]

Deciding which of your writes are idempotent is not busywork you do to satisfy the helper. It is the design. Once you can answer “is it safe to run this twice?” for every write, the retry strategy falls out of it almost mechanically.

What I took away from it

A missing timeout is not a small omission. It coupled a slow payment provider to unrelated database errors and kept invocations alive far past anything reasonable. Every call to something you do not control needs a deadline.

Model timeouts as their own failure. A 504 you can reason about beats a 500 you have to investigate. The classification is worth the few extra lines, and keying it off the abort signal keeps it robust.

Bound the operation, not the leg. Per-call deadlines silently add up across sequential calls. Share one budget when the steps are part of one logical action.

“Add retries” is half an instruction. The other half is “to the operations that are safe to repeat.” Backoff and jitter keep retries from becoming the outage; idempotency is what makes them correct in the first place.