incident report · sev-2

Export queue backlog, 12 June

A retry storm from one malformed workspace held the export queue for 94 minutes. No exports were lost; the oldest was delayed 81 minutes.

Detected
09:14, queue-depth alert
Resolved
10:48, poison job quarantined
Blast radius
exports only; imports and sync unaffected
1,204exports delayed+1,204 81 minoldest delay 0exports lost

Timeline

09:14 Queue-depth alert fires Depth crosses 500; normal peak is 60. 09:26 On-call confirms a single hot job The same export id retrying at the head of the queue, every 45 seconds. 09:51 First fix doesn't hold Raising worker count clears depth briefly; the retry storm refills it because the head job still fails first. 10:32 Poison job quarantined The workspace's export is parked in a dead-letter table; the queue drains at normal rate. 10:48 Backlog cleared Depth back under 60; delayed exports delivered.

Root cause

The export worker treats any failure as retryable. One workspace carried an attachment with a declared size of −1, which fails serialization every time; with retries capped by attempt count but not by queue position, the job returned to the head on each attempt and starved everything behind it.

serialize fails

starved

queue head

worker

retry in 45s

1,204 jobs behind

Follow-ups

  1. Classify worker failures as permanent vs. retryable; dead-letter permanents on first failure.
  2. Re-enqueue retries at the tail.
  3. Alert on a single job id exceeding five attempts, not only on queue depth.