Skip to content
All posts

1 min readcetus, providers, debugging

The provider account that had quietly died

Uploads kept failing for a week. The retry code was fine the whole time.

8 July 2026

Uploads kept failing. Not all of them, never loudly, and nothing in the logs said why. A job would go in, sit there, and never come back out.

I spent a week inside the retry code. Read the backoff twice. Added logging around the queue, took it out, put it somewhere else. Rewrote the handler. Every time, the retries were doing exactly what I had told them to do.

One of the provider accounts had gone dead. Every job routed to it queued behind a worker that was never going to answer, and the queue waited, because waiting is the entire job of a queue.

The provider layer rotates off an account now once it stops answering, and everything behind it moves to the next one. That change is nine lines. The week in front of it was not.