How do you safely retry background jobs that call paid APIs?

0
0
Asked By MellowCedar42 On

I'm building background jobs that make paid API calls, mostly to LLM services. When a job fails partway through and gets retried, I can end up paying again for work that already succeeded. Some actions, such as sending an email, may also happen twice. How do you usually handle this? Do you save progress after each step, use idempotency keys, or structure the job another way so completed work is not repeated? I'm especially interested in how you handle a crash that occurs after the paid API call succeeds but before the job records that progress, and in examples of failures that ended up costing money.

4 Answers

Answered By QuietOrbit7 On

Persist durable state after each meaningful step, including the job status, step results, retry count, and any provider request IDs. On retry, inspect that state and skip steps already known to have completed. A queue helps with delivery and retries, but it does not by itself prevent duplicate side effects.

Answered By SilverMaple28 On

Break the workflow into small, durable steps and make the boundaries explicit. Store inputs and outputs where they can be audited, retry only the failed step, and set limits such as maximum attempts and spending alerts. This does not eliminate every duplicate, but it turns a whole-job replay into a controlled recovery process.

Answered By BriskHarbor19 On

Make each operation idempotent whenever possible. Generate a stable idempotency key from the job and step, then send it with the paid API request if the provider supports idempotency. For emails or other side effects, keep an idempotency record or an outbox entry keyed to the same operation so a retry does not send the message twice.

MellowCedar42 -

The tricky case for me is a crash immediately after the API succeeds but before I save the result. I’m looking at provider request IDs and reconciliation checks for that gap.

Answered By AmberWillow63 On

There is an unavoidable uncertainty window if the remote call succeeds and your process dies before recording the result. A database transaction cannot make an external API call atomic with your local save. Use provider-side idempotency or lookup-by-request-key when available; otherwise mark the step as uncertain and reconcile it before retrying instead of blindly calling again.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.