I'm building features where background jobs make paid API calls, mostly to language models. When a job fails partway through and retries, it can repeat work that already succeeded, which means paying twice or triggering side effects twice—for example, sending the same email again. How do you handle this in practice? Do you persist progress after each step, make operations idempotent, use request keys, or rely on another pattern to prevent completed work from running again?
3 Answers
Persist the job’s state and progress in durable storage, usually after each meaningful step. On retry, load that state and resume from the first unfinished step instead of starting over. A queue with retry counts and backoff also helps keep failures controlled.
Design each step to be idempotent whenever possible. Store a stable operation or idempotency key, check whether that key has already completed, and make side effects such as email delivery tolerate duplicate attempts. Saving intermediate results is useful, but it cannot completely remove the small failure window between an external call and recording the result.
Break the workflow into durable, independently retryable steps. Cache or save the result of expensive calls, track retry counts, and only retry the exact step that failed. For actions that cannot safely be repeated, use an outbox or deduplication record so the worker can confirm whether the side effect was already issued before sending it again.

That’s where I’m leaning too. The tricky case is a crash immediately after the paid call succeeds but before the progress update is saved.