At 02:17 UTC, our finance dashboard showed that daily LLM spending had jumped from an expected $1,200 to $4,080. Usage volume looked flat compared with the previous week, and nobody had knowingly changed the model or configuration.
We investigated model and config changes first, then compared request volume and per-request costs. The outlier turned out to be a background retry job. The cost of each individual call was unchanged, but the average cost per completed task had nearly tripled because retries were being recorded as separate normal calls rather than grouped under one task.
We fixed the reporting by attaching a stable task ID to every invocation and rolling costs up by task instead of by raw call. Has anyone else tracked down a similar AI cost spike? Was it a retry storm, an overlapping batch job, a model change, or something else?
4 Answers
Raw-call dashboards can hide this kind of problem. Track cost per completed task alongside retry counts, and alert when either moves well above its normal range. A task that costs three times as much may still look normal if the dashboard only counts individual requests.
For prevention, set separate budgets and hard caps for automated workloads, keep token-level audit logs by key or service, and alert around twice the normal daily spend. Background jobs, agent loops, and overlapping schedules are usually worth checking before blaming the model.
Human subscriptions generally don’t cap programmatic API usage, so automated calls still need their own pay-as-you-go controls.
Instrument the full call chain with OpenTelemetry or an equivalent tracing system. Put a consistent task ID on every span so a retry storm appears as several attempts within one trace instead of unrelated calls. I’d record model and token counts on the spans, then calculate dollars downstream because provider pricing, caching, and batch discounts can change.
It’s also worth reconciling the estimate with the provider invoice. Timed-out requests may still consume tokens and can sometimes be missing or represented differently in application traces.
We saw something similar with an evaluation job scheduled by cron. Each run lasted longer than its interval, so multiple copies overlapped, and timeouts caused additional retries. Per-call cost stayed flat while the bill climbed.
Grouping spend by service account exposed the problem faster than grouping by model. We eventually tagged every invocation with a run ID so retries were treated as one unit. Stable task or run IDs are much more useful than request counts for this kind of investigation.

That’s essentially where we landed, although we added the task IDs after the incident. Having retries appear as multiple spans under one trace would have made the issue obvious much sooner.