I'm trying to understand the billing for Claude Opus hosted through Microsoft Foundry on Azure. In one day, the usage report showed 359.52K input tokens, 434.39K output tokens, and €46.32 in Claude Consumption Unit charges. Using published rates of roughly $5 per million input tokens and $25 per million output tokens, I calculate about $1.80 for input and $10.86 for output, or approximately $12.66 total. That seems closer to €10–€11 than the amount Azure reported. Could extended thinking or reasoning tokens, prompt-cache reads, or prompt-cache writes account for the difference? A fourfold increase seems unusually high. Has anyone seen a similar discrepancy with Claude Consumption Units in Azure AI Foundry?
3 Answers
If the detailed usage metrics still do not reconcile, open a billing case with the Foundry team. They can inspect the underlying consumption records and provide a category-by-category breakdown, which is especially useful when CCUs do not line up with the basic token totals.
Extended reasoning tokens may not appear in the regular input/output split, but they can still be billed at output-token rates. In some cases, extended thinking has made the effective cost three or four times higher than a calculation based only on visible input and output tokens. Check the detailed trace or usage metrics for reasoning tokens. Prompt-cache writes are typically charged at a multiple of the input rate, while cache reads are much cheaper, so writes alone probably would not explain the full difference.
Before assuming the model is responsible, check whether the bill includes other Foundry services such as content understanding, document processing, agent runtime, memory, or evaluations. Those can appear alongside model usage and make the total look much higher than the token calculation.
In this case, the Anthropic models were deployed in a separate Foundry project and resource group specifically to keep the proof of concept isolated. The usage was limited to model inference.

I checked the detailed metrics for that day. They showed about 33.25 million prompt tokens read from cache, 804K tokens written with a one-hour TTL, and 2.66 million written with a five-minute TTL, alongside 359.5K input tokens, 434.4K output tokens, and 2,129 calls. The heavy cache activity explains the higher charge. I’m glad we caught it before expanding the proof of concept.