How a request is charged
Each request is metered into lines, one per kind of usage: prompt, completion, reasoning, cache reads, cache writes, images and the like. Every line carries its quantity, the unit price, any discount, and the price version it was billed at. The request’s cost is the sum of its lines, computed exactly and rounded once.
The unit price is copied onto the line when the request is metered. Re-reading an old charge therefore reproduces it even after the catalogue’s prices change.
Which endpoint served a request, at which price version, and every line with its cost: enough to recompute the charge yourself.
On your own provider key
With your own key, your provider bills you for the tokens directly. Those lines are recorded at zero cost with the provider’s list price alongside, so spend ceilings, budgets and Activity measure what the request cost while nothing is charged twice. Cached tokens are valued at the provider’s cache rate where it publishes one.
Nothing produced, nothing charged
When a response produced no output and gave no finish reason, or ended in a provider error, its prompt, completion and reasoning lines are not charged. Auxiliary services the provider billed regardless, such as web search, still appear as their own lines.
Holds
Before dispatch, a hold for the request’s largest possible cost is placed against the balance and against every ceiling that applies. At settlement the hold becomes the metered cost and the rest is released at once; a request that fails before a provider answers releases its hold in full. This is what stops concurrent requests from together spending past a limit that each of them alone would have respected.
Exports
A month of requests as CSV, for one client workspace or the whole organisation, with the same costs the Activity page shows. Formula-like cells are escaped, so the file opens safely in a spreadsheet.