A run that never happens costs nothing
Most cost tools estimate a saving after the fact. CoverOps guardrails decide before a run is billed, so there is no model assumption to argue about. The run simply did not happen.
| Model | Input | Cached | Output |
|---|---|---|---|
| Claude Opus 4.8 | $5.00 | $0.50 | $25.00 |
| Claude Sonnet 5 | $3.00 | $0.30 | $15.00 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
Cached input is a tenth of fresh input on every model, which is why caching is usually the biggest line in a savings report.
Where token money goes
On the example estate in the console, these three patterns add up to 40% of agent spend.
- Caching off
Every run pays for the same system prompt again.
Cached input costs a tenth of fresh input. An agent with a stable prompt should hit the cache on nearly every run. Most miss, because caching is off or the prefix shifts by a few bytes each time.
- Over-modelled
A frontier model doing work a cheaper one handles.
An agent succeeding 96% of the time on a frontier model is telling you the task does not need one. The price gap between tiers is most of the bill.
- Looping runs
Billed in full for an answer nobody got.
A run that loops to its step cap, or times out waiting on a tool, still bills every token it spent getting there.
Seven guardrails, checked before billing
A run that trips one is recorded and shown to you. It costs nothing, because it never reached the model.
budget_cap- A hard monthly ceiling. The agent stops before it overruns.
token_cap- A per-run token ceiling, so one strange input cannot cost a fortune.
step_cap- The most reasoning steps a run gets before it is ended.
rate_limit- Runs per interval, so a retry storm does not turn into a bill.
latency_cap- Abandon a run that has clearly stalled instead of paying it out.
pii_filter- Block runs carrying data that should never reach a model provider.
tool_allowlist- Limit which tools an agent can call at all.
Tokens, metered like any other resource
Fresh, cached and output tokens are priced separately on every run, against the model that actually served it. Agent cost stops being one line on a provider invoice.
incident-summarizer Opus 4.8, 82.7% cache miss $327/mo recoverable fix: stabilise the prompt prefix, enable caching
Training your own models? GPU waste runs on the same engine, with bigger numbers.