A run that never happens costs nothing

Most cost tools estimate a saving after the fact. CoverOps guardrails decide before a run is billed, so there is no model assumption to argue about. The run simply did not happen.

Model catalog, USD per million tokens
ModelInputCachedOutput
Claude Opus 4.8$5.00$0.50$25.00
Claude Sonnet 5$3.00$0.30$15.00
Claude Haiku 4.5$1.00$0.10$5.00

Cached input is a tenth of fresh input on every model, which is why caching is usually the biggest line in a savings report.

Where token money goes

On the example estate in the console, these three patterns add up to 40% of agent spend.

  1. Caching off

    Every run pays for the same system prompt again.

    Cached input costs a tenth of fresh input. An agent with a stable prompt should hit the cache on nearly every run. Most miss, because caching is off or the prefix shifts by a few bytes each time.

    How it is caught
    Cache hit rate tracked per agent. Off, or under 25%, and the recoverable share of that agent's spend is priced.
    What fixes it
    Make the prompt prefix byte-identical between runs, then switch caching on.
  2. Over-modelled

    A frontier model doing work a cheaper one handles.

    An agent succeeding 96% of the time on a frontier model is telling you the task does not need one. The price gap between tiers is most of the bill.

    How it is caught
    Success rate compared against model tier. High success on a frontier model raises a routing finding, priced from the catalog difference.
    What fixes it
    Route to the balanced model and keep the frontier model for the runs that actually fail.
  3. Looping runs

    Billed in full for an answer nobody got.

    A run that loops to its step cap, or times out waiting on a tool, still bills every token it spent getting there.

    How it is caught
    Runs ending in a loop or timeout are counted over a rolling week and priced against the agent's spend.
    What fixes it
    Lower the step cap and give the agent loop an explicit stopping condition.

Seven guardrails, checked before billing

A run that trips one is recorded and shown to you. It costs nothing, because it never reached the model.

budget_cap
A hard monthly ceiling. The agent stops before it overruns.
token_cap
A per-run token ceiling, so one strange input cannot cost a fortune.
step_cap
The most reasoning steps a run gets before it is ended.
rate_limit
Runs per interval, so a retry storm does not turn into a bill.
latency_cap
Abandon a run that has clearly stalled instead of paying it out.
pii_filter
Block runs carrying data that should never reach a model provider.
tool_allowlist
Limit which tools an agent can call at all.

Tokens, metered like any other resource

Fresh, cached and output tokens are priced separately on every run, against the model that actually served it. Agent cost stops being one line on a provider invoice.

incident-summarizer Opus 4.8, 82.7% cache miss
$327/mo recoverable
fix: stabilise the prompt prefix, enable caching

Console opening soon

The console isn't open to the public yet.

We are letting teams in a few at a time while design partners run it on real estates. Leave your details and you will hear from a person, not a drip campaign, when your seat is ready.

Opens your mail app with this filled in. Nothing sends until you do.

Training your own models? GPU waste runs on the same engine, with bigger numbers.