Four bills, one place to read them

A cost tool, a Kubernetes dashboard, an ML platform, and whatever you built for agents. Four tools that never see each other, so nobody can say what the whole estate costs. CoverOps sits over all four.

  1. “Your job is running. Your GPUs are idle.”

    Teams watch whether a job is alive. Nobody watches accelerator utilisation. So eight H100s sit at 30% for a fortnight and the bill looks completely normal.

    See the arithmetic
    • Runs under 35% utilisation get flagged as starved, not compute-bound
    • Only the idle share is priced as recoverable
    • Spot runs without checkpointing, where one preemption bins the epochs you paid for
    • Endpoints under 25% utilisation, holding a GPU warm between requests
  2. “A blocked run costs nothing. That is the whole saving.”

    Cached tokens cost a tenth of fresh ones. Most teams have caching off, or a prompt prefix that shifts just enough to miss it every time.

    See the arithmetic
    • Every run metered against the catalog. Fresh, cached and output priced apart
    • Frontier models doing work a cheaper tier handles at 96% success
    • Caching off, or hitting under a quarter of the time
    • Looping and timed-out runs that billed in full for nothing
    • Guardrails run before billing. A blocked run is a clean saving
  3. “The familiar one. Real money, but everybody already sells it.”

    Unattached volumes. Oversized instances. Node pools still scaled for last year's launch. Every finding arrives with the change attached.

    • Read-only connect to AWS, Azure, GCP, DigitalOcean and OCI
    • Cost and utilisation per resource, on the schedule you set
    • Findings that carry the fix, not just the observation
    • Scale, cordon and drain node pools from the console
  4. “One outage should page once. Not three times.”

    Every metric is measured against its own rolling baseline, so a signal has to be unusual to register. Related anomalies collapse into one incident with a probable cause attached.

    • A rolling baseline, not a threshold nobody tuned
    • Anomalies inside five minutes become one incident
    • Playbooks propose the fix, or apply it with the same audit trail
    • Automation can only do what an operator could

Every screen in the console

These are the console's own menu labels, so a demo uses the same words.

Overview
Spend under management, what is still recoverable, live CPU, memory, latency and GPU vitals over one stream, and an activity rail.
Cloud accounts
Read-only connect to AWS, Azure, GCP, DigitalOcean, OCI. Inventory reconciled on a schedule, per-resource cost and utilisation, idle findings with the change attached.
Kubernetes
Register clusters, scale node pools, cordon and drain nodes, deploy, scale, restart and delete workloads, live capacity meters.
Containers
Every pod with live CPU and memory, restart counts, restart, stop and evict controls, and streaming logs.
Alerts
A 26-metric rule builder spanning infrastructure, cost, security, GPU and agent metrics, with duration and cooldown windows and routing to email, Slack, webhooks, PagerDuty or SMS.
MLOps
Model registry with drift tracking, GPU training runs priced per accelerator-hour, inference endpoints with scale-to-zero, and a waste report that names the recoverable dollars.
AIOps
Baseline-relative anomaly detection, correlation into single incidents with a probable cause, and playbooks that suggest or auto-apply the remediation.
AgentOps
Agent registry, per-run token metering priced from the model catalog, budget and token guardrails that block before billing, and a savings report.
Delivery
Apps with verified GitOps deployments and one-click rollback, plus a security centre for supply-chain, container, IAM and configuration findings.
Account
Cost and carbon reporting, team and role management, API keys, and a full audit log of everything the platform did.

Three rules it is built around

Your account, not ours
A control plane, not a host. Every resource stays in your cloud account, reached through a role you wrote and can delete.
One live stream
Sync results, metric ticks, firing alerts and deployments all arrive on one event stream. Nothing polls, nothing goes stale.
Automation cannot outrank you
Playbooks call the same functions the console does, so automation can never do something you could not do by hand.

Where the build actually is: discovery synthesises inventory from a catalog today. Everything downstream of it runs for real. Swapping in live SDK clients is the next big job, and it touches one function per domain. The roadmap says where that sits.

See what it finds on your estate

We are onboarding a few design partners on real production workloads. Tell us what you run and we will show you what it flags.

Console opening soon

The console isn't open to the public yet.

We are letting teams in a few at a time while design partners run it on real estates. Leave your details and you will hear from a person, not a drip campaign, when your seat is ready.

Opens your mail app with this filled in. Nothing sends until you do.