The most expensive thing nobody measures
Every team has a dashboard that says whether the training job is running. Almost none has one that says whether the accelerators are doing anything. The money sits in that gap.
| Accelerator | On demand | Spot |
|---|---|---|
| T4-16GB | $0.53 | $0.16 |
| L4-24GB | $0.98 | $0.31 |
| A10G-24GB | $1.21 | $0.38 |
| A100-40GB | $3.13 | $0.98 |
| A100-80GB | $4.09 | $1.31 |
| H100-80GB | $9.89 | $3.42 |
The price list the control plane costs every run against.
Under 35%: flagged as starved on input, not compute-bound.
Recoverable per year
$339,615
₹2.82 crore
- Node run rate
- $57,758/mo
- Idle accelerator share
- 70%
- Unavoidable-idle discount
- ×0.7
- Recoverable
- $28,301/mo
Arithmetic on our own accelerator price list, using the formula wasteReport() runs in the product. It is not a measured customer outcome. Rupees at ₹83 to the dollar.
Where GPU money goes
Three patterns cover most accelerator overspend, so these are the three the control plane models.
- Starved runs
The job is alive and the accelerators are waiting.
A run under 35% utilisation is not compute-bound. It is waiting on the input pipeline, and the GPUs stay allocated and billed between batches.
- Spot without checkpoints
One preemption throws away every epoch you paid for.
Spot costs about a third of on-demand, which makes it tempting for long runs. Without checkpointing, a single preemption restarts the run from zero.
- Idle serving
Paying for a warm GPU between requests.
An endpoint under 25% utilisation is holding an accelerator for traffic that is not arriving. Overnight and at weekends that is close to pure loss.
A finding names the run and the money
Not “GPU utilisation is low”, which is a ticket someone has to go and investigate. This:
churn-propensity 8× A100-80GB, 31% utilised $8,412/mo recoverable fix: raise batch size or dataloader workers
Shipping agents rather than training models? The same engine prices LLM agent spend.