Writing · Cloud infrastructure · 5 min read
Spot instances without the 3am page
Spare cloud capacity is cheap until it's taken back. Interruptions, checkpoints, retries and the spot/on-demand mix behind a batch pipeline bill that came down by 28%.
AWS sells its spare capacity as Spot Instances, at up to 90% off the on-demand price. The catch is in the name: when AWS needs the capacity back, it takes it, with two minutes' warning. For a web server that's a problem. For a batch job it's a trade, and whether it's a good one depends almost entirely on what happens to the work in progress.
I designed AWS Batch pipelines that run more than ten thousand jobs a month on a mix of spot and on-demand capacity, and cut their infrastructure cost by 28%. The discount is what saves the money. Making interruptions boring is what lets you keep it.
1What an interruption actually costs
When a spot instance is reclaimed, the job on it dies, and whatever it computed since it last saved its progress is gone. The cost of an interruption isn't the instance. It's the wasted work, plus the time it takes to get going again. A six-hour job that saves nothing, interrupted in its fifth hour, has thrown away five hours.
That makes checkpointing the most important property of any job you want to run on spot. A job that writes its progress somewhere durable every few minutes, and can resume from it, loses at most those few minutes. Interruptions go from catastrophic to mildly annoying.
work redone –last job done at –cost vs on-demand –
The figure also shows the less obvious cost. Lost work has to be run again, so an interrupted job finishes later than its runtime suggests. If the deadline matters, that tail is what you plan for, not the average.
2Retry the right failures
AWS Batch will retry a failed job, but a blanket retry is the wrong tool. A job that fails because its input is malformed will fail the same way three more times, at full price. A job that fails because its host was reclaimed should go straight back in the queue. Batch lets you tell those apart by the reason a job ended:
"retryStrategy": {
"attempts": 3,
"evaluateOnExit": [
{ "onStatusReason": "Host EC2*", "action": "RETRY" },
{ "onReason": "*", "action": "EXIT" }
]
}
That retries jobs whose host went away and fails everything else immediately, where a person can look at it. On the capacity side, a job queue can list a spot compute environment first and an on-demand one second, so that work overflows to on-demand when spot capacity runs short instead of waiting for it.
3Don't put every job in one pool
Spot capacity isn't one big pool. It's many: each instance type in each availability zone is its own pool, with its own supply and demand. Interruptions are correlated inside a pool, because when AWS needs that capacity back it tends to need a lot of it. Ask for exactly one instance type, and a single reclamation can take a large part of your fleet at once.
Asking for many interchangeable types spreads that risk. Batch's capacity-optimised allocation strategies choose from the pools with the most spare capacity, which are the least likely to be reclaimed. Most batch jobs don't care whether they run on one instance family or its neighbour, so the flexibility costs almost nothing.
last reclamation took –worst so far –
4The mix
None of this means everything belongs on spot. Some work is better off on-demand: long jobs that can't checkpoint, anything on a hard deadline, and the last attempt of a job that has already been interrupted twice. The right mix is a portfolio, and the price of the cheap part is engineering: checkpoints, outputs that are safe to write twice, and retry rules that know a reclaimed host from a bug.
cheapest mix –saving there –
The curve has the same shape for most batch workloads. Moving the first jobs to spot saves money almost for free. Each further step saves a little less, because the jobs left are the ones that hurt most to interrupt, and past some point the expected rework costs more than the discount. Where that point sits depends on the jobs and on how often capacity is reclaimed, which is why it's worth measuring rather than guessing.
5Making it boring
The test of a spot setup isn't the bill. It's whether anyone notices when capacity is reclaimed at 3am. If jobs checkpoint, retries know what to retry, the fleet is spread across pools and the deadline-critical work sits on capacity that can't be taken away, an interruption is a line in a log rather than an incident. That's when the discount is really yours.