Rohit Swami
India Resume ↗

Writing · Cloud infrastructure · 5 min read

Spot instances without the 3am page

Spare cloud capacity is cheap until it's taken back. Interruptions, checkpoints, retries and the spot/on-demand mix behind a batch pipeline bill that came down by 28%.

AWS sells its spare capacity as Spot Instances, at up to 90% off the on-demand price. The catch is in the name: when AWS needs the capacity back, it takes it, with two minutes' warning. For a web server that's a problem. For a batch job it's a trade, and whether it's a good one depends almost entirely on what happens to the work in progress.

I designed AWS Batch pipelines that run more than ten thousand jobs a month on a mix of spot and on-demand capacity, and cut their infrastructure cost by 28%. The discount is what saves the money. Making interruptions boring is what lets you keep it.

1What an interruption actually costs

When a spot instance is reclaimed, the job on it dies, and whatever it computed since it last saved its progress is gone. The cost of an interruption isn't the instance. It's the wasted work, plus the time it takes to get going again. A six-hour job that saves nothing, interrupted in its fifth hour, has thrown away five hours.

That makes checkpointing the most important property of any job you want to run on spot. A job that writes its progress somewhere durable every few minutes, and can resume from it, loses at most those few minutes. Interruptions go from catastrophic to mildly annoying.

interruptions

work redone –last job done at –cost vs on-demand –

Fig. 1 Six jobs, six hours of work each, on spot capacity at 35% of the on-demand price. Red marks an interruption, hatching is work lost and redone, and grey is waiting for a new instance. The interruptions are random but repeatable: the same rate replays the same day.

The figure also shows the less obvious cost. Lost work has to be run again, so an interrupted job finishes later than its runtime suggests. If the deadline matters, that tail is what you plan for, not the average.

2Retry the right failures

AWS Batch will retry a failed job, but a blanket retry is the wrong tool. A job that fails because its input is malformed will fail the same way three more times, at full price. A job that fails because its host was reclaimed should go straight back in the queue. Batch lets you tell those apart by the reason a job ended:

"retryStrategy": {
  "attempts": 3,
  "evaluateOnExit": [
    { "onStatusReason": "Host EC2*", "action": "RETRY" },
    { "onReason": "*", "action": "EXIT" }
  ]
}

That retries jobs whose host went away and fails everything else immediately, where a person can look at it. On the capacity side, a job queue can list a spot compute environment first and an on-demand one second, so that work overflows to on-demand when spot capacity runs short instead of waiting for it.

3Don't put every job in one pool

Spot capacity isn't one big pool. It's many: each instance type in each availability zone is its own pool, with its own supply and demand. Interruptions are correlated inside a pool, because when AWS needs that capacity back it tends to need a lot of it. Ask for exactly one instance type, and a single reclamation can take a large part of your fleet at once.

Asking for many interchangeable types spreads that risk. Batch's capacity-optimised allocation strategies choose from the pools with the most spare capacity, which are the least likely to be reclaimed. Most batch jobs don't care whether they run on one instance family or its neighbour, so the flexibility costs almost nothing.

last reclamation took –worst so far –

Fig. 2 Rows are instance types, columns are availability zones, and every cell is a separate spot pool. Every few seconds one pool is reclaimed. The fleet is the same size either way; only how much of it shares a fate changes.

4The mix

None of this means everything belongs on spot. Some work is better off on-demand: long jobs that can't checkpoint, anything on a hard deadline, and the last attempt of a job that has already been interrupted twice. The right mix is a portfolio, and the price of the cheap part is engineering: checkpoints, outputs that are safe to write twice, and retry rules that know a reclaimed host from a bug.

interruptions

cheapest mix –saving there –

Fig. 3 A model, not a measurement. Jobs are moved to spot in order of how cheaply they recover from an interruption, and each one costs the spot price plus its expected rework. Past the lowest point, the next job's rework costs more than the discount saves.

The curve has the same shape for most batch workloads. Moving the first jobs to spot saves money almost for free. Each further step saves a little less, because the jobs left are the ones that hurt most to interrupt, and past some point the expected rework costs more than the discount. Where that point sits depends on the jobs and on how often capacity is reclaimed, which is why it's worth measuring rather than guessing.

5Making it boring

The test of a spot setup isn't the bill. It's whether anyone notices when capacity is reclaimed at 3am. If jobs checkpoint, retries know what to retry, the fleet is spread across pools and the deadline-critical work sits on capacity that can't be taken away, an interruption is a line in a log rather than an incident. That's when the discount is really yours.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.