Rohit Swami
India Resume ↗

Writing · Cloud infrastructure · 8 min read

Spot instances without the 3am page

Spare cloud capacity is cheap until it's taken back. Interruptions, checkpoints, retries and the spot/on-demand mix behind a batch pipeline bill that came down by 28%.

AWS sells its spare capacity as Spot Instances, at up to 90% off the on-demand price. The catch is in the name: when AWS needs the capacity back, it takes it, with two minutes' warning. For a web server that's a problem. For a batch job it's a trade, and whether it's a good one depends almost entirely on what happens to the work in progress.

I designed AWS Batch pipelines that run more than ten thousand jobs a month on a mix of spot and on-demand capacity, and cut their infrastructure cost by 28%. The discount is what saves the money. Making interruptions boring is what lets you keep it.

1What an interruption actually costs

When a spot instance is reclaimed, the job on it dies, and whatever it computed since it last saved its progress is gone. The cost of an interruption isn't the instance. It's the wasted work, plus the time it takes to get going again. A six-hour job that saves nothing, interrupted in its fifth hour, has thrown away five hours.

That makes checkpointing the most important property of any job you want to run on spot. A job that writes its progress somewhere durable every few minutes, and can resume from it, loses at most those few minutes. Interruptions go from catastrophic to mildly annoying.

interruptions

work redone –last job done at –cost vs on-demand –

Fig. 1 Six jobs, six hours of work each, on spot capacity at 35% of the on-demand price. Red marks an interruption, hatching is work lost and redone, and grey is waiting for a new instance. The interruptions are random but repeatable: the same rate replays the same day.

The figure also shows the less obvious cost. Lost work has to be run again, so an interrupted job finishes later than its runtime suggests. If the deadline matters, that tail is what you plan for, not the average.

2Two minutes' warning

A reclaim doesn't arrive completely unannounced. Two minutes before it takes an instance back, EC2 posts an interruption notice, which shows up in the instance's metadata and as an EventBridge event. A job that watches for it can spend those two minutes writing one last checkpoint, and lose seconds of work instead of everything since its last scheduled save.

But two minutes is a deadline, and a save that doesn't fit inside it can do real damage if it was writing over the previous checkpoint. A checkpoint cut off halfway is neither the old state nor the new one, and the attempt that resumes from it either crashes or, worse, carries on from garbage. So every save goes to a new file first, and only a complete one takes the old one's place. A rename within one file system is atomic, and an S3 object only becomes visible once its upload has finished, so whatever reads the checkpoint sees the old version or the new one, never half of each.

work lost –next attempt resumes from –

Fig. 2 One job, checkpointing every twenty minutes, on an instance reclaimed at 14:35. Blue is work, grey is writing a checkpoint, and hatching is work the next attempt will have to do again. Try saving on the warning with the slow save, and then the other way of writing it.
def save_checkpoint(state, path="checkpoint.pkl"):
    tmp = path + ".tmp"
    with open(tmp, "wb") as f:
        pickle.dump(state, f)
        f.flush()
        os.fsync(f.fileno())        # really on disk before it replaces anything
    os.replace(tmp, path)           # atomic: a reader gets the old file or the new one
    s3.upload_file(path, BUCKET, f"checkpoints/{job_id}.pkl")   # visible only once complete

EC2 sometimes sends an earlier and vaguer signal too, a rebalance recommendation, when an instance is at elevated risk of being reclaimed. It's a good moment for a checkpoint that isn't due yet. Neither signal replaces the schedule, though: the warning only helps if the final save fits in it, and the schedule is what bounds the loss when it doesn't.

3Retry the right failures

AWS Batch will retry a failed job, but a blanket retry is the wrong tool. A job that fails because its input is malformed will fail the same way three more times, at full price. A job that fails because its host was reclaimed should go straight back in the queue. Batch lets you tell those apart by the reason a job ended:

"retryStrategy": {
  "attempts": 3,
  "evaluateOnExit": [
    { "onStatusReason": "Host EC2*", "action": "RETRY" },
    { "onReason": "*", "action": "EXIT" }
  ]
}

That retries jobs whose host went away and fails everything else immediately, where a person can look at it. On the capacity side, a job queue can list a spot compute environment first and an on-demand one second, so that work overflows to on-demand when spot capacity runs short instead of waiting for it.

4Safe to run twice

A retry runs the job again, from the top or from its last checkpoint, and that has a consequence that's easy to miss: everything the first attempt wrote before it died is still there. A job that appends to its output appends the same rows a second time. A job that names its output after the attempt, with a timestamp or a random ID, leaves the first attempt's half-finished files lying next to the second attempt's complete ones. Either way, whatever reads the output next sees some rows twice and has no way to tell.

The fix is to make writing idempotent. Name each output after what it contains rather than after the attempt that produced it, so a second attempt overwrites the first instead of adding to it. Then write a marker last, an empty _SUCCESS object, and have readers wait for it. Until the marker exists a reader sees nothing at all, rather than a partial result it could mistake for the whole thing.

a reader sees –duplicates –reading halfway –

Fig. 3 Four parts of 250 rows each. The first attempt writes two parts and is reclaimed; the retry writes all four. Red marks rows the bucket now holds twice. "Reading halfway" is what a reader would have got between the two attempts.

5Don't put every job in one pool

Spot capacity isn't one big pool. It's many: each instance type in each availability zone is its own pool, with its own supply and demand. Interruptions are correlated inside a pool, because when AWS needs that capacity back it tends to need a lot of it. Ask for exactly one instance type, and a single reclamation can take a large part of your fleet at once.

Asking for many interchangeable types spreads that risk. Batch's capacity-optimised allocation strategies choose from the pools with the most spare capacity, which are the least likely to be reclaimed. Most batch jobs don't care whether they run on one instance family or its neighbour, so the flexibility costs almost nothing.

last reclamation took –worst so far –

Fig. 4 Rows are instance types, columns are availability zones, and every cell is a separate spot pool. Every few seconds one pool is reclaimed. The fleet is the same size either way; only how much of it shares a fate changes.

In Batch, this is a property of the compute environment. List instance families rather than single sizes, give it subnets in several availability zones, and let the allocation strategy choose among the pools:

"computeResources": {
  "type": "SPOT",
  "allocationStrategy": "SPOT_PRICE_CAPACITY_OPTIMIZED",
  "instanceTypes": ["c6i", "c6a", "c5", "c5a", "m6i", "m6a", "m5", "m5a"],
  "subnets": ["subnet-zone-a", "subnet-zone-b", "subnet-zone-c"],
  "minvCpus": 0,
  "maxvCpus": 2048
}

SPOT_PRICE_CAPACITY_OPTIMIZED looks for pools that are both unlikely to be reclaimed and cheap, where the older SPOT_CAPACITY_OPTIMIZED considers only the first. And with minvCpus at zero, the environment shrinks to nothing when the queue is empty, which is a saving of its own.

6Smaller pieces

A checkpoint limits how much work one interruption can destroy. So does the size of the job. A run that processes a thousand samples in a single job is one long thing to lose. As an array job of a thousand children, each child processes one sample in a few minutes, and an interruption costs one child its few minutes. Batch tells each child its position in an environment variable, so picking the work is a single line:

i = int(os.environ["AWS_BATCH_JOB_ARRAY_INDEX"])   # 0 … size − 1, one per child job
sample = manifest[i]

There's a floor, because every job pays a fixed price to start: being placed, pulling its image, fetching its inputs. Cut the pieces so small that start-up is a large share of their runtime and the overhead eats the discount. The sweet spot is a job long enough that starting it is cheap and short enough that losing it isn't.

7The mix

None of this means everything belongs on spot. Some work is better off on-demand: long jobs that can't checkpoint, anything on a hard deadline, and the last attempt of a job that has already been interrupted twice. The right mix is a portfolio, and the price of the cheap part is engineering: checkpoints, outputs that are safe to write twice, and retry rules that know a reclaimed host from a bug.

interruptions

cheapest mix –saving there –

Fig. 5 A model, not a measurement. Jobs are moved to spot in order of how cheaply they recover from an interruption, and each one costs the spot price plus its expected rework. Past the lowest point, the next job's rework costs more than the discount saves.

The curve has the same shape for most batch workloads. Moving the first jobs to spot saves money almost for free. Each further step saves a little less, because the jobs left are the ones that hurt most to interrupt, and past some point the expected rework costs more than the discount. Where that point sits depends on the jobs and on how often capacity is reclaimed, which is why it's worth measuring rather than guessing.

8What wakes someone up

On spot, interruptions are weather, and paging someone about the weather is how alerts end up ignored. What deserves a page is what a retry can't fix: a job that has used up its attempts, a rise in jobs failing with errors of their own rather than lost hosts, or the oldest job in the queue waiting far longer than usual, which means capacity has run short and the on-demand fallback isn't absorbing it. Everything else belongs on a dashboard that someone reads over coffee.

9Making it boring

The test of a spot setup isn't the bill. It's whether anyone notices when capacity is reclaimed at 3am. If jobs checkpoint, retries know what to retry, the fleet is spread across pools and the deadline-critical work sits on capacity that can't be taken away, an interruption is a line in a log rather than an incident. That's when the discount is really yours.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.