Writing · Cloud infrastructure · 8 min read
Spot instances without the 3am page
Spare cloud capacity is cheap until it's taken back. Interruptions, checkpoints, retries and the spot/on-demand mix behind a batch pipeline bill that came down by 28%.
AWS sells its spare capacity as Spot Instances, at up to 90% off the on-demand price. The catch is in the name: when AWS needs the capacity back, it takes it, with two minutes' warning. For a web server that's a problem. For a batch job it's a trade, and whether it's a good one depends almost entirely on what happens to the work in progress.
I designed AWS Batch pipelines that run more than ten thousand jobs a month on a mix of spot and on-demand capacity, and cut their infrastructure cost by 28%. The discount is what saves the money. Making interruptions boring is what lets you keep it.
1What an interruption actually costs
When a spot instance is reclaimed, the job on it dies, and whatever it computed since it last saved its progress is gone. The cost of an interruption isn't the instance. It's the wasted work, plus the time it takes to get going again. A six-hour job that saves nothing, interrupted in its fifth hour, has thrown away five hours.
That makes checkpointing the most important property of any job you want to run on spot. A job that writes its progress somewhere durable every few minutes, and can resume from it, loses at most those few minutes. Interruptions go from catastrophic to mildly annoying.
work redone –last job done at –cost vs on-demand –
The figure also shows the less obvious cost. Lost work has to be run again, so an interrupted job finishes later than its runtime suggests. If the deadline matters, that tail is what you plan for, not the average.
2Two minutes' warning
A reclaim doesn't arrive completely unannounced. Two minutes before it takes an instance back, EC2 posts an interruption notice, which shows up in the instance's metadata and as an EventBridge event. A job that watches for it can spend those two minutes writing one last checkpoint, and lose seconds of work instead of everything since its last scheduled save.
But two minutes is a deadline, and a save that doesn't fit inside it can do real damage if it was writing over the previous checkpoint. A checkpoint cut off halfway is neither the old state nor the new one, and the attempt that resumes from it either crashes or, worse, carries on from garbage. So every save goes to a new file first, and only a complete one takes the old one's place. A rename within one file system is atomic, and an S3 object only becomes visible once its upload has finished, so whatever reads the checkpoint sees the old version or the new one, never half of each.
work lost –next attempt resumes from –
def save_checkpoint(state, path="checkpoint.pkl"):
tmp = path + ".tmp"
with open(tmp, "wb") as f:
pickle.dump(state, f)
f.flush()
os.fsync(f.fileno()) # really on disk before it replaces anything
os.replace(tmp, path) # atomic: a reader gets the old file or the new one
s3.upload_file(path, BUCKET, f"checkpoints/{job_id}.pkl") # visible only once complete
EC2 sometimes sends an earlier and vaguer signal too, a rebalance recommendation, when an instance is at elevated risk of being reclaimed. It's a good moment for a checkpoint that isn't due yet. Neither signal replaces the schedule, though: the warning only helps if the final save fits in it, and the schedule is what bounds the loss when it doesn't.
3Retry the right failures
AWS Batch will retry a failed job, but a blanket retry is the wrong tool. A job that fails because its input is malformed will fail the same way three more times, at full price. A job that fails because its host was reclaimed should go straight back in the queue. Batch lets you tell those apart by the reason a job ended:
"retryStrategy": {
"attempts": 3,
"evaluateOnExit": [
{ "onStatusReason": "Host EC2*", "action": "RETRY" },
{ "onReason": "*", "action": "EXIT" }
]
}
That retries jobs whose host went away and fails everything else immediately, where a person can look at it. On the capacity side, a job queue can list a spot compute environment first and an on-demand one second, so that work overflows to on-demand when spot capacity runs short instead of waiting for it.
4Safe to run twice
A retry runs the job again, from the top or from its last checkpoint, and that has a consequence that's easy to miss: everything the first attempt wrote before it died is still there. A job that appends to its output appends the same rows a second time. A job that names its output after the attempt, with a timestamp or a random ID, leaves the first attempt's half-finished files lying next to the second attempt's complete ones. Either way, whatever reads the output next sees some rows twice and has no way to tell.
The fix is to make writing idempotent. Name each output after what it contains rather than after the attempt that produced it, so a second attempt overwrites the first instead of adding to it. Then write a marker last, an empty _SUCCESS object, and have readers wait for it. Until the marker exists a reader sees nothing at all, rather than a partial result it could mistake for the whole thing.
a reader sees –duplicates –reading halfway –
5Don't put every job in one pool
Spot capacity isn't one big pool. It's many: each instance type in each availability zone is its own pool, with its own supply and demand. Interruptions are correlated inside a pool, because when AWS needs that capacity back it tends to need a lot of it. Ask for exactly one instance type, and a single reclamation can take a large part of your fleet at once.
Asking for many interchangeable types spreads that risk. Batch's capacity-optimised allocation strategies choose from the pools with the most spare capacity, which are the least likely to be reclaimed. Most batch jobs don't care whether they run on one instance family or its neighbour, so the flexibility costs almost nothing.
last reclamation took –worst so far –
In Batch, this is a property of the compute environment. List instance families rather than single sizes, give it subnets in several availability zones, and let the allocation strategy choose among the pools:
"computeResources": {
"type": "SPOT",
"allocationStrategy": "SPOT_PRICE_CAPACITY_OPTIMIZED",
"instanceTypes": ["c6i", "c6a", "c5", "c5a", "m6i", "m6a", "m5", "m5a"],
"subnets": ["subnet-zone-a", "subnet-zone-b", "subnet-zone-c"],
"minvCpus": 0,
"maxvCpus": 2048
}
SPOT_PRICE_CAPACITY_OPTIMIZED looks for pools that are both unlikely to be reclaimed and cheap, where the older SPOT_CAPACITY_OPTIMIZED considers only the first. And with minvCpus at zero, the environment shrinks to nothing when the queue is empty, which is a saving of its own.
6Smaller pieces
A checkpoint limits how much work one interruption can destroy. So does the size of the job. A run that processes a thousand samples in a single job is one long thing to lose. As an array job of a thousand children, each child processes one sample in a few minutes, and an interruption costs one child its few minutes. Batch tells each child its position in an environment variable, so picking the work is a single line:
i = int(os.environ["AWS_BATCH_JOB_ARRAY_INDEX"]) # 0 … size − 1, one per child job
sample = manifest[i]
There's a floor, because every job pays a fixed price to start: being placed, pulling its image, fetching its inputs. Cut the pieces so small that start-up is a large share of their runtime and the overhead eats the discount. The sweet spot is a job long enough that starting it is cheap and short enough that losing it isn't.
7The mix
None of this means everything belongs on spot. Some work is better off on-demand: long jobs that can't checkpoint, anything on a hard deadline, and the last attempt of a job that has already been interrupted twice. The right mix is a portfolio, and the price of the cheap part is engineering: checkpoints, outputs that are safe to write twice, and retry rules that know a reclaimed host from a bug.
cheapest mix –saving there –
The curve has the same shape for most batch workloads. Moving the first jobs to spot saves money almost for free. Each further step saves a little less, because the jobs left are the ones that hurt most to interrupt, and past some point the expected rework costs more than the discount. Where that point sits depends on the jobs and on how often capacity is reclaimed, which is why it's worth measuring rather than guessing.
8What wakes someone up
On spot, interruptions are weather, and paging someone about the weather is how alerts end up ignored. What deserves a page is what a retry can't fix: a job that has used up its attempts, a rise in jobs failing with errors of their own rather than lost hosts, or the oldest job in the queue waiting far longer than usual, which means capacity has run short and the on-demand fallback isn't absorbing it. Everything else belongs on a dashboard that someone reads over coffee.
9Making it boring
The test of a spot setup isn't the bill. It's whether anyone notices when capacity is reclaimed at 3am. If jobs checkpoint, retries know what to retry, the fleet is spread across pools and the deadline-critical work sits on capacity that can't be taken away, an interruption is a line in a log rather than an incident. That's when the discount is really yours.