Rohit Swami
India Resume ↗

Writing · Part 107 · ASVAB · CitizenOz · 6 min read

Marked like the real thing

Three exam-prep apps, for the FAA drone licence, the US military's enlistment test and the Australian citizenship test, run on one engine. What it takes for practice to predict the real exam honestly, and why the readiness score is hard to flatter.

Part 107 Test prepares people for the FAA's drone pilot knowledge test. ASVAB Test is for the US military's enlistment test, and CitizenOz is for the Australian citizenship test. Three very different exams, one engine. The engine is pure, tested TypeScript: it builds tests, marks them, schedules reviews, predicts a score and merges progress between devices. A new exam is mostly a new content pack, its format, its questions and its study notes, and everything else is shared.

Practice apps have a common way of failing. They make you feel ready. Score 90% on questions you've already seen and the app congratulates you, and then the real test, with its own weighting and its own rules, disagrees. Most of the engine exists to not do that.

1A mock exam shaped like the exam

The FAA publishes how its 60 questions are spread: 15 to 25% on regulations, 15 to 25% on airspace, 11 to 16% on weather, 7 to 11% on loading and performance, and 35 to 45% on operations. A mock exam should have the same shape, so the engine gives each area a share of the questions in proportion to the middle of its range.

The shares come out fractional, and rounding each one on its own doesn't add up. For the 20-question diagnostic the exact shares are 3.9, 3.9, 2.6, 1.8 and 7.8, and rounding them gives 21 questions. So the engine uses largest-remainder rounding, the method some countries use to share out parliamentary seats: every area gets the whole part of its share, and the questions left over go to the areas with the largest fractions. The total is always exact.

questions

rounded one by one –largest remainder –

Fig. 1 Each area's exact share is the thin bar; squares are the questions it gets. Blue squares are leftovers handed to the largest fractions. A red outline is a question that rounding each share on its own would add, and a red cross one it would drop.

Within each area the least-practised questions come first, so taking mock exams again and again walks you through the whole bank instead of showing you the same favourites. The answer choices are shuffled as well, unless a question's choices only make sense in their order, so nobody passes by remembering that the answer was the second one.

2Pass or fail the way the real one decides

Marking has to follow the real rules, including the unusual ones. The Australian citizenship test has 20 questions and a pass mark of 75%, but five of the questions are about Australian values, and you have to get all five right. Nineteen out of twenty, with the one mistake on a values question, is a fail. An app that only reports a percentage would call that a comfortable pass.

score –result –

Fig. 2 Click any answer to flip it between right and wrong, or pick an example sheet. The verdict comes from the same two checks CitizenOz runs: the score against the pass mark, and the values questions.

A smaller detail in the same function: the pass check compares whole numbers, correct answers times 100 against the pass mark times the number of questions, instead of comparing a rounded percentage. A percentage that displays as 70 can be 69.6, and rounding should never be what turns a fail into a pass.

3A readiness score you can't flatter

The number people look at most is the predicted score, so it's the one that has to be hardest to fool. Two things keep it honest.

The first is smoothing. One correct answer is not 100% mastery, so an area's accuracy is (correct + 1) / (answered + 2), which starts in the middle and only moves to the extremes with evidence. The second is coverage. Questions you haven't studied count at 45%, a little better than guessing one of three choices, since some answers can be reasoned out. Each area's prediction mixes your accuracy on what you've answered with that baseline for what you haven't, in proportion to how much of the area you've covered, and the areas are then weighted like the real exam.

accuracy
answered

predicted score –ready after –

Fig. 3 The Part 107 bank of 460 questions, treated as one area to keep the picture simple; the app does this per area and weights the areas like the exam. The dashed line is smoothed accuracy on the questions answered, and the solid line is the predicted score.

The effect is that being right isn't enough. You also have to have seen the material. And "ready" asks for a predicted 80% even though the Part 107 pass mark is 70%, because a prediction built from a question bank, however good the bank, is still a prediction.

4Short intervals, on purpose

Reviews are scheduled with Leitner boxes, the simplest kind of spaced repetition. A correct answer moves a question up a box and a miss sends it back to the first. The usual intervals stretch out to months, but exam prep lasts weeks, so these stop at a month:

// a miss comes back within minutes, a mastered question in a month
export const BOX_INTERVALS = [10 * MINUTE, 1 * DAY, 3 * DAY, 7 * DAY, 14 * DAY, 30 * DAY];
export const MAX_BOX = BOX_INTERVALS.length - 1;

export function review(stat, correct, now) {             // shortened
  const box = correct ? Math.min((stat?.box ?? 0) + 1, MAX_BOX) : 0;
  return { ...stat, box, lastSeen: now, due: now + BOX_INTERVALS[box] };
}

5Estimates that say they're estimates

The ASVAB's headline number, the AFQT, is a percentile against a national sample, built from four subtests: arithmetic reasoning, mathematics knowledge, and twice a verbal score that pools word knowledge and paragraph comprehension. The real calculation runs on standard scores an app can't see. So ASVAB Test weights your answers the same way, AR + MK + 2×VE, and maps the share you got right onto a percentile through anchor points taken from published score distributions: guessing sits at the bottom, about 65% correct is the median, and a clean sheet reads 99. The screen calls it an estimate for study, not an official score, because that is what it is.

6Sync that doesn't double count

Progress follows a signed-in learner between devices, and merging it is full of small traps. The study calendar is one. Each device counts the questions answered per day, but once they've synced, a device can't tell its own answers from the ones it was sent, so adding the counts would double them on every sync. The merge keeps the larger count for each day instead, which never overstates a day and never loses one. A saved question leaves a tombstone when it's un-saved, so the latest choice wins on every device instead of the question coming back. Neither of these shows up in a demo. Both would show up in the first week of real use.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.