Writing · YAPD · 15 min read
A WhatsApp export is not a file format
YAPD turns an exported chat into a Wrapped-style recap. Reading a format nobody designed, defining what a conversation is, letting a language model quote people without misquoting them, and sealing everything that leaves the device so that not even the server can read it.
YAPD takes the file WhatsApp gives you when you export a chat and turns it into a recap in the style of Spotify Wrapped: a deck of more than twenty animated cards with totals, top emojis, reply speed, streaks, the moments that mattered and a playful archetype for everyone in the chat. It's live at yapd.in. A paid tier, Encore, adds a seven-chapter essay about the same chat.
Two rules shaped every part of it. The free recap is built entirely in the browser: the chat is never uploaded, so there's nothing on a server to leak, hand over or misuse. And anything that does have to leave the device, a saved recap or an Encore essay, leaves sealed, under a key the server never holds. The first rule made the parser and the statistics harder, because there's no backend to tidy things up afterwards. The second made the product slower to build and much easier to trust. This is how both work, from the first byte of the file to the last row in the database.
1Four formats wearing one name
"Export chat" sounds like a format. It isn't one. An iPhone wraps the chat in a zip archive as _chat.txt, next to any media; Android hands over the text file itself, sometimes renamed. So the first step happens before any parsing. YAPD looks at what was dropped, sniffs it (a zip file starts with the bytes PK, which matters when a platform has stripped the extension), and picks the chat out of the archive: _chat.txt if it's there, otherwise the largest text file. Anything over 50 MB is turned away with a message, rather than with a frozen tab.
Inside the file, the lines differ by platform and by the phone's language settings:
[12/01/24, 9:41:22 PM] Rohit: Bro where are you iPhone
12/01/24, 9:41 PM - Rohit: Bro where are you Android
[12/01/2024, 21:41:22] Rohit: Bro where are you iPhone, 24-hour clock
01.12.2024, 21:41 - Rohit: Bro where are you Android, much of Europe
Then there are the characters nobody can see. Exports scatter right-to-left marks, directional embeddings and isolates (U+200F, U+202A to U+202E, U+2066 to U+2069) and byte-order marks around the names, colons and timestamps, plus non-breaking spaces where you'd expect ordinary ones. None of them show up in a text editor, and every one of them breaks a regular expression that looks perfectly correct. They're stripped before anything is matched. One invisible character, U+200E, is deliberately kept, for a reason that comes up below.
Messages with line breaks in them carry on over lines that have no header of their own. So the parser works line by line: a line that starts with something shaped like a date and a time begins a new message, and any other line belongs to the message before it.
messages –joined to the one before –dropped –
2Is 05/06 the fifth of June?
The hardest line in that list is the one that looks easiest. 12/01/24 is the 12th of January in India and the UK, and the 1st of December in the US, and nothing on the line says which. Guess wrong and every date in the chat is wrong: the busiest month moves, the streak breaks, the anniversary lands in the wrong season.
The file does hold the answer, just not on any single line. A date whose first number is above 12 can only be day-first, and one whose second number is above 12 can only be month-first. So the parser samples the headers and lets the unambiguous ones vote, while the ambiguous ones abstain. If nobody votes, it falls back to day-first, WhatsApp's default in most of the world. And if a single date still makes no sense in the chosen order, that date alone is tried the other way round.
read as –dates right –
That fallback is a guess, and the figure doesn't hide it: a short American chat from the first days of a month really is ambiguous, and nothing in the file can settle it. What matters is that the guess is only made when the evidence is genuinely missing. Anything that runs past the 12th of a month settles itself.
3Messages nobody sent
An export also contains lines WhatsApp wrote itself: the encryption notice, "you were added", someone changing the group's subject. On an iPhone those lines are attributed to the group, so if one of them were read as a message, the group's name would turn up in the recap as a person, with a message count and a personality of its own. Most of them carry a quiet signal: an invisible left-to-right mark, U+200E, at the start of the text, which nobody types by accident. That's the one invisible character the parser keeps until it has used it.
The catch is that media placeholders carry the same mark, "image omitted" and "sticker omitted", and those are real messages from real people with the content left out. Dropping them would make the chattiest sticker-sender look quiet. So a body that starts with the mark is dropped unless it's a media placeholder, and a short list of tightly scoped patterns catches system lines in the exports where the mark is missing.
function splitSenderAndText(rest) {
const idx = rest.indexOf(": ");
if (idx === -1) return null;
const sender = rest.slice(0, idx).trim();
const text = rest.slice(idx + 2);
if (!sender || sender.length > 80) return null;
const injected = text.startsWith(""); // WhatsApp wrote this body itself
if (injected && !isMediaMessage(text)) return null; // a system line, not a person
if (isSystemBody(text)) return null; // the same, for exports without the mark
return { sender, text: text.replace(//g, "") };
}
4What counts as a conversation
Once the messages are clean, the recap is a single pure function from messages to statistics. It runs in the browser on chats of up to about a hundred thousand messages, so every number comes from one linear pass, with no nested loops. The arithmetic is easy. The definitions are where the judgement goes, and each one is a small decision about what people mean.
Take reply speed. A reply is a message whose previous message came from someone else, so a burst of five messages from one person is one turn, not five replies. The time between the two is a reply latency, but only up to 24 hours, because "good morning" the next day isn't really an answer to "good night", and a few of those would wreck any average. Then YAPD reports the median rather than the mean, because one weekend away shouldn't define how fast someone texts back.
Conversations need a definition too. A silence of more than six hours ends one, and that single threshold answers three of the questions people care about most: who starts conversations (the first message after a silence), who ends them (the last message before one), and which days count towards a streak. Night owls are whoever writes most between 11pm and 5am.
conversations –opened by Asha –by Dev –median reply –
None of these definitions is right in any absolute sense. What matters is that each one is written down, applied the same way to every chat, and chosen so that the number on the card means what a person reading it would assume it means.
5A deck that fits the chat
A fixed deck would give a chat with no emojis an emoji slide. So a curator scores about fifteen candidate slides against the chat's own numbers and keeps the most relevant ones, in a fixed order. There's no randomness in it. Open the same recap a year later and you get the same deck, while two different chats come out meaningfully different.
The archetypes are rules over each person's behaviour, not a model's opinion of their messages. That makes them free, instant and identical every time, and it means YAPD jokes about patterns, like yapping, ghosting and late nights, never about what anyone said. The "moments that mattered" slide works the same way. Declarations of love, engagements, pregnancies, new jobs and losses are found with patterns written in each language's own script, for nine languages including Hindi, Spanish, Tagalog and Mandarin, so detection runs on the device and gives the same answer every time. It's tuned to prefer a false positive to a missed moment, and the slide frames what it finds as things that were said, not as verdicts.
6Quoting people exactly
The paid tier, Encore, writes a seven-chapter essay about a chat, with the exchanges that defined it quoted in full. A language model writes the prose. It doesn't get to write the quotes.
Models are far more reliable at choosing than at copying. Asked to reproduce a message, they tend to tidy it: fix the spelling, smooth the punctuation, swap a word. In a recap of someone's own conversations that's the worst possible failure, because people remember exactly what was said, and how. So the pipeline turns copying into choosing. It cuts candidate windows of real messages around the landmarks it found, numbers them, and asks the model to pick the strongest four or five by number and write a caption for each. The quoted messages are then rebuilt from the local data by that number. The model can choose badly. It can't misquote, because it never writes the quote.
words changed –can it misquote –
for (const pick of picks) {
const cand = byIndex.get(pick.index);
if (!cand || seen.has(pick.index)) continue; // an index we never offered is ignored
seen.add(pick.index);
moments.push({
caption: pick.caption,
messages: cand.messages, // verbatim: from local data, not from the model
});
}
7Four kinds of call, one shared prefix
Encore is several model calls, not one: the moments; an essay of about 900 words; one call for the relationship's arc, its private vocabulary and a few predictions; and a profile of every participant, packed five people to a call. Every one of them needs the same long context: the rules, the shape of the chat, and its statistics as JSON.
That shared context is the expensive part, and it's identical across calls. So it goes first, as two system blocks marked for prompt caching, and the call that picks the moments runs on its own before the others to write the cache. Every call after it reads that prefix at about a tenth of the normal input price, and pays full price only for its own short instructions and its output. Profiles go one batch at a time, so even a thirty-person group never has more than one profile call in flight. Five to a batch is a deliberate middle: fewer, bigger calls repeat the prefix less, but one failed call shouldn't take a whole group's profiles down with it.
calls –input cost –whole report –
Cost shaped one more decision. The essay used to come from the largest model available. It now comes from the same mid-sized model as everything else, at about a fifth of the price, because that model writes a perfectly good keepsake essay, and the price of an essay is the price of the product. If any call fails, the whole report fails and the credits go back: a paid report is never saved half-done, or quietly patched with a fallback. Because generation runs in the reader's tab, closing the tab halfway sends a beacon that refunds the charge.
8Private by construction
Everything so far runs in the browser, and that takes care of the free recap: it is never uploaded, so there's nothing on a server to protect. Two features do need a server, saving a recap to come back to later and Encore. The goal for those was easy to state and hard to build: the database should hold nothing that the person running the database could read.
Sealed before it leaves
When a signed-in user saves a recap, the browser serialises it (the statistics, a sampled slice of messages and any AI insights) and encrypts it with AES-256-GCM through the Web Crypto API before a byte is sent. Every ciphertext gets its own random 96-bit IV, so saving the same recap twice never produces the same bytes and nobody watching can tell two saves are related. GCM is authenticated encryption: a ciphertext that has been tampered with doesn't decrypt to garbage, it fails loudly. The key that does all of this, the master key, is 256 random bits generated on the device.
One key, several envelopes
The obvious place to keep the master key is the server, and that's exactly where it can't go. Instead the server keeps envelopes: copies of the master key, each encrypted under a different secret that only the user has.
- A passphrase. It's stretched with PBKDF2-HMAC-SHA-256 over a random 16-byte salt into a 256-bit wrapping key. New envelopes use 1,000,000 iterations, so every guess at a stolen envelope costs an attacker a million rounds of hashing, and the browser marks the derived key as non-extractable: it can wrap and unwrap, but it can never be read out.
- A passkey. Through WebAuthn's PRF extension, Face ID, Touch ID or Windows Hello unlocks a passkey, the passkey computes a stable secret from a salt specific to YAPD, and that secret wraps the master key. There's one envelope per enrolled device, and nothing biometric ever leaves the device.
Any one secret unwraps the master key, which is what makes recovery work. A new phone uses the passphrase once, unwraps the master key, then enrolls its own passkey as another envelope. Each envelope records its own parameters, so the iteration count can go up over time without re-encrypting anything that already exists. The one thing no design can offer is a way back without any secret at all. Lose the passphrase and every enrolled device, and the recaps are gone, for the user and for YAPD alike. That isn't an oversight: it's what "only you can read it" means.
recaps you can open –
Once unwrapped, the master key lives in a JavaScript variable for the session, never in localStorage, a cookie or anything sent to the server. So that every restart doesn't mean typing the passphrase again, the device also keeps one more wrapped copy, under a device key: a non-extractable key the browser stores but won't hand to any script, which only works in that browser profile on that machine. Anyone holding the unlocked device gets in, which is the point of remembering a device. A copy of its disk doesn't, because the copy can't use the key. Locking or signing out deletes it. To keep the library fast, the decrypted statistics of recaps you've already opened are cached on the device too: never the key, never the messages.
What the server can see
This is the complete list of what YAPD's database learns about a saved recap: the ciphertext and its IV, a format version, the number of messages, the number of participants, the kind of relationship and the visual theme (so the library can draw placeholder cards before anything is unlocked), and a fingerprint. For the account itself it knows the name and email Google provides at sign-in. It never sees a message, a participant's name, the chat's title, an Encore report, or any key that could open them.
The fingerprint exists so that a purchase can follow a chat. It's a SHA-256 hash of a version tag, the participants' names in sorted order, and the month the chat began, cut to 32 hex characters. Export the same chat again next week, with more messages in it, and the fingerprint doesn't change, so Encore doesn't have to be bought twice. The hash only goes one way: it can't be turned back into names, and a guess can only be confirmed by someone who already knows everyone in the chat and when it started.
can read your messages –can read names –
One limit applies to every web app that encrypts in the browser, and it's worth saying out loud: the encryption is only as trustworthy as the code the site sends. That's part of why the most sensitive piece of server code YAPD runs is published in full.
A proxy that never reads
Encore's essay is generated by a model whose API can't be called straight from a browser without publishing the API key. So there's exactly one place where chat content passes through YAPD's servers: a pass-through route that forwards the request to Anthropic and streams the answer back. It's written so that it can't look at what it carries. It never calls req.json() or req.text(); it hands the request body to the upstream call as a stream, and everything it needs to decide anything arrives in headers. Its full source is on yapd.in, comments and all. Shortened, it reads:
export async function POST(req: Request): Promise<Response> {
const userId = await requireActiveUserId();
if (!userId) return error(401, "Not signed in");
// the grant arrives in a header, so the body is never opened
const grant = req.headers.get("x-encore-grant") ?? "";
const used = await prisma.$queryRaw`
UPDATE "EncoreProxyGrant" SET "callsUsed" = "callsUsed" + 1
WHERE "id" = ${grant} AND "userId" = ${userId}
AND "callsUsed" < "callsMax" AND "expiresAt" > now()
RETURNING "callsUsed"`; // if the store is down: 503, fail closed
if (used.length === 0) return error(402, "Encore session expired or exhausted.");
// ...a per-account limit, a global ceiling, and a 4 MB cap read from Content-Length...
const upstream = await fetch(ANTHROPIC_URL, {
method: "POST",
headers: { "Content-Type": "application/json", "x-api-key": apiKey, "anthropic-version": "2023-06-01" },
body: req.body, // a stream, handed on unread
duplex: "half",
});
return new Response(upstream.body, { status: upstream.status, headers: { "Cache-Control": "no-store" } });
}
Before a byte goes anywhere, the request has to clear a series of gates, each one answered without touching the body. The account must be signed in and active. A grant, minted when the credits were charged, is consumed by a single atomic UPDATE that only succeeds for the grant's owner, within its call budget and before it expires, so an account that never paid can't use YAPD as a free relay and one purchase can't exceed its ceiling. Then a per-account rate limit, keyed on a hash of the account id; a ceiling across every account; and a 4 MB cap read from the Content-Length header. If the grant store can't be reached, the proxy fails closed: no proof of payment, no call. The finished essay comes back to the browser and is encrypted there before it's saved, like everything else.
response –body read by the proxy never
The one deliberate exception
Sharing a recap publicly is the only time plaintext is stored, and it only happens because the user asked. The browser, which can decrypt, sends a separate plaintext copy of the deck for the public link, capped in size; the private, encrypted recap stays the source of truth. The copy can be set to expire, and turning the link off deletes it.
9Free, but hard to farm
Every new device gets a couple of free recaps, which is an invitation to farm them: clear the cookies, turn on a VPN, disguise the browser. So YAPD recognises a device by three signals instead of one: a random cookie, a hash of the IP address and a hash of a browser fingerprint. A device counts as known if any one of the three matches, so getting fresh credits means changing all three at once.
still recognised –
The hashes are salted with a secret, and in production the code refuses to run without one. That isn't decoration. There are only about four billion IPv4 addresses, so an unsalted hash of one, or one salted with a publicly known value, can be reversed by simply trying them all. The hash protects anything only because the salt is private.
Credits are charged before work starts and refunded if it fails, and refunds are where money quietly leaks. Each charge mints a single-use refund ticket, and a refund claims it with a guarded delete: the row disappears for exactly one caller, so a retried or duplicated refund pays out once. The rate limits on the AI routes live in Postgres as fixed-window counters, bumped by a single atomic statement so that concurrent requests can't race past them. They're in the database rather than in memory because the app runs as serverless functions, which don't share memory: an in-process counter would reset on every cold start.
None of this shows up in the recap, which is the point. People see their year in a group chat. Underneath it is a parser that expects the unexpected, statistics whose definitions can be defended, a model that can't misquote anyone, and a server that, by construction, has nothing to tell.