Rohit Swami
India Resume ↗

Writing · YAPD · 6 min read

A WhatsApp export is not a file format

YAPD turns an exported chat into a Wrapped-style recap, and the chat never leaves the browser. Most of the engineering went into reading a format nobody designed, and into letting a language model quote people without ever misquoting them.

YAPD takes the file WhatsApp gives you when you export a chat and turns it into a recap in the style of Spotify Wrapped: a deck of more than twenty animated cards with totals, top emojis, reply speed, the longest streak and the moments that mattered. It's live at yapd.in. The free recap is parsed and analysed entirely in the browser, so the messages never reach a server. That was the first decision, and it made everything after it harder in a useful way: there's no backend to tidy things up later, so the parser has to get it right on the device.

1Four formats wearing one name

"Export chat" sounds like a format. It isn't one. An iPhone wraps the chat in a zip file as _chat.txt; Android hands over the text file itself. The lines inside differ too, by platform and by the phone's language settings:

[12/01/24, 9:41:22 PM] Rohit: Bro where are you      iPhone
12/01/24, 9:41 PM - Rohit: Bro where are you         Android
[12/01/2024, 21:41:22] Rohit: Bro where are you      iPhone, 24-hour clock
01.12.2024, 21:41 - Rohit: Bro where are you         Android, much of Europe

Some exports scatter invisible direction-control characters around the colons and timestamps, and a message with line breaks in it carries on over lines with no header at all. So the parser works line by line. The invisible characters are stripped and Unicode spaces become plain ones first, or none of the patterns would match. Then a line that starts with something shaped like a date and a time begins a new message, and any other line belongs to the message before it.

2Is 05/06 the fifth of June?

The hardest line in that list is the one that looks easiest. 12/01/24 is the 12th of January in India and the UK, and the 1st of December in the US, and nothing on the line says which. Guess wrong and every date in the chat is wrong: the busiest month moves, the streak breaks, the anniversary lands in the wrong season.

The file does hold the answer, just not on any single line. A date whose first number is above 12 can only be day-first, and one whose second number is above 12 can only be month-first. So the parser samples the headers and lets the unambiguous ones vote, while the ambiguous ones abstain. If nobody votes, it falls back to day-first, WhatsApp's default in most of the world. And if a single date still makes no sense in the chosen order, that date alone is tried the other way round.

chat starts
lasts

read as –dates right –

Fig. 1 One date per day of the chat, as the export writes it. Filled dates have a number above 12 and vote; outlined ones could be read either way. Shorten a month-first chat until it ends before the 13th and the evidence disappears.

That fallback is a guess, and the figure doesn't hide it: a short American chat from the first days of a month really is ambiguous, and nothing in the file can settle it. What matters is that the guess is only made when the evidence is genuinely missing. Anything that runs past the 12th of a month settles itself.

3Messages nobody sent

An export also contains lines WhatsApp wrote itself: the encryption notice, "you were added", someone changing the group's subject. If one of those were read as a message, the group's name would turn up as a person, with a message count and a personality of its own. Most of them carry a quiet signal, an invisible left-to-right mark, U+200E, at the start of the text, which nobody types by accident. The catch is that media placeholders carry it too, "image omitted" and "sticker omitted", and those are real messages from real people with the content left out. So a body that starts with the mark is dropped unless it's a media placeholder, and a short list of tightly scoped patterns catches system lines in exports where the mark is missing.

function splitSenderAndText(rest) {
  const idx = rest.indexOf(": ");
  if (idx === -1) return null;
  const sender = rest.slice(0, idx).trim();
  const text = rest.slice(idx + 2);
  if (!sender || sender.length > 80) return null;

  const injected = text.startsWith("‎");             // WhatsApp wrote this body itself
  if (injected && !isMediaMessage(text)) return null;     // a system line, not a person
  if (isSystemBody(text)) return null;                     // the same, for exports without the mark
  return { sender, text: text.replace(/‎/g, "") };
}

4Quoting people exactly

The paid tier, Encore, writes a seven-chapter essay about a chat, with the exchanges that defined it quoted in full. A language model writes the prose. It doesn't get to write the quotes.

Models are far more reliable at choosing than at copying. Asked to reproduce a message, they tend to tidy it: fix the spelling, smooth the punctuation, swap a word. In a recap of someone's own conversations that's the worst possible failure, because people remember exactly what was said, and how. So the pipeline turns copying into choosing. It cuts candidate windows of real messages around likely landmarks, numbers them, and asks the model to pick the strongest four or five by number and write a caption for each. The quoted messages are then rebuilt from the local data by that number. The model can choose badly. It can't misquote, because it never writes the quote.

words changed –can it misquote –

Fig. 2 An illustration with a made-up chat. The tidying is the typical way models drift when they copy text, not a measurement of any one model. Asked for a number, the worst it can do is pick a dull moment.
for (const pick of picks) {
  const cand = byIndex.get(pick.index);
  if (!cand || seen.has(pick.index)) continue;   // an index we never offered is ignored
  seen.add(pick.index);
  moments.push({
    caption: pick.caption,
    messages: cand.messages,                      // verbatim: from local data, not from the model
  });
}

5Private by construction

The same idea runs through the rest of YAPD: make the safe behaviour structural rather than a promise. The free recap never touches a server. A signed-in user's saved recaps are encrypted in the browser with AES-256-GCM before upload, under a key derived on the device from a passphrase or a passkey and kept only in memory, so the database holds ciphertext and a few counts. Encore runs in the browser too. Its calls pass through a proxy that doesn't log, because the model's API can't be called from a browser directly, and a grant issued by the server caps how many calls one paid essay can make. If the essay fails, the credit goes back. Nobody pays for a report they didn't get.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.