Writing · Kaggle · 8 min read
A dataset is a product
Five datasets and six notebooks on Kaggle, downloaded nearly 18,000 times and built on in 63 public notebooks. What it takes to make data that strangers can trust, from every NeurIPS paper to two million Indian companies, and the notebooks I built on top of my own datasets.
On Kaggle I'm "Ragnar". Most of my work there happened in 2020: five public datasets and six notebooks, which took me to the Expert tier in both, in the top 0.05% of dataset contributors and the top 0.07% of notebook authors. As I write this, the datasets have been viewed about 135,000 times and downloaded nearly 18,000 times, and 63 public notebooks have been built on them, most of them by other people.
The numbers aren't the point, but they taught me something I've used in every job since. A dataset isn't a file. It's a product, with users who decide within a minute whether to trust it. They read the title, skim the description, look for the columns they need and a licence that lets them use them, and either open a notebook or close the tab. Almost everything that makes that minute go well has nothing to do with the data itself.
1Finishing someone else's dataset
My second most-voted dataset started as a gap. Ben Hamner's "NIPS Papers" dataset on Kaggle covered the conference from 1987 to 2017: 7,241 papers. By the end of 2019 NeurIPS had published 2,439 more, and anyone studying the field's recent years had to scrape them on their own. So I did it once, for everyone: a scraper in Python with BeautifulSoup that walks nips.cc, published on my GitHub, and a dataset with the year, title, authors, abstract and full text of every NeurIPS paper from 1987 to 2019.
Full text is what made it useful, and also what made it heavy: the papers file alone is 325 MB. Authors live in a separate file, one row per author per paper, so the main table stays one row per paper and nobody has to split a comma-separated field to count co-authors. The description says where the data came from, credits the dataset it extends, says what's in each file and lists what it's good for: topic modelling, keyword extraction, exploration, a semantic search engine. Then I built two of those myself, and they became my best notebooks.
2Two million companies, thirteen statuses
My most-voted dataset, at 171 votes, is a different kind of problem. India's Open Government Data platform publishes the master data the Registrar of Companies keeps on every company: here 1,992,170 of them, registered between 1857 and 2020, each with 17 columns, from the corporate identification number and the date of registration to authorised and paid-up capital, the registered state and the principal business activity.
The raw columns are terse codes, and a code nobody can read is a column nobody uses. Nobody can analyse "STOF" without knowing it means struck off. So most of the work on this dataset was its documentation: every column explained, every class and category listed, and every status decoded, so that a stranger's first minute goes on the data rather than on deciphering it.
| Status | Means | Status | Means |
|---|---|---|---|
| ACTV | active | DISD | dissolved |
| NAEF | not available for e-filing | CLLD | converted to LLP and dissolved |
| ULQD | under liquidation | UPSO | under process of striking off |
| AMAL | amalgamated | CLLP | converted to LLP |
| STOF | struck off | LIQD | liquidated |
| DRMT | dormant | MLIQ | vanished |
| D455 | dormant under section 455 |
The thirteen company statuses, as the dataset's description decodes them.
The notebook I published with it is a plain exploration: where the missing values are, how companies divide by status, class and business activity, how registrations grew year by year, with a sharp dip in 2001, and which companies hold the most authorised capital. The most useful lesson in it was the least glamorous. To draw companies per state on a map of India, the dataset's state names had to match the map's, and they didn't.
places missing from the map –
Orissa had become Odisha and Uttaranchal Uttarakhand; Pondicherry is Puducherry now; "Chattisgarh" and "Nagra" were simply misspelt; the map writes "Andaman & Nicobar" with an ampersand. Dadra and Nagar Haveli and Daman and Diu had merged into one union territory in 2020, so their counts had to be added together, and Ladakh, a union territory since 2019, had no companies under its own name, so it went in with zero to complete the map. None of that is analysis. All of it decides whether the analysis is right.
3Keywords without a model
My most popular notebook, with 188 votes, extracts keywords from NeurIPS papers with TF-IDF, and it runs entirely on two of my own datasets: the papers, and a list of English stopwords. TF-IDF is old, but it's worth understanding properly, because it's still the backbone of keyword search and the first thing to reach for before anything heavier.
It multiplies two numbers. Term frequency is how often a word appears in this document. Inverse document frequency is how rare the word is across all the documents: scikit-learn's smoothed version is ln((1 + n) / (1 + df)) + 1, where n is the number of documents and df the number that contain the word. A word that's frequent here and rare elsewhere scores high, and that's roughly what a keyword is. Each document's vector is then scaled to length one, so long papers don't win just by being long.
The notebook's choices were all about cleaning. Text is lowercased, punctuation stripped and whitespace collapsed before anything is counted. The stopwords are the 733 from the Terrier information-retrieval platform, which I'd published as their own dataset because every text project needs a list like it. And the vectoriser learns its vocabulary and document frequencies from every paper except the first ten, which are held back to test it on, so the keywords for those ten come from a model that has never seen them.
vectorizer = TfidfVectorizer(stop_words=stopwords, smooth_idf=True, use_idf=True)
vectorizer.fit_transform(corpora[10::]) # learn from all but the first ten papers
feature_names = vectorizer.get_feature_names()
def get_keywords(vectorizer, feature_names, doc):
tf_idf_vector = vectorizer.transform([doc]) # score one held-out paper
sorted_items = sort_coo(tf_idf_vector.tocoo()) # highest TF-IDF first
return extract_topn_from_vector(feature_names, sorted_items, TOP_K_KEYWORDS)
4Asking nine thousand papers a question
Keywords tell you what a paper is about. The next notebook, with 36 votes, tried to answer questions from all of them. It's the retrieve-then-read pattern that sits under every retrieval-augmented system today, built in 2020 with Haystack. Elasticsearch indexes every paper and, given a question, returns the ten most relevant by BM25, a refined TF-IDF. Then a reader, ALBERT xxlarge fine-tuned on SQuAD 2.0, reads those ten and marks the exact spans most likely to answer, returning the best five with their confidence.
The interesting decision was about length. A reader model looks at a few hundred tokens at a time, and a NeurIPS paper runs to tens of thousands of characters. Sliding a large reader over every word of ten whole papers, for every question, would be slow on a notebook's hardware, so each paper was cleaned and trimmed to about 10,000 characters, cut at a sentence boundary so the last sentence kept is a whole one. That trades recall for speed, and it's the same trade I describe in "Retrieval is the hard part": whatever the reader never sees, it can't answer from.
5Small files travel furthest
The smallest thing I published is a 6 KB text file: those 733 stopwords, one per line. It has been downloaded 5,738 times, more than the 325 MB of NeurIPS papers, and roughly one in five people who look at it download it. Big datasets get looked at. Small, boring, correct ones get used.
The other two were made for questions of my own. One is a month of US flight data from the Bureau of Transportation Statistics, for February 2020: scheduled and actual departure and arrival times from every US airline with at least one percent of domestic passenger revenue, published alongside a flight-delay prediction project on my GitHub. The other is the TOP500 and Green500 lists of the world's fastest and most energy-efficient supercomputers, with a notebook that maps them by continent, by maker and by operating system, and draws how every system's rank moved between two editions of the list.
moved up –moved down –held –
6What a dataset needs
Looking back across the five, the ones people used shared the same few things, and Kaggle's own usability score, which rates three of them at a perfect 1.0, rewards most of them:
- A title and subtitle that say what's inside, and when. "1857 to 2020" and "1987 to 2019" answer the first question anyone has.
- Provenance. Where the data came from, with a link, and credit for anything it builds on.
- A licence. Without one, careful people can't use it at all.
- A column dictionary, with every code decoded. The thirteen statuses are worth more than the two million rows without them.
- Files shaped for use. One row per thing, and related things in their own file, joined by an id.
- A first notebook. Showing that the data loads, and one interesting thing it can do, is the best documentation there is.
None of it is specific to Kaggle. Every pipeline I've built since has had readers: the next team, the dashboard, the model, the person on call at 3am. The habits are the same. Say where the data came from, say what every field means, shape it for the question people will ask, and make the first minute easy.