Rohit Swami
India Resume ↗

Writing · Kaggle · 8 min read

A dataset is a product

Five datasets and six notebooks on Kaggle, downloaded nearly 18,000 times and built on in 63 public notebooks. What it takes to make data that strangers can trust, from every NeurIPS paper to two million Indian companies, and the notebooks I built on top of my own datasets.

On Kaggle I'm "Ragnar". Most of my work there happened in 2020: five public datasets and six notebooks, which took me to the Expert tier in both, in the top 0.05% of dataset contributors and the top 0.07% of notebook authors. As I write this, the datasets have been viewed about 135,000 times and downloaded nearly 18,000 times, and 63 public notebooks have been built on them, most of them by other people.

The numbers aren't the point, but they taught me something I've used in every job since. A dataset isn't a file. It's a product, with users who decide within a minute whether to trust it. They read the title, skim the description, look for the columns they need and a licence that lets them use them, and either open a notebook or close the tab. Almost everything that makes that minute go well has nothing to do with the data itself.

Fig. 1 My five public datasets, as Kaggle's API reported them in October 2026. "Notebooks" counts every public notebook that uses the dataset, mine included. The smallest file has the best conversion by far.

1Finishing someone else's dataset

My second most-voted dataset started as a gap. Ben Hamner's "NIPS Papers" dataset on Kaggle covered the conference from 1987 to 2017: 7,241 papers. By the end of 2019 NeurIPS had published 2,439 more, and anyone studying the field's recent years had to scrape them on their own. So I did it once, for everyone: a scraper in Python with BeautifulSoup that walks nips.cc, published on my GitHub, and a dataset with the year, title, authors, abstract and full text of every NeurIPS paper from 1987 to 2019.

Full text is what made it useful, and also what made it heavy: the papers file alone is 325 MB. Authors live in a separate file, one row per author per paper, so the main table stays one row per paper and nobody has to split a comma-separated field to count co-authors. The description says where the data came from, credits the dataset it extends, says what's in each file and lists what it's good for: topic modelling, keyword extraction, exploration, a semantic search engine. Then I built two of those myself, and they became my best notebooks.

2Two million companies, thirteen statuses

My most-voted dataset, at 171 votes, is a different kind of problem. India's Open Government Data platform publishes the master data the Registrar of Companies keeps on every company: here 1,992,170 of them, registered between 1857 and 2020, each with 17 columns, from the corporate identification number and the date of registration to authorised and paid-up capital, the registered state and the principal business activity.

The raw columns are terse codes, and a code nobody can read is a column nobody uses. Nobody can analyse "STOF" without knowing it means struck off. So most of the work on this dataset was its documentation: every column explained, every class and category listed, and every status decoded, so that a stranger's first minute goes on the data rather than on deciphering it.

StatusMeansStatusMeans
ACTVactiveDISDdissolved
NAEFnot available for e-filingCLLDconverted to LLP and dissolved
ULQDunder liquidationUPSOunder process of striking off
AMALamalgamatedCLLPconverted to LLP
STOFstruck offLIQDliquidated
DRMTdormantMLIQvanished
D455dormant under section 455

The thirteen company statuses, as the dataset's description decodes them.

The notebook I published with it is a plain exploration: where the missing values are, how companies divide by status, class and business activity, how registrations grew year by year, with a sharp dip in 2001, and which companies hold the most authorised capital. The most useful lesson in it was the least glamorous. To draw companies per state on a map of India, the dataset's state names had to match the map's, and they didn't.

places missing from the map –

Fig. 2 The names that needed work, straight from the notebook. Old names, a misspelling, a territory that merged with its neighbour in 2020 and one created in 2019: joined as they are, every one of these places is blank on the map.

Orissa had become Odisha and Uttaranchal Uttarakhand; Pondicherry is Puducherry now; "Chattisgarh" and "Nagra" were simply misspelt; the map writes "Andaman & Nicobar" with an ampersand. Dadra and Nagar Haveli and Daman and Diu had merged into one union territory in 2020, so their counts had to be added together, and Ladakh, a union territory since 2019, had no companies under its own name, so it went in with zero to complete the map. None of that is analysis. All of it decides whether the analysis is right.

3Keywords without a model

My most popular notebook, with 188 votes, extracts keywords from NeurIPS papers with TF-IDF, and it runs entirely on two of my own datasets: the papers, and a list of English stopwords. TF-IDF is old, but it's worth understanding properly, because it's still the backbone of keyword search and the first thing to reach for before anything heavier.

It multiplies two numbers. Term frequency is how often a word appears in this document. Inverse document frequency is how rare the word is across all the documents: scikit-learn's smoothed version is ln((1 + n) / (1 + df)) + 1, where n is the number of documents and df the number that contain the word. A word that's frequent here and rare elsewhere scores high, and that's roughly what a keyword is. Each document's vector is then scaled to length one, so long papers don't win just by being long.

Fig. 3 TF-IDF computed live, the way scikit-learn's vectoriser does it, on eight short made-up abstracts standing in for 9,680 papers. The stopword list is the real one from my dataset. Raw counts surface "we", "the" and "of"; the stopword list removes those; IDF then discounts words most abstracts share, like "model", so a shared word only ranks high by being very frequent.

The notebook's choices were all about cleaning. Text is lowercased, punctuation stripped and whitespace collapsed before anything is counted. The stopwords are the 733 from the Terrier information-retrieval platform, which I'd published as their own dataset because every text project needs a list like it. And the vectoriser learns its vocabulary and document frequencies from every paper except the first ten, which are held back to test it on, so the keywords for those ten come from a model that has never seen them.

vectorizer = TfidfVectorizer(stop_words=stopwords, smooth_idf=True, use_idf=True)
vectorizer.fit_transform(corpora[10::])            # learn from all but the first ten papers
feature_names = vectorizer.get_feature_names()

def get_keywords(vectorizer, feature_names, doc):
    tf_idf_vector = vectorizer.transform([doc])     # score one held-out paper
    sorted_items = sort_coo(tf_idf_vector.tocoo())  # highest TF-IDF first
    return extract_topn_from_vector(feature_names, sorted_items, TOP_K_KEYWORDS)

4Asking nine thousand papers a question

Keywords tell you what a paper is about. The next notebook, with 36 votes, tried to answer questions from all of them. It's the retrieve-then-read pattern that sits under every retrieval-augmented system today, built in 2020 with Haystack. Elasticsearch indexes every paper and, given a question, returns the ten most relevant by BM25, a refined TF-IDF. Then a reader, ALBERT xxlarge fine-tuned on SQuAD 2.0, reads those ten and marks the exact spans most likely to answer, returning the best five with their confidence.

Fig. 4 An illustration of the notebook's pipeline, with its real settings: ten papers retrieved, five answers read out. Retrieval alone hands back documents; the reader turns them into answers, each a span of the original text with a score.

The interesting decision was about length. A reader model looks at a few hundred tokens at a time, and a NeurIPS paper runs to tens of thousands of characters. Sliding a large reader over every word of ten whole papers, for every question, would be slow on a notebook's hardware, so each paper was cleaned and trimmed to about 10,000 characters, cut at a sentence boundary so the last sentence kept is a whole one. That trades recall for speed, and it's the same trade I describe in "Retrieval is the hard part": whatever the reader never sees, it can't answer from.

5Small files travel furthest

The smallest thing I published is a 6 KB text file: those 733 stopwords, one per line. It has been downloaded 5,738 times, more than the 325 MB of NeurIPS papers, and roughly one in five people who look at it download it. Big datasets get looked at. Small, boring, correct ones get used.

The other two were made for questions of my own. One is a month of US flight data from the Bureau of Transportation Statistics, for February 2020: scheduled and actual departure and arrival times from every US airline with at least one percent of domestic passenger revenue, published alongside a flight-delay prediction project on my GitHub. The other is the TOP500 and Green500 lists of the world's fastest and most energy-efficient supercomputers, with a notebook that maps them by continent, by maker and by operating system, and draws how every system's rank moved between two editions of the list.

moved up –moved down –held –

Fig. 5 The real data from my TOP500 dataset: 456 systems that were on both the June and November 2020 lists, from their previous rank on the left to their new one on the right. Hover or drag across the chart to pick one out. Forty-four new systems arrived, so nearly everyone else slid down.

6What a dataset needs

Looking back across the five, the ones people used shared the same few things, and Kaggle's own usability score, which rates three of them at a perfect 1.0, rewards most of them:

None of it is specific to Kaggle. Every pipeline I've built since has had readers: the next team, the dashboard, the model, the person on call at 3am. The habits are the same. Say where the data came from, say what every field means, shape it for the question people will ask, and make the first minute easy.

I'm Rohit Swami. I build the unglamorous machinery real products run on: data pipelines, real-time services, open-source tools, and products of my own. More about me, or write to me.

The figures on this page are simulations written for it. They run in your browser, and the numbers in them are illustrative unless the text says otherwise.