ImageNet: How Fei-Fei Li's Team Built the Dataset

The 2009 project used WordNet, web search and Mechanical Turk voting to build millions of labeled images, enabling the annual benchmark AlexNet later won.

Fifty million images, one synset at a time

WordNet groups English nouns into roughly 80,000 synsets, its term for a single concept. The Princeton team, where Fei-Fei Li was on the faculty before moving to Stanford later in 2009, set out to attach 500 to 1,000 photographs to most of those synsets, aiming for on the order of 50 million labeled images altogether. By the time their paper appeared in June 2009, they had reached 5,247 synsets across twelve subtrees, among them mammal, bird, vehicle and furniture, holding 3.2 million images: about ten percent of WordNet's noun vocabulary. [2][4]

Borrowing a hierarchy to kill ambiguous words

ImageNet's design choice was to inherit WordNet's sense distinctions rather than build its own vocabulary. The rival ESP dataset, built from a labeling game, collected words with no such disambiguation: a label of bank could mean a riverbank or a financial institution, a real cost at this scale. Candidate images were instead crawled using each synset's WordNet synonyms, expanded with parent-synset terms and queries translated into Chinese, Spanish, Dutch and Italian, since ordinary search results were only about ten percent accurate for a given query. [2]

A vote threshold that moves with the synset

Cleaning ran on Amazon Mechanical Turk. Workers saw a candidate image, the synset's definition and a Wikipedia link, then voted on whether the image showed that concept. Some categories are harder to agree on than others, telling a Burmese cat from a Siamese cat is not the same task as telling a cat from a truck, so the team built a confidence score from an initial batch of at least ten votes per synset, then kept voting on the rest until each image crossed that synset's threshold. Across an 80-synset sample, the result held to 99.7 percent precision. [2]

A two-year target, and a three-year wait for proof

Atlas interpretation: The summary's two years is best read as the paper's own forward-looking target, not a look back. In June 2009 the team wrote that their goal was to finish the roughly 50-million-image database in the next two years. The 3.2 million images that already existed were the product of crawling and labeling up to that point. [2]

That target became an annual benchmark, the ImageNet Large Scale Visual Recognition Challenge, run every year from 2010. Progress was incremental at first: winning classification error was 28.2 percent in 2010 and 25.8 percent in 2011. In 2012 a deep convolutional network cut that figure to 16.4 percent, and nearly every serious entry afterward used the same kind of network. [3]

Atlas interpretation: Three years, then, separates Li's bet that data rather than algorithms was the field's bottleneck from the year that bet stopped being arguable. ImageNet kept growing past its own two-year plan besides: by 2014 it held more than 14 million images across nearly 22,000 synsets, short of the 50-million target but far past any comparable dataset of the time. [3]

Sources

  1. ImageNet: A large-scale hierarchical image database

    IEEE CVPR · Jun 2009

  2. ImageNet: A Large-Scale Hierarchical Image Database

    IEEE CVPR · Jun 20, 2009

  3. ImageNet Large Scale Visual Recognition Challenge

    arXiv · Jan 30, 2015

  4. 'The Worlds I See' by AI visionary Fei-Fei Li '99 selected as Princeton Pre-read

    Princeton University · Feb 23, 2024