Why did AI need 14 million labelled photos to finally see?
Most AI researchers were chasing cleverer algorithms. One of them decided to chase data instead, and hired 49,000 strangers to help.
▶ Start the storyAI needed millions of labelled photos because a program learns to see from examples, and good examples were in short supply. In 2006, most AI research focused on cleverer models and algorithms. Researcher Fei-Fei Li bet on the other half of the problem. She set out to expand and improve the data that vision programs learn from.
She started from WordNet's roughly 22,000 English nouns. She was also inspired by an estimate that an average person recognizes about 30,000 kinds of objects. The original plan called for 400 million images, each checked three times. But a person can sort at most two images per second. At that pace, the job would have taken an estimated 19 human-years without rest.
19 yrs
So Li's team turned to Amazon Mechanical Turk, a website for hiring remote workers to do small tasks that computers could not yet do cheaply. From July 2008 to April 2010, 49,000 workers in 167 countries filtered and labelled more than 160 million candidate images. What remained was 14 million images in over 20,000 categories, each labelled three times.
ImageNet debuted as a modest poster at a 2009 conference. From 2010, it ran a yearly contest where programs compete to recognize its images. In 2012, a deep neural network called AlexNet won it by more than 10 percentage points over the runner-up. It had been trained on two off-the-shelf graphics cards in Alex Krizhevsky's bedroom at his parents' house. Yann LeCun called it a turning point for computer vision, and the whole tech industry started to pay attention.
Quiz me
0/3
Recap
AlexNet's 2012 win, more than ten points ahead of the runner-up, came from a deep network trained on graphics cards and on the dataset Fei-Fei Li had spent years building.
Surprising fact · 49,000 workers in 167 countries filtered and labelled over 160 million candidate images down to ImageNet's 14 million, because the original plan would have taken about 19 human-years of nonstop labelling.
Sources (4)
No source, no claim. Every fact in this lesson (15 claims) cites at least one of these.