ImageNet
यह क्यों मायने रखता है
गहन अध्ययन
ImageNet वास्तव में दो चीज़ें हैं जिन्हें mix किया जाता है: dataset और competition। 2009 में release dataset में web से scraped तथा WordNet lexical hierarchy से लिए 20,000 से अधिक categories में manually sorted करीब 14 million images हैं, specific dog breeds से cheese types तथा household appliances तक। Competition ILSVRC 2010 से 2017 तक smaller, curated subset — 1,000 classes में करीब 1.2 million training images — पर annually चला और teams से images classify तथा objects localize करने को कहा। उस competition में modern डीप लर्निंग प्रभावी रूप से पैदा हुई, और उसका subset, जिसे अक्सर ImageNet-1K कहते हैं, AI के सबसे cited बेंचमार्क में है।
WordNet और Crowdsourcing से निर्मित
Fei-Fei Li और collaborators उस समय के लिए contrarian bet लेकर निकले: जब अधिकतर computer vision research cleverer algorithms पर focused थी, उनका argument था कि progress data के लिए starved थी। उन्होंने categories को WordNet noun hierarchy से map किया, search engines से candidate images collect की और हर label verify करने के लिए Amazon Mechanical Turk पर workers hire किए, scale पर एनोटेशन (Annotation) reliable रखने के लिए redundancy तथा quality checks बनाकर। Result उस समय vision research में common datasets से orders of magnitude बड़ा था, जिनमें अधिकतम tens of thousands images होती थीं। Dataset 2009 में release हुआ और field को वास्तव में इस्तेमाल करने के लिए annual challenge 2010 में आया। Early results uninspiring थे — traditional methods केवल incrementally improve हुए — जिससे 2012 outcome और dramatic बना।
2012 Breakthrough
2012 में AlexNet — Alex Krizhevsky, Ilya Sutskever तथा Geoffrey Hinton द्वारा बनाया deep CNN — ने 15.3% top-5 error rate से classification task जीता, जबकि traditional hand-engineered features वाला runner-up करीब 26% पर था। इतना बड़ा margin benchmark competitions में essentially कभी नहीं होता और field ने weeks में notice किया। Few years में हर serious entry deep neural network थी और error rates गिरती रहीं: Residual Connection वाले architectures ने 2015 तक top-5 error को 4% से नीचे push किया, trained human labeler के करीब 5% error rate से आगे। Community ने double lesson लिया — deep networks काम करते हैं और ImageNet scale का data मिलने पर ही इतना अच्छा काम करते हैं।
इसने Vision को वास्तव में Solve नहीं किया
एक आम गलतफ़हमी है कि superhuman ImageNet scores का अर्थ computer vision finished है। Benchmark narrow skill test करता है: ऐसी web photo को एक clean label assign करना जहाँ subject आम तौर पर frame भरता है। Cluttered scenes में Object Detection, fine-grained counting, unusual camera angles के विरुद्ध robustness या real-world imagery की long tail के बारे में यह कम कहता है। ImageNet ace करने वाले models अक्सर shortcut cues — object shape के बजाय textures तथा backgrounds — पर depend करते हैं और उनकी accuracy exact brittleness expose करने वाले stress-test variants पर noticeably गिरती है। Labels single और कभी ambiguous होने के कारण, border collie वाली beach photo dog तथा beach scene दोनों है, top-5 metric partially छिपाता है कि Classification वास्तव में कितनी messy है। ImageNet वह ladder था जिसे field ने climb किया और फिर उतरना पड़ा।
Challenge के बाद का जीवन
ILSVRC 2017 के बाद wind down हुआ, लेकिन ImageNet workflow से कभी नहीं गया। इसका most durable contribution Transfer Learning निकला: ImageNet पर trained networks general-purpose visual features सीखते हैं और small task-specific dataset पर उन pre-trained weights की fine-tuning अधिकतर कंप्यूटर विज़न applications, medical imaging से industrial inspection तक, में scratch से training को beat करती है। ImageNet-1K standard proving ground भी बना रहता है जहाँ Vision Transformer जैसे architectures ने पहले demonstrate किया कि वे convolutions से compete कर सकते हैं, और larger 21,000-category version अब भी scale पर pre-training के लिए इस्तेमाल होता है। Field तब से web-scale, weakly labeled तथा self-supervised data की ओर बढ़ा है — ImageNet की champion की वही data-first philosophy, इस point तक push जहाँ few million hand-checked images अब small लगती हैं। यही real legacy है: ऐसा dataset नहीं जिस पर लोग अब भी train करते हैं, बल्कि proof कि better data वाला आम तौर पर जीतता है।