मुख्य सामग्री पर जाएँ
Zubnet AIसीखेंWiki › Turing Test
मूल तत्व

Turing Test

इसे भी कहा जाता है: Imitation Game
Alan Turing द्वारा अपने 1950 paper “Computing Machinery and Intelligence” में proposed machine intelligence का behavioral test। Hidden interlocutors के साथ text messages exchange करता human judge machine को human से reliably अलग न बता सके तो machine को intelligent behavior exhibit करने वाला कहा जाता है। Turing ने इसे machines सोच सकती हैं या नहीं जैसे vague question का concrete replacement दिया था।

यह क्यों मायने रखता है

Turing Test ने machine intelligence क्या मानी जाए इस पर seventy years से अधिक arguments shape किए और chatbot द्वारा pass करने के हर claim पर अब भी public debate anchor करता है। Test वास्तव में क्या measure करता है — time pressure में convincing imitation, understanding नहीं — यह जानना उन headlines को critically पढ़ने और face value पर लेने के बीच difference है।

गहन अध्ययन

Turing का move philosophy को पूरी तरह sidestep करना था। “Thinking” define करने के बजाय उन्होंने parlor game describe किया: interrogator दो hidden players, एक human तथा एक machine, को questions type करता है और केवल answers से पता लगाने की कोशिश करता है कि कौन कौन है। Interrogator chance से better न करे तो machine जीतती है। Setup conversational behavior के अलावा सब हटा देता है — कोई नहीं पूछता machine अंदर से कैसे काम करती है, conscious है या किससे बनी है। वही deliberate narrowness test के popular होने और तब से criticize होने, दोनों का कारण है: यह mechanism नहीं, performance measure करता है। Large language models के modern era में, जिनका पूरा training objective plausible text produce करना है, test ठीक वही चीज़ extraordinarily अच्छी करने वाले systems से सीधे टकराया है जो वह measure करता है।

Original Imitation Game

Turing का 1950 paper तीन players वाला game describe करता है: man, woman और किसी gender का interrogator, सभी typed messages से communicate करते हैं ताकि voice तथा appearance कुछ reveal न करें। Interrogator का काम तय करना है कौन-सा player woman है, जबकि man उसे mislead करने की कोशिश करता है। फिर Turing ने पूछा कि machine man की जगह ले तो क्या होगा — क्या interrogator लगभग उतनी बार गलत होगा। Framing महत्वपूर्ण है क्योंकि test facts recite करने के बारे में कभी नहीं था; यह इस बारे में था कि machine ऐसे flexible, open-ended conversation sustain कर सकती है या नहीं जिसके लिए mind आवश्यक लगता था। Turing ने predict किया कि year 2000 तक machines game इतनी अच्छी खेलेंगी कि average interrogator के पास five minutes questioning बाद correct identification की 70% से अधिक chance नहीं होगी, remark जो बाद में अक्सर quote होने वाला “five minutes के लिए 30% judges” passing bar बना।

ELIZA से Loebner Prize तक

Test trickery reward करता है इसकी first warning जल्दी आई। 1966 में Joseph Weizenbaum ने ELIZA बनाया, ऐसा चैटबॉट जो simple pattern matching से users के statements को questions के रूप में reflect करके psychotherapist की parody करता था — और कुछ users को विश्वास हुआ कि वह उन्हें समझता है। Reaction ने Weizenbaum को इतना disturb किया कि career का बाकी हिस्सा simulation को understanding समझने के विरुद्ध argue करने में बिताया। Pattern Loebner Prize में decades तक repeat हुआ, 1990 से 2019 की annual competition जो most human-like program को medals देती थी; कोई entry unrestricted long-form test कभी pass नहीं कर पाई। 2014 में imperfect English वाले 13-year-old Ukrainian boy के रूप में pose करता Eugene Goostman नामक program five-minute event में third judges को convince कर पाया और pass घोषित हुआ — claim widely dismiss हुआ क्योंकि persona उसकी evasions excuse करता और conversations short थीं।

LLMs ने Test तोड़ दिया

Modern बड़ा भाषा मॉडल ने picture पूरी तरह बदल दी। Enormous text corpora पर pretraining उन्हें construction से human conversation के fluent mimics बनाती है, ठीक वह skill जिसे imitation game score करता है। 2024 तथा 2025 की controlled studies में three-party imitation games के judges ने GPT-4-class systems को लगभग half time या अधिक human identify किया, और एक 2025 study में persona-prompted model को actual human participant से अधिक human rate किया गया। Crucially outcome deep reasoning से नहीं बल्कि surface style से tip हुआ: typing speed, typos, emoji use, brevity और believable persona। Five-minute text conversation pass करना, जिसे कभी AGI के रास्ते का distant milestone माना गया, ऐसी चीज़ निकली जिसे well-tuned प्रॉम्प्ट largely achieve कर सकता था।

Intelligence Test नहीं, Deception Test

Deepest misconception है कि Turing Test pass करना intelligence demonstrate करता है। Test measure करता है कि judge fool हो सकता है या नहीं, और यह machine की abilities जितना judge की expectations, time limit तथा format पर depend करता है — इसीलिए skeptical AI researcher तथा casual participant के बीच results wildly vary करते हैं। John Searle के Chinese Room thought experiment ने 1980 में point press किया: system कुछ समझे बिना symbols manipulate करके perfect conversational answers produce कर सकता है। Modern chatbots के critics same argument का version बनाते हैं, LLMs को training text remix करने वाले स्टोकेस्टिक तोता कहते हैं। व्यवहार में field quietly आगे बढ़ गई है: serious मूल्यांकन अब capability benchmarks, मानव मूल्यांकन (Human Evaluation) protocols और Chatbot Arena जैसे preference platforms पर निर्भर है, जबकि Turing Test मुख्यतः historical landmark और headline generator के रूप में बचा है।

← सभी शब्द
ESC