मुख्य सामग्री पर जाएँ
Zubnet AIसीखेंWiki › AI Energy Consumption
बुनियादी ढांचा

AI Energy Consumption

इसे भी कहा जाता है: AI Power Consumption, AI Energy Use
AI systems के lifecycle में consumed total electricity, जो two phases से dominated है: model बनाने वाला one-time training run और हर user request answer करने वाला ongoing inference। Load में accelerators स्वयं और उनके आसपास cooling, networking तथा power-delivery overhead शामिल हैं। Models और usage scale होने के साथ demand इतनी बड़ी हुई है कि कई countries में power-grid planning reshape कर रही है।

यह क्यों मायने रखता है

Energy अब AI progress की hard constraints में है: hundreds of megawatts draw करने वाला training cluster केवल वहाँ exist कर सकता है जहाँ grid उतना power deliver करे, और power availability increasingly तय करती है कि new capacity कहाँ build होगी। Practitioners के लिए energy serving costs और sustainability claims में directly दिखती है जिन्हें हर major AI lab को अब defend करना पड़ता है।

गहन अध्ययन

पहली चीज़ training तथा inference का split समझना है। Frontier model training single enormous burst है: tens of thousands accelerators weeks या months तक lockstep में चलते हैं, और largest runs की energy estimates tens of gigawatt-hours में आती हैं — roughly small town की annual electricity use। Inference का profile opposite है: हर request cheap है लेकिन billions requests a day accumulate होती हैं, और deployed usage explode होने के साथ inference ने AI की total electricity demand में larger share के रूप में training को overtake किया है। दोनों phases अंततः same place — डेटा सेंटर — में आते हैं, जहाँ GPU और TPU power को heat में बदलते हैं जिसे फिर remove करना होता है, इसीलिए modern AI campuses silicon जितने substations तथा cooling के आसपास engineer होते हैं।

Training बनाम Inference

Frontier training run अपनी energy single window में concentrate करता है। उसके पीछे cluster — high-bandwidth networking से connected tens of thousands accelerators — run की duration में 100 MW या अधिक continuous IT load draw कर सकता है, इसलिए training sites available substations तथा cheap, firm power के आसपास चुने जाते हैं। Smaller jobs भी मायने रखते हैं: existing model की फ़ाइन-ट्यूनिंग scratch से training के मुकाबले orders of magnitude cheap है, लेकिन thousands of organizations का daily ऐसा करना जुड़ता जाता है।

इन्फ़ेरेंस energy को spread करती है। हर generated टोकन tiny amount compute cost करता है, लेकिन popular model round the clock millions of users serve करता है, इसलिए meter कभी नहीं रुकता। Google की अपनी 2025 measurement ने median Gemini text prompt करीब 0.24 watt-hours बताया, और comparable chatbots के independent estimates वही few-tenths-of-a-watt-hour range में हैं — per request small, aggregate में enormous। Reasoning models ने per-request energy sharply बढ़ाई है क्योंकि visible answer से पहले long hidden chains of thought generate करते हैं। Rule of thumb: training energy model तथा dataset size के साथ स्केलिंग नियम के अनुसार scale करती है, जबकि inference energy users, context length और output length के साथ।

Gigawatt Campuses और Nuclear Deals

Industry का response gigawatts में measured buildout रहा है, modern AI इन्फ्रास्ट्रक्चर का core हिस्सा। Announced AI campuses single site पर अब 1 GW या अधिक capacity plan करते हैं — large nuclear reactor का output — और IEA global data-center consumption को 2024 में roughly 415 TWh, total electricity का करीब 1.5%, estimate करती है, जो 2030 तक double से अधिक और AI main driver होगा। Strain locally already visible है: US data centers ने 2023 में national electricity का करीब 4.4% draw किया, 2028 तक 7–12% federal projections के साथ, और Ireland में वे national demand का roughly fifth हैं, जिससे Dublin के आसपास new connections पर restrictions लगीं।

Grids इतनी fast expand नहीं कर सकते, इसलिए biggest buyers सीधे source पर गए। Microsoft ने Three Mile Island के undamaged reactor को restart करने के लिए 20-year agreement sign किया, roughly 835 MW round-the-clock carbon-free power secure करके; Google ने 2035 तक कुल करीब 500 MW के small modular reactors के लिए Kairos Power से deal की; Amazon ने अपने nuclear-backed campus deals के साथ follow किया। ये agreements green branding से कम physics के बारे में हैं: training तथा serving clusters को firm, 24/7 power चाहिए और nuclear तथा handful other sources ही gigawatt scale पर देते हैं।

एक Query Grid को Crash नहीं करेगी

Common claim है कि हर chatbot query environmental disaster है और per-query numbers इसे support नहीं करते। Few tenths watt-hour पर text prompt LED bulb के couple minutes के comparable और उन earlier viral estimates से बहुत नीचे है जो हर query के massive fresh computation trigger करने की assumption करते थे। Models तथा serving stacks improve होने पर per-query figure भी लगातार गिरती है।

लेकिन opposite claim — AI energy non-issue है — equally wrong है। Impact aggregate में है: billions daily requests, image तथा video generation जैसी heavier modalities जिनकी cost text से कहीं अधिक है, और search, office software तथा customer support में AI embed होने के साथ compound demand। Doom headlines और dismissal दोनों real picture miss करते हैं, जो per-query guilt problem के बजाय grid-level planning problem है।

Efficiency Gains और Jevons Rebound

Efficiency curve genuinely steep है। क्वांटाइज़ेशन, डिस्टिलेशन, स्पेक्युलेटिव डिकोडिंग और mixture-of-experts architectures प्रत्येक compute per token को large factors से cut करते हैं, और NVIDIA तथा competitors की हर accelerator generation performance per watt substantially improve करती है। Model Serving stacks भी smarter हुई हैं, KV-cache reuse तथा continuous batching utilization high रखते हुए, और कुछ inference edge devices पर shift हो रही है जहाँ small models locally चलते हैं।

फिर भी total consumption climb करती रहती है, जो Jevons paradox in action है: cost per token गिरने पर demand flat नहीं रहती — previously uneconomical applications switch on होते और usage budget fill करने के लिए expand होती है। Token prices roughly order of magnitude per year गिरी हैं जबकि total inference volume और fast बढ़ी है। Realistic outlook यह नहीं कि efficiency grid बचाती है, बल्कि efficiency तय करती है कि given grid कितना AI deliver कर सकती है।

← सभी शब्द
ESC