भविष्य को आकार देने वाली तकनीक पर गहन लेख।

Python का JIT कंपाइलर आखिरकार आ रहा है

Python 3.15 में copy-and-patch JIT कंपाइलर आ रहा है। यह कैसे काम करता है, कितनी speedup मिलेगी, और CPython को इतना समय क्यों लगा।

पटरी पर तेज़ दौड़ने को तैयार, सर्किट-ट्रेस जेटपैक लगा हुआ साँप

Python जब से बना है, तब से 'धीमा' कहलाता आ रहा है। Python समुदाय का आम जवाब — 'गर्म लूप्स के लिए C extensions इस्तेमाल करो' — हमेशा इस बात की स्वीकृति रहा है कि भाषा का डिफ़ॉल्ट execution model मूल रूप से सीमित है। CPython हर बार एक instruction करके bytecode interpret करता है, और हर instruction एक switch statement के ज़रिए dispatch होती है। यह सरल है, portable है, और debug करना आसान है। लेकिन compute-heavy काम में यह compiled C से लगभग 100 गुना धीमा है।

Python 3.15 इसे बदल रहा है। कुछ साल के experimental काम के बाद, copy-and-patch JIT compiler अब डिफ़ॉल्ट रूप से चालू होकर आ रहा है। यह Python को C जितना तेज़ नहीं बनाएगा — static compilation के बिना कुछ भी नहीं बनाएगा — लेकिन शुरुआती benchmarks असली दुनिया के कोड पर 15-30% speedup दिखा रहे हैं, और कुछ खास patterns में सुधार इससे कहीं ज़्यादा है। 30 सालों से जिस भाषा की परफ़ॉर्मेंस की कहानी 'hot path को C में फिर से लिख दो' रही है, उसके लिए एक JIT जो शुद्ध Python कोड को सचमुच तेज़ करे, एक वास्तविक मील का पत्थर है।

CPython में JIT अब तक क्यों नहीं आया

ऐसा इसलिए नहीं कि किसी ने कोशिश नहीं की। PyPy में एक दशक से ज़्यादा से JIT है और वह Python कोड को CPython से नियमित रूप से 5-10 गुना तेज़ चलाता है। लेकिन PyPy एक अलग implementation है जिसका अपना runtime है, और वह कभी CPython जैसी बाज़ार हिस्सेदारी हासिल नहीं कर पाया, क्योंकि C extension इकोसिस्टम — NumPy, pandas, scikit-learn, वह सब कुछ जिससे Python data science की भाषा बनता है — CPython के C API से जुड़ा है।

CPython में JIT बनाने की बात कई बार हो चुकी है और कोशिशें भी हुई हैं। चुनौतियाँ अच्छी तरह दस्तावेज़ित हैं। CPython का architecture JIT compilation को मुश्किल बनाता है: bytecode dynamically typed है (JIT को efficient code बनाने के लिए type की जानकारी चाहिए, और Python वह static रूप में नहीं देता), C API C code को Python objects में सीधे ऐसे बदलाव करने देता है जो JIT की धारणाओं को तोड़ देते हैं, और interpreter का reference counting garbage collector ऐसा bookkeeping overhead पैदा करता है जिसे JIT आसानी से खत्म नहीं कर सकता।

पिछली कोशिशें — Unladen Swallow (Google, 2009), Pyston (Dropbox, 2014) — LLVM-based JIT compilation को CPython पर जोड़ने की थीं। दोनों ने पाया कि Python के आम workloads के लिए LLVM का compilation overhead बहुत ज़्यादा है। LLVM बड़े codebases के ahead-of-time compilation के लिए बना है; छोटे Python functions को JIT करने में इसे इस्तेमाल करने पर माइक्रोसेकंड की execution के लिए मिलीसेकंड का compilation time लगता है। compilation का खर्च speedup से ज़्यादा पड़ रहा था।

Copy-and-Patch: JIT का एक अलग तरीका

copy-and-patch तकनीक, जो 2021 के एक research paper में आई थी, JIT compilation के लिए मूल रूप से अलग रास्ता अपनाती है। bytecode को intermediate representation में बदलकर optimization passes चलाने (LLVM वाला तरीका) की बजाय, copy-and-patch पहले से compile किए गए code templates के साथ काम करता है।

idea यह है: हर bytecode instruction (LOAD_FAST, BINARY_ADD, CALL_FUNCTION, आदि) के लिए compiler एक C implementation को machine code में पहले से compile कर लेता है, और variable-specific data — register allocations, constant values, memory offsets — के लिए placeholder 'holes' छोड़ देता है। runtime पर JIT compilation बस इतना है: template को copy करो, holes में इस function के specific values भरो। कोई optimization pass नहीं, कोई register allocation algorithm नहीं, कोई instruction selection नहीं। बस memcpy और patch।

Traditional JIT pipeline:
Python bytecode
→ Parse into IR
→ Type inference
→ Optimization passes (CSE, DCE, loop unrolling, inlining...)
→ Register allocation
→ Instruction selection
→ Machine code
Total: 1-100ms per function
Copy-and-patch pipeline:
Python bytecode
→ For each instruction, copy pre-compiled template
→ Fill in holes (constants, offsets)
→ Done — machine code ready
Total: 1-100μs per function (1000x faster compilation)

trade-off code quality में है। LLVM बहुत optimized machine code बनाता है। Copy-and-patch ऐसा machine code बनाता है जो मूल रूप से interpreter loop का compiled रूप है — हर bytecode instruction अभी भी एक अलग template है, instructions के बीच न्यूनतम cross-instruction optimization के साथ। बना हुआ code interpretation से बेहतर है (कोई dispatch overhead नहीं, कोई switch statement नहीं, branch prediction बेहतर), लेकिन एक पूरे optimizing compiler के उत्पाद से कमतर है।

Python के लिए यह trade-off बहुत अच्छा है। Python functions आम तौर पर छोटे होते हैं, कई बार call होते हैं, और हर एक माइक्रोसेकंड में चलता है। जो JIT माइक्रोसेकंड में compile होकर 20-30% speedup दे, वह उस JIT से ज़्यादा कीमती है जो मिलीसेकंड में compile होकर 200% speedup दे — क्योंकि पहले वाले का compilation overhead लगभग तुरंत amortize हो जाता है।

क्या तेज़ होगा

JIT सारे Python कोड को एक जैसी रफ़्तार से तेज़ नहीं करता। कौन सा हिस्सा सबसे ज़्यादा फ़ायदा पाएगा, यह समझने के लिए यह जानना ज़रूरी है कि interpreter अपना समय किस चीज़ पर खर्च करता है।

Bytecode dispatch overhead. interpreter में हर bytecode instruction के लिए चाहिए: अगला opcode fetch करो, उसे decode करो, switch statement के ज़रिए handler पर जाओ। tight loops में यह dispatch overhead कुल execution time का 30-50% तक हो सकता है। JIT इसे पूरी तरह खत्म कर देता है — instructions सीधे jumps के साथ sequential machine code में compile हो जाती हैं।

Type-specialized operations. Python 3.11 ने specializing adaptive interpreter पेश किया, जो असली types देखने के बाद generic operations को type-specific operations से बदल देता है। जब दो integers दिखते हैं तो BINARY_ADD बदलकर BINARY_ADD_INT बन जाता है। JIT इन specialized instructions को efficient machine code में compile करता है — integer addition एक single add instruction बन जाता है, function call नहीं।

Branch prediction. interpreter का central dispatch loop — सैकड़ों cases वाला switch — CPU के branch predictor के लिए बुरा सपना है। JIT इसकी जगह direct control flow देता है जिसे CPU सटीक तरीके से predict कर लेता है। आधुनिक CPUs पर जहाँ branch misprediction 15-20 cycles लगाती है, अकेले इसी से speedup का बड़ा हिस्सा आता है।

क्या तेज़ नहीं होगा: C extension calls (NumPy, pandas), I/O operations (network, disk), और वे operations जिनमें memory allocation हावी है (लाखों छोटे objects बनाना)। अगर आपका Python program 95% समय C extensions में और 5% शुद्ध Python में बिताता है, तो JIT उस 5% को तेज़ करता है — मापने लायक, लेकिन बदलाव लाने वाला नहीं।

Specialization Pipeline

JIT अकेले काम नहीं करता। यह उस performance pipeline का आखिरी चरण है जो Python 3.11 के specializing interpreter से शुरू हुआ और 3.12-3.14 के क्रमिक सुधारों के साथ आगे बढ़ा।

  1. Tier 0: Interpreter. सारा code यहीं से शुरू होता है। adaptive specialization के साथ standard bytecode interpretation। कोई function कई बार call होने के बाद, hot instructions को type-specialized versions से बदल दिया जाता है।
  2. Tier 1: JIT-compiled bytecode. copy-and-patch JIT specialized bytecode को machine code में compile करता है। इससे dispatch overhead खत्म होता है और compiled templates के भीतर constant folding और dead code elimination जैसे basic optimizations संभव होते हैं।
  3. Tier 2 (भविष्य): Trace-based optimization. hot code paths के execution traces रिकॉर्ड करके पूरे traces — function boundaries के पार — को optimized machine code में compile करना। यह योजना में है, पर अभी ship नहीं हुआ है।

यह tiered approach Ruby के YJIT और दूसरे आधुनिक language runtimes जैसा ही है। तेज़ interpretation से शुरू करो, code hot होने पर तेज़ compilation पर जाओ, और महंगे optimization को सबसे hot paths के लिए बचाओ। मूल सोच एक ही है: ज़्यादातर code optimize करने लायक नहीं होता, इसलिए compilation का बजट वहाँ खर्च करो जहाँ सबसे ज़्यादा code चलता है।

Memory और Startup पर असर

JIT compilers compiled code के लिए memory लेते हैं। copy-and-patch JIT का memory overhead मामूली है — compiled code bytecode से बड़ा होता है, पर LLVM-based JITs के output से छोटा (क्योंकि optimization का bloat नहीं है)। मौजूदा implementation जिस bytecode की जगह लेता है उसकी लगभग 1.5-3 गुना memory इस्तेमाल करता है, और केवल उन्हीं functions को compile करता है जो इतनी बार call होते हैं कि फ़ायदा हो।

छोटे समय तक चलने वाली Python scripts के लिए startup time चिंता का विषय है। JIT template library लोड करने और compilation infrastructure सेट करने में overhead लेता है। जो scripts एक सेकंड से कम चलती हैं, उनमें JIT का overhead speedup से ज़्यादा पड़ सकता है। CPython इसका हल ऐसे निकालता है कि functions सिर्फ़ एक निश्चित संख्या में call होने के बाद JIT होते हैं — छोटी scripts interpreter में ही रहती हैं और JIT का कोई टैक्स नहीं देतीं।

यह configurable है। -X jit flag JIT का व्यवहार नियंत्रित करता है, और environment variables compilation threshold ट्यून करते हैं। serverless functions और CLI tools के लिए, जहाँ startup मायने रखता है, आप threshold बढ़ा सकते हैं या JIT पूरी तरह बंद कर सकते हैं। लंबे समय तक चलने वाले servers और data processing scripts के लिए, जहाँ steady-state performance मायने रखती है, defaults ठीक काम करते हैं।

Python इकोसिस्टम के लिए इसका मतलब

JIT Python की performance hierarchy में उसकी जगह नहीं बदलता — compute-heavy काम में C, Rust, Go और Java अभी भी काफ़ी तेज़ हैं। यह जो बदलता है वह वह सीमा है जिस पर Python developers को उन विकल्पों की ओर जाना पड़ता है।

शुद्ध Python कोड पर 20-30% speedup का मतलब है कि कुछ workloads, जिनके लिए पहले C extensions या rewrite चाहिए था, अब शुद्ध Python में पर्याप्त तेज़ चलेंगे। जो data processing scripts 10 मिनट लेती थीं, अब 7 मिनट लेंगी। जो web server 1000 requests per second संभालता था, अब 1300 संभालेगा। ये क्रांतिकारी आंकड़े नहीं हैं, लेकिन यही फ़र्क है 'Python काफ़ी तेज़ है' और 'इसे Go में rewrite करना पड़ेगा' के बीच।

और अहम बात, JIT infrastructure भविष्य के optimization की नींव रखता है। copy-and-patch तरीके को बेहतर templates, ज़्यादा specialization, और अंततः trace-based compilation से आगे बढ़ाया जा सकता है। Python का हर नया version JIT के मूल architecture को बदले बिना बेहतर templates ship कर सकता है। 3.15 में 20-30% speedup एक floor है, ceiling नहीं।

दुनिया की सबसे लोकप्रिय और सबसे धीमी भाषाओं में से एक बनने के तीन दशक बाद, CPython आखिरकार performance में गंभीरता से निवेश कर रहा है। JIT 'Python बहुत धीमा है' वाले लोगों को संतुष्ट नहीं करेगा — कुछ भी नहीं करेगा, क्योंकि कुछ workloads के लिए Python सचमुच धीमा है, JIT हो या न हो। लेकिन Python कोड के बड़े हिस्से के लिए, जहाँ execution speed 'ठीक है पर बढ़िया नहीं' थी, JIT उसे 'सचमुच अच्छा' के करीब ले जाता है। यह लगता है उससे कहीं बड़ी बात है।