India AI Compute Buyers Face 36-To-52-Week GPU Lead Times
Indian AI infrastructure buyers are still reserving GPU capacity months ahead. NeevCloud cofounder Narendra Sen said next-generation enterprise AI GPU lead times now range between 36 and 52 weeks, while named Indian startups are shifting training and inference around scarce capacity.

India’s AI infrastructure market is still operating against a large supply gap: global data centres added 8.9 GW of capacity in 2025, compared with nearly 21.1 GW of demand, leaving a shortfall of about 12 GW, according to cited Jefferies figures.
For Indian buyers, that imbalance now reaches beyond GPUs.
Delivery schedules depend on memory, networking, packaging and power systems as well as the chips themselves.
The early generative-AI shortage has eased from its most severe point, but the latest accelerators remain difficult to secure.
NeevCloud cofounder Narendra Sen said lead times for next-generation enterprise AI GPUs now range from 36 to 52 weeks, with some new orders pushed into 2027.
Dedicated cluster setups that once took about two months have stretched to roughly four months.
Yotta cofounder and chief executive Sunil Gupta said large deployments took between 6 and 15 months to deliver in 2024.
The process is now more predictable, he said, but large builds still require early reservations and close coordination among original equipment manufacturers, system integrators and GPU vendors.
Supply Gap And Delivery Timelines
India Electronics and Semiconductor Association president Ashok Chandak described the situation as a structural imbalance rather than a temporary squeeze.
Delivery timelines have improved since the peak of the shortage, he said, while global data-centre demand continues to outpace supply.
The hardware available to buyers is also changing.
Neysa chief product officer Karan Kirpalani said NVIDIA is concentrating on its newest architectures and retiring the Hopper generation, with the H100 already declared end-of-sale.
Older-generation chips are therefore becoming easier to source, while newer high-performance GPUs remain constrained.
Hyperscalers are adding to that pressure.
Microsoft, Amazon, Google and Meta are committing hundreds of billions of dollars to AI infrastructure and locking up much of NVIDIA’s latest shipment through long-term purchase agreements.
The Jefferies figures also put expected hyperscaler investment at $770 Bn in 2026, up 74% year on year.
The Components Behind The Queue
The constraint has spread into the rest of the AI server.
Sen pointed to CoWoS chip packaging, high-bandwidth memory and the power infrastructure required to run large clusters.
For NVIDIA’s B200 and B300 accelerators, he said, memory and networking components can now take longer to arrive than the chips they support.
That means a processor can be available while the rest of the system is not.
New factories planned by SK Hynix, Samsung and Micron are not expected to add substantial memory capacity until 2027 or 2028, limiting expectations of a quick recovery.
Geopolitics is reinforcing the hierarchy of access.
Nearly 90% of advanced logic-chip production at the 2 nm, 3 nm and 5 nm nodes is concentrated in Taiwan, while China dominates critical minerals.
Export controls, tariffs and compliance requirements are pushing chipmakers to diversify capacity into the US, Europe, India and other Southeast Asian nations, while vendors prioritise strategic markets, sovereign AI programmes and hyperscalers.
How Indian Companies Ration Compute
The queue is changing how companies plan workloads.
Gupta said large enterprises, model builders and AI-platform companies are planning compute requirements several quarters ahead.
Indian cloud providers are reserving capacity years in advance and keeping older hardware in service while they wait for newer clusters, moving away from a purely pay-as-you-go model.
Murf.AI separates scheduled training from continuous real-time inference for its live voice traffic, according to cofounder and chief executive Ankur Edkie.
The company reserves guaranteed capacity before training runs instead of relying on the unpredictable spot market.
Nurix operates a fleet of 15 to 22 GPUs, running real-time inference during high-throughput hours and shifting fine-tuning to off-peak windows.
It uses NVIDIA H100 GPUs for heavier inference and older T4 and L4 architectures for lighter models, while spreading workloads across hyperscalers and neoclouds.
CoRover uses a composite AI architecture that routes only its most demanding workloads to large language models.
Cofounder Ankush Sabharwal said nearly 80% of the company’s tasks avoid GPU-heavy inference.
CoRover sources capacity from Google Cloud, Yotta, NxtGen and the IndiaAI Mission, with peak usage reaching around 1,200-1,250 GPUs.
India’s Strategic Exposure
India imports almost all of its high-end chips, making access to compute a strategic issue as the country expands private AI investment and the IndiaAI Mission.
Chandak called GPU allocation a sovereign security concern and argued that India must manufacture its own compute.
For now, the practical response is earlier procurement and tighter utilisation.
Indian startups are scheduling training, prioritising inference, combining new and old hardware, and routing workloads across providers while the supply chain catches up.
The shortage is no longer only a question of how many GPUs a company can buy; it determines how much work that company can reliably run on the capacity it secures.




















