AI Compute Shortage Shifts Chip Strategy Towards Efficiency
Semiconductor Engineering analysed how AI data-centre growth is being shaped by foundry, memory, power and optical-link bottlenecks as chip buyers look for more tokens from scarce capacity.

Semiconductor Engineering analysed AI data-centre growth through the bottlenecks now shaping chip supply, with efficiency presented as the main lever for the next two to five years while foundry capacity, memory, power and optical links remain constrained.
The piece is less a demand forecast than a map of workarounds.
Hyperscalers and model developers are trying to extract more tokens from scarce infrastructure through custom silicon, smaller models, routing tools, improved GPUs, optical compute experiments and supply deals for unused capacity.
Scarce Compute Makes Efficiency The First Constraint
On Amazon’s July 30 earnings call, chief executive Andy Jassy said demand remained strong into 2028, while returns on capital would become clearer once revenue growth outpaced incremental capital expenditure.
The analysis also treated spare capacity as a near-term pressure valve.
It cited AI infrastructure agreements linked to xAI and SpaceX, including a $45 billion Anthropic deal for access to 220K Nvidia GPUs and 300MW, plus a $920 million-a-month Google deal through mid-2029 for ~110K GPUs.
Model selection is another route around shortages.
SemiAnalysis reports that Anthropic’s annualised revenue run rate per megawatt rose fourfold in nine months, and OpenRouter’s late-July leaderboard had Chinese models among eight of the 10 most-used models by tokens.
Custom Chips And Memory Bandwidth Become Capacity Levers
Semiconductor Engineering placed Nvidia’s Vera Rubin generation, Google and Amazon custom silicon, Broadcom-designed accelerators, optical compute startups and small language models in the same efficiency discussion.
The common aim is more work from limited wafers, memory, power and data-centre capacity.
Amazon’s Trainium revenue run rate exceeds $25 billion a year in the Semiconductor Engineering text, while Google’s reported Frozen v2 project targets six to 10 times more tokens per watt than current hardware by hard-coding parts of Gemini into silicon.
Those figures show why model owners are moving design choices closer to their own workloads.
Memory supply remains just as important.
The analysis describes strong demand and years-long backlogs for DRAM companies, while suggesting that customised HBM for specific AI accelerators could move the business away from a pure commodity cycle.
The supply-chain view keeps the focus on chips rather than general cloud demand.
Semiconductor Engineering framed faster accelerators as useful only when packaging, HBM, networking, power delivery and wafer starts arrive in enough volume for production systems.
Near-term scarcity can coexist with later buyer pressure for lower-cost, workload-specific hardware once supply improves.
Shortage Boom Leaves Capacity-Cycle Questions
The current shortage benefits companies with compute, memory, power and optical-networking capacity.
The harder question is what happens if supply catches up with demand around 2028, 2029 or 2030, especially for providers whose valuations assume rented capacity remains scarce.
Differentiation is the unresolved test for chip suppliers, neoclouds and memory vendors.
The public gaps are actual optical-compute system performance, future wafer allocation, sustained customer demand for rented GPU capacity, whether specialised HBM becomes a durable product category, binding supply volumes for smaller accelerator startups and customer commitments after the shortage eases.




















