Research

GPU Shortage 2026: How to Secure AI Compute When GPUs Are Sold Out

Back to BlogWritten by Published Apr 6, 2026Updated
GPU ShortageGPU AvailabilityGPU Supply ChainHBM ShortageCoWoSGPU Lead TimesAI Compute ShortageGPU Price IncreaseGPU CloudAI InfrastructureSpot GPU
GPU Shortage 2026: How to Secure AI Compute When GPUs Are Sold Out

H100 SXM5 nodes are sitting at 36-52 week lead times from resellers right now. That is not a supply blip. It is a structural problem with two root causes: CoWoS packaging capacity at TSMC is fully allocated, and HBM production from SK Hynix cannot keep pace with demand. For AI teams that did not lock in compute in 2025, the practical reality is this: training jobs are queuing, inference costs are rising, and planning horizons are collapsing to weeks instead of quarters.

This post covers what is driving the shortage, what the supply picture looks like through 2027, and four strategies that let you keep workloads running despite constrained hardware availability. Export policy is the other lever on supply: NVIDIA's 2026 H200 shipments to China show how quickly allocation can shift.

TL;DR: GPU Shortage Status, Lead Times, and Availability in 2026

  • Still short, but narrower than 2025. As of 09 Sep 2026 the squeeze sits in Blackwell and top-end Hopper SXM parts. A100 and L40S are broadly available.
  • Buying takes months, renting takes minutes. Neither NVIDIA nor the OEMs publish lead times, and circulating estimates for H100 range from 6 to 52 weeks. Cloud rental skips the queue entirely.
  • The cause is packaging and memory, not silicon. TSMC CoWoS throughput and HBM output cap NVIDIA shipments, not GPU die yield.
  • Cloud rates now: Spheron H100 SXM5 is $2.65/hr on-demand and $2.10/hr spot, H200 $4.80/hr on-demand, as of 09 Sep 2026, with no reserved contract.
  • Easing, not over. Analysts expect the CoWoS supply-demand gap to narrow from roughly 20% to 10% by end of 2026, with Rubin absorbing much of the new capacity. Check live GPU rates.

Is There Still a GPU Shortage in 2026?

Yes, and the honest version of that answer has a qualifier attached: the shortage is real but it is no longer general. In 2023 and 2024, "GPU shortage" meant almost any accelerator was hard to get. In 2026 it means something narrower and more useful to know.

The constraint now lives almost entirely in the top-end SXM parts. H100 SXM5 and H200 SXM5 are the SKUs where CoWoS packaging slots and HBM3e stacks are the gating input, and those are the parts hyperscalers pre-committed to years in advance. Below that tier, the picture is unrecognisably different. A100 80GB is broadly obtainable, L40S has genuine surplus in most regions, and RTX-class inventory is limited more by GDDR7 supply than by anything structural.

That distinction changes the practical question. It is rarely "can I get a GPU" and almost always "can I get this GPU, in this region, this week, without signing a twelve-month commitment". Three teams can hit three completely different answers on the same day depending on which of those three variables they are unwilling to move.

There is a second thing worth separating out, because it gets conflated constantly. Scarcity of hardware and scarcity of bookable cloud capacity are not the same problem. A hyperscaler showing no on-demand H100 capacity is usually not out of H100s. It is prioritising reserved enterprise commitments and first-party AI workloads ahead of on-demand requests. The GPUs are racked and running; they are just not available to you. That is an allocation policy, not a supply shortage, and it is the reason a GPU-first provider can show H100 SXM5 at $2.65/hr on-demand as of 09 Sep 2026 on the same day AWS returns an insufficient-capacity error.

So the useful framing for planning: treat top-end SXM supply as constrained and forecast it as a managed risk, treat mid-tier GPUs as available, and treat hyperscaler on-demand availability as a policy variable you do not control rather than a market signal.

The 2026 GPU Shortage Explained

The shortage is not primarily about GPU die production. It is about the memory and packaging that surrounds the die.

HBM supply chain bottleneck. NVIDIA H100 SXM5 uses HBM3 (the PCIe variant uses HBM2e). H200 and the Blackwell lineup use HBM3e. SK Hynix supplies the majority of HBM stacked memory for NVIDIA's data center products. TSMC's CoWoS (Chip on Wafer on Substrate) packaging process is required to bond HBM dies onto the GPU substrate, and CoWoS capacity is fully allocated through at least mid-2027. Samsung and Micron are ramping HBM capacity, but neither will meaningfully ease the shortage before late 2026 at the earliest.

Hyperscaler reservation activity. Microsoft, Google, Meta, and Amazon placed multi-billion-dollar forward orders for Blackwell GPUs (GB200, B200) in 2025, consuming most of NVIDIA's available allocation capacity through the end of 2026 and into 2027. This crowded out mid-market and enterprise customers who previously purchased through standard channels or direct resellers. The circular financing that backs many neoclouds means new capacity is typically committed before it reaches the spot market, compounding access constraints for teams without forward agreements.

Consumer GPU production cuts. NVIDIA reportedly cut RTX 5000-series production by 30-40%, according to industry reports, driven primarily by GDDR7 memory shortages and a strategic shift toward data center SKUs. The result: the consumer-grade secondary market that smaller AI teams have historically relied on when cloud supply is tight is now thinner than usual.

Commonly quoted purchase lead times, with the caveat that these are channel estimates rather than published figures and the spread between sources is wide. The lead times section below covers why:

GPUCommonly quoted lead timeCloud availability
H100 SXM536-52 weeks via resellersLimited on hyperscalers; available spot on neo-clouds
H200 SXM5Months, estimates vary widelyReserved pools mostly sold out
B200Allocation-gated, booked well aheadLimited to select providers
A100 80GB8-16 weeksMore available; watch for constrained VRAM configs
L40S2-8 weeksGood availability; strong for inference

HBM and CoWoS: The Two Bottlenecks Behind the GPU Shortage

If you only track one thing about GPU supply, track these two inputs. Neither is a GPU. Both cap how many GPUs exist.

CoWoS is the packaging bottleneck. CoWoS stands for Chip on Wafer on Substrate, and it is the TSMC 2.5D packaging process that mounts the GPU die and its HBM stacks side by side on a shared silicon interposer. NVIDIA has used it on every HBM-equipped data center part since P100, AMD's MI300X and MI350X use CoWoS-S, and there is no way to wire memory to compute at these bandwidths without it. TSMC dominates supply rather than monopolising it: its own lines are projected to reach roughly 120,000 to 140,000 wafers per month by the end of 2026, with OSAT partners ASE and Amkor adding perhaps another 50,000 to 60,000. Adding a line means new front-end-grade cleanroom space, and industry estimates put the cycle from decision to volume at roughly 18 to 24 months. The consequence is the counterintuitive part of this whole story: packaging capacity, not wafer fabrication, is the binding constraint on how many accelerators exist. TSMC can etch more GPU dies than it can package.

HBM is the memory bottleneck. High Bandwidth Memory is DRAM dies stacked vertically and bonded with through-silicon vias. Only three companies make it at volume, and the concentration is real: in Q2 2026 SK Hynix held about 50% of HBM revenue and Samsung about 33%, with Micron third. Each generation is harder to yield than the last because the stacks get taller and the tolerances tighter; test-and-yield analysis puts the loss at roughly 5-8% for each die added beyond an 8-layer stack. H100 SXM5 uses HBM3 and the PCIe variant HBM2e. H200, B200 and B300 use HBM3e. Rubin moves to HBM4. Every step up that ladder consumes more wafer and yields worse per wafer.

InputWho supplies itWhat it gatesWhy it is tight
CoWoS packagingTSMC (effectively sole volume source)Total NVIDIA and AMD DC GPU outputInterposer line capacity, 18-24 month expansion cycle
HBM3SK Hynix, Samsung, MicronH100 SXM5, MI300XDisplaced by HBM3e as fabs convert lines
HBM3eSK Hynix (lead), Samsung, MicronH200, B200, B300, MI350XTaller stacks, lower yield per wafer than HBM2e
HBM4SK Hynix, Samsung, Micron (ramping)Rubin generation, 2027 onwardPre-production, absorbs capacity before it eases anything
GDDR7SK Hynix, Samsung, MicronRTX 5000 series, some inference partsCompetes with HBM for the same DRAM fab capacity

The row that matters most is the last one. GDDR7 and HBM are made in the same DRAM fabs and compete for the same wafer starts, and HBM consumes far more capacity per usable gigabyte. When a memory maker shifts a line to HBM because the margin is better, consumer GDDR7 supply tightens as a direct result. Micron alone is adding around 60,000 HBM wafers per month, reaching roughly 100,000 by the end of 2026, and that capacity comes from somewhere. That is the mechanism connecting a data center memory shortage to RTX 5000-series production cuts, and it is why "the GPU shortage" and "the consumer graphics card shortage" are the same shortage viewed from two ends.

There is a compounding effect here that gets missed. These two constraints multiply rather than add. A GPU needs both a CoWoS slot and its HBM stacks. Solving one alone moves nothing: extra HBM with no packaging capacity is inventory, and extra packaging capacity with no HBM is idle line time. Any forecast that promises relief from a single announcement, whether a Samsung capacity ramp or a TSMC capex release, is only describing half the equation.

For teams renting rather than buying, the practical read-through is that this cost structure follows the hardware into rental rates. Memory subsystem cost is now a large and rising share of GPU bill of materials, which is why per-hour rates for H200 SXM5 availability and H100 have stayed firm even at providers holding real inventory. Cheaper rental comes from supply of the input, not from a provider deciding to discount.

The Memory Crisis: HBM and GDDR Demand

The binding constraint is memory, not the GPU die itself. AMD, Intel, and NVIDIA all compete for HBM allocation from the same three suppliers: SK Hynix, Samsung, and Micron. AMD's MI300X uses HBM3 and MI350X uses HBM3e, Intel's Gaudi 3 uses HBM2e, and NVIDIA's entire data center stack now requires HBM3 or HBM3e. They are all pulling from the same limited supply.

HBM3e production is more demanding than HBM2e. Higher die stacks and tighter tolerances mean lower yield per wafer. H200 and B200 both require HBM3e, so as Blackwell ramps, it compounds the same bottleneck that is already constraining H100 H200 supply. The result is predictable: less supply available for the installed base while demand from new architectures increases.

The downstream effect on cloud pricing is less obvious but real. HBM shortages raise GPU memory subsystem costs, which flows into lease and rental prices even when a cloud provider does have inventory. This is why H100 spot prices have not collapsed to historical lows despite the fact that neo-cloud providers have maintained meaningful availability. The hardware itself costs more to source. For current H100 price trends and how Blackwell supply is shifting the market, see our H100 price tracker.

GDDR7, used in the RTX 5000 series and some inference-optimized GPUs, is also constrained. This has pushed RTX 5090 prices above levels that make sense for most AI workloads, limiting the consumer GPU secondary market as an overflow valve.

GPU Lead Times in 2026: Direct Purchase vs Cloud

Lead time is the number most procurement teams anchor on, and it is also the number most likely to be quoted without saying which channel it applies to. A 40 week figure and a 4 minute figure can both be true for the same GPU on the same day.

Start with a warning about the numbers themselves, because it matters more than any figure below. Neither NVIDIA nor the major OEMs publish lead times. Every number in circulation is a channel estimate, and the published estimates disagree with each other badly: for H100 alone you can find sourced-to-nobody figures ranging from 6 weeks to 52 weeks, sometimes on the same day, sometimes attributed to the same channel. Much of that spread is real (lead time genuinely depends on order size, region, configuration and whether you are an existing account) and some of it is content-mill repetition. Anyone quoting you a single confident lead-time number for a GPU is telling you about their sourcing relationship, not about the market.

What is reliable is the shape of the difference between buying and renting.

Buying, through any channel. Months, with wide variance. You are joining an allocation queue and your position depends on order size and existing relationship: a first-time buyer ordering four nodes sits behind a hyperscaler ordering forty thousand regardless of when each order was placed. Quoted dates are forecasts, not commitments, and they slip. The one dependable rule is ordinal rather than cardinal: Blackwell parts are hardest to get, Hopper SXM parts next, A100 and L40S easiest by a wide margin.

Renting. Minutes. The hardware is already installed, powered, networked and running. You are not buying a GPU, you are buying time on one that already exists.

That gap is the actual decision, and it is not close. Buying makes sense at sustained high utilisation over multiple years, where amortisation beats rental and you can absorb months of uncertainty in the schedule. Below that threshold, a procurement queue is simply the wrong instrument: if a training run is blocked for two or three quarters waiting on a purchase order, the constraint was never really hardware. Teams working through that tradeoff properly should start with our GPU capacity planning guide, which covers the utilisation point where buying starts to win.

If you do need a procurement estimate, get it as a written quote from the specific distributor you would actually buy from, for your specific configuration and quantity. Treat every published figure, including any you find on this site, as an order of magnitude rather than a number you can plan a budget against.

GPU Availability in 2026 by Model: H100, H200, B200, A100, L40S

Availability is per SKU, per tier, per provider type. A single "is there a shortage" answer hides all three. Here is the model-by-model picture as of 09 Sep 2026.

H100 SXM5. The part most teams ask for by name, and no longer the scarcest thing on the market: Blackwell has taken that title as frontier demand moved up a generation. H100 is now best described as available but uneven. Hyperscaler on-demand remains unreliable during peak periods and reserved pools are largely spoken for, while GPU-first providers generally have it bookable. H100 on Spheron lists at $2.65/hr on-demand with spot at $2.10/hr. If your workload is flexible on interconnect, PCIe variants clear faster than SXM.

H200 SXM5. Tight, and reported to be improving as Blackwell pulls frontier labs upward and frees Hopper inventory behind them. The 141GB of HBM3e at 4.8TB/s is the reason to want one, though the popular version of that pitch needs a correction. You will read that a 70B model at FP16 fits on one H200 where it needs two H100s. Weights-only, that is arithmetically true: 70B at two bytes per parameter is 140GB, which clears 141GB and does not clear 80GB. In practice it is not a serving configuration, because it leaves roughly 1GB for KV cache, activations and framework overhead, and a single 128K-context request alone wants tens of GB of KV cache. The real single-GPU win on H200 is quantized: at FP8 you get the weights down near 70GB and have genuine headroom to serve. Listed at $4.80/hr on-demand, $3.36/hr spot.

B200. Allocation-gated at the manufacturer level, so availability depends entirely on whether a provider secured an allocation. Most did not. Where it exists it is frequently spot-tier only, because providers hold dedicated capacity for reserved customers. Spheron B200 instances list at $5.37/hr spot as of 09 Sep 2026.

A100 80GB. Not scarce, and consistently the most underrated answer to a shortage problem. Displaced from frontier training by Hopper and Blackwell, it remains entirely adequate for fine-tuning, most inference, and any memory-bound workload where bandwidth matters more than raw FLOPS. On-demand from $1.43/hr.

L40S. Genuine surplus. 48GB of GDDR6 with strong FP8 inference throughput, and it is the correct default for quantized serving workloads that do not need HBM bandwidth. On-demand L40S access starts at $0.96/hr.

RTX-class (4090, 5090). Constrained by GDDR7 rather than by data center allocation, so the scarcity has a different cause and a different shape. Still the cheapest route to running quantized models under 30B. RTX 4090 on Spheron is $0.72/hr on-demand.

The pattern across all six: the shortage is a top-of-stack problem, and most workloads are not top-of-stack workloads. Teams that specify H100 by reflex rather than by measurement are competing for the scarcest part in the market to run jobs that an A100 or an L40S would serve at a fraction of the rate. Sizing the workload before sourcing the GPU is the cheapest availability strategy there is, and our GPU performance buyer's guide walks through how to benchmark that properly.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 09 Sep 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

Why GPU Prices Are Going Up, and What Would Bring Them Down

Cloud GPU rates have not fallen the way a normal hardware cycle would predict, and the reason is that three cost inputs are rising at once underneath them.

Memory is a growing share of what a GPU costs to build. Each HBM generation carries a price premium over the one before it, and each adds stack height that costs yield. On a top-end accelerator the memory subsystem is a large fraction of bill of materials rather than a line item. That floor moves with DRAM contract pricing, not with GPU demand. Precise percentages are not disclosed by NVIDIA or the memory makers, so treat any specific figure you see as an estimate.

Packaging is priced as a scarce good. While advanced packaging lines are allocated, the slots price accordingly, and that premium is embedded in every unit shipped. This is observable rather than theoretical: OSAT provider ASE raised advanced packaging quotes by more than 20% in July 2026 in an AI-driven price increase.

Forward orders remove supply before it reaches the open market. Hyperscaler and neocloud commitments take inventory out of circulation months before it is racked, so what reaches spot and on-demand is the residual. A market clearing on residual supply prices higher than one clearing on total supply, even when total supply grows.

Together these explain the thing that confuses people most: why rental rates hold firm even at providers with real inventory. The provider is not withholding capacity. Their replacement cost went up.

So what would actually bring prices down? Not demand cooling, and not a single provider deciding to compete on price. Three specific things, in rough order of impact:

  1. HBM3e supply outrunning HBM4 pre-production. New Samsung and Micron capacity helps only to the extent it is not immediately consumed by Rubin.
  2. CoWoS expansion landing faster than the next architecture absorbs it. The historical pattern is that it does not.
  3. Hopper displacement at scale. As frontier labs migrate to Blackwell and Rubin, H100 and H200 inventory returns to the general market. This is the mechanism most likely to deliver real relief, and it is already visible in H200 availability.

The near-term implication for anyone budgeting compute: model your costs on rates holding roughly flat for top-end SXM parts through 2026, with softening more likely in Hopper than in Blackwell. If that exposure is material to your unit economics, the serverless vs on-demand vs reserved GPU breakdown covers which billing model absorbs rate volatility best.

The AWS rates cited elsewhere in this post were checked against published instance pricing on 7 Sep 2026. Cost estimates for HBM and packaging as a share of accelerator bill of materials are widely reported but not disclosed by NVIDIA or its suppliers, so treat the direction as sound and any specific percentage you see quoted, here or anywhere, as an estimate.

Impact on AI Teams

The shortage hits AI teams in three distinct ways.

Training delays. Teams that planned Q2 2026 training runs expecting to reserve H100 nodes on AWS or GCP have found reserved pools locked behind existing customers. The fallback is on-demand pricing, which runs 2-3x more expensive and is often throttled or unavailable during peak demand periods. A training run budgeted at $40,000 on reserved capacity is now looking at $80,000-120,000 on on-demand, if capacity is accessible at all.

Inference cost spikes. As H100 on-demand rates climb with demand, inference teams running production APIs face rising per-token costs with no clear path to reserved discounts. Workloads that were profitable at a given per-hour H100 rate are now priced above their unit economics ceiling. Switching to smaller models or alternative GPU classes becomes a financial necessity, not an architectural preference. That kind of unhedged rate exposure is exactly what's pushing CME Group and Silicon Data to launch cash-settled compute futures contracts on H100 and B200 rental pricing in October 2026, giving finance teams a way to hedge the swing even when reserved capacity isn't available.

Capacity planning breakdowns. Twelve-month planning cycles assume predictable GPU availability. With 36-52 week lead times for physical hardware and reserved cloud capacity booked 6-plus months ahead, teams that did not commit compute in 2025 are now reacting to scarcity rather than executing against a plan. Roadmaps built around "we'll add compute when we need it" are running into walls. For a practical framework on sourcing and right-sizing GPU capacity under these conditions, see our GPU capacity planning guide. EU teams face the steepest lead times because hyperscalers prioritize US-East capacity: the impact on B200 and H200 availability in European regions is covered in our European GPU cloud guide for 2026. The Middle East is adding further demand pressure: Saudi Arabia's HUMAIN placed the largest sovereign GB300 order of 2026, and the UAE-US AI Campus is accelerating UAE GPU procurement. For a full breakdown of GPU availability in the Gulf region, see our Middle East GPU cloud guide.

There is an important asymmetry at work. Hyperscalers and well-funded frontier labs locked in supply via forward contracts a year or two before the shortage became acute. Everyone else is competing for spot and on-demand capacity that was not pre-reserved. The strategies below are how teams in the second group are staying operational. Deciding whether to chase reserved capacity, ride out on-demand, or lean on spot is really a billing-model question; see our serverless vs on-demand vs reserved GPU breakdown for the tradeoffs under exactly this kind of scarcity.

How to Get GPU Capacity Right Now

Everything above is diagnosis. If you have a job blocked today, this is the order to work through, cheapest and fastest first.

1. Check whether you actually need the GPU you asked for. Run the memory maths before you run the search. A 70B model at FP16 is 140GB of weights before KV cache and activations, so in practice it is a two-GPU deployment on 80GB cards and still wants headroom on a 141GB H200. At FP8 the same weights drop to roughly 70GB, which fits a single H100 80GB with about 10GB left for KV cache: tight, but a genuine single-GPU serving configuration and the standard practitioner setup for 70B on Hopper. Quantizing before sourcing turns a two-scarce-GPU problem into a one-scarce-GPU problem, and it is the only step here that costs nothing.

2. Widen the SKU. If the workload runs on A100 80GB or L40S, source those instead. Both are available, both are cheaper, and neither is competing with a hyperscaler's frontier training queue. This solves more blocked pipelines than any other single move.

3. Widen the region. Capacity is not uniform. US-East gets new generations first and clears fastest; European and APAC regions run behind on Blackwell and H200 specifically. If your data residency requirements allow movement, a region change often resolves an availability error immediately. Where they do not, our European GPU cloud guide covers what is realistically bookable in-region.

4. Move off hyperscaler on-demand. This is the step teams delay longest and benefit from most. An insufficient-capacity error on AWS or GCP is an allocation decision, not a market condition. A GPU-first provider running the same silicon has no first-party AI workload competing for it.

5. Take the spot tier with checkpointing. For training, spot plus checkpointing every 15-30 minutes converts a preemption risk into at most a 30 minute loss. For inference, run a warm on-demand replica as primary and spot as burst behind a load balancer. Do not assume spot is cheaper than on-demand: the two tiers track different pools and the ordering does invert, so compare both rates at the moment you book.

6. Add a second provider before you need one. A fallback script is an afternoon of work and it is the difference between a full spot pool being an inconvenience and being an outage.

A note on what not to do. Waiting for a reserved allocation to open up is the most common failure mode we see, because it feels like the responsible procurement path while quietly costing weeks. Reserved capacity is genuinely the cheapest tier at sustained high utilisation, and it is the wrong instrument for an unblocking problem. Book on-demand or spot now, negotiate reserved on your own timeline. Teams that need the reserved conversation anyway should read the GPU cluster reservation and contract negotiation guide before entering it.

The four strategies below go deeper on the structural versions of steps 4 through 6.

Strategy 1: GPU-First Cloud Providers vs Hyperscalers

AWS, GCP, and Azure are general-purpose cloud platforms. When GPU capacity is tight, they prioritize their own first-party AI workloads and their largest enterprise customers. On-demand H100 availability on these platforms has become genuinely unreliable for teams without pre-existing reserved capacity. For more context on this tradeoff, see our guide to AWS, GCP, and Azure GPU alternatives.

Neo-clouds (Spheron, CoreWeave, Lambda, Hyperstack, and others) source GPU inventory from vetted data center partners globally and run GPU-first infrastructure. They are not diverting GPU capacity to internal workloads, because they do not have internal workloads competing for the same GPUs. This structural difference matters a lot when supply is tight. A hyperscaler will always allocate reserved capacity to a $100M enterprise customer before making on-demand inventory available to a startup. A GPU-focused provider's entire business model depends on keeping that on-demand capacity accessible. For a comparison of the top GPU rental platforms and what to look for when evaluating them, see our GPU rental marketplace overview, GPU performance buyer's guide, and the live Spheron GPU rental catalog with availability across H100, H200, A100, B200, L40S, and consumer Blackwell.

Spheron aggregates supply from data center partners across multiple regions. H100, H200, A100, L40S, and RTX-class inventory is available through on-demand instances and spot, with no reserved contracts required.

Current pricing comparison (on-demand and spot, per GPU per hour):

GPUAWS instanceAWS on-demand, per GPUSpheron on-demandSpheron spot
H100 SXM5p5.48xlarge$6.88 ($55.04 / 8)$2.65/hr$2.10/hr
A100 40GBp4d.24xlarge$2.74 ($21.96 / 8)$1.43/hr$1.19/hr
L40Sg6e.48xlarge$3.77 ($30.13 / 8)$0.96/hr$1.07/hr

Two caveats on that table. AWS per-GPU figures are derived by dividing the published instance rate by GPU count, which is how you compare a bare instance price to a per-GPU marketplace rate, and they exclude storage, networking and egress. The A100 row is not quite like for like either: AWS p4d carries the 40GB A100 while the Spheron rate is for the 80GB part, so the real gap on equivalent hardware is wider than the row shows.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 09 Sep 2026; AWS rates were checked against published instance pricing on 7 Sep 2026 and may have changed. Check current GPU pricing → for live rates.

The real advantage with neo-cloud providers is not always on-demand price, it is availability and access to a spot tier at all. H100 SXM5 spot runs $2.10/hr against $2.65/hr on-demand as of 09 Sep 2026; the two tiers track different pools, so compare both at the moment you book rather than assuming an ordering. More practically, hyperscalers routinely show on-demand capacity as unavailable during peak periods and deprioritize teams without existing reserved commitments. GPU-focused providers do not have internal workloads competing for inventory, so on-demand availability is more consistent. For spot-capable workloads, the economics are clear. For on-demand, compare current rates and factor in actual availability when the supply is this tight.

Strategy 2: Spot GPU Instances and Preemptible Compute

Spot instances provide access to GPUs that are not reserved, at a discount that varies sharply by provider (see our spot GPU instances buying guide for the 2026 breakdown by cloud). On Spheron, spot prices vary dynamically based on pool availability. The tradeoff is preemptibility: your instance can be reclaimed with short notice when the underlying capacity is needed elsewhere.

For training, preemption is manageable with automated checkpointing. Save model state, optimizer state, and data loader position every 15-30 minutes. When a spot instance gets reclaimed, you lose at most 30 minutes of training. You saved 40-70% on compute across the entire run. A 12-person team used this exact approach to train a 70B model for $11,200 using this approach, documented in full detail here.

Here is a minimal PyTorch checkpoint save and resume pattern:

python
import torch
import os

# Save checkpoint every N steps
def save_checkpoint(model, optimizer, step, path):
    tmp_path = path + ".tmp"
    torch.save({
        'step': step + 1,  # store next step to execute, not the one just completed
        'model_state_dict': model.state_dict(),
        'optimizer_state_dict': optimizer.state_dict(),
    }, tmp_path)
    os.rename(tmp_path, path)  # atomic on POSIX filesystems; verify NAS rename semantics if using NFS

# Resume from checkpoint if one exists
def load_checkpoint(model, optimizer, path):
    if os.path.exists(path):
        checkpoint = torch.load(path, weights_only=True, map_location='cpu')
        model.load_state_dict(checkpoint['model_state_dict'])
        optimizer.load_state_dict(checkpoint['optimizer_state_dict'])
        return checkpoint['step']
    return 0

# In your training loop
start_step = load_checkpoint(model, optimizer, "/persistent-storage/checkpoint.pt")
for step in range(start_step, total_steps):
    # ... training logic ...
    if step % checkpoint_interval == 0:
        save_checkpoint(model, optimizer, step, "/persistent-storage/checkpoint.pt")

The key requirement: checkpoint to persistent network storage, not local NVMe. Local storage disappears when the instance is reclaimed. Network-attached storage survives.

For inference on spot, the approach is different. Run a warm on-demand instance as your primary serving replica. Use spot instances as burst capacity behind a load balancer. If a spot instance is preempted, the load balancer routes requests to on-demand. Over-provision spot by 20% to absorb preemption lag without dropping requests.

For the full cost breakdown on a spot training setup, see the GPU cost optimization playbook.

Strategy 3: Model Optimization to Reduce GPU Requirements

When you cannot get more GPUs, reduce how many GPUs your model needs. Three techniques have real traction in 2026.

Quantization

FP8 is native on H100 and H200. It reduces model memory roughly 50% compared to FP16. A 70B model that requires 8x H100 80GB at FP16 fits on 4x H100 80GB at FP8 with minimal quality degradation on most benchmarks. Halving your H100 requirement effectively doubles your access to available inventory.

INT4 via GPTQ or AWQ is more aggressive. A 13B model at INT4 fits on a single 24GB GPU (RTX 4090 class), and RTX 4090 availability is far better than H100 right now. Quality degradation is task-dependent, but for many production inference workloads it is within acceptable bounds. Test against your task before committing.

FP4 on Blackwell (B200, GB200): NVIDIA's MXFP4 support cuts memory requirements further still. If your team has Blackwell access, this is worth evaluating for inference-heavy workloads. See our FP4 quantization guide for Blackwell GPUs for the details.

Mixture of Experts Inference

MoE architectures activate a fraction of total parameters per token. DeepSeek V3.2 has 671B total parameters but only 37B active per token. That is 37B active parameter compute cost on a model that delivers near-dense-model quality. The catch: you still need to load all 671B parameters into VRAM for weight storage, which means multiple GPUs. But the compute requirement per token is dramatically lower than the total parameter count implies.

The GPU selection calculus changes for MoE: you need memory-rich GPUs for weight storage, but the compute throughput requirement is lower than the total size implies. H200 and A100 80GB become more attractive for MoE inference than for dense models, and those have better availability right now than H100. See our MoE inference optimization guide for serving strategies.

Knowledge Distillation

Train a smaller student model on outputs from a larger teacher model. A 7B student trained on GPT-4 or Llama 4 Maverick outputs often reaches 85-95% of teacher quality on specific task domains. The GPU requirement for the student at inference time is 10-20x lower than the teacher.

For shortage conditions, this is a viable medium-term strategy: run the teacher on whatever expensive reserved capacity you can access to generate training data, train the student (which is significantly cheaper), then serve the student on widely available, lower-cost GPUs. The upfront cost to generate training data is real, but the ongoing inference savings on widely available hardware are substantial. See our 7B student from 70B teacher distillation guide for a concrete implementation.

Strategy 4: Multi-Provider GPU Orchestration and Failover

Single-provider dependency is a risk in a constrained market. If your primary provider fills its spot pool and cannot spin up new instances, your pipeline stops. This happens. Planning for it is not paranoia.

Multi-provider orchestration routes workloads across two or more GPU cloud providers based on live availability and price. It is operationally more complex, but the availability hedge is worth the overhead when supply is tight. Implementation approaches range from full Kubernetes multi-cloud node pools (high engineering lift, maximum flexibility) to simple provider fallback scripts (an afternoon's work). Distributed training across providers raises a separate question: whether you can get an InfiniBand-connected cluster at all during a shortage. Our guide to multi-node GPU training without InfiniBand covers the fallback pattern for teams that can only source Ethernet-connected nodes right now.

Here is a simple two-provider fallback pattern in Python:

python
import requests
import time

SPHERON_API = "https://api.spheron.ai/v1"
PROVIDER_B_API = "https://api.provider-b.com/v1"

def try_provision_gpu(api_base, headers, config):
    """Attempt to provision a GPU instance. Returns instance ID or None."""
    try:
        resp = requests.post(
            f"{api_base}/instances",
            json=config,
            headers=headers,
            timeout=10
        )
        if resp.ok:
            return resp.json().get("instance_id")
    except (requests.RequestException, ValueError):
        pass
    return None

def provision_with_fallback(primary_config, fallback_config):
    """Try primary provider first, fall back to secondary if unavailable."""
    instance_id = try_provision_gpu(
        SPHERON_API,
        {"Authorization": "Bearer YOUR_SPHERON_TOKEN"},
        primary_config
    )
    if instance_id:
        print(f"Provisioned on Spheron: {instance_id}")
        return "spheron", instance_id

    print("Spheron capacity unavailable, trying fallback provider...")
    time.sleep(2)
    instance_id = try_provision_gpu(
        PROVIDER_B_API,
        {"Authorization": "Bearer YOUR_PROVIDER_B_TOKEN"},
        fallback_config
    )
    if instance_id:
        print(f"Provisioned on Provider B: {instance_id}")
        return "provider_b", instance_id

    raise RuntimeError("No GPU capacity available from any provider")

This is illustrative, not production-ready. A real implementation needs idempotency, retry logic, and provider-specific error codes. But the core pattern is achievable in an afternoon and gives your team an availability hedge with minimal ongoing maintenance.

Spheron's REST API is well-suited as either the primary or fallback provider in this pattern: on-demand availability is strong, spot instances are available on most GPU types, and there are no reserved commitments required to get started. See the Spheron documentation for API reference and configuration details.

When Will the GPU Shortage End? The 2026-2027 Supply Forecast

Short answer: the acute phase is easing, and the market is not going back to 2023 conditions either. The most concrete evidence is on the packaging side. TSMC's CoWoS capacity is estimated to have roughly doubled, from around 70,000 to 80,000 wafers per month in 2025 to a projected 120,000 to 140,000 by the end of 2026, and analysts now expect the CoWoS supply-demand gap to narrow from about 20% to about 10% by the end of 2026. A gap that is halving is still a gap, which is the honest summary of where this goes: tight, not desperate, and improving unevenly by SKU.

Note that those are analyst estimates rather than TSMC guidance. TSMC does not publish a CoWoS capacity figure, so every number in circulation, including the ones above, is a modelled estimate.

The reason relief keeps arriving slower than the headlines suggest is structural. Capacity increases land at roughly the same time as architectures that consume them. HBM3e output rises just as HBM4 pre-production begins. CoWoS lines expand just as Rubin needs packaging slots. Supply is genuinely growing; it is also being absorbed close to the rate it appears.

MilestoneStatusWhat it changesWhat it does not change
CoWoS capacity expansionRamping through 2026Raises total NVIDIA and AMD DC GPU outputRubin absorbs a large share on arrival
Samsung and Micron memory rampLate 2026Eases H200 and B200 memory constraintMuch of the new output is HBM4, not HBM3e
Hopper displacement to BlackwellThrough 2027Returns H100 and H200 to general marketBlackwell supply stays allocation-gated
Blackwell Ultra (B300)Shipping since late 2025Already ramped, 288GB HBM3e per GPUConsumed allocation rather than easing it
Rubin launchRamping into production nowNew performance tierConsumes new CoWoS and HBM4 capacity

Read that table by column rather than by row. The middle column is what gets announced. The right-hand column is what determines whether you can book a GPU, and it is why "new capacity coming online" headlines have translated into availability more slowly than each announcement implied.

The mechanism most likely to deliver real relief is the third row, and it is the least discussed: as frontier labs migrate to Blackwell and then Rubin, their Hopper fleets return to the general market. A team that can run on H100 or H200 rather than insisting on the newest part available will find 2027 considerably easier than a team that has standardised on Blackwell.

Here is what the supply picture looks like input by input:

HBM3e capacity ramp. Samsung and Micron are ramping HBM3e production. Meaningful new capacity is expected online in late 2026, which should begin to ease on-demand price premiums for H200 and B200. This will not clear the existing backlog immediately, but the pressure should start to reduce in Q4 2026.

CoWoS packaging expansion. TSMC announced capital expenditure for CoWoS capacity expansion in 2024-2025. Meaningful new packaging capacity is expected online in H2 2026. This is the gating factor for GPU production volume, so when CoWoS capacity expands, output can ramp more quickly than raw HBM yield improvements alone would allow.

NVIDIA Rubin (R100). NVIDIA's next architecture after Blackwell is scheduled for H2 2026, with Rubin Ultra following in 2027. Pre-production allocations are already underway among hyperscalers. Rubin will not ease 2026 supply pressure. It will likely absorb most new CoWoS capacity as it comes online. Teams that want to secure a slot early can pre-order on the Spheron R100 page; see our Rubin GPU guide for architecture details.

Blackwell Ultra (B300). This one is worth correcting if you are working from an older roadmap: B300 is not a future part. Volume shipments began in late 2025, and NVIDIA now lists HGX B300 and HGX B200 as shipping, with 288GB of HBM3e and 8TB/s of bandwidth per GPU. It consumed allocation on the way in rather than easing it, which is the pattern every Blackwell-generation part has followed. See our B300 Blackwell Ultra guide if you are evaluating it.

The practical implication: GPU supply should be treated as a managed risk through at least 2026, not a procurement commodity. Multi-provider strategies and spot-plus-checkpoint patterns should be standing operating procedures, not emergency responses to shortage events. Teams that build these patterns now will be structurally better positioned when the next shortage cycle arrives, which, given the trajectory of AI compute demand, is when, not if.

On a related front, AI data center power constraints have emerged as the binding bottleneck as hardware availability improves. The power approval backlog is now the longer tail for teams planning new on-prem or colocation capacity.


If your team is hitting GPU availability walls on AWS, GCP, or Azure, Spheron gives you on-demand access to H100, H200, A100, and L40S inventory sourced from data center partners globally, with spot pricing available on most GPU types. No reserved contracts, no minimums, no egress fees.

On-demand H100 → | On-demand A100 → | View all GPU pricing →

Get started on Spheron →

FAQ / 12

Frequently Asked Questions

Yes, but it is narrower than it was in 2025 and it is easing. The constraint now sits in Blackwell parts and top-end Hopper SXM, where CoWoS packaging and HBM supply gate how many units NVIDIA can ship. A100 and L40S are broadly available in cloud, and RTX-class supply is limited by GDDR7 rather than data center allocation. The practical test is not whether GPUs exist but whether the specific SKU you want is bookable this week, in your region, without a reserved contract.

Nobody publishes a reliable figure. Neither NVIDIA nor the major OEMs disclose lead times, and circulating channel estimates for H100 alone span roughly 6 to 52 weeks depending on order size, region, configuration and whether you are an existing account. The dependable part is ordinal, not cardinal: Blackwell parts are hardest to source, Hopper SXM next, A100 and L40S easiest. Cloud rental sidesteps the question entirely, provisioning in minutes because the hardware is already racked and running. For a procurement budget, get a written quote from the distributor you would actually buy from.

On Spheron, H100 SXM5 is listed at $2.65/hr on-demand and $2.10/hr spot, H200 at $4.80/hr on-demand, A100 80GB at $1.43/hr, and L40S at $0.96/hr, as of 09 Sep 2026. Availability moves with partner inventory, so check the live catalog rather than a cached number. Hyperscaler on-demand H100 capacity is the least reliable tier during peak demand because reserved enterprise commitments are served first.

CoWoS (Chip on Wafer on Substrate) is the TSMC packaging process that bonds HBM memory stacks onto the GPU die on a shared silicon interposer. Every NVIDIA data center GPU from A100 onward requires it. TSMC's CoWoS lines are the narrowest point in the chain, so total GPU output is capped by packaging throughput rather than by how many GPU dies TSMC can etch. Expanding CoWoS takes new cleanroom capacity and roughly 18-24 months, which is why the bottleneck persists even as wafer supply improves.

It is easing rather than ending. TSMC CoWoS capacity is estimated to roughly double from around 70,000-80,000 wafers per month in 2025 to 120,000-140,000 by the end of 2026, and analysts expect the CoWoS supply-demand gap to narrow from about 20% to about 10% over the same period. That is a shrinking gap, not a closed one, and NVIDIA's Rubin generation is ramping into production now and will absorb much of the new capacity. Plan for top-end parts staying tight into 2027, with A100, L40S and RTX-class GPUs broadly available throughout. Those capacity numbers are analyst estimates; TSMC does not publish a CoWoS figure.

Three inputs push cloud GPU prices up at once: each HBM generation carries a price premium over the last and memory is a growing share of accelerator bill of materials, CoWoS packaging capacity is priced at a premium while it stays allocated, and hyperscaler forward orders remove inventory from the open market before it reaches spot. That raises the cost of sourcing hardware for every provider, so rental rates hold up even where a provider has inventory. Prices ease when supply of the input, not demand for the output, changes.

The four patterns that work are: rent from GPU-first providers instead of hyperscalers, since they have no internal workloads competing for the same inventory; run training on spot with automated checkpointing every 15-30 minutes; cut GPU count per workload through FP8 or INT4 quantization, which can halve the requirement; and route across two or more providers with a fallback script so a full spot pool at one vendor does not stop the pipeline.

NVIDIA H100 and H200 lead times are running 36-52 weeks due to constrained CoWoS packaging capacity at TSMC, surging HBM demand that exceeds SK Hynix and Micron production capacity, and a spike in hyperscaler reservation activity following the LLM arms race of 2024-2025. Consumer GPU production cuts of 30-40% have compounded supply pressure across the entire NVIDIA stack.

Spot GPU instances from cloud providers are typically 40-70% cheaper than on-demand rates and are available even when reserved capacity is fully booked. Providers like Spheron aggregate supply from multiple data centers, giving access to GPUs that are sold out on AWS, GCP, and Azure. Spot instances are preemptible, so pair them with automated checkpointing for training workloads.

Supply constraints from HBM shortages and CoWoS packaging bottlenecks are expected to persist through at least H1 2027. NVIDIA's Blackwell ramp is absorbing most new production capacity. Prices for H100 and H200 are unlikely to drop significantly before new HBM capacity from Samsung and Micron comes online in late 2026 to early 2027.

FP8 quantization reduces memory footprint by roughly 50% compared to FP16, meaning a model that required 8x H100 80GB can often run on 4x H100 80GB. INT4 quantization via GPTQ or AWQ cuts that further. For inference workloads with latency tolerance, quantized models on L40S or A100 GPUs can replace H100 deployments at a fraction of the cost, and H100 supply is where the shortage is most acute.

Multi-provider GPU orchestration means routing workloads across two or more GPU cloud providers based on price, availability, and latency rather than committing exclusively to one. This hedges against single-provider capacity constraints and spot preemptions. Tools like Kubernetes with multi-cloud node pools, or marketplaces that aggregate provider inventory, handle the routing layer.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute with no contracts and no minimum. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min