← The Bala Gazette

Field Notes · Investing

AI Infrastructure Investment Cycle
(2023 → 2030)

Understanding the bottlenecks, capital rotation, and next AI opportunities.

AI infrastructure investment cycle illustration

Executive Summary

The AI boom is not just a software story. It is a chain reaction across:

Every phase of AI scaling exposes a new bottleneck. Capital rotates toward:

  1. the most constrained layer,
  2. the highest pricing power,
  3. the hardest supply chain to expand.

This explains why the market moved:

The key insight:

AI investing is fundamentally "bottleneck investing."

Phase 1: Compute Bottleneck (2022–2023)"AI needs massive compute"

What triggered the boom?

The launch of:

created explosive demand for AI training infrastructure. Training models required:

Traditional CPUs were insufficient.

Why GPUs won

GPUs were ideal because they could:

This created a compute bottleneck.

Key winners

GPU leaders

Semiconductor manufacturing

Equipment suppliers

Why NVIDIA dominated

NVIDIA had:

The market realized:

AI compute = NVIDIA GPUs

This caused historic demand spikes.

Bottlenecks during this phase

1. GPU supply shortage

Demand exceeded production capacity.

2. Foundry capacity shortage

Advanced nodes were limited.

3. Advanced packaging shortage

Especially:

Result

GPU prices surged. Datacenter capex exploded. AI infrastructure became the dominant tech investment theme.

Enterprise infrastructure players
Enterprise Infra Players
Compute bottleneck supporting chart
Neo-cloud infrastructure players
Neo-cloud Infra Players

Phase 2: Storage Bottleneck (2023–2024)"AI training consumes enormous storage"

After GPUs became the focus, the next issue emerged:

AI models needed massive amounts of data storage and ultra-fast retrieval.

AI training required:

This suddenly increased demand for:

The bottleneck shifted to:

storage capacity + storage bandwidth.

Why storage stocks surged:

Major beneficiaries:

Storage bottleneck supporting chart

Key insight:

AI is fundamentally a data explosion cycle, not just a GPU cycle.

Phase 3: Memory Bottleneck (2024–2025)"Compute alone is not enough"

The problem discovered

As AI models became larger:

The limitation shifted from:

compute power

to:

memory bandwidth.

GPUs became starved for data.

Why HBM became critical

HBM (High Bandwidth Memory):

Without HBM:

Industry realization

The real AI accelerator stack became: GPU + HBM + packaging — not just GPU alone.

Key winners

Memory leaders

Memory bottleneck supporting chart

Why memory stocks exploded

HBM had:

Supply could not scale fast enough. This created: pricing power, margin expansion, multi-year demand visibility.

Hidden bottleneck: Packaging

Even if HBM supply improved, advanced packaging remained constrained. Critical technologies:

became essential.

Industry lesson

AI scaling is a systems problem, not just a chip problem.

Phase 4: Networking Bottleneck (Emerging Now)"How do 100,000 GPUs communicate?"

Why networking became critical

AI clusters evolved from single servers to hyperscale GPU fabrics. Large AI systems require:

The bottleneck shifted to:

interconnect bandwidth.

The core problem

Moving data between GPUs consumes:

At massive scale, communication itself becomes expensive.

Current industry focus

Optical networking

The next major transition is toward:

because copper networking is approaching physical limits.

Why optics matter

Optics offer:

Companies gaining attention

Networking / optics

SMH ETF covers some of them in this layer + chip layer.

Lumentum stock chart
Lumentum
Intel stock chart
Intel
Coherent stock chart
Coherent
Marvell stock chart
Marvell
Broadcom (AVGO) stock chart
AVGO
Intel stock chart, second view
Intel
Potential Future Upside
SegmentCurrent StagePotential Re-ratingCompanies
Optical NetworkingEarly-middle inningsStrongCiena, Nokia, Juniper
Optical ComponentsCorning, Coherent
Silicon PhotonicsEarly adoptionVery highCoherent, Lumentum, Intel
AI switching fabricsRapid hyperscaler demandStrongBroadcom, Cisco, Arista, Networks, NVIDIA Spectrum-X
AI Interconnect StartupsAyar Labs, Astera Labs, Credo
Credo positioning chart
Credo — solve future problems inside rack/system. Marvell for system to system.

Why this could be the next "HBM moment"

Characteristics are similar:

This often creates:

Phase 5: Power Bottleneck (Current + Future)"AI runs on electricity"

Biggest industry realization

AI datacenters consume enormous energy. Earlier datacenter racks: 5–15 kW. Modern AI racks: 100–150+ kW. Future AI systems may require gigawatt-scale campuses.

New constraints emerging

Hyperscalers now face:

This changes AI from a:

semiconductor problem

to a:

national infrastructure problem.

Where spending is increasing

Electrical infrastructure

Demand rising for:

Key beneficiaries

Important insight

Power infrastructure scales slower than semiconductors. This creates:

Phase 6: Cooling Bottleneck"AI generates enormous heat"

The issue

Dense AI clusters produce:

Traditional air cooling is becoming inadequate.

Industry shift

Datacenters are moving toward:

Why cooling becomes critical

Without proper cooling:

Cooling becomes essential infrastructure.

Why investors care

This creates new demand for:

This area remains relatively under-owned compared to semiconductors.

Phase 6 (Parallel): Agentic / Orchestration Bottleneck (2026+)"Running millions of AI agents efficiently"

AWS × Meta Graviton AI partnership announcement

Agentic orchestration bottleneck illustration

After solving:

a new constraint is emerging:

how to manage and scale AI systems in real-world usage.

What changed?

AI is moving from model training to continuous execution (agents, copilots, workflows). Instead of one model run, systems now involve:

The new bottleneck

Not GPU. Not memory. It is:

orchestration compute (CPU + system layer)

Why CPU becomes critical again

These workloads are:

Handled mainly by CPUs, system memory (DDR), and backend infrastructure.

Where pressure builds

1. Agent concurrency

Thousands to millions of agents running simultaneously.

2. Context management

Frequent reads/writes from memory systems.

3. Tool execution

API calls, DB queries, external integrations.

4. Scheduling & coordination

Managing workflows across systems.

Resulting bottlenecks

Key beneficiaries

Compute / CPU

Cloud / orchestration layer

Why this matters

The next wave of AI demand is not:

"train bigger models"

It is:

"run AI continuously at scale"

This shifts optimization toward: cost per request, latency, efficiency, orchestration.

Key insight

GPUs power intelligence.
CPUs coordinate intelligence.

As AI moves into real-world deployment, coordination becomes the bottleneck.

Investor takeaway

This phase may not create a sudden spike like GPUs or HBM. But it can drive:

If earlier phases were about building intelligence, this phase is about operating intelligence at scale.

Phase 7: Inference Era (Future)"Inference may become larger than training"

What changes?

Today's AI boom is training-heavy. Future AI demand may come primarily from:

New optimization goals

Focus shifts from raw FLOPS toward:

Likely future winners

Custom AI chips

AI devices

Full Chain Reaction of AI Infrastructure

  1. AI models grow larger
  2. Need more GPUs
  3. GPU shortage emerges
  4. Need more HBM memory
  5. Memory shortage emerges
  6. Packaging capacity becomes constrained
  7. GPU clusters become enormous
  8. Networking bandwidth becomes bottleneck
  9. Power consumption explodes
  10. Cooling infrastructure becomes critical
  11. Inference efficiency becomes dominant

The AI Infrastructure Stack

LayerBottleneckMajor Cost DriverKey Companies
ComputeGPU shortageAI acceleratorsNVIDIA, AMD
ManufacturingFoundry capacityAdvanced nodesTSMC
MemoryHBM shortageBandwidth scalingSK hynix, Micron
PackagingCoWoS limitsIntegration complexityTSMC ecosystem
NetworkingData movementOptical interconnectsBroadcom, Arista
PowerElectricity demandGrid infrastructureEaton, Vertiv
CoolingThermal densityLiquid coolingVertiv ecosystem
InferenceCost/queryEfficient AI chipsFuture ASIC leaders

Key Framework for Retail Investors

The biggest AI opportunities usually emerge where there is:

1. Demand explosion

AI adoption accelerates rapidly.

2. Limited suppliers

Only a few companies can manufacture the solution.

3. Long expansion cycles

Capacity takes years to build.

4. High switching costs

Customers cannot easily replace suppliers.

5. Mission-critical infrastructure

The AI ecosystem cannot function without it.

Most Important Insight

The AI investment cycle is not random. Capital rotates toward:

the next infrastructure bottleneck.

The sequence so far:

  1. Compute
  2. Memory
  3. Networking
  4. Power
  5. Cooling
  6. Inference efficiency

Understanding this rotation early is where asymmetric investment opportunities emerge.