The Five-Layer Cake

A Cal Bay AI℠ Essay

Introduction: The Alchemy of Electrons to Tokens

The global AI industry is no longer a collection of “computer companies.” It has reorganized into a massive, interconnected AI Factory — one that performs a modern form of alchemy: transforming raw electricity (electrons) into digital intelligence (tokens).

A token is not merely a unit of text. Just as one molecule can be more valuable than another, the value of a token is determined by the artistry and science of its transformation. Nvidia’s stated philosophy is to do “as much as necessary, but as little as possible,” partnering with an entire ecosystem so the most valuable intelligence is produced with the least friction.

“The input is electrons, the output is tokens. In the middle is Nvidia.” — Jensen Huang

To understand how this factory works, we peel back the five layers of the “AI cake” — from the physical bedrock of energy to the digital workers of the future.

The Five Layers at a Glance

LayerNameRole in the AI Factory
5ApplicationsDigital workers — AI agents that use tools and solve problems
4SoftwareThe ecosystem that gives hardware its instructions
3SystemsRack-scale AI factories built through extreme co-design
2HardwareThe engines that convert power into computation
1EnergyThe physical bedrock and absolute limit

Layer 1: Energy — The Physical Bedrock

Energy is the absolute physical limit of the AI revolution. While the world focuses on chips, energy is the real constraint for the United States, which operates in a regime of energy scarcity. China, by contrast, has adopted a strategy of energy abundance, using its vast infrastructure to overcome technological hurdles.

FeatureEnergy Scarcity (U.S. Approach)Energy Abundance (Chinese Approach)
Primary GoalMaximize performance per watt.Maximize total throughput per area.
Hardware StrategyCutting-edge chip nodes (N3/N2) to save power.Older nodes (7nm) ganged together at scale.
InfrastructureHighly optimized, dense data centers.Fully powered “ghost data centers” and “ghost cities.”
Key AdvantageHigh throughput per limited watt.Energy is cheap or abundant, so older chips remain viable.

Energy is the fuel for the revolution, but it requires specialized engines to convert that power into calculations — which leads to the hardware layer.

Layer 2: Hardware — The Engines of Intelligence

General-purpose computing (the CPU) is a Cadillac: a comfortable, easy-to-drive cruiser that simply cannot reach the velocities AI requires. Intelligence at scale demands accelerated computing — the F1 racer of the digital world. It is harder to drive and requires expert tuning, but it is the only vehicle capable of breaking through the limits of Moore’s Law.

This engine is built from three essential sub-components. They are no longer just “parts” but scarce resources, secured through Nvidia’s purchase commitments reported at more than $100 billion:

ComponentWhat It IsKey Suppliers
1. Logic DiesThe primary processors, built on the most advanced manufacturing nodes (N3/N2).TSMC
2. High-Bandwidth Memory (HBM)Specialized memory that eliminates the data bottleneck.SK Hynix, Micron, Samsung
3. Advanced Packaging (CoWoS)The technology that links logic and memory into a single functional unit.TSMC

As AI models grew, the problem became too large for a single engine. That forced the jump from individual chips to massive, integrated systems.

Layer 3: Systems — The Rise of the AI Factory

The industry has shifted from chip-scale to rack-scale design. We no longer just build computers; we build AI factories. Modern systems such as Nvidia’s NVLink 72-GPU racks and the Vera Rubin pod are feats of extreme co-design, where the physical layout is dictated by the needs of the algorithm.

System Vital Signs: Vera Rubin Pod (as described by Nvidia)

Vital SignFigure
Transistor Count1.2 quadrillion transistors
Component Density1.3 million components per rack
Raw Compute60 exaflops (FP4)
Bandwidth10 petabytes per second
Infrastructure RequirementLimited by plumbers and electricians

This massive wall of hardware is a silent giant until it is given instructions. That is where the software ecosystem — the true treasure of the industry — comes into play.

Layer 4: Software — The Treasure of the Ecosystem

CUDA is the industry’s “great treasure.” Its value lies in being the language of researchers and in its massive installed base of hundreds of millions of GPUs. Developers target CUDA first because it offers the largest possible audience, creating a self-reinforcing flywheel: more users lead to more tools, which lead to more users.

The software moat rests on three pillars:

PillarWhat It Means
RichnessThousands of domain-specific libraries — such as cuLitho for chip lithography or cuDF for data processing — that solve problems others haven’t yet identified.
VersatilityThe ability to run on any cloud (AWS, Azure, Google, Oracle), on-premises, or inside a robot.
LongevityTwenty years of backward compatibility, so the software is thoroughly proven and trusted by the world’s most demanding labs.

Once software is running on the system, we reach the final layer — where people interact with AI.

Layer 5: Applications — The Era of Digital Workers

The application layer is moving from simple chatbots to AI agents — the “iPhone of tokens.” This is the moment AI becomes a digital worker rather than a search tool. Agents do not just output text; they use tools, access file systems, and perform research to solve complex problems.

Legacy Computing (Warehouses)AI Computing (Factories)
Retrieval-based: Finding pre-recorded files.Generative-based: Creating new tokens in real time.
Static storage: A place to store information.I/O subsystems: Agents using tools (for example, OpenClaw).
Tool users: Humans using software.Digital workers: AI accessing ground truth to solve tasks.

Together, these five layers create a new economic reality that transforms the nature of human productivity.

Conclusion: The Inevitable Scale

The AI revolution is not about replacing humans; it is about elevating them.

The Radiology Analogy: When AI vision became superhuman at reading scans, many predicted the end of radiologists. Instead, it helped address a global shortage — allowing radiologists to focus on their true purpose, patient care, while AI handled the repetitive work of scanning.

The “Speed of Light” Heuristic: To thrive in this era, test every problem against the physical limits of physics and reason from first principles. If a process takes 74 days, ask why it doesn’t take six. By breaking problems down to their speed-of-light potential, we can navigate this industrial revolution with clarity.

Final Learning Insight

The ultimate goal of the five-layer cake is to drive the cost of a token toward zero. As intelligence becomes a commodity, the humanity we bring — our compassion, our character, and our determination — becomes the ultimate premium.

Intelligence is now a utility; humanity is the frontier.