The dominant story about Nvidia went like this: for the first years of the AI boom it was the only source of state-of-the-art GPUs, which became immensely profitable as the industry scaled. In recent years hyperscalers like Amazon and Google started building their own chips, so Nvidia is no longer the only game in town — and investors began wondering how durable its advantage really is.
It is a compelling story and mostly true. After growing its market cap tenfold between the start of 2023 and mid-2025, Nvidia shares have been on a more modest trajectory for the past year, driven by concerns about GPU competition.
Since the company's latest earnings, a new narrative has taken shape: Nvidia's advantage goes far beyond GPUs. As AI compute grows into the gigawatt scale, orchestration has become an increasingly complex task — and Nvidia has already built much of the hardware needed to handle it.
The system, sold rack by rack
You can see it by looking at what the company actually sells. Nvidia is currently rolling out its Vera Rubin architecture, which pairs the Rubin GPU with a Vera CPU, an inference accelerator and similar racks for storage and networking.
What ships in the Vera Rubin package:
- Rubin — the GPU itself
- Vera CPU — the processor responsible for orchestrating data
- Inference accelerator — a separate unit that handles token generation
- Storage and networking racks — the paths data arrives on and leaves by
Those systems are as specialised as the GPU itself, but their job is not churning through tokens: it is making sure everything outside the GPU works as efficiently as possible. If the GPU is the engine, these are the rest of the car.
The real bottleneck: data at the right moment
The Vera CPU focuses specifically on orchestrating data. In the words of Jason Hardy, Nvidia's VP of storage technology: "Vera is important because there's only so much memory that you can put in a single server or any sort of compute platform."
As data centres scaled up compute, memory capacity scaled up too — which is why memory makers grew rich in the second wave of the infrastructure boom. But getting that data to the GPU at the right time is not straightforward. As companies push tokens-per-watt lower, they are realising how important that kind of traffic direction is.
Hardy's figure: "We saw upwards of 3x improvement in these operations, where the Vera CPU is allowing for acceleration. So now we can use our flash to its fullest potential, because we can get all that performance out of it without bottlenecking."
Same problem, different solution
Versions of the same problem show up outside Nvidia. When OpenAI developed its Jalapeño chip, a major focus was avoiding these challenges entirely by minimising the amount of data that needs to be moved.
From the company's blog post: "We designed Jalapeño to minimize data movement and communication delays. Its large domain allows the entire workload to remain within one connected system."
A different approach — avoiding data movement by keeping the workload inside one integrated chip — but the same logic: increasing efficiency with smarter traffic control instead of just more processor cycles.
What changes
This new focus is not an automatic win for Nvidia. The company will have to compete with rival chipmakers and hyperscalers just as it has on GPUs. But the competition has moved to a new layer, where building a rival GPU matters less than being able to make the entire system work efficiently.
And at least in these early stages, Nvidia looks to have a commanding lead.