The Vera Rubin NVL72 from NVIDIA has ceased being a roadmap part. On September 30, 2026, the product went into production by CoreWeave to serve an actual customer using real workloads.
The company called Cognition, which developed an AI coding assistant called Devin, is the said customer. CoreWeave revealed it during its Fully Connected event held in San Francisco.
There are tons of articles describing the event as a milestone achievement. Very few of them try to distinguish facts from marketing claims, and almost none of them discuss its pricing.
It’s more than just another customer story because Vera Rubin is the upcoming generation of GPUs by NVIDIA after Blackwell. The rate at which it will make its way into actual workloads will dictate the pricing and availability of GPU clouds in the coming year.
The AI Compute Threshold Report
We analyzed pricing from 150+ GPU cloud providers to find the exact threshold where an AI startup's OpenAI API bill eclipses the cost of a dedicated H100 cluster.
Read the Full ReportThat’s precisely what this article is going to do.
What’s actually being offered
Vera Rubin is not one component either. The NVIDIA platform is made up of Rubin GPUs, Vera CPUs, and even the whole NVL72 system containing both.
The NVL72 system consists of 72 Rubin GPUs and 36 Vera CPUs connected via NVLink. This is what CoreWeave has just brought into production, rather than one GPU only.
And this will become relevant further down this article. The “price of Vera Rubin GPU” and “price of Vera Rubin NVL72” are different issues. Mixing them up leads to incorrect pricing information being reported.
As reported by NVIDIA in May 2026, Vera Rubin had entered full production mode. NVIDIA’s investor relations news report indicated that it expected to start shipments in the fall.
The Rubin GPU itself is an impressive leap on paper. TSMC manufactured it using 3-nanometer process technology, incorporating 336 billion transistors each and 288GB of HBM4 memory per GPU.
NVLink 6, which connects the chips together within the rack, provides a bandwidth of 3.6 TB/s for each GPU. This represents a 50% increase compared to the previous generation. And it’s one of the features that allow 72 GPUs to operate as one unified system.
Also, NVIDIA launched its NVIDIA Spectrum-X Ethernet Photonics in production, alongside the GPUs. This product is a combination of co-packaged optics and Spectrum-X switching. NVIDIA states that this product specifically enables AI factories with one million GPUs or more.
These specifications aren’t under debate–NVIDIA provides them in its technical information about the product. The only thing that has changed since September 30 is that the rack, which uses all these components, is now being used by one of the customers.
The announcement itself
CoreWeave announced this at their Fully Connected Conference in San Francisco on September 30, 2026: Vera Rubin NVL72 is now available to a select few on CoreWeave Cloud.
Fully Connected is CoreWeave’s first-ever user conference. It took place from September 29 to October 1 in Moscone South. The conference program states that more than 2,000 AI leaders and engineers attended.
Back in early September, Cognition joined forces with CoreWeave and launched the Vera Rubin NVL72 cluster. This configuration helps Devin to be trained, reinforced, and produce inference on the CoreWeave platform.
Cognition states that they went from bridge capacity to thousands of GPUs for both training and inference in less than nine months. CoreWeave engineers completed the first-ever benchmark of Vera Rubin, comparing it to GB200 NVL72.
It wasn’t the first time that CoreWeave had shown such leadership in this industry. They were one of the first cloud service providers to install Dell PowerRack with NVIDIA GB200 and GB300 NVL72 families.
And Vera Rubin follows the trend. CoreWeave considers itself to be the provider that puts the latest generation of NVIDIA up and running before the competition.
The performance numbers
CoreWeave and Cognition have provided two sets of numbers along with their announcement. Neither independent laboratory provided either set of numbers, so we must regard both as assertions.
It is important to note that Devin, Cognition’s AI that develops and distributes software autonomously, has some relevance in the case. Cognition runs the benchmark SWE-2 inference to target agentic coding tasks, rather than serving as a general-purpose benchmark.
The first set of numbers is from Cognition itself. According to Silas Alberti, the company’s SVP of research, there is up to 4.8x token throughput on SWE-2 inference relative to a GB200 NVL72 baseline. Besides that, Cognition reports a 3.8x improvement in output token throughput for reinforcement learning, with matched interactivity.
CoreWeave provides the second set of numbers. It claims that Vera Rubin NVL72 provides 10x the token throughput per megawatt of GB200 NVL72, according to tests with DeepSeek R1.
Both sets of numbers are based on actual measurements, not marketing. However, both originate from the side with an obvious interest in making Vera Rubin look good, and neither of them has independent confirmation at this point.
What else CoreWeave announced that day
Vera Rubin isn’t the only story coming out of Fully Connected. CoreWeave used the event to launch a standalone Vera CPU product that doesn’t depend on the entire NVL72 rack.
CoreWeave also launched the CoreWeave Partner Network, which is a network of technology partners that have validated integrations with the CoreWeave platform. The confirmed members of this network include CrowdStrike, VAST Data, Reflection, and ClickHouse.
There’s also a newly announced customer contract by CoreWeave with Ennoble Care. This isn’t as big a story as the Vera Rubin news. But it shows that the announcement was more of a platform play than just one product.
Where else Rubin is headed
But Vera Rubin isn’t just found on CoreWeave’s website either. NVIDIA’s Q2 fiscal 2027 earnings filing highlights four additional customers already using Vera Rubin racks.

These are Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius. The SEC filing provides this information directly, rather than a press release or blog post.
Third parties have also pointed to Lambda and Nscale as Rubin’s other two partners. NVIDIA’s SEC filing excludes both, so consider this list as reported rather than confirmed.
All of this doesn’t mean that every partner currently has Vera Rubin in production, though. CoreWeave is currently the only company with a named customer using Vera Rubin.
The industry analysts tracking NVIDIA’s rollout schedule have the broader deployment around Q4 2026. This date is based on third-party reporting, not an official promise from NVIDIA.
This time, context comes to the rescue. NVIDIA has said that its previous generation, Blackwell, has had the “fastest product ramp” in NVIDIA’s history.
The industry has already deployed huge numbers of GPUs worldwide. There are close to twice as many partners with data centres above 10 megawatts as there were just one year ago, at more than 80 facilities. Rubin is entering an already vast user base that is skyrocketing.
There is even greater demand than that. At GTC 2026, NVIDIA CEO Jensen Huang quoted figures. He stated that the company expects total orders worth $1 trillion of Blackwell and Vera Rubin GPUs by 2027.
What it costs
There have been no published prices on Vera Rubin NVL72. Neither NVIDIA nor CoreWeave nor any of the other listed partners.
This was true even before this launch. As one tracker stated: Any current price per hour for Vera Rubin is nothing but an estimate.
This is quite common practice with new hardware. There have been no published prices on Vera Rubin NVL72 by NVIDIA, CoreWeave, or any of the other listed partners. Our DGX Cloud pricing article explains this in more detail.
This should be the case with Vera Rubin as well. The public price will become available once this product becomes more widely available.
Experience tells us that this expectation will prove correct. AWS reduced H100 pricing by 44% in June 2025, months after the launch of this chip, because of increasing supply and declining demand.
The reduction in A100 prices was 33%. Vera Rubin’s price should follow a similar trajectory.
How to get access
CoreWeave calls it a limited availability release and not a full release. This means customers can access the service directly through CoreWeave based on engagement with certain clients, such as Cognition.
CoreWeave also offers CoreWeave ARENA (AI Ready, Native Applications). ARENA is an evaluation process and not self-service immediate access to any GPU, including Vera Rubin. At present, ARENA is available only to current customers of CoreWeave, while CoreWeave will begin processing new team applications in Q2 2026.
“Test” has a special definition–the team makes success criteria with CoreWeave and runs the workload within a guided notebook interface. The standard benchmark timeframe is about 14 days, according to CoreWeave’s FAQ. Teams can pay for evaluations, and CoreWeave refunds this money in case of moving to the production phase.
The number of allocated GPUs is determined by the region, capacity, and goals of the evaluation and offered by CoreWeave during the qualification process. Information on the accessibility of Vera Rubin NVL72 via ARENA from the materials of CoreWeave is not available.
FAQ: Vera Rubin NVL72 questions
Does the 4.8x throughput claim mean Vera Rubin is 4.8x faster overall?
No, the 4.8x refers to Cognition’s SWE-2 Inference workload only. Another model or workload might yield a lower gain, a higher gain, or even something in between.
Will CoreWeave’s benchmark results carry over to a different workload?
False. That’s what CoreWeave ARENA is for, but it currently serves only existing customers, not the public.
Is Vera Rubin NVL72 the same as DGX Rubin?
Close, but not quite. DGX Rubin is actually a branded server from NVIDIA for direct sale to consumers. NVL72 is the name of the architecture that partners can implement and sell.
How fast did CoreWeave go from bring-up to production?
Quickly. Cognition built its cluster in early September, and CoreWeave proclaimed availability on September 30–less than a month later.
Was the Vera Rubin news the only thing CoreWeave announced that day?
No, CoreWeave also introduced a Partner Network and an independent Vera CPU service, and a contract for a customer named Ennoble Care.
Conclusion
Vera Rubin NVL72 is real, live, and running production workloads at one provider. That’s confirmed, not speculation.
The performance numbers are genuinely impressive. But they’re self-reported by the two companies with the most reason to make them look good. Treat them as a starting point, not a verdict.
Pricing is the one question nobody can answer yet, including this article. Watch for that to change once Rubin moves past this first limited rollout.
The next signal to watch is simple: does a second cloud provider announce its own named customer on Vera Rubin. That would confirm this isn’t a one-provider story.
Get personalised, no-commitment quotes from top AI infrastructure providers in under 2 minutes.



