Nvidia's Vera Rubin reaches customer production as AMD buys an AI lab

CoreWeave says Cognition is the first customer running production work on Vera Rubin NVL72. Two days earlier, AMD agreed to buy World Labs for about $8.2 billion in stock.

ByShajanthanFounder & Editor
Published
Reading4 MIN
Diagram of the six chips that make up Nvidia's Vera Rubin NVL72 rack.
What’s new, in 20 seconds
  1. On September 30, 2026, CoreWeave made Nvidia's Vera Rubin NVL72 available. Cognition reports up to 4.8x more inference throughput than on GB200 NVL72, a customer benchmark that has not been independently reproduced.
  2. On September 28, AMD agreed to acquire Fei-Fei Li's World Labs in an all-stock deal valued at about $8.2 billion, to steer its chip roadmap with in-house model research.
  3. Microsoft's Maia 300, reported in August as possibly launching as early as September, had not been announced by Microsoft as of October 8.
Contents

Nvidia's next-generation Vera Rubin platform is now running paying production work. On September 30, 2026, CoreWeave announced that Vera Rubin NVL72 is available on its cloud. It named Cognition, the company behind the Devin coding agent, as the first customer anywhere running production workloads on the system. Two days earlier, AMD made a different kind of chip bet: it agreed to acquire World Labs, the AI model lab led by Fei-Fei Li, for about $8.2 billion in stock. Both moves matter for the AI datacenters now being built at gigawatt scale (see our tracker of those projects).

What CoreWeave announced

In its September 30 release, issued at its Fully Connected event in San Francisco, CoreWeave said customers can run Vera Rubin NVL72 under the same operating model and tooling as their GB200 and GB300 NVL72 fleets.

Cognition stood up a Vera Rubin NVL72 cluster in early September, according to the release. Its engineers then ran benchmarks against a GB200 NVL72 baseline. Cognition reports:

  • up to 4.8x higher total token throughput on its SWE-2 inference workloads
  • 3.8x higher output-token throughput for reinforcement learning

"Our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads," said Silas Alberti, Cognition's SVP of research. Cognition says this means more concurrent Devin sessions per GPU and a lower cost per session.

These are a customer's numbers on its own workload, published in a vendor's press release. CoreWeave calls them "independent benchmarks". They are useful evidence of real-world gains, but nobody outside these companies has reproduced them. CoreWeave also claims 10x the token throughput per megawatt of GB200 NVL72 on DeepSeek R1 at matched interactivity, which is its own measurement.

The release doesn't say whether access is general or limited, or which regions have capacity. According to DCD, CoreWeave also plans to offer Nvidia's Vera CPU as standalone bare-metal racks, with 128 CPUs and 11,264 cores per rack, and some customers will begin testing in the coming weeks.

Why it matters

Vera Rubin has been "in full production" since at least Nvidia's August 26, 2026 earnings. Nvidia said then that racks were running at partners including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius. That quarter, Nvidia reported revenue of $96.2 billion, $89.0 billion of it from data centers. It guided to about $108 billion for the current quarter, assuming no data center compute revenue from China.

"Racks running at partners" and "a customer serving production traffic" are different milestones. The CoreWeave announcement is the first public, named example of the second that we found. It matters for three reasons:

  1. Inference economics. If Cognition's gains hold for other agentic workloads, the cost of serving long-running coding agents falls. That cost feeds directly into API and subscription pricing (see why token pricing isn't the whole story).
  2. The 2H 2026 commitments. OpenAI's ≥10 GW deal with Nvidia, announced September 22, 2025, said the first gigawatt would arrive in the second half of 2026 on Vera Rubin. Customer production use on a neocloud doesn't confirm that milestone, but it shows the platform is past bring-up.
  3. Pressure on AMD. AMD says its Helios racks (72 MI455X GPUs each) are in production, and that OpenAI expects to bring Helios online starting in Q4 2026. AMD's own estimate claims up to 30% more tokens per dollar than Vera Rubin NVL72. That claim will now be tested against real Rubin deployments, not spec sheets.

AMD buys a model lab

On September 28, 2026, AMD signed a definitive agreement to acquire World Labs in an all-stock deal valued at about $8.2 billion. Founded in 2024 and based in San Francisco, World Labs builds "spatial intelligence" models that generate and simulate interactive 3D environments. After closing, Li will join AMD as executive vice president and chief scientist, reporting to CEO Lisa Su. AMD expects the deal to close by the end of 2026, subject to regulatory approval.

"Building the compute platforms for the next generation of AI requires a deep understanding of how models are evolving," Su said. AMD's release names no Instinct product. It says World Labs' research will inform AMD's future hardware, software and systems roadmaps.

What's confirmed: the price, structure, leadership role and expected close. What's uncertain: how a 3D world-model lab changes AMD's GPU designs, and whether regulators will review the deal closely. It is still an unusual move: a GPU vendor buying a model lab to shape its roadmap, rather than only investing in or partnering with model developers.

Other chip news in the window

  • September 28: Nvidia added $150 billion to its share buyback authorization, leaving $235 billion available.
  • September 30: Synopsys and Amazon signed a $1 billion multi-year IP agreement supporting Amazon's custom chips, including Trainium and Graviton, according to DCD.
  • September 22: Alibaba unveiled its Zhenwu V900 AI chip and set a target of 20 GW of Alibaba Cloud data centre capacity by 2032, according to Bloomberg (via The Star). In May, DCD reported that the V900 was planned for the third quarter of 2027. Alibaba hasn't published full specifications or a shipping date in the coverage cited here.
  • Not yet: in August, TrendForce, citing a Reuters report on The Information's reporting, said Microsoft could unveil Maia 300 "as early as September" and was eyeing capacity for more than 300,000 chips for 2027 delivery. We found no Microsoft announcement as of October 8. We also found no new Google TPU, AWS Trainium or OpenAI–Broadcom milestones in this period.

What to watch next

  • First confirmed Helios deployment at OpenAI (AMD says from Q4 2026).
  • Hyperscaler, not just neocloud, production announcements on Vera Rubin.
  • Independent benchmarks of Rubin against GB300 and MI455X on the same workloads.

For how OpenAI's chip deals with Nvidia, AMD and Broadcom fit its business, see OpenAI explained.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on AI datacenters & hardware

The Week in AI

Get the cluster, not just the headline.

0