CoreWeave announced the availability of NVIDIA Vera Rubin NVL72 systems on its cloud platform, powered by Spectrum-X 102.4T Ethernet networks, in a move aimed at shifting agentic AI workloads from training and experimentation into production operation. The company says Cognition, the developer of the Devin agentic software engineer, has become the first customer to run production workloads on Vera Rubin.
The announcement comes as part of the CoreWeave Fully Connected event in San Francisco. It also includes plans to make the NVIDIA Vera processor available, designed for AI agent workloads, and the launch of the CoreWeave Forge environment to connect the training, evaluation, and continuous optimization of models and agents.
Higher Performance for Devin Workloads
Cognition uses CoreWeave infrastructure to train Devin, conduct reinforcement learning, and run inference in production. After CoreWeave received the first production Vera Rubin NVL72 racks, Cognition compared performance using a set of software engineering tasks taken from FrontierCode and deployed agents to solve them.
Initial tests showed up to a 4.8x increase in aggregate token generation rate for SWE-2 inference workloads compared with a GB200 NVL72 baseline. In practice, according to the source, the result means faster code generation and better responsiveness during Devin’s multi-step reasoning phases. The text did not provide sufficient details about the full measurement methodology or the test conditions determining how comparable the figure is across different environments.
One Platform for the AI Lifecycle
CoreWeave enables Vera Rubin capacity through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes, and CoreWeave Inference. The company also announced that CoreWeave Forge will bring together Weights & Biases tools, OpenPipe’s post-training expertise, and the open-source marimo project in a single environment, while remaining open to different models, frameworks, and clouds.
Forge includes the generally available CoreWeave ARIA service for analyzing experiments, suggesting modifications, and saving them to GitHub, as well as the new CoreWeave Agent Lens service for turning production agent monitoring data into inputs for optimization. CoreWeave says Agent Lens improved failure detection by 20% and cut the cost of fixing failures in half, without providing independent details on how these percentages were calculated.
CoreWeave Sandboxes is also now generally available for running agents, tool calls, reinforcement learning, and evaluation inside isolated environments on CPU or GPU units. The company says serverless reinforcement learning runs 1.4 times faster and at 40% lower cost than a self-managed setup.
What Changes in Practice?
The most important shift is not merely the addition of new hardware, but an attempt to narrow the gap between running an agent, collecting its signals, and then reusing them in training. NVIDIA Dynamo, an open-source inference framework, supports the managed inference service and the RL Rollouts feature currently available in private preview, which can load new checkpoints into a live deployment without redeployment.
CoreWeave says the Vera processor can place 128 CPU units and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments when allocating one core to each environment. In testing, the startup process for agent environments accelerated by more than three times, while the processor achieved a 1.7x gain across successful tasks in Terminal-Bench.
Canva, Capital One, and MasterClass use Forge, while Ennoble Care selected reserved capacity of NVIDIA RTX PRO 6000 to run clinical AI inference. However, most of the performance and cost figures cited here were provided by CoreWeave or its customers, and therefore require independent review before being considered general indicators for all agentic AI workloads.