CoreWeave said on September 30, 2026, that it will add the NVIDIA Vera CPU to its cloud platform, one of the first large commercial rollouts of a processor NVIDIA built specifically for AI agents rather than general-purpose computing. The announcement landed at Fully Connected, CoreWeave’s AI cloud conference in San Francisco, where the company also detailed a rack-scale system packing 128 Vera CPUs and 11,264 cores into a single cabinet. CoreWeave says that configuration supports more than 11,000 concurrent agent environments, and that agent sandboxes start more than three times faster on Vera than on an x86 processor in its own tests.
The move reframes part of what “AI infrastructure” has meant through 2026. For three years the industry conversation around AI compute centered almost entirely on GPUs, how many H100s, B200s, or now Rubin chips a cloud provider could rack up. NVIDIA Vera CPU targets something different: the unglamorous but growing work that surrounds a model’s actual reasoning step, tasks like spinning up code sandboxes, running reinforcement-learning environments, and moving data fast enough to keep expensive GPUs fed with work.
What Is the NVIDIA Vera CPU, and Why Build One for Agents
According to NVIDIA’s Vera CPU product page, the chip is built around 88 custom Olympus cores and uses a technique NVIDIA calls Spatial Multithreading to expose 176 threads with partitioned core resources. A second-generation NVIDIA Scalable Coherency Fabric ties the cores, cache, memory, IO, and NVLink-C2C together across a single compute die, delivering 3.4 TB/s of bisectional bandwidth, the company states. NVLink-C2C then links Vera to NVIDIA GPUs at up to 1.8 TB/s of coherent bandwidth, letting CPU and GPU share memory space instead of shuttling data over a slower PCIe connection.
On memory, NVIDIA says Vera reaches up to 1.2 TB/s of bandwidth on LPDDR5X with as much as 1.5 TB of capacity per chip. The company claims that setup delivers twice the bandwidth, and three times the bandwidth per core, of what it calls leading x86 CPUs running DDR5. For the compilation, code-analysis, and Python tool-chain work that make up much of an AI agent’s reasoning loop, NVIDIA says Vera runs up to 1.8x faster than those x86 rivals. Those figures come from NVIDIA’s own measured data, baselined against current-generation x86 chips, and the company notes they are subject to change as software matures. NVIDIA’s engineering team detailed the core design in a separate developer blog post focused on single-threaded performance for agentic workloads.
Inside CoreWeave’s 11,264-Core Rack
CoreWeave’s rack-scale Vera deployment fits 128 CPUs and 11,264 cores into one rack, connected with NVIDIA BlueField-4 DPUs and Spectrum-X Ethernet switching. The company says that density supports more than 11,000 concurrent agent environments, assuming roughly one core per environment, and that it runs Vera on bare metal under the same platform, consumption model, and pricing structure used across the rest of its fleet. CoreWeave executive vice president of product and engineering Chen Goldberg said the company’s platform natively supports Vera through products such as CoreWeave Sandboxes, letting teams spin up thousands of isolated environments without added operational overhead, according to the company’s announcement.
CoreWeave launched CoreWeave Sandboxes on May 14, 2026, as an execution layer for reinforcement learning, agent tool use, and model evaluation, available through CoreWeave Kubernetes Service or as a serverless runtime through Weights & Biases. Each sandbox runs in its own isolated virtual environment by default, so a runaway process in one cannot spill into another, the company has said.
The Agent Loop CoreWeave Is Building Around
CoreWeave frames agentic AI as a continuous loop: run, observe, curate, improve, and evaluate. GPUs handle the training and reasoning steps inside that loop, while CPUs run the surrounding work, isolated sandboxes, reinforcement-learning environments, tool calls, code execution, and data pipelines. The company describes demand for that CPU-side work as bursty and hard to predict. A single step in the loop can call for thousands of environments for an hour, then almost none until the next training run starts. As post-training and agentic workloads grow, CoreWeave argues, that CPU-intensive work increasingly sets the pace of the whole AI development cycle, which is the gap Vera is meant to close.
CoreWeave Sandboxes and the 3x Startup Claim
CoreWeave says one of the clearest performance measures for agentic AI is how many isolated agent environments a team can run at once, and whether per-sandbox performance holds steady as that count rises. “General-purpose infrastructure bottlenecks agentic AI; Vera is the first CPU explicitly designed to accelerate it,” Chen Goldberg said in the company’s announcement. In CoreWeave’s own testing, agent sandbox startup was more than three times faster on Vera than on an x86 CPU, though the company did not name the comparison chip or share enough benchmark detail to judge how broadly the result applies, a gap flagged in TheEnergyMag’s coverage of the announcement.
NVIDIA Vera CPU: Key Specifications
The table below pulls together the specifications NVIDIA and CoreWeave have disclosed so far. Several figures, including pricing and a firm launch date, remain unconfirmed.
| Specification | Figure |
|---|---|
| Custom CPU cores | 88 Olympus cores |
| Threads (Spatial Multithreading) | 176 |
| Memory type | LPDDR5X |
| Memory bandwidth | Up to 1.2 TB/s |
| Memory capacity | Up to 1.5 TB per CPU |
| Coherency fabric bandwidth | 3.4 TB/s bisectional |
| NVLink-C2C bandwidth to GPU | Up to 1.8 TB/s coherent |
| Code/compile workload speedup vs. x86 | Up to 1.8x |
| Memory bandwidth per core vs. x86 DDR5 | 3x |
| CoreWeave rack: CPUs / cores | 128 CPUs / 11,264 cores |
| CoreWeave rack: concurrent environments | 11,000+ |
| NVIDIA Vera CPU Rack (MGX): max CPUs | Up to 256 |
| NVIDIA Vera CPU Rack: concurrent environments | 22,500+ |
| Price / exact launch date | Not disclosed (“coming soon”) |
From Standalone CPU to Vera Rubin NVL72
Vera is not only a standalone rack product. NVIDIA also positions it as the host CPU for accelerated systems including Vera Rubin NVL72 and HGX Vera Rubin NVL8, where 72 Rubin GPUs pair with 36 Vera CPUs, ConnectX-9 SuperNICs, and BlueField-4 DPUs in one rack-scale platform. CoreWeave said separately on September 30 that a multi-rack Vera Rubin NVL72 cluster had entered limited availability on its cloud, with Cognition, the developer of the coding agent Devin, becoming the first customer to run production workloads on the platform, according to TheEnergyMag. That deployment is distinct from the standalone Vera CPU service aimed at customers who want CPU-heavy agent capacity without provisioning a full GPU rack alongside it.
CoreWeave’s CPU Portfolio: Vera Joins EPYC and Xeon
Vera does not replace CoreWeave’s existing CPU lineup. It joins AMD EPYC and Intel Xeon processors already available on the platform, giving customers a third option built specifically for agent-adjacent, GPU-coupled workloads rather than general-purpose computing, per TheEnergyMag’s reporting. That lets teams scale CPU capacity for sandboxes and data pipelines independently of their GPU fleet, rather than over-provisioning GPU racks just to get more CPU cores attached to them.
| Compute Option | Primary Role | Notable Data Point |
|---|---|---|
| NVIDIA Vera (CoreWeave) | Agent sandboxes, RL environments, data pipelines | 11,264 cores/rack, claimed 3x faster sandbox starts |
| AMD EPYC (CoreWeave) | General-purpose host CPU | Existing standard across CoreWeave’s fleet |
| Intel Xeon (CoreWeave) | General-purpose host CPU | Existing standard across CoreWeave’s fleet |
| AWS AgentCore Runtime v2 | Managed agent runtime on AWS | Claimed 93% cut in cold-start times |
| Google Cloud / GKE agentic tooling | Agentic workload migration on GKE | Positioned against AWS’s roughly 28% cloud-market share lead |
The contrast with hyperscalers is instructive. AWS has pushed its own agent-runtime improvements through AgentCore Runtime v2, while Google has leaned on GKE’s agentic migration tooling to chip at AWS’s cloud lead. CoreWeave’s approach is narrower but more vertically integrated: rather than optimizing a general-purpose runtime, it is pairing a purpose-built CPU with its own bare-metal fleet and its Sandboxes product, betting that hardware-software fit matters more than breadth of managed services for agent-heavy customers.
Fully Connected 2026 and the Nvidia-CoreWeave Relationship
CoreWeave made the announcement at Fully Connected, its AI cloud conference held September 29 through October 1, 2026, at Moscone South in San Francisco. The company said the event brought together more than 4,500 customers, partners, developers, and AI leaders. The Vera news arrives alongside a broader financial relationship between the two companies. NVIDIA invested roughly $2 billion in CoreWeave shares in January 2026 at $87.20 apiece, as the pair outlined plans to accelerate construction of more than 5 gigawatts of shared AI infrastructure capacity, according to TheEnergyMag’s reporting. That investment sits alongside NVIDIA’s broader capital moves this year, including the $150 billion share buyback the company expanded through 2028.
No Price Tag Yet: What CoreWeave Hasn’t Said
What’s missing from Sunday’s announcement is as notable as what’s in it. CoreWeave’s blog post describes the Vera CPU service as “coming soon” without naming a launch date or pricing, according to TheEnergyMag. NVIDIA’s own Vera specifications are described as based on internal, measured comparisons rather than independently audited benchmarks. Readers should treat the 3x sandbox-startup figure and the 1.8x code-workload figure as vendor-reported claims until third-party testing, such as an MLPerf submission, becomes available. Neither company has said which specific customers, beyond Cognition’s early Vera Rubin NVL72 use, will get first access to standalone Vera CPU capacity, or how it will be priced relative to CoreWeave’s existing EPYC and Xeon instances.
Market Reaction and What It Means for CRWV
CoreWeave trades on the Nasdaq under the ticker CRWV, and the Vera announcement lands in a stretch of active capital activity for the company. TheEnergyMag separately reported that CoreWeave opened a $3 billion debt sale and a 35-million-share at-the-market equity program on September 17, 2026, roughly two weeks before the Vera news. Taken together, the debt raise, the equity program, and the NVIDIA-backed capacity buildout point to a company still spending aggressively to secure compute ahead of demand, a pattern that has defined CoreWeave’s growth since its 2025 IPO. For investors, the Vera CPU line is less about immediate revenue, since no pricing has been disclosed, and more about differentiation: a CPU product no hyperscaler currently offers gives CoreWeave a pitch that is harder for AWS, Azure, or Google Cloud to match without their own NVIDIA-designed silicon partnership.
Historical Context: From GPU Clouds to Agent-Native CPUs
CoreWeave’s path to this announcement runs through a rapid shift in identity. The company started as a cryptocurrency-mining operation before pivoting to GPU cloud rental in the late 2010s, then grew into one of the largest NVIDIA-aligned AI clouds by the time it listed on Nasdaq in 2025. NVIDIA’s own CPU ambitions follow a similar arc: the Grace CPU and Grace Hopper Superchip introduced the idea of a data-center CPU tuned for GPU-adjacent workloads, and Grace Blackwell systems extended that pairing through 2025. Vera represents the next step in that lineage, a CPU designed not primarily for training throughput but for the CPU-bound scaffolding around agents, code sandboxes, evaluation loops, and reinforcement learning, that NVIDIA and CoreWeave argue is becoming its own bottleneck as agentic AI scales past simple chatbot use cases into autonomous coding, research, and operations tools.
Competitive Landscape: AWS, Azure, Google Cloud, and Agent Compute
Every major cloud provider is now racing to own some part of the agent-infrastructure stack, but each is attacking it from a different angle. AWS has focused on managed runtime software with AgentCore, aiming to cut the operational overhead of running agents rather than redesigning the underlying silicon. Google has pushed agentic tooling into GKE, framing it as a way to close the gap with AWS’s roughly 28% share of the cloud infrastructure market. Microsoft Azure has leaned on AI Foundry and its existing enterprise relationships. CoreWeave’s bet with Vera is different in kind: instead of software-layer optimization on commodity CPUs, it is offering hardware co-designed with NVIDIA specifically for the CPU-bound half of agent workloads, a niche none of the three hyperscalers currently fill with a comparable purpose-built chip. Pricing pressure elsewhere in the GPU cloud market, including the roughly 79% jump in B200 rental rates reported earlier this year, adds urgency for customers looking to extract more throughput per dollar from CPU-GPU pairs rather than simply renting more GPUs.
What the Companies Are Saying
NVIDIA has kept its public language focused on the workload gap Vera is meant to fill. “NVIDIA Vera CPU is purpose-built for agentic workloads,” the company said in a blog post detailing the CoreWeave rollout. The same post confirmed that “CoreWeave will also offer NVIDIA Vera, the first CPU built for AI agents,” tying the standalone CPU service to the broader Vera Rubin platform rollout. On the performance side, NVIDIA and CoreWeave jointly stated that “in testing, CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPU compared to an x86 CPU,” the clearest quantified claim either company has made about real-world gains so far. CoreWeave’s Chen Goldberg tied the hardware directly to the company’s broader agentic strategy, saying, “General-purpose infrastructure bottlenecks agentic AI; Vera is the first CPU explicitly designed to accelerate it,” in the company’s announcement.
5 Predictions for Agent-Native Compute Through 2027
- Pricing lands within weeks, not months. CoreWeave rarely leaves a “coming soon” product undetailed for long once it has publicly quantified rack-scale capacity, and investor pressure from its recent debt and equity raises favors a quick monetization path.
- At least one hyperscaler responds with its own CPU-silicon push. AWS and Google have both invested in custom silicon before, with Graviton and Axion, so an agent-specific CPU announcement from one of them within a year would not be a surprise.
- Independent benchmarks narrow the gap between vendor claims and reality. Expect MLPerf-style or third-party sandbox-startup benchmarks to appear within two to three quarters, testing whether the 3x figure holds outside CoreWeave’s own environment.
- Vera Rubin NVL72 customer names multiply beyond Cognition. Coding-agent and RL-heavy startups are the most likely early adopters, given the direct fit between Vera’s design goals and their workloads.
- CPU-GPU ratio becomes a standard line item in cloud pricing comparisons. As agent workloads mature, buyers will increasingly evaluate clouds on cores-per-GPU and sandbox-density metrics, not GPU count alone, pushing competitors to publish comparable figures.
Frequently Asked Questions
What is the NVIDIA Vera CPU?
NVIDIA Vera is a CPU built around 88 custom Olympus cores that NVIDIA describes as the first processor designed specifically for AI agent workloads, including code sandboxes, reinforcement-learning environments, and data pipelines that support GPU-based model training and inference.
When will CoreWeave offer NVIDIA Vera CPU?
CoreWeave has not given a firm launch date. The company’s announcement described the service as “coming soon” without specifying pricing or availability, according to reporting from TheEnergyMag.
How many cores does CoreWeave’s Vera rack have?
CoreWeave’s rack-scale Vera deployment packs 128 CPUs and 11,264 cores into a single rack, which the company says supports more than 11,000 concurrent agent environments.
Is NVIDIA Vera the same as Vera Rubin NVL72?
No. Vera is a standalone CPU product, while Vera Rubin NVL72 is a combined rack-scale system pairing 72 Rubin GPUs with 36 Vera CPUs. CoreWeave offers both, and said its Vera Rubin NVL72 cluster entered limited availability separately, with Cognition as its first production customer.
How much faster is Vera than x86 CPUs for AI agents?
CoreWeave reported agent sandbox startup times more than three times faster on Vera than on an x86 CPU in its own testing. NVIDIA separately claims up to 1.8x faster performance on compilation and code-analysis workloads compared with leading x86 chips. Both figures are vendor-reported and have not yet been independently benchmarked.
Does Vera replace AMD EPYC and Intel Xeon on CoreWeave?
No. Vera joins CoreWeave’s existing AMD EPYC and Intel Xeon CPU options rather than replacing them, giving customers a specialized choice for agent-heavy, GPU-coupled workloads alongside general-purpose compute.
What is Fully Connected 2026?
Fully Connected is CoreWeave’s AI cloud conference, held September 29 through October 1, 2026, at Moscone South in San Francisco. CoreWeave said the event drew more than 4,500 customers, partners, developers, and AI leaders, and served as the venue for the Vera CPU announcement.




