OpenAI put real numbers behind its custom silicon bet this week, and the figures land squarely in Nvidia’s backyard. The company’s first in-house inference chip, codenamed Jalapeño, posted benchmark results showing 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than the comparison systems OpenAI tested against, according to the company’s own August 25, 2026 results post. For a market that has run almost entirely on Nvidia GPUs since the generative AI boom began, that is a loaded claim, and it is the clearest evidence yet that OpenAI intends to build, not just buy, the chips that run ChatGPT.

Jalapeño was first announced in partnership with Broadcom on June 24, 2026, as OpenAI’s first custom inference chip, what the company calls an “Intelligence Processor.” The design goal was narrow on purpose: rather than chase the general-purpose flexibility that makes Nvidia’s GPUs useful for training, fine-tuning, and inference alike, Jalapeño is built to do one thing, serve model responses, as efficiently as possible. Eight months after the announcement, OpenAI has published its first public performance data, and the numbers are prompting a fresh round of scrutiny over whether the AI industry’s reliance on a single GPU vendor is starting to crack at the edges.

What OpenAI Actually Announced

The headline fact is simple: OpenAI and Broadcom co-designed an inference-only chip as part of what the companies describe as a multi-generation compute platform, not a one-off part. OpenAI President Greg Brockman told CNBC that the chip was designed end-to-end in nine months, a timeline he credited partly to OpenAI’s own AI models assisting with the engineering work. That is a notable detail on its own: a frontier AI lab says its models helped speed up the design of the chips meant to run those same models.

OpenAI has said it plans to begin deploying Jalapeño inside its own compute infrastructure by the end of 2026. That is a tight runway for a first-generation ASIC, and it suggests the chip has already cleared whatever internal validation gate OpenAI uses before committing capacity to production traffic. The company has been careful, though, to frame Jalapeño as additive rather than a replacement. OpenAI has repeated, in multiple statements, that it will keep buying and deploying Nvidia accelerators, plus hardware from other partners, for both training and inference. Jalapeño is aimed at a specific slice of the workload: serving responses to users and AI agents at scale, where throughput per watt and latency per token matter more than raw training flexibility.

The Benchmark Numbers, Explained

OpenAI’s August 25 results used InferenceX, a public inference benchmark, and ran three open models through it: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Those are deliberate choices. All three are widely available open-weight models that outside researchers can test on their own hardware, which gives OpenAI’s claims a degree of independent verifiability that a purely internal benchmark would lack. Across that test set, OpenAI reported Jalapeño delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than the systems it was measured against.

Those two metrics point at different parts of the inference problem. Throughput per watt is an economics number, it tells a cloud operator how many tokens they can generate per dollar of electricity, which matters enormously at OpenAI’s scale given how much of ChatGPT’s and the API’s operating cost is power and cooling. Latency is a product number, it is the gap between when a user hits enter and when the first token of a response appears. OpenAI has leaned on the latency figure specifically to argue that Jalapeño makes AI agents, not just chat responses, feel more responsive, since agentic workloads chain together many inference calls in sequence and latency compounds across that chain.

What Broadcom and OpenAI’s Hardware Chief Are Saying

Broadcom CEO Hock Tan did not hedge when asked how Jalapeño compares to the competition. Tan said the chip was as good as Nvidia’s Blackwell chips and Google’s tensor processing units, putting Jalapeño in the same conversation as the two most established AI accelerator lines on the market. That is a bold comparison for a first-generation inference ASIC to invite, and it signals how much confidence Broadcom has riding on this partnership continuing into future chip generations.

Richard Ho, OpenAI’s hardware vice president, has given the most detailed public account of how Jalapeño performs and why OpenAI built it. Ho told reporters that Jalapeño offers “best of both worlds” with lower latency and higher throughput, a claim reported by The Verge from a briefing on the chip’s benchmark results against Nvidia superchips. Ho went further on the product impact, saying the efficiency and latency gains mean the chip can give users “faster responses, more responsive agents, and more reliable access as the demand grows,” according to the same Verge briefing.

Ho also addressed a detail that matters more than it might look: the CPUs sitting alongside Jalapeño in OpenAI’s racks. Tom’s Hardware reported that Jalapeño ASICs are deployed alongside AMD EPYC Turin CPUs as hosts. Ho described the decision to use Turin as “pragmatic,” and when asked about Nvidia’s new Vera CPU as an alternative, he called it “a little bit behind… on that maturity level.” That is a pointed, specific knock on Nvidia’s non-GPU silicon, not just a hedge about chip availability, and it is one of the more candid technical assessments a major OpenAI hire has given about a direct Nvidia product.

OpenAI’s own framing of the results, as covered by CNBC, called the testing results “a significant performance advance.” The company has stopped short of claiming Jalapeño beats Nvidia outright across every workload, but the language is confident enough to read as a direct challenge rather than a modest incremental update.

Homogeneous vs. Heterogeneous: What’s Actually Different About This Design

The core design bet behind Jalapeño is specialization. Nvidia’s GPUs, including the Blackwell and now-shipping Vera Rubin data center platforms, are built to handle training and inference on the same silicon, which is a major reason they have become the default choice across the industry: one chip, one software stack, one supply chain to manage for every workload a lab runs. That flexibility carries a cost, though. A GPU built to be good at everything from scratch-training a trillion-parameter model to serving a single chat response is not optimized for either job the way a purpose-built chip could be.

Jalapeño gives up that flexibility entirely. It is not designed to train models at all, only to serve them. OpenAI’s public material describes a purpose-built inference ASIC with local model-state placement and compute, memory, and networking tuned to the specific phases of serving a large language model response. That is a meaningfully different approach from Nvidia’s one-chip-does-everything strategy, though it is worth being precise about the limits of what has been confirmed: OpenAI’s own materials do not use the exact phrase “homogeneous inference design” to describe the chip, so that specific label should be read as a characterization of the approach rather than an official OpenAI term.

How Jalapeño Compares on Paper

Factor OpenAI Jalapeño Nvidia Blackwell / Vera Rubin Google TPU
Primary job Inference only Training and inference Training and inference
Manufacturing partner Broadcom Nvidia (in-house design) Google (in-house design)
Reported efficiency gain (OpenAI’s own test) 1.5x-1.9x more work per watt vs. comparison systems Baseline in OpenAI’s comparison Baseline in OpenAI’s comparison
Reported latency gain (OpenAI’s own test) 1.7x-3.6x lower end-to-end latency vs. comparison systems Baseline in OpenAI’s comparison Baseline in OpenAI’s comparison
Host CPU pairing reported AMD EPYC Turin Nvidia Vera (standalone), per Ho’s comments Not disclosed
Availability for other customers Not confirmed; built for OpenAI’s own infrastructure Sold broadly across cloud and enterprise Cloud-only, via Google Cloud
Announcement date June 24, 2026 Ongoing product line Ongoing product line

Every figure in that OpenAI-vs-comparison row comes from OpenAI’s own InferenceX testing, not from an independent third-party benchmark, which is an important caveat. Broadcom’s CEO backed the comparison in public remarks, but no outside lab has yet published a head-to-head re-run of those exact numbers.

Why Inference Economics Are the Real Battlefield

Training gets the headlines, inference drains the budget. Once a model like GPT-5 or GPT-6 ships, it runs a response to millions of prompts a day, every single day, for as long as the product stays live. That is where OpenAI’s bills actually accumulate, and it is why a chip that only serves inference, rather than one built to do everything, can still move the needle on the company’s bottom line even if it never touches a training job. OpenAI has framed Jalapeño as solving exactly that tension: processing more AI workloads per unit of power while delivering faster responses, which the company has described as a trade-off baked into existing hardware systems.

That framing lines up with a broader shift already visible across the market. OpenAI’s $8 billion Nvidia chip leaseback deal with Amazon shows the company juggling multiple financing structures just to secure enough Nvidia capacity, and HPE’s stock jump on its own Nvidia Vera CPU server bet shows how much the rest of the hardware industry is still betting on Nvidia’s roadmap holding steady. A custom inference chip gives OpenAI a lever that is not subject to Nvidia’s allocation decisions or its pricing, at least for the specific workload Jalapeño is built to handle.

Nvidia’s Position Isn’t Going Anywhere Soon

None of this means Nvidia is suddenly vulnerable. OpenAI has said explicitly, more than once, that it will continue to deploy Nvidia accelerators and hardware from other partners for both training and inference. Jalapeño augments OpenAI’s hardware mix, it does not replace the GPUs the company already runs at scale. Nvidia also still dominates the training side of the market almost entirely, and training is where the largest capital commitments in AI infrastructure continue to flow. The Vera Rubin NVL72’s results in the latest MLPerf inference run show Nvidia is not standing still on the inference side either, even as custom silicon from OpenAI, Google, and Amazon chips away at specific corners of the workload.

The more interesting comparison may be the CPU layer rather than the accelerator layer. Ho’s comment that Nvidia’s standalone Vera CPU is “a little bit behind… on that maturity level” is a specific, named criticism of a specific Nvidia product line, not a broad dismissal of Nvidia’s GPUs. That distinction matters for anyone trying to read how seriously to take the Jalapeño announcement: OpenAI is picking apart Nvidia’s platform piece by piece, keeping the GPUs it still needs while building around the parts it thinks it can do better or source more pragmatically, like pairing Jalapeño with AMD’s EPYC server chips rather than Nvidia’s own CPU line.

The Custom Silicon Trend OpenAI Just Joined

OpenAI is late to custom AI silicon relative to its biggest rivals, not early. Google has run its own TPUs since 2015 and is years into deploying them at scale for both internal workloads and Google Cloud customers, a track record Hock Tan referenced directly when he compared Jalapeño’s performance to Google’s TPUs. Amazon has its own Trainium and Inferentia chip lines. Microsoft has its Maia accelerator. What sets Jalapeño apart is less the idea of custom silicon and more the specific scope: an inference-only design from a company whose core business is still, nominally, a software and API business rather than a hardware maker.

That context also explains why OpenAI leaned on Broadcom rather than building a chip design team from scratch. Broadcom has spent years as the quiet manufacturing and design partner behind Google’s TPUs, which gives the Jalapeño partnership a credible technical pedigree even though OpenAI itself has no public history of chip design before this project. It is the same playbook Google used a decade ago, applied to a company racing to catch up rather than one that pioneered the approach.

Market and Competitive Impact

The immediate market reaction to custom AI silicon announcements has become a predictable pattern over the past two years: Nvidia’s stock dips briefly on the headline, then recovers once investors register that the chip in question is narrow in scope and the company announcing it is still buying Nvidia GPUs in bulk. Jalapeño fits that pattern. It is a genuine technical achievement and a real statement of intent, but it does not replace the GPUs OpenAI needs for training its next generation of models, and OpenAI has said as much directly.

Where the competitive pressure actually lands is on Nvidia’s inference pricing and on the broader narrative that GPUs are the only credible way to run AI workloads at scale. Every lab that stands up a credible, benchmarked alternative, even a narrow one, gives cloud buyers and enterprise customers a data point to use in price negotiations with Nvidia. That dynamic already played out with Google’s TPUs, and it is part of why Broadcom’s stock has become a proxy for how investors think about the custom-silicon trend broadly, alongside Nvidia’s own moves like the export controls and enforcement actions around GPU smuggling that have made some Nvidia hardware harder to source in certain markets.

Historical Context: How We Got Here

OpenAI’s hardware ambitions did not appear overnight. The company has spent more than two years signing a stack of compute deals, with Microsoft, Oracle, CoreWeave, Amazon, and now Broadcom, each one adding a different piece to its infrastructure puzzle. The common thread across all of them has been capacity anxiety: OpenAI has consistently needed more compute than it could secure through any single vendor relationship, which is part of why CoreWeave’s buildout around Nvidia’s Vera CPU for AI agents and OpenAI’s own custom chip effort have developed in parallel rather than as substitutes for each other.

The nine-month design timeline Brockman described to CNBC is fast by chip industry standards, where custom ASICs have historically taken much longer to move from spec to production. If that pace holds up for future Jalapeño generations, it would give OpenAI the ability to iterate on inference hardware faster than a conventional chip design cycle allows, which would matter enormously if AI model architectures keep shifting as quickly as they have over the past three years.

The Memory and Power Bottleneck Nobody’s Chip Solves Yet

Jalapeño’s efficiency gains address compute and latency, not the memory supply crunch that has been squeezing the entire AI hardware industry through most of 2026. High-bandwidth memory remains scarce and expensive regardless of which logic chip sits next to it, a constraint that has already forced tradeoffs on Nvidia’s own roadmap and that shows up in deals like Volantis’s $88 million raise to attack the AI memory wall through photonic interconnects. Nothing in OpenAI’s public Jalapeño disclosures addresses memory supply directly, which means even a successful inference chip rollout does not insulate OpenAI from the same HBM shortages affecting Nvidia, AMD, and every other accelerator maker.

What a Jalapeño-Class Chip Needs to Prove Next

Open question Why it matters Status as of October 2026
Independent benchmark verification All current figures are OpenAI’s own InferenceX results Not yet independently reproduced
Production deployment at scale Lab benchmarks don’t always hold under live traffic Deployment targeted for end of 2026
Fabrication process and yield Determines real-world cost and supply reliability Not disclosed
Availability beyond OpenAI’s own infrastructure Would signal a shift from internal tool to a market product Not confirmed; no stated plans for third-party sales
Next-generation chip roadmap Tests whether the Broadcom partnership compounds across generations Described as “multi-generation” but no second chip detailed yet

Five Predictions for Where This Goes Next

  • OpenAI will publish a second round of Jalapeño benchmarks once the chip is live in production, likely tied to a specific GPT model’s serving infrastructure rather than another open-model comparison.
  • Nvidia will respond with inference-specific tuning or pricing moves rather than a direct rebuttal of OpenAI’s numbers, following the same playbook it used after Google’s TPU benchmarks years ago.
  • Broadcom’s AI chip design business will keep expanding its customer list beyond Google and OpenAI, with at least one more major lab or cloud provider confirming a custom inference ASIC partnership within the next two to three quarters.
  • OpenAI will keep its Nvidia GPU orders flat or growing for training even as inference capacity gradually shifts toward Jalapeño, since the two chips serve different jobs rather than competing for the same budget line.
  • Independent researchers or rival labs will attempt to reproduce the InferenceX comparison on their own hardware, and the results will likely show a narrower gap than OpenAI’s headline figures, as has happened with most vendor-published AI hardware benchmarks over the past two years.

What This Means If You Build on OpenAI’s API

For developers and companies building on top of GPT models, Jalapeño’s most immediate relevance is latency, not price. If OpenAI’s reported 1.7x to 3.6x latency improvements translate into production traffic, API response times for high-volume applications could drop noticeably over the next two to three quarters, which matters most for agentic applications that chain multiple model calls together. Pricing changes are a separate question entirely. OpenAI has not linked Jalapeño’s rollout to any announced API price cuts, and inference cost savings at the infrastructure level do not automatically pass through to customer-facing rates on any fixed timeline.

Frequently Asked Questions

What is OpenAI’s Jalapeño chip?
Jalapeño is OpenAI’s first custom chip, co-designed with Broadcom and announced June 24, 2026. It is built specifically for large language model inference, serving model responses, rather than training.

How does Jalapeño compare to Nvidia’s GPUs?
In OpenAI’s own InferenceX benchmark results, published August 25, 2026, Jalapeño delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than the comparison systems tested, using the open models GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. These figures come from OpenAI’s own testing and have not yet been independently reproduced.

Is OpenAI dropping Nvidia?
No. OpenAI has stated repeatedly that it will continue deploying Nvidia accelerators and hardware from other partners for both training and inference. Jalapeño is designed to handle a specific slice of inference workloads, not to replace Nvidia GPUs across the board.

When will Jalapeño be deployed?
OpenAI has said it plans to begin deploying Jalapeño in its own compute infrastructure by the end of 2026.

What CPUs run alongside Jalapeño?
OpenAI’s hardware vice president Richard Ho has said Jalapeño ASICs are deployed alongside AMD EPYC Turin CPUs as hosts, a decision he called “pragmatic” given the relative maturity of alternatives like Nvidia’s standalone Vera CPU.

Will Jalapeño be available to other companies, not just OpenAI?
That has not been confirmed. Public statements so far describe Jalapeño as built for OpenAI’s own infrastructure, with no announced plans to sell the chip or related cloud instances to third parties.

Who else makes custom AI inference chips?
Google has used custom TPUs since 2015, Amazon has its Trainium and Inferentia lines, and Microsoft has its Maia accelerator. Broadcom, which co-designed Jalapeño, has also been the manufacturing partner behind Google’s TPUs for years.

Does Jalapeño solve the AI memory shortage?
No. Jalapeño addresses compute efficiency and inference latency, not high-bandwidth memory supply, which remains constrained across the entire AI hardware industry regardless of which logic chip a system uses.