OpenAI has a new AI chip, and it’s built to hit Nvidia where it hurts most: profit margin. The chip, named Jalapeño, was unveiled by OpenAI and Broadcom on June 24, 2026, marking OpenAI’s entry into a custom-silicon race that Google, Amazon, and Microsoft already joined years earlier. CNBC reported that OpenAI president Greg Brockman said the chip moved from initial design to a finished, working prototype in about nine months, with AI models handling much of the engineering work along the way.

The stakes go beyond one company’s hardware roadmap. Nvidia currently earns a data-center gross margin that trade publication InformationWeek puts between 75% and 78%. Every hyperscaler that swaps a Nvidia GPU for a custom accelerator captures a slice of that spread for itself. Jalapeño is OpenAI’s attempt to do exactly that, and its arrival adds a fresh data point to a shift that’s been building since Google shipped its first Tensor Processing Unit nearly a decade ago.

This piece breaks down what the Jalapeño AI chip actually is, what named outlets have reported about it, how custom-silicon economics compare with buying merchant GPUs, and what the shift could mean for Nvidia’s margins through the rest of 2026.

What Is Jalapeño? Inside OpenAI’s First Custom AI Chip

Jalapeño is OpenAI’s first self-designed AI accelerator, and unlike Nvidia’s general-purpose GPUs, it targets one job specifically: running inference on large language models as cheaply and efficiently as possible. In its June 2026 announcement, OpenAI said the chip is meant to deliver performance per watt substantially better than current state-of-the-art accelerators, with the goal of cutting the cost of serving every token a model outputs.

Broadcom, TSMC, and a Nine-Month Design Sprint

Broadcom co-designed the chip, and manufacturing runs through TSMC on the 3-nanometer (N3) process node, paired with HBM3E high-bandwidth memory, the same memory class used in Nvidia’s newest Blackwell GPUs. CNBC reported that Brockman credited AI-assisted design tools for compressing what’s normally a multi-year chip program into roughly nine months, something OpenAI and Broadcom describe as among the fastest ASIC development cycles achieved for an advanced chip at this scale. Celestica is building the servers and racks that will house the processor, based on details in OpenAI’s announcement materials.

Built for Inference, Not Training

Jalapeño won’t replace Nvidia GPUs for training OpenAI’s next frontier models. It’s an inference chip, meaning it handles the job of running an already-trained model to answer prompts, write code, or process requests at scale. That’s a deliberate choice. Inference is where OpenAI’s compute bill grows fastest as usage climbs, and it’s also the workload where custom ASICs tend to show the biggest cost edge over general-purpose GPUs. TechCrunch reported in August 2026 that OpenAI’s head of hardware, Richard Ho, said early benchmark results show “a very, very significant performance advance over state of the art,” though OpenAI has not published full, independently verified production benchmarks.

Why Nvidia’s Margins Are the Real Target

Nvidia doesn’t just sell chips, it sells them at a markup few hardware companies have ever sustained at this scale. InformationWeek’s coverage of the custom-silicon trend quoted analyst Alexander Harrowell laying out the math plainly: when a hyperscaler buys a merchant Nvidia GPU, Nvidia’s 75-78% gross margin effectively comes out of the buyer’s own margin. Replace that GPU with an ASIC built by an outsourcer like Broadcom, which typically earns a 30-35% gross margin on manufacturing, and the buyer recaptures roughly half that difference for itself.

That’s the entire logic behind the Jalapeño AI chip, and it’s the same logic that pushed Google, Amazon, and Microsoft to build their own silicon years earlier. OpenAI, unlike those three companies, doesn’t run its own cloud platform to fall back on. It has been posting large losses tied heavily to compute spending, which makes trimming inference costs by double-digit percentages a real lever on its finances, not a nice-to-have efficiency project. Broadcom CEO Hock Tan has said the chip cuts inference cost per token by roughly 50% compared with current GPU setups, a figure that, if it holds up once Jalapeño runs at production volume, represents a direct transfer of value away from Nvidia’s pricing power.

The Numbers: Nvidia’s Market Share and Gross Margin Under Pressure

Nvidia’s dominance in AI accelerators remains real. Estimates of its exact share vary by methodology, with some trackers putting Nvidia’s slice of AI accelerator revenue as high as 86% in 2025 and others closer to 70-80%. What’s more consistent across estimates is the direction of travel: most trackers show Nvidia’s share easing from a peak near 86-87% in 2024 toward roughly 75% by 2026, as custom ASICs and AMD’s GPUs pick up the difference.

On margins, Nvidia’s most recent reported figures carry some nuance. Its GAAP gross margin came in at 60.5% in a recent quarter, with non-GAAP margin at 61.0%, both depressed by a charge tied to U.S. export restrictions on AI chip sales to China involving the H20 chip. Strip that charge out and non-GAAP gross margin would have landed closer to 71.3%, still well above what most chipmakers post and still the figure custom-silicon programs are chasing. AMD, Nvidia’s closest merchant-GPU rival, runs a company-wide gross margin closer to 54-57%, a reminder of how unusual Nvidia’s data-center profitability has been.

Custom Silicon vs Merchant GPU Economics

The gap between what a merchant GPU vendor charges and what an ASIC partner charges is the entire commercial case for a chip like Jalapeño. The table below lines up how the main custom programs compare on structure, though exact unit economics for unreleased chips remain private.

Chip / ProgramTypeManufacturing PartnerProcess NodeReported Cost Edge vs Nvidia GPU
OpenAI JalapeñoCustom inference ASICBroadcom + TSMC3nm (N3)~50% lower cost per token (Broadcom, reported)
Google TPU v7 “Ironwood”Custom ASICBroadcom + TSMCAdvanced node (undisclosed)~65-67% (industry estimate)
AWS Trainium2Custom training/inference ASICMarvell + TSMCAdvanced node (undisclosed)~50% (industry estimate)
Microsoft Maia 200Custom inference ASICMarvell + TSMC3nm (N3)~3x cost-performance vs Trainium Gen3 (industry estimate)
Nvidia Blackwell / H100 / H200Merchant GPUTSMC (Nvidia in-house design)Advanced nodeBaseline, premium pricing, no discount

Figures for competing chips come from industry hardware trackers rather than the companies’ own audited disclosures, so treat the percentages as directional estimates rather than exact numbers. Even with that caveat, the pattern holds across every program: custom ASICs are being pitched, consistently, as roughly half the cost of Nvidia’s GPUs for comparable inference work.

How OpenAI’s Compute Costs Are Fueling the Shift

OpenAI’s compute bill is the backdrop for all of this. Financial trackers covering OpenAI’s 2025 results point to revenue near $13 billion against operating expenses around $34 billion, with a large share of that spend going toward research and development tied to compute infrastructure. Separate analyses have put OpenAI’s 2025 operating loss at roughly $20.9 billion, citing high inference costs as a central driver. OpenAI is privately held and doesn’t publish audited GAAP results, so these figures should be read as estimates from financial analysts rather than confirmed company disclosures, but they explain why a 50% cut to inference cost per token matters enormously to OpenAI’s bottom line even before Jalapeño ships at meaningful volume.

Locking in a chip that shaves inference costs also reduces OpenAI’s dependence on Nvidia’s supply allocation, which matters when demand for top-tier GPUs regularly outstrips supply. Owning more of its compute stack gives OpenAI leverage in future GPU negotiations too, even for the training workloads Jalapeño isn’t built to handle.

The Broader Custom Silicon Landscape

Jalapeño didn’t arrive in a vacuum. It’s the newest entry in a custom-silicon trend that Google started building nearly ten years ago and that Amazon and Microsoft have since joined at meaningful scale.

Google’s TPU Program and Broadcom’s Role

Google’s Tensor Processing Units are the longest-running hyperscaler chip program, and its newest generation, TPU v7 “Ironwood,” was announced in November 2025. Like Jalapeño, it’s co-designed with Broadcom. Industry estimates put its inference cost savings versus Nvidia GPUs in the 65-67% range inside Google Cloud, among the largest gaps reported for any custom chip so far. Google mostly keeps TPUs for internal workloads and Google Cloud customers rather than selling them as standalone hardware.

Amazon Trainium, Inferentia, and Marvell

Amazon Web Services runs two custom chip lines: Trainium for training and Inferentia for inference, both developed with Marvell rather than Broadcom. Trainium2 deployment has reportedly passed 500,000 chips, and AWS has pitched the line as offering roughly 50% inference cost savings over Nvidia-based instances. AWS has also struck deals to expand Nvidia GPU capacity in its own datacenters even while building out Trainium, a sign that hyperscalers currently treat custom silicon and merchant GPUs as complementary rather than an either-or choice.

Nvidia vs Hyperscaler Custom Chips: The Market Share Picture

YearNvidia (merchant GPU), est. shareBroadcom-designed ASICs (Google, OpenAI), est. shareMarvell-designed ASICs (AWS, Microsoft), est. shareAMD + others
2024~86-87%~7-10%~1-2%Remainder
2025 (est.)~80-81%~10-12%~2-3%Remainder
2026 (projected)~75%~12-15%~3-4%Remainder

These are estimates compiled from industry chip-market trackers rather than company-reported figures, and different research firms produce different numbers depending on how they define “AI accelerator.” Directionally, though, nearly every tracker shows the same pattern: Nvidia’s share is easing as custom silicon and AMD grow underneath it.

Microsoft Maia and AMD’s MI Series: The Rest of the Field

Microsoft’s custom chip, Maia 200, is built on TSMC’s 3nm process and is reported to deliver more than 10 PFLOPS of 4-bit compute, with Microsoft targeting roughly 3x the cost-performance of AWS’s third-generation Trainium chip, according to industry hardware analyses on Microsoft’s Azure blog and elsewhere. Even so, an estimated 70% of Azure’s AI workloads still run on Nvidia GPUs, which shows how far custom silicon still has to go before it challenges Nvidia on the training side of the market, where GPU flexibility and Nvidia’s CUDA software layer remain hard to replace.

AMD takes a different path entirely, competing head-on with Nvidia as a merchant GPU vendor rather than building bespoke silicon for one customer. Its MI300 and newer MI-series chips are the most direct alternative to Nvidia’s Blackwell and H-series data-center lineup. AMD’s company-wide gross margin, in the mid-50% range, sits well below Nvidia’s, which shows how much of Nvidia’s advantage comes from pricing power rather than manufacturing cost alone.

Historical Context: From Merchant GPUs to In-House Silicon

For most of the 2010s, AI compute meant buying Nvidia GPUs, full stop. CUDA, Nvidia’s software layer, locked developers into the ecosystem, and no other chipmaker matched Nvidia’s combination of hardware performance and tooling. That gave Nvidia room to post gross margins most semiconductor companies could only envy, often in the 70%-plus range even as rivals like AMD ran margins twenty points lower.

Google broke from that pattern first, introducing its original TPU in 2016 for internal workloads. Amazon followed with Inferentia in 2019 and Trainium in 2021. Microsoft came last among the major hyperscalers, unveiling Maia in 2023. Each company reached the same conclusion on a different timeline: at hyperscaler volume, designing a chip and paying an ASIC partner a lower margin beats paying Nvidia’s premium indefinitely, provided the software and performance gap closes enough to make the switch worthwhile. Jalapeño puts OpenAI, a company that doesn’t operate its own cloud platform the way its rivals do, into that same club roughly three years after Microsoft.

What Named Outlets and Executives Are Saying

The public record on Jalapeño so far comes largely from OpenAI’s own announcement and a handful of interviews given around its unveiling and subsequent benchmark disclosures.

“The degree to which our models have been able to accelerate it was very surprising to us.”

Greg Brockman, OpenAI President, via CNBC

“The bottom line is that the results show a very, very significant performance advance over state of the art.”

Richard Ho, OpenAI Head of Hardware, via TechCrunch

“While OpenAI is still measuring final performance, early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art.”

OpenAI company statement, via CNN

Broadcom CEO Hock Tan has separately said the chip performs on par with Nvidia’s Blackwell GPUs and Google’s TPUs, and told reporters that Jalapeño cuts inference cost per token by roughly 50% versus current Nvidia-based setups, a claim first widely circulated through Bloomberg’s coverage of the announcement.

Analyst Reactions and the Margin-Compression Thesis

Wall Street hasn’t reacted to Jalapeño with the kind of single-day stock swings that accompany major product launches, and none of the reporting gathered here ties a specific Nvidia or Broadcom share-price move directly to the June announcement. The reaction has instead centered on a slower-moving thesis: every workload that shifts to custom silicon compounds against Nvidia’s long-term pricing power, even as Nvidia keeps growing in absolute dollar terms.

InformationWeek’s coverage captures the core argument well. Nvidia’s premium pricing works as long as buyers have no credible alternative. Jalapeño, alongside TPU, Trainium, and Maia, chips away at that lack of alternative one hyperscaler at a time. None of these chips individually threatens Nvidia’s near-term revenue, since Nvidia’s order backlog for Blackwell-class GPUs remains large, but the cumulative effect on margin mix over several years is what analysts following the custom-silicon trend are watching most closely.

Predictions: Where the Custom Silicon Race Goes Next

The following points are analysis and forecasting, not confirmed fact, based on the reporting and trend data above.

  • Jalapeño’s first production deployment, expected internally at OpenAI by the end of 2026, will likely stay limited to inference workloads for a year or more before OpenAI risks custom silicon for any training runs, a much harder engineering problem.
  • OpenAI’s reported second-generation chip on TSMC’s more advanced A16 node suggests a multi-year custom-silicon roadmap rather than a one-off cost-cutting project, echoing Google’s decade-long TPU investment.
  • Nvidia’s headline revenue growth will likely continue in the near term given the size of its GPU order backlog, even as its share of the overall AI accelerator market keeps drifting toward the roughly 75% level several trackers project for 2026.
  • Broadcom and Marvell stand to be among the biggest financial winners of the custom-silicon shift regardless of which hyperscaler’s chip wins, since both companies collect design and manufacturing fees across nearly every major program in this space.
  • Expect at least one more hyperscaler-adjacent AI company, beyond OpenAI, Google, Amazon, and Microsoft, to announce a custom AI chip program within the next twelve months, following the same playbook of partnering with an established ASIC vendor rather than building an in-house design team from scratch.

What This Means for Developers and Enterprise AI Buyers

For companies buying AI compute rather than building chips, the immediate impact of the Jalapeño AI chip is limited. It’s an internal OpenAI chip, not a product enterprises will provision through an API the way they pick GPU instance types on AWS or Azure today. What matters more for buyers is the downstream effect: if Jalapeño and its peers cut OpenAI’s inference costs meaningfully, that pressure could eventually show up as more competitive API pricing, since inference cost is one of the largest line items behind what OpenAI charges per token.

It’s also a signal worth tracking for anyone planning long-term AI infrastructure strategy. The assumption that Nvidia GPUs are the only serious option for production AI workloads is eroding gradually, not overnight. Enterprises locked into Nvidia-based infrastructure aren’t at risk of losing hardware support anytime soon, but the growing menu of custom-silicon options, especially through cloud providers offering TPU-, Trainium-, or Maia-backed instances, gives buyers more real leverage in pricing conversations than they had two years ago. That’s the underlying story behind OpenAI’s custom AI chip push as much as any single benchmark number.

Frequently Asked Questions

What is OpenAI’s Jalapeño chip?

Jalapeño is OpenAI’s first custom AI accelerator, co-designed with Broadcom and built for running inference on large language models rather than training them.

Who manufactures the Jalapeño chip?

TSMC manufactures Jalapeño on its 3-nanometer (N3) process node, paired with HBM3E memory, with Celestica handling server and rack integration.

When will OpenAI start using Jalapeño?

OpenAI has targeted internal deployment by the end of 2026, according to its own announcement, with a second-generation chip reportedly planned for TSMC’s newer A16 node.

How much does Jalapeño save compared with Nvidia GPUs?

Broadcom CEO Hock Tan has said the chip cuts inference cost per token by roughly 50% compared with current Nvidia GPU setups, though OpenAI has not published independently verified production benchmarks yet.

Why does custom silicon threaten Nvidia’s margins?

Nvidia’s data-center gross margin runs in the 75-78% range according to InformationWeek, while ASIC manufacturing partners like Broadcom typically earn 30-35%. Hyperscalers building their own chips can recapture roughly half that spread for themselves.

Will Jalapeño replace Nvidia GPUs at OpenAI?

No. Jalapeño is designed specifically for inference, not training, so OpenAI is expected to keep using Nvidia GPUs for training its frontier models for the foreseeable future.

What other companies have built custom AI chips?

Google (TPU), Amazon (Trainium and Inferentia), and Microsoft (Maia) all run custom AI silicon programs, with Google’s TPU line dating back to 2016 and the others following between 2019 and 2023.

Is Nvidia’s AI chip market share actually shrinking?

Various industry trackers estimate Nvidia’s AI accelerator market share has eased from a peak near 86-87% in 2024 toward roughly 75% by 2026, as custom ASICs and AMD’s GPUs gain ground, though estimates vary by source and methodology.