Nvidia spent most of 2026 selling chips. Now it wants to sell the models that run on them. On August 11, Nvidia released Nemotron 3.5 Lightning, a 30-billion-parameter open-weight model built for AI agents, and reporting from Stocktwits, citing The Information, says a far bigger successor, Nemotron 4, is already in development and could land at roughly 1 trillion parameters. The story resurfaced hard this week, and for good reason: if the trillion-parameter figure holds, Nemotron 4 would be the largest open model Nvidia has ever shipped, and it would put the GPU maker in a stranger position than it has ever occupied, competing directly with the same AI labs that buy its chips.

None of this is happening in a vacuum. Nvidia’s stock trades on the assumption that demand for AI compute keeps climbing, and a bigger, better open model from Nvidia itself is one of the cleanest ways to keep that demand curve pointed up. The company’s calculation seems straightforward: the more open models get used, the more GPUs get rented, bought, and racked, regardless of which lab’s model wins any particular benchmark. That logic is why this story, even though the base model launched two months ago, keeps climbing news feeds and search results in October as the “late autumn” release window for the trillion-parameter model draws closer.

What Nvidia Actually Shipped on August 11

Nemotron 3.5 Lightning is confirmed, documented, and already downloadable. Nvidia built it as a mixture-of-experts (MoE) model with 30 billion total parameters but only 3 billion active at any given time, a design choice that keeps inference costs low while still giving the model access to a much larger pool of specialized “expert” sub-networks. Nvidia published the weights openly, per reporting from Silicon, meaning companies can download, inspect, and modify the model without paying a license fee or asking Nvidia’s permission first.

The pitch is narrow and specific: long-running AI agents. Nvidia designed Nemotron 3.5 Lightning for tasks where a model has to act over extended stretches of time, calling tools, checking its own work, and making dozens of small decisions in a row, rather than answering a single prompt well. According to Nvidia, cited in reporting from ChosunBiz, the model generates tokens up to four times faster than comparable open models and lets agents finish entire multi-step tasks about 30% quicker. Nvidia told CNBC, per the same ChosunBiz report, that Nemotron 3.5 Lightning was distilled from a larger Nemotron model, a common technique where a smaller model is trained on the outputs of a bigger one to shrink size while preserving most of the performance.

Nvidia didn’t stop at the model itself. It also shipped a free routing tool called Nemo Switchyard that reads an incoming prompt, estimates how hard the task actually is, and sends it to whichever model, small or large, can handle it most cheaply. Nvidia says that routing layer can cut costs to roughly a third of what a company would pay by sending every request to its most expensive model, while keeping output quality close to that top-tier baseline. That’s a practical, cost-driven feature, and it tells you something about who Nvidia is building for: not hobbyists, but enterprises running agents at scale who care about the bill at the end of the month.

The Trillion-Parameter Rumor: What’s Confirmed and What Isn’t

Here’s where the story shifts from announcement to speculation, and the distinction matters. Nvidia has not announced a model called Nemotron 4. There is no confirmed parameter count, no confirmed release date, and no confirmed price. What exists is a report, originally from The Information and relayed by Reuters and multiple outlets since, claiming Nvidia is training a next-generation Nemotron model that could reach or exceed 1 trillion total parameters.

Reporting attributed to The Information, as summarized by Ground News and by Stocktwits, says training on the larger model had not finished as of the reports and that Nvidia has not locked a release date. Two people described as familiar with the project told The Information the model could be ready as early as this fall, but that’s a target, not a commitment. If Nemotron 4 does land near 1 trillion parameters, it would be roughly double the size of Nemotron 3 Ultra, the model that held the title of Nvidia’s largest Nemotron release after it shipped in June 2026, according to the Stocktwits report.

Scale alone isn’t the interesting part. The Information’s reporting, relayed by Stocktwits, also says more than 570 authors contributed to Nvidia’s most recent major model research paper, and the team working on Nemotron 4 is expected to be larger still. Nvidia has reportedly capped its cloud-compute budget for Nemotron development at $7 billion through fiscal 2028, a number that signals this isn’t a side project squeezed out of spare GPU cycles. It’s a funded, multi-year bet with a ceiling attached, which suggests Nvidia is treating its own model lineup as a real product line rather than a marketing exercise to sell more chips.

DetailNemotron 3.5 Lightning (confirmed)Nemotron 4 (reported, unconfirmed)
StatusReleased August 11, 2026In development, not announced
Total parameters30 billionReportedly up to ~1 trillion
Active parameters3 billionNot disclosed
ArchitectureMixture-of-experts (MoE)Not disclosed
AvailabilityOpen weights, no license feeNot confirmed
Target use caseLong-running AI agentsNot confirmed; reported to target frontier-level tasks
Release timingShippedReportedly possible as early as late autumn 2026
Primary sourceNvidia / NIM documentationThe Information, via Reuters and Stocktwits

Why Nvidia Wants to Be a Model Company, Not Just a Chip Company

It’s worth asking why a company that already dominates AI hardware would bother building its own frontier-scale language model at all. The answer, based on the reporting around Nemotron 3.5 Lightning, comes down to ecosystem gravity. Closed models from OpenAI, Anthropic, and Google run on infrastructure those companies control tightly. Open models, by contrast, get downloaded, copied, fine-tuned, and deployed across thousands of private data centers and cloud accounts that Nvidia doesn’t own but very much wants to sell GPUs into.

Every additional open model that runs well on Nvidia hardware is, functionally, another reason for a company to buy more Nvidia chips rather than a competitor’s. That’s the logic Stocktwits’ reporting on The Information’s findings points to directly: Nvidia’s push into open models is partly aimed at building a wider pool of high-quality models optimized specifically for its own silicon, which supports demand for its AI infrastructure even in scenarios where Nvidia’s own model doesn’t end up being anyone’s first choice. Whether Nemotron 4 becomes a top-five frontier model or a solid also-ran barely matters to that strategy, as long as developers keep running Nemotron variants on Nvidia GPUs.

Nvidia has been building the surrounding ecosystem for months. The company organized what reporting describes as a “Nemotron Coalition,” reportedly including Mistral, Cursor, Cognition, and Prime Intellect, and has separately backed open-source AI efforts at Reflection AI and Thinking Machines Lab. That’s a deliberate pattern: fund or partner with multiple open-model players, rather than betting everything on one in-house release. It also mirrors how Nvidia already committed roughly $60 million toward AI factories built around open-source models, a move that lines up with the same goal of making open weights the default choice for enterprise AI deployments running on Nvidia hardware.

Nvidia’s Growing Open-Source Partner Network

The Nemotron Coalition is the clearest evidence that Nvidia sees this as a long game rather than a one-off launch. Bringing in Mistral, Cursor, Cognition, and Prime Intellect gives Nvidia partners who already have developer mindshare in coding tools and agent frameworks, two areas where Nemotron 3.5 Lightning is explicitly aimed. Pairing that with direct backing of Reflection AI and Thinking Machines Lab spreads Nvidia’s bet across several independent model-building efforts instead of concentrating everything into a single internal team’s output.

That diversification matters because open-model development is unusually risky to centralize. If Nvidia’s internal Nemotron 4 effort stalls or underperforms on release, the company still has equity in the broader open-model ecosystem through its coalition partners and portfolio bets, so the GPU-demand thesis survives even if one specific model release disappoints.

The Awkward Part: Competing With Your Own Customers

If Nemotron 4 does arrive at trillion-parameter scale, Nvidia walks into a genuinely uncomfortable position. OpenAI, Microsoft, and SpaceX, three of the companies most dependent on Nvidia’s chips to train and run their own frontier models, would suddenly find Nvidia fielding a direct competitor in the same weight class. Nvidia has roughly $30 billion invested in OpenAI directly, on top of an infrastructure partnership announced in September 2025 that could eventually involve deploying up to $100 billion in AI infrastructure, according to the Stocktwits report. Competing with a company you’ve put tens of billions of dollars into is not a normal position for a supplier to be in.

Microsoft and OpenAI are both already working on their own custom AI silicon, which flips the tension in the other direction too: Nvidia’s biggest customers are trying to need Nvidia less, at the exact moment Nvidia is trying to need them less. It’s a standoff built on mutual hedging. Neither side wants to fully cut the cord, because Nvidia still makes the best training and inference hardware on the market today, and OpenAI, Microsoft, and the rest still generate the demand that keeps Nvidia’s order books full. But a trillion-parameter Nemotron model would be the clearest signal yet that Nvidia isn’t content to just watch that demand from the sidelines as a hardware vendor.

The Geopolitical Layer: China, Export Rules, and Open Weights

There’s a geopolitical layer here too. Reporting from the South China Morning Post, cited in the Stocktwits coverage, found that some of China’s top AI models are still being trained on Nvidia chips despite Beijing’s own restrictions on foreign hardware dependence. That detail matters for Nemotron’s open-weight strategy specifically, because an openly licensed Nvidia model that runs best on Nvidia GPUs gives both US and Chinese AI developers another reason to stay inside Nvidia’s hardware ecosystem, regardless of which country’s export rules are in effect at any given moment.

An open, freely licensed model adds a layer of neutrality that a closed, subscription-gated model doesn’t have. Export restrictions tend to target hardware shipments and, increasingly, specific closed-model access, not open weights that can already be mirrored and redistributed globally once published. That structural detail may be part of why Nvidia is comfortable releasing Nemotron openly even as broader US-China AI tensions remain unresolved.

How Nemotron 3.5 Lightning Stacks Up Against Other Small Open Models

Nvidia isn’t the only company betting that smaller, cheaper, agent-focused models are where the next wave of adoption happens. Microsoft shipped its own small open model, Decision-1, built on a 9-billion-parameter base with claims of a 35x speed improvement on targeted decision tasks, and JetBrains released Mellum 2.1, a 12-billion-parameter coding model that reportedly hit 47% on the SWE-Bench coding benchmark. Both are smaller than Nemotron 3.5 Lightning’s 30 billion total parameters, but they chase the same basic idea: don’t make the model bigger, make it good enough at a narrow job and cheap enough to run constantly.

Where Nemotron 3.5 Lightning differs is the mixture-of-experts architecture paired with full weight openness and no license fee. That combination, low active-parameter count for cost, large total-parameter count for capability, open weights for flexibility, isn’t unique to Nvidia, but Nvidia is pairing it with Nemo Switchyard’s routing layer in a way few competitors currently offer as a bundled, free tool. On the other end of the spectrum, OpenAI has been pushing its own next-generation release cadence hard, including GPT-6 Sol, which the company claims reaches 99.99% prompt injection defense, a very different kind of flagship bet focused on safety hardening rather than agent speed.

ModelMakerTotal parametersActive parametersFocus
Nemotron 3.5 LightningNvidia30B3BFast, low-cost AI agents
Decision-1Microsoft9BNot disclosedDecision-task speed (35x claim)
Mellum 2.1JetBrains12BNot disclosedCoding (47% SWE-Bench)
Nemotron 3 UltraNvidiaNot disclosed (Nvidia’s prior largest)Not disclosedGeneral frontier-scale tasks
Nemotron 4 (reported)Nvidia~1 trillion (unconfirmed)Not disclosedReportedly frontier-scale, unconfirmed by Nvidia

Historical Context: Nvidia’s Slow Walk From Chips to Models

Nvidia didn’t wake up in August 2026 and decide to become a model company overnight. The Nemotron family has existed as Nvidia’s open-model line for a while, positioned explicitly as a set of models built to help companies and researchers construct their own specialized AI agents rather than compete head-on with ChatGPT or Claude as a general consumer assistant. Nemotron 3 Ultra, released in June 2026, was the largest entry in that family before Nemotron 3.5 Lightning arrived two months later as a smaller, faster sibling rather than a direct replacement.

That build-out has tracked closely with Nvidia’s broader hardware roadmap. Around the same period, Nvidia has been navigating its own supply-chain balancing act, reportedly shifting GB202 chip supply toward RTX Pro GPUs at the expense of RTX 5090 availability, a sign that enterprise and AI-factory demand is pulling production priority away from consumer gaming cards. The Nemotron strategy and the GPU allocation strategy point the same direction: Nvidia is increasingly organizing its entire product line, chips and models alike, around enterprise AI agent workloads rather than any single market segment.

The Distillation Debate Hanging Over the Release

It’s also worth remembering that distillation, the technique Nvidia reportedly used to build Nemotron 3.5 Lightning from a larger internal model, has become a contentious topic across the industry in 2026. The same technique has drawn scrutiny after allegations that some Chinese AI firms used distillation to closely mimic the outputs of leading US models without the underlying research investment. Nvidia using distillation on its own larger model to produce Nemotron 3.5 Lightning is a different situation entirely, since it’s distilling from its own IP, but the timing puts Nvidia’s release inside a conversation the industry is already having about where distillation crosses into imitation.

For Nvidia, the practical upside of distillation is clear: a 30-billion-parameter model that performs close to a much larger original, at a fraction of the inference cost, is exactly what agent-focused customers are asking for. The reputational risk is that every distilled open model released in 2026 now gets measured against the same imitation debate, regardless of whether the company doing the distilling owns the source model outright.

Market Impact: What This Means for Nvidia’s AI Infrastructure Bet

Nvidia’s hardware business doesn’t actually need Nemotron 4 to succeed as a standalone product to benefit from its existence. Every dollar Nvidia spends building a credible open model is a dollar spent making its own GPUs the default home for a wider slice of the open-source AI ecosystem. That’s the same logic behind Nvidia’s reported $60 million commitment to open-source-model AI factories, and it’s consistent with how the company has spent 2026 expanding the Nemotron Coalition to include partners like Mistral, Cursor, Cognition, and Prime Intellect rather than trying to out-build every rival lab solo.

The $7 billion cloud-compute budget cap through fiscal 2028, reported by The Information, is also a useful signal for how seriously Nvidia is treating Nemotron as a line item rather than a research curiosity. That’s real money set aside specifically for training runs, and it implies Nvidia expects multiple more Nemotron releases between now and 2028, not just one trillion-parameter model and then a pause. For customers evaluating whether to build on Nemotron versus a rival open model, a multi-year funded roadmap is a meaningfully different signal than a single splashy release with no stated follow-up plan.

For competitors in the open-model space, the pressure is now twofold. They’re competing against Nvidia’s model quality on one axis, and against Nvidia’s hardware-plus-software bundling on the other, since a model optimized for Nvidia silicon paired with a free routing tool like Nemo Switchyard is a hard combination for a smaller lab to match without owning the underlying hardware stack. That dynamic echoes other recent infrastructure plays in the space, including Google’s push with its Ironwood TPU, reportedly priced to undercut Nvidia’s B200 by 18%, which shows rival hardware makers are also trying to bundle silicon and software advantages rather than compete purely on raw chip specs.

What Enterprises Building AI Agents Should Actually Do Right Now

For teams already building agent workflows, Nemotron 3.5 Lightning is available today and worth testing against existing small-model deployments, particularly if agent task latency or per-task cost is the current bottleneck. The open license and lack of upfront fees make it low-risk to pilot. Nemo Switchyard is worth a separate evaluation on its own, since routing logic that cuts costs by roughly two-thirds while preserving top-tier accuracy is a meaningful operational win independent of which underlying models get routed to.

Nemotron 4 is a different story. Nothing about it is confirmed, including whether the 1-trillion-parameter figure survives to release, so building a product roadmap around a model that doesn’t exist yet, with no price and no date, is premature. The more reasonable move is to track the story the way Nvidia itself seems to be building it: as a slow, funded, multi-year commitment to open models rather than a single make-or-break launch. Companies that plan around that longer arc, rather than around speculative benchmark numbers from an unreleased model, will be better positioned whenever Nemotron 4 (or whatever Nvidia eventually calls it) does ship.

Five Predictions for Where This Goes Next

  • Nvidia will likely confirm more details about its next large Nemotron model before the end of 2026, even if a full release slips past the reported “late autumn” window, given the scale of the reported $7 billion compute budget already committed.
  • Expect the Nemotron Coalition to add more members over the next two quarters, following the same pattern as Nvidia’s existing partnerships with Mistral, Cursor, Cognition, and Prime Intellect, as Nvidia works to make its model family the default open option across more developer tool stacks.
  • Tension between Nvidia and major customers like OpenAI and Microsoft over competing models will stay mostly unspoken in public, since both sides have too much financial entanglement, including Nvidia’s roughly $30 billion OpenAI investment, to risk an open rift over model competition.
  • Rival chipmakers and cloud providers will keep pairing hardware releases with their own open-model or routing-tool announcements, following the same bundling logic Nvidia is using with Nemo Switchyard, rather than competing on raw chip specifications alone.
  • If Nemotron 4 does ship near 1 trillion parameters, expect immediate scrutiny over whether its training data or methodology overlaps with the distillation controversy already surrounding smaller open models in 2026, given that Nemotron 3.5 Lightning was itself built using distillation from a larger in-house model.

Frequently Asked Questions

Is Nemotron 4 officially confirmed by Nvidia?

No. Nvidia has not announced a product called Nemotron 4. The 1-trillion-parameter figure comes from reporting attributed to The Information, relayed by outlets including Reuters and Stocktwits, describing a model reportedly in development. Nvidia has not confirmed parameters, pricing, or a release date.

What is Nemotron 3.5 Lightning?

It’s an open-weight mixture-of-experts language model Nvidia released on August 11, 2026, with 30 billion total parameters and 3 billion active parameters, designed specifically for running long-running AI agents efficiently.

Can companies use Nemotron 3.5 Lightning for free?

Yes. Nvidia released the model’s weights publicly, and companies can download and use it without a separate license fee or prior approval from Nvidia, according to reporting from Silicon and ChosunBiz.

Why would Nvidia build its own AI model when it already sells the chips that run competing models?

Reporting from The Information, relayed by Stocktwits, suggests Nvidia wants to grow the broader pool of high-quality open models optimized for its own hardware, which supports GPU demand even when Nvidia’s models aren’t the top choice for a given task.

Does a trillion-parameter Nemotron model compete with OpenAI and Microsoft?

Potentially, yes. Both companies are major Nvidia customers and major AI model developers. If Nemotron 4 reaches frontier scale, it would put Nvidia in more direct competition with the same companies that rely heavily on its chips, a tension reporting has already flagged as notable.

What is Nemo Switchyard?

It’s a free tool Nvidia released alongside Nemotron 3.5 Lightning that analyzes how difficult a given prompt or task is and routes it to the most cost-appropriate model, which Nvidia says can cut expenses to roughly a third of always using the top-tier model.

When might Nemotron 4 actually be released?

Reports citing people familiar with the project say it could be ready as early as late autumn 2026, but Nvidia has not set or confirmed a formal release date, and training reportedly had not finished as of the available reporting.

How big is Nemotron 4 compared to Nvidia’s current largest model?

If the reported figure of roughly 1 trillion parameters holds, Nemotron 4 would be about twice the size of Nemotron 3 Ultra, the Nvidia model that held the title of largest in the Nemotron family after its release in June 2026.