Mistral AI opened a public preview of Mistral Large 4 on October 6, 2026, putting a 1 trillion-parameter model into developers’ hands through the Mistral Studio API. The French AI lab’s own announcement describes the system as a Mixture-of-Experts model with 49 billion active parameters, a 1-million-token context window, and native multimodal support, and it carries an internal nickname that has already spread across developer forums: “Le Chonk.”
The timing matters. Mistral has spent the back half of 2026 fighting to stay relevant against Anthropic, OpenAI, and a wave of Chinese open-weight labs, and a $24 billion valuation only raises the stakes on whether Large 4 can actually compete. This piece breaks down what Mistral has confirmed, what’s still murky, how Large 4 stacks up against the rest of the field, and why the model’s parameter count has already become a small controversy of its own.
Mistral Large 4 enters public preview
Mistral’s announcement opens with a simple line: “Today, we’re launching a public preview of Mistral Large 4,” Mistral AI said in its official post. That preview runs through the Mistral Studio API under the model ID mistral-large-4, giving developers a way to test the model before any weights become publicly downloadable.
Previews are a familiar move for frontier labs trying to generate buzz and collect usage data before a full release. What’s different here is the gap between preview and open weights: Mistral says in its own post that “weights drop end of this month,” which puts a rough three-week window between the API going live and the model becoming downloadable. That’s a notably tighter turnaround than some rivals have offered, and it fits Mistral’s broader identity as the lab most committed to shipping open weights alongside commercial products.
The announcement lands less than a year after Mistral’s previous flagship release, and it follows closely on the heels of earlier reporting on the Large 4 debut that circulated different parameter figures than what Mistral has now confirmed in writing. We’ll get to that discrepancy in detail below, because it’s one of the more interesting threads in this story.
Inside Le Chonk: a 1 trillion-parameter Mixture of Experts
Mistral’s own description is direct: “ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters,” the company wrote in its launch post. That’s the core spec sheet, and it’s worth sitting with for a moment because of what it implies architecturally.
A trillion total parameters with only 49 billion active per token tells you Mistral built Large 4 as a sparse Mixture-of-Experts (MoE) system. Instead of running every parameter on every request, an MoE model routes each token through a small subset of specialized expert sub-networks. The result is a model that stores a trillion parameters’ worth of learned knowledge but only pays the compute cost of a much smaller, roughly 49-billion-parameter model at inference time. That’s the entire appeal of MoE architecture: scale on paper, speed in practice.
This isn’t a new trick. Mixture-of-Experts has become the default architecture for frontier-scale models over the past two years, and other recent releases chase the same math. Tencent’s Hy4 model, for example, uses a similar total-versus-active split to keep a 770-billion-parameter model affordable to run. What stands out about Large 4 is the ratio: 49 billion active out of 1 trillion total is a roughly 20-to-1 split, a notably sparse design even by current MoE standards.
Why the nickname matters more than it sounds
Mistral’s announcement makes a point of naming the model twice: “Unofficially ML4, very officially: le Chonk,” the company wrote. Model nicknames aren’t just marketing flourish. Developers search for them, communities rally around them, and in Mistral’s case the self-deprecating joke about size signals confidence that the trillion-parameter count is a selling point, not something to downplay. Compare that to labs that quietly bury exact parameter counts in technical appendices. Mistral is leading with the number.
The 1-million-token context window
Beyond parameter count, Large 4 ships with a 1-million-token context window, according to reports on the announcement. That puts Mistral in the same bracket as a small group of models built to handle entire codebases, full-length books, or massive document sets in a single pass without chunking.
Long context windows have become one of the clearest competitive battlegrounds in the model market this year. Anthropic shipped a 1M-token context window with Claude Opus 5.5 specifically to court coding tool developers who need a model to hold an entire repository in memory. If Mistral is matching that figure with Large 4, it’s a direct signal about which market the company is chasing: enterprise and developer tooling, not just chatbot use cases.
A large context window only matters if a model can actually use it, though. Long-context benchmarks like “needle in a haystack” retrieval tests exist precisely because raw window size and effective recall often diverge. Mistral has not published benchmark results for Large 4’s long-context performance as of this writing, so that remains an open question rather than a confirmed strength.
Getting access: the Mistral Studio preview API
Mistral’s post is explicit about where to find the model right now: “You can try the preview API today on Mistral Studio,” the company said. For developers already building on Mistral’s platform, that means pointing an existing integration at the new model ID rather than waiting for a separate rollout.
A typical request against the preview follows the same shape as Mistral’s existing chat completion endpoints, just with the new model identifier swapped in:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{"role": "user", "content": "Summarize this quarterly report."}
]
}'
Because this is a preview rather than a stable release, expect the usual caveats: rate limits, possible latency spikes under load, and a model that may change behavior before the final weights ship. Mistral has not published preview-specific pricing, so teams testing the API should treat current usage as developer-tier access rather than a production-ready commercial tier.
Natively multimodal: what that means in practice
Mistral describes Large 4 as “natively multimodal,” a phrase that distinguishes it from models that bolt vision or audio encoders onto a text-only backbone after the fact. A natively multimodal model is trained from the start on mixed data (text, images, and potentially other modalities) rather than having that capability stitched on later through fine-tuning.
The practical upside of native multimodality is usually better grounding between modalities. A model trained end-to-end on image-text pairs tends to reason about what it sees more coherently than one where vision was added as an afterthought. Mistral hasn’t published specific multimodal benchmark numbers for Large 4, so this is a design claim worth tracking once independent testing starts, rather than a verified performance edge today.
The parameter count confusion that preceded this launch
Here’s where the story gets genuinely messy. In the days before Mistral’s own confirmation, multiple outlets were already reporting on Large 4 using different numbers entirely. Some of that earlier coverage, including reporting on this site citing 675 billion total and 41 billion active parameters, and separate coverage of a 1.05 trillion-parameter claim that conflicted with Mistral’s own official figure, painted a picture that doesn’t match what Mistral’s announcement text now states.
Mistral’s own post is unambiguous on the figure that matters: 1 trillion total parameters, 49 billion active. That’s the number from the company directly, and it should be treated as the authoritative figure going forward. The 675 billion, 41 billion, and 1.05 trillion figures that circulated earlier appear to have come from leaks, early benchmarking teasers, or speculation ahead of the official post, and none of them match what Mistral itself has now put in writing.
This kind of pre-launch number confusion isn’t unique to Mistral. Coverage of the broader AI industry shows frontier labs increasingly tease models through partial leaks, cherry-picked benchmark screenshots, and researcher social media posts well before an official announcement, and those early numbers rarely survive contact with the real spec sheet. The lesson for readers tracking frontier AI releases: treat any parameter count as provisional until it appears in a company’s own announcement, not a leak or a third-party estimate.
Mistral Large 4 versus the rest of the field
Placing Large 4 next to other recent releases shows both where it fits and where the public record still has gaps. Several labs have not disclosed full parameter breakdowns for their latest models, so the table below marks those cells honestly rather than guessing.
| Model | Developer | Total parameters | Active parameters | Context window | Status as of Oct. 2026 |
|---|---|---|---|---|---|
| Mistral Large 4 ("Le Chonk") | Mistral AI | 1 trillion | 49 billion | 1M tokens | Public preview; weights due end of October |
| Tencent Hy4 | Tencent | 770 billion | Undisclosed | 1M tokens | Preview, per earlier reporting |
| Reflection AI Beam | Reflection AI | 501 billion | 23 billion | Undisclosed | Launched, per earlier reporting |
| Claude Opus 5.5 | Anthropic | Undisclosed | Undisclosed | 1M tokens | Generally available |
| GLM-5.3 | Z.ai | Undisclosed | Undisclosed | Undisclosed | Generally available |
| DeepSeek V4.1-Flash | DeepSeek | Undisclosed | Undisclosed | Undisclosed | Generally available |
Two things jump out. First, among labs that do disclose a total parameter count, Mistral’s trillion-parameter figure is now the largest publicly confirmed total on the list, ahead of Tencent’s 770 billion and Reflection AI’s 501 billion. Second, the 1-million-token context window is quickly becoming table stakes rather than a differentiator. Mistral, Tencent, and Anthropic have all converged on roughly the same window size, which means the real competition is shifting toward active-parameter efficiency and real-world benchmark performance rather than headline specs.
Mistral has not published head-to-head benchmark results against any of these models for Large 4, so none of this table should be read as a performance ranking. It’s a spec comparison, and Mistral’s own framing (that ML4 “pushes the frontier of open-weight performance,” per the company’s announcement) is a claim the industry hasn’t yet had a chance to independently verify through tools like Artificial Analysis or LMArena.
Market impact: open-weight momentum and Mistral’s bet
Mistral’s entire identity rests on being the loudest advocate for open weights among well-funded frontier labs. A trillion-parameter model with a committed weights-release window inside the same month as the preview is a direct statement to that market: Mistral isn’t retreating into a closed, API-only strategy the way some competitors have leaned toward.
That bet carries real financial weight behind it. Mistral’s business has already drawn a $24 billion valuation, a figure that depends partly on the company proving it can keep shipping frontier-class models without abandoning the open-weight commitments that built its reputation. If Large 4’s eventual open release lands on schedule and the model performs competitively, it reinforces the thesis that open weights and commercial viability aren’t mutually exclusive. If the weights slip past “end of this month” or underperform expectations once independent testing starts, it hands ammunition to critics who argue open-weight strategies can’t keep pace with closed frontier labs funding ever-larger training runs.
There’s also a downstream effect on the broader open-weight ecosystem. Hugging Face hosts the bulk of Mistral’s prior model releases, and the company’s existing presence there gives a strong signal for where Large 4’s weights will eventually land once they drop. Developers building on open-weight infrastructure, from fine-tuning pipelines to self-hosted inference stacks, have a direct stake in whether a trillion-parameter model actually becomes practical to self-host, even with a sparse 49-billion active-parameter footprint.
Historical context: how Mistral Large got here
Mistral’s flagship Large line has moved fast by industry standards, with each generation arriving well inside a year of the last. The company built its reputation on being small and nimble relative to OpenAI, Google, and Anthropic, shipping smaller, efficient models that punched above their weight on benchmarks. Large 4 represents a clear break from that early positioning: a trillion total parameters is squarely in frontier-lab territory, not the lean, efficiency-first category Mistral built its name on in its earliest releases.
That shift tracks with Mistral’s broader trajectory over the past two years, including its rising valuation and expanding enterprise ambitions. The company has also had to manage its share of security scrutiny along the way, including an incident where it had to patch a prompt-injection exploit in its chat product shortly after a major funding raise. Scaling up model size while also fielding security and safety pressure is the balancing act every frontier lab now faces, and Mistral is no exception.
What the industry is saying
Because Large 4 only became public within the last day, independent commentary from outside analysts and researchers hasn’t had time to accumulate yet. The clearest record of intent still comes from Mistral’s own announcement, which doubles as both a technical disclosure and a statement of strategy. The company’s framing that the preview is live “today” on Mistral Studio, paired with its direct claim that the model “pushes the frontier of open-weight performance,” both come straight from Mistral’s official launch post, and they’re worth reading in full context rather than as isolated soundbites. Expect third-party benchmark commentary, developer reaction threads, and analyst notes to accumulate over the coming days as more people get preview access through Mistral Studio.
Confirmed specs versus unconfirmed claims
Given how much pre-launch speculation preceded this announcement, it’s worth laying out exactly what’s nailed down versus what’s still floating around as rumor.
| Detail | Status | Source |
|---|---|---|
| 1 trillion total parameters | Confirmed | Mistral’s official announcement |
| 49 billion active parameters | Confirmed | Mistral’s official announcement |
| Mixture-of-Experts, natively multimodal | Confirmed | Mistral’s official announcement |
| Public preview via Mistral Studio | Confirmed | Mistral’s official announcement |
| 1-million-token context window | Reported | Secondary reporting on the announcement |
| Weights released "end of this month" | Confirmed (exact date not given) | Mistral’s official announcement |
| Weights release on October 27 | Unconfirmed by Mistral | Secondary reports only |
| 1.05 trillion-parameter figure | Unconfirmed, conflicts with official 1T figure | Earlier secondary claims |
| Benchmark scores | Not yet published | None available |
| Preview or production pricing | Not yet published | None available |
That split matters for anyone deciding how to cover or act on this story. The architecture and headline parameter counts are locked down by Mistral’s own words. The exact weight-release date, training details, benchmark scores, and pricing are not, and treating any of those as settled fact right now would be getting ahead of what’s actually been published.
What’s still unconfirmed
Three gaps stand out. First, training data composition and compute budget for Large 4 haven’t been disclosed, which makes it hard to independently assess how the model was built relative to its trillion-parameter peers. Second, no benchmark results have been published by Mistral or by independent evaluators, so every claim about Large 4’s quality relative to GPT, Gemini, Claude, or Chinese open-weight rivals remains untested. Third, pricing for both the preview and any eventual production tier is absent from the announcement, leaving enterprise buyers without the information they’d need to budget for adoption.
Each of these gaps is normal for a same-day preview announcement. Labs routinely hold back benchmarks and pricing until closer to general availability, both to control the narrative and because preview-stage models are still subject to change. The open question is simply how long that gap lasts before Mistral fills it in.
Predictions: what happens next
- Independent benchmark results will start appearing on platforms like LMArena and Artificial Analysis within one to two weeks of preview access opening, giving the first real signal on whether Large 4’s performance matches its parameter count.
- Mistral will publish pricing for Large 4 closer to the open-weight release rather than during the preview window, following the pattern most labs use to avoid committing to numbers before demand is clear.
- The weights release will most likely land within the final week of October 2026, consistent with Mistral’s “end of this month” language, though an exact date announcement should be treated as unconfirmed until Mistral states it directly.
- Rival labs, particularly Chinese open-weight developers like Tencent and Z.ai, will respond with their own parameter-count disclosures or model updates within weeks, continuing the pattern of competitive one-upmanship on trillion-parameter-class MoE models.
- Expect continued confusion in secondary coverage between the 1 trillion/49 billion figures Mistral has confirmed and the earlier 675 billion/41 billion and 1.05 trillion figures that circulated pre-launch, until enough outlets correct the record against Mistral’s own announcement.
Frequently asked questions
What is Mistral Large 4?
Mistral Large 4, nicknamed “Le Chonk,” is Mistral AI’s newest flagship model, announced October 6, 2026. It’s a natively multimodal Mixture-of-Experts model with 1 trillion total parameters and 49 billion active parameters, currently available as a public preview through the Mistral Studio API.
How many parameters does Mistral Large 4 have?
According to Mistral’s own announcement, Large 4 has 1 trillion total parameters and 49 billion active parameters per token. Earlier reports citing 675 billion, 41 billion, or 1.05 trillion parameters predate Mistral’s official confirmation and don’t match the figures in the company’s launch post.
How can I access the Mistral Large 4 preview?
The preview is available now through the Mistral Studio API using the model ID mistral-large-4. Developers with existing Mistral API access can point requests at the new model identifier directly.
When will Mistral release the Large 4 weights?
Mistral’s announcement says weights “drop end of this month,” without giving an exact date. Some secondary reports have cited October 27, 2026, but that specific date has not been confirmed by Mistral directly.
Is Mistral Large 4 open source?
Mistral has confirmed that open weights are coming by the end of October 2026, following the public API preview. The exact license terms for those weights haven’t been published yet.
What does “natively multimodal” mean for this model?
It means Large 4 was trained from the start on mixed data types, rather than having image or other modality support added after text-only pretraining. Mistral has not yet published specific multimodal benchmark results to verify how well this performs in practice.
How does Mistral Large 4 compare to Claude, Gemini, and GPT models?
No independent benchmark comparisons exist yet. Large 4’s 1-million-token context window matches what Anthropic has shipped with Claude Opus 5.5, but head-to-head performance data isn’t available until third-party evaluators run their own tests.
Why did earlier reports cite different parameter counts for Mistral Large 4?
Pre-launch leaks and early estimates circulated figures like 675 billion and 1.05 trillion parameters before Mistral’s official post went live. Those numbers don’t match the 1 trillion total and 49 billion active figures Mistral has now confirmed directly, which should be treated as the authoritative source going forward.




