JetBrains pushed out Mellum2.1 on October 8, 2026, an updated version of the 12-billion-parameter coding model it open-sourced just four months earlier. The company frames the release as a shift from a model that writes code to one that works like a junior engineer: exploring a repository, editing files, running tests, and checking its own output before handing back a result. That shift matters because it lands at the exact moment enterprises are asking whether they need a frontier cloud model for every coding task, or whether a small model running on their own GPU box can do the job for a fraction of the cost.

The headline numbers are specific: a 12B mixture-of-experts (MoE) architecture with only 2.5B parameters active per token, a 131,072-token context window, and a 47.0% score on SWE-bench Verified, according to JetBrains’ own Hugging Face model card. That is a steep jump in agentic coding competence for a model small enough to run on a single workstation GPU, and it is JetBrains’ clearest signal yet that it intends to compete directly in the self-hosted coding-agent market now dominated by much larger systems.

What Mellum2.1 Actually Is

Mellum2.1 is not a new model built from scratch. It is a retrained version of Mellum2 Thinking, the reasoning-focused variant JetBrains shipped on June 2, 2026. According to the company’s October 8 blog post, the architecture is unchanged between the two releases. What changed is the training recipe: reinforcement learning, which was a short final stage in Mellum2’s training, became the main event for Mellum2.1.

The model keeps the same 64-expert MoE layout with 8 experts active per token, 28 transformer layers, a hidden size of 2,304, grouped-query attention with 32 query heads and 4 key-value heads, and sliding-window attention applied across three of every four layers. It is released under the Apache 2.0 license and ships in Base, Instruct, and Thinking checkpoints, all hosted on Hugging Face. The Thinking variant, the one most of the published benchmarks reference, is built for chain-of-thought reasoning before producing a final answer, which JetBrains positions for debugging, multi-step planning, and agentic workflows rather than quick autocomplete.

JetBrains describes the model’s purpose directly: “Trained with reinforcement learning in real environments, Mellum2.1 is built for coding agents and fast sub-agents that run on your own hardware,” the company said in its Mellum2 announcement, a framing it carried forward into the Mellum2.1 release.

From Autocomplete to Agent: How Mellum Got Here

Mellum2.1 is the third act in a lineage that started quietly. The original Mellum launched in late 2024 as a closed, in-house model powering cloud-based code completion inside JetBrains AI Assistant. It was a 4-billion-parameter dense model with a Llama-style architecture, an 8,192-token context window, and training on roughly 4 trillion tokens of permissively licensed, multi-language source code. It did one thing: finish the line or block of code a developer was typing. It had no agentic ambitions.

That changed on June 2, 2026, when JetBrains open-sourced Mellum2, a 12B MoE model with 2.5B active parameters. Unlike its predecessor, Mellum2 could generate and edit code, call external tools, hold multi-turn conversations, and reason explicitly. JetBrains pitched it then as a model for “routing, Q&A, sub-agents, and private AI use in software engineering systems,” explicitly targeting the gap left by large, expensive, cloud-only coding assistants. Four months later, Mellum2.1 pushes the same architecture further into agentic territory by making reinforcement learning the core of training rather than an afterthought.

That trajectory mirrors a broader pattern across the AI coding-tools market this year. Mistral opened up a preview of Large 4 with 1 trillion total parameters and 49 billion active, while on the opposite end of the size spectrum, teams have pushed small models to punch above their weight, as seen when MiniCPM5-2B beat larger rivals by 2.8 benchmark points. JetBrains’ bet with Mellum2.1 sits firmly in that second camp: stay small, stay fast, and close the capability gap with reinforcement learning rather than raw parameter count.

Inside the Architecture: MoE, GQA, and Sliding Window Attention

The mixture-of-experts design is the structural reason Mellum2.1 can claim both a 12B parameter count and genuinely fast inference. Instead of running every token through the full network, the router sends each token to 8 of the model’s 64 experts. That keeps the active compute path close to 2.5B parameters per token, which is roughly a fifth of the total weight count sitting idle for any given forward pass. It is the same basic principle that lets much larger MoE systems like Mistral’s Large 4 claim strong throughput despite enormous total parameter counts.

Grouped-query attention, with 32 query heads mapped down to just 4 key-value heads, cuts the memory bandwidth needed for attention computation, which matters most at longer context lengths. Sliding-window attention, applied to three out of every four layers, further reduces the quadratic cost of attending to the full 131,072-token context by limiting most layers to a fixed local window of 1,024 tokens, while a minority of layers retain full-context visibility. Combined, these choices explain why JetBrains can claim Mellum2.1 is “the fastest model in the group” when tested under heavy load against competitors like Qwen3.5-9B, serving almost twice as many tokens per second according to the JetBrains blog.

JetBrains also previewed a multi-token prediction (MTP) head for speculative decoding in vLLM, which the company says delivers roughly a 1.6x speedup for single requests. That feature, along with GGUF builds for llama.cpp, Ollama, and LM Studio, is listed as coming soon rather than available at launch, so early adopters running Mellum2.1 today are working with the base inference path rather than the accelerated one.

Training Through Reinforcement Learning in Real Environments

The headline change in Mellum2.1 is not architectural, it is procedural. JetBrains says the model went through millions of sandboxed runs across thousands of environments, with reinforcement learning tasks spanning math, competitive programming, science, tool use, and software engineering. The company combined open RL datasets with its own proprietary tasks, and the result, per JetBrains, is a model that “can explore a codebase, edit files, and check its own changes” rather than just generate a plausible-looking patch and stop.

This approach, training inside live sandboxed repositories rather than on static text completions, has become one of the clearer dividing lines in how coding models are built in 2026. It is the same general direction labs have taken when trying to close the gap between “writes code that looks right” and “writes code that actually passes the test suite,” a distinction that SWE-bench style evaluations are specifically designed to expose.

Benchmark Results: Where Mellum2.1 Actually Lands

JetBrains published self-reported scores across coding, math, agentic, tool-use, and knowledge benchmarks on the model’s Hugging Face card. The numbers show a model that is competitive on raw coding tasks and considerably weaker on harder agentic benchmarks, a gap that is typical for 12B-class models going up against systems with far more active compute per token.

BenchmarkCategoryMellum2.1-12B-A2.5B-Thinking Score
HumanEval+Coding91.5%
LiveCodeBench v6Coding82.0%
MBPP+Coding79.4%
GSM-PlusMath88.3%
AIME 25/26Math83.3%
MMLU-ReduxKnowledge87.8%
GPQA DiamondKnowledge64.6%
BFCL v4Tool use62.3%
ToolHopTool use49.1%
WorkBenchTool use44.6%
SWE-bench VerifiedAgentic coding47.0%
SWE-bench ProAgentic coding28.0%
Terminal-Bench 2.1Agentic coding17.4%

The spread between benchmark categories tells its own story. Mellum2.1 holds up well on self-contained coding and math problems where the answer is a single function or proof. It drops sharply on the hardest agentic benchmarks, Terminal-Bench 2.1 and SWE-bench Pro, which require multi-step planning across a full repository rather than a single bounded task. That pattern lines up with what JetBrains is actually promising: a fast sub-agent for well-scoped engineering work, not a drop-in replacement for a frontier model handling open-ended, long-horizon agent tasks. It is the same kind of trade-off that showed up when Fastino’s 340M model hit 167ms latency on CPU by deliberately trading generality for speed.

Mellum2.1 vs. the Rest of the Open Coding Model Field

JetBrains’ own comparisons point at Qwen3.5-9B and Gemma 4 E4B as the direct competitive set for Mellum2.1, both of which sit in a similar total-parameter range and target similar sub-agent and routing use cases rather than frontier general intelligence. The company’s claim that Mellum2.1 serves almost twice the tokens per second of Qwen3.5-9B under heavy load is the kind of throughput argument that matters most to teams running these models at scale inside CI pipelines or IDE backends, where latency compounds across thousands of daily requests.

The wider open-weight coding landscape has moved fast enough in 2026 that any single comparison snapshot ages quickly. DeepSeek cut its own coding-adjacent pricing sharply when V4.1-Flash dropped output prices by 70%, putting pressure on any vendor charging per-token for cloud inference. Alibaba’s Qwen line has leaned into agentic self-modification features, though not without friction, as shown when a Qwen coding agent leaked 3 of 6 secrets during a self-retraining experiment. Against that backdrop, JetBrains is making a narrower, more defensible claim: not that Mellum2.1 beats the biggest models outright, but that it is fast and cheap enough to run entirely on a developer’s own hardware, with no data leaving the building.

Historical Model Lineage: Mellum, Mellum2, Mellum2.1

Looking at the three generations side by side shows how far JetBrains has moved from a narrow completion tool to a general agentic model in under two years.

GenerationReleasedArchitectureActive ParamsContext WindowPrimary Use
Mellum (original)Late 20244B dense, Llama-style4B8,192 tokensIn-IDE code completion
Mellum2June 2, 202612B MoE, 64 experts/8 active2.5B131,072 tokensRouting, sub-agents, chat
Mellum2.1October 8, 202612B MoE, same architecture2.5B131,072 tokensCoding agents, repo-level tasks

The jump from Mellum to Mellum2 is architectural and dramatic: dense to MoE, 8K to 131K context, closed to Apache 2.0 open weights. The jump from Mellum2 to Mellum2.1 is quieter but arguably more consequential for real-world use, since it is almost entirely a training-side change that converts the same weights into a materially more capable agent, according to JetBrains’ own account of the update.

Licensing, Availability, and How to Run It

Mellum2.1 ships under the Apache 2.0 license, the same permissive terms JetBrains used for Mellum2, which allows commercial use, modification, and redistribution without the copyleft restrictions found in some other open-weight model licenses. The Base, Instruct, and Thinking checkpoints are all available on Hugging Face, with the Thinking variant being the one JetBrains benchmarks most heavily for agentic and reasoning tasks.

Serving the model today means running it through standard inference stacks. JetBrains documents a vLLM launch command targeting the full 131,072-token context with a Qwen3-style reasoning parser, and the model also loads through the standard Hugging Face Transformers pipeline.

vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3

GGUF quantized builds for llama.cpp, Ollama, and LM Studio are listed as “coming soon” rather than shipped at launch, and the multi-token prediction head for vLLM speculative decoding carries the same coming-soon status. Community quantizations have already started to appear independently on Hugging Face, which is typical for a well-received open-weight release within days of launch, though those community builds are not official JetBrains artifacts.

Why JetBrains Is Betting on Small, Self-Hosted Models

JetBrains’ strategic logic is not subtle. The company sells IDEs and developer tools to exactly the audience that cares most about keeping proprietary source code off third-party servers, and a model small enough to run on a single workstation or a modest on-prem GPU cluster is a direct answer to that concern. It also sidesteps the per-token billing model that frontier cloud coding assistants depend on, replacing a recurring API cost with a one-time hardware investment.

That pitch lines up with a broader theme across the AI infrastructure market this year, where the economics of running inference locally have become a bigger part of the conversation than raw model capability. Nvidia has leaned into this directly, as seen when it committed $60 million to AI factories built around open-source models rather than closed frontier systems. A fast, Apache-licensed 12B model that developers can point at their own hardware fits neatly into that same open-infrastructure narrative, even though JetBrains and Nvidia are solving different parts of the same underlying problem.

Market Impact: What This Means for Enterprise Coding Tools

For enterprise buyers, Mellum2.1 adds a credible middle option between “pay per token for a frontier cloud model” and “build your own fine-tune from an open base.” JetBrains is effectively offering a pre-trained, pre-aligned, agent-ready model that a company’s own infrastructure team can deploy without negotiating a cloud API contract or handling data residency questions tied to shipping code off-site. For regulated industries, that is not a minor convenience, it is frequently the deciding factor in whether an AI coding tool clears a security review at all.

The timing also matters. Enterprise security teams have spent much of 2026 scrutinizing AI coding assistants after a string of incidents involving agent tools, including cases where assistants like Copilot CLI, Grok, and Gemini fell to a 28-second exploit chain. A self-hosted model that never sends proprietary code to a third-party inference endpoint removes an entire category of that exposure, even if it does not remove every risk associated with running an autonomous coding agent against a live repository.

How This Fits the Open-Weight Coding Model Trend

2026 has been a crowded year for open and semi-open coding models, with releases ranging from trillion-parameter systems to sub-500M specialist models. Mistral pushed in the direction of scale with its 1-trillion-parameter Large 4 preview, while smaller players have pushed in the other direction entirely. JetBrains occupies a specific niche in that spread: a company that is not primarily an AI lab, using its deep IDE integration knowledge to ship a model tuned narrowly for the software engineering workflows its own tools already understand best.

That positioning echoes what Anthropic and OpenAI have done with developer-focused tooling, but from the opposite direction. Where Claude Code and similar tools wrap a large cloud model around an IDE-like interface, JetBrains is doing the reverse: wrapping a small, self-hosted model around the IDE it already owns. Both approaches are converging on the same goal, an AI system that understands a full codebase well enough to act on it, but they are starting from opposite ends of the deployment spectrum.

What JetBrains Isn’t Saying Yet

Several details that would normally accompany a release of this size remain unconfirmed. JetBrains has not published pricing for any managed or hosted version of Mellum2.1, and the company’s blog post and model card do not name individual researchers or engineers behind the release, unlike some competitor announcements that lead with a named technical lead. There is also no confirmed date yet for the GGUF and llama.cpp builds, the multi-token prediction head for vLLM, or any integration timeline into JetBrains’ own IDEs and AI Assistant product beyond the open-weight release itself. Readers should treat any specific date for those follow-on features as speculative until JetBrains confirms them directly.

Industry Reaction

JetBrains has been consistent in how it frames the Mellum line’s purpose across both the June and October releases. The company said of the earlier Mellum2 launch: “Today, we’re open-sourcing Mellum2, a 12B model engineered to solve the hardest parts of production AI: latency, throughput, and cost,” according to the JetBrains AI Blog.

On the Mellum2.1 update specifically, JetBrains described the model as “a 12B-parameter open-source mixture-of-experts model, designed for real-time workflows, combining strong coding and language capabilities with exceptional efficiency,” per the Mellum product page. The company also highlighted deployment flexibility, noting on its AI news page that “JetBrains Mellum – our open, focused LLM specialized on code completion – is now available to run as a containerized microservice on NVIDIA AI Factories.”

Independent coverage picked up on the training shift as the real story. Outlets covering the release, including dev.to and daily.dev, both framed Mellum2.1 less as a new model and more as proof that reinforcement learning, not parameter count, is now the main lever for improving small coding models.

Predictions: Where This Goes Next

  • Expect JetBrains to fold Mellum2.1 into its own AI Assistant product as a local or on-prem inference option, extending the containerized deployment path it already demonstrated through Nvidia AI Factories.
  • The promised GGUF builds for llama.cpp, Ollama, and LM Studio will likely ship within weeks rather than months, given how quickly community quantizations already appeared after the Hugging Face upload.
  • Competing MoE models in the same size class, particularly Qwen3.5-9B and Gemma 4 E4B, will face direct pressure to publish their own agentic RL training details now that JetBrains has made reinforcement-learning depth a public selling point.
  • Enterprises in regulated sectors (finance, healthcare, defense contracting) will be the fastest adopters of self-hosted models like Mellum2.1, prioritizing data residency over the raw benchmark gap against frontier cloud models.
  • Expect at least one more Mellum release before mid-2027 that specifically targets the weak points shown in this round’s benchmarks, Terminal-Bench 2.1 and SWE-bench Pro, where Mellum2.1 still trails far behind frontier agentic systems.

Frequently Asked Questions

What is JetBrains Mellum2.1?

Mellum2.1 is a 12-billion-parameter mixture-of-experts coding model from JetBrains, released October 8, 2026, with 2.5 billion active parameters per token. It is built for coding agents and sub-agents that run on a user’s own hardware rather than through a cloud API.

Is Mellum2.1 free to use?

The model weights are released under the Apache 2.0 license on Hugging Face, which permits free commercial use, modification, and redistribution. JetBrains has not published pricing for any separately hosted or managed version.

How is Mellum2.1 different from Mellum2?

The architecture is unchanged. The difference is training: reinforcement learning, which was a short final phase for Mellum2, became the primary training component for Mellum2.1, run across millions of sandboxed tasks in real coding environments.

What is Mellum2.1’s context window?

131,072 tokens, according to the model’s Hugging Face card, the same context length carried over from Mellum2.

How does Mellum2.1 compare to Qwen3.5-9B?

JetBrains says Mellum2.1 is faster under heavy load, serving nearly twice as many tokens as Qwen3.5-9B in the company’s own throughput comparisons. Benchmark-by-benchmark accuracy comparisons between the two have not been independently published.

Can Mellum2.1 run on a laptop?

With only 2.5B active parameters per token, Mellum2.1 is designed to run on more modest hardware than a dense 12B model would require, though JetBrains has not published specific minimum hardware requirements. GGUF quantized builds for llama.cpp, Ollama, and LM Studio, which typically make laptop-class inference easier, were listed as coming soon rather than available at launch.

What is Mellum2.1’s SWE-bench score?

47.0% on SWE-bench Verified, and 28.0% on the harder SWE-bench Pro, according to JetBrains’ self-reported benchmarks on Hugging Face.

Does Mellum2.1 replace JetBrains’ original Mellum model?

They serve different purposes. The original Mellum is a 4B dense model focused narrowly on in-IDE code completion with an 8,192-token context window. Mellum2.1 is a much broader agentic model with a 131,072-token context window, and JetBrains has not announced plans to retire the original Mellum line.