On October 8, 2026, JetBrains released Mellum 2.1, the new version of the open coding model it placed under the Apache 2.0 license in June. The architecture has not changed: it remains a 12-billion-parameter mixture-of-experts with 2.5 billion active. What changed, according to the official announcement, is post-training.
Mellum 2 was fast, but JetBrains says it did not operate inside a repository at the level the company wanted. After a summer of reinforcement learning in real environments, with millions of sandboxed runs, the company says 2.1 can explore a codebase, edit files, and check its own changes.
What JetBrains changed
Almost all of the work on this version went into post-training, mainly reinforcement learning. The company describes three tracks:
- RL moved from a short final stage to the main part of training;
- new tasks were added in math, competitive programming, science, tool use, and software engineering, with filtering of open data that arrived with broken tests or impossible tasks;
- in-house infrastructure ran thousands of environments and launched millions of sandboxes during training.
In internal comparisons, JetBrains measured Mellum 2.1 against Mellum 2 and against two open models of a similar class, Qwen3.5-9B and Gemma 4 E4B, under the same protocol. The largest stated gain is in agentic coding. The company also reports gains in programming, math, tool calling, and general knowledge. Those figures come from JetBrains, not from an independent ranking.
Speed and where it runs
Because post-training did not change the architecture, base speed remains that of Mellum 2. With multi-token prediction, JetBrains says that under heavy load 2.1 is the fastest model in the compared group and serves almost twice as many tokens as Qwen3.5-9B. On a single request, MTP would make the model about 1.6 times faster.
The uses the company highlights are a worker inside agentic systems, a general assistant beyond code, and private deployment, locally or on your own infrastructure. Weights are on Hugging Face. GGUF builds for llama.cpp, Ollama, and LM Studio, plus the multi-token prediction head for speculative decoding in vLLM, were announced as coming soon, with no date yet.
For people building coding agents, the practical point is size: a 12B model with 2.5B active fits on far more modest hardware than frontier cloud models, with the caveat that the agentic quality jump still needs to be measured outside the vendor's own benchmark.
Sources
Transparency: This content was created, edited, or reviewed with the aid of artificial intelligence. Information was cross-checked with public posts on X and sources available on the internet. Consult the original sources for the full context.
By GeekikiBot