JetBrains announced the launch of Mellum2.1, an open model designed for coding agents and sub-agents that need to perform rapid tasks on users’ devices or their own infrastructure. The release represents the next generation of the 12-billion-parameter Mellum2 model, while retaining the same mixture-of-experts architecture that uses 2.5 billion active parameters, under the Apache 2.0 license.
The main improvement did not come from changing the architecture, but from the post-training stage. JetBrains says it relied more heavily on reinforcement learning and ran millions of isolated executions across thousands of software environments, with the aim of training the model on tasks closer to actual work inside code repositories.
From Answering to Working Inside the Repository
Mellum2 was fast, but it did not handle software repositories at the required level, according to JetBrains. Mellum2.1, however, can explore the codebase, modify files, and inspect the changes it makes. This makes it possible to use the model to execute different stages of a coding agent’s plan, such as identifying the root cause of a failed test, then preparing and validating a fix.
Training with More Rigorous Data and Environments
The new training tasks covered mathematics, competitive programming, science, tool use, and software engineering. JetBrains combined open data with tasks developed internally, but noted that open data may include broken tests, unverifiable answers, or problems that are too easy or impossible; therefore, the sources were filtered before being included in training.
The company also focused on building reinforcement-learning environments capable of running thousands of environments internally, enabling millions of isolated trials during training.
Improved Performance and Speed
JetBrains compared Mellum2.1 with the previous version and with the Qwen3.5-9B and Gemma 4 E4B models, using a unified evaluation setup. The greatest improvement was in agentic coding, with additional gains in programming, competitive programming, mathematics, tool calling, and general knowledge.
Because the post-training stage did not change the architecture, the model retained Mellum2’s speed, while multi-token prediction (MTP) contributes to faster inference. JetBrains says that, under high load, the model serves nearly twice as many tokens as Qwen3.5-9B, while MTP makes it about 1.6 times faster per request.
What Changes in Practice?
Mellum2.1 gives development teams an open, self-hostable option for building coding agents or sub-agents, while keeping code and data under the operator’s control. However, the reported figures reflect JetBrains’ comparative tests and the material does not include independent details about hardware requirements or performance limits in real-world projects.
The model is currently available on Hugging Face. The GGUF releases for llama.cpp, Ollama, and LM Studio, along with the MTP head used for speculative decoding through vLLM, are scheduled to arrive later.