JetBrains released Mellum2.1 on Wednesday, the second major version of its open coding model in four months, and the headline is a single number: SWE-bench Verified went from 2.0 to 47.0. The 12B-parameter mixture-of-experts model keeps the exact architecture of Mellum2 - 12 billion total parameters with 2.5 billion active per token - and the company attributes nearly all of the gain to one change: reinforcement learning moved from a short final stage to the center of post-training.
The training recipe is the interesting part. JetBrains put the model to work in real software repositories inside sandboxes, giving it a shell and file-editing tools and rewarding it when tests pass, with math, competitive programming, science and tool-use tasks mixed in. The company says it filtered open RL datasets for broken tests and unverifiable answers, and launched millions of sandboxes across thousands of environments. The released checkpoint, Mellum2.1-12B-A2.5B-Thinking, emits its chain of thought before answering and is aimed at three uses: agent worker, general reasoning assistant, and private self-hosted deployment.
The numbers - all JetBrains self-reported, measured with one pipeline in thinking mode - show both the strength and the limits of that recipe. SWE-bench Pro rose from 0.0 to 28.0 and Terminal-Bench 2.1 from 0.6 to 17.4. The model leads its comparison group on LiveCodeBench v6 at 82.0, HumanEval+ at 91.5 and MBPP+ at 79.4, and beats its own predecessor on 15 of 17 listed benchmarks. But Qwen3.5-9B still leads on the hardest agentic software tasks, taking SWE-bench Verified at 50.0, SWE-bench Pro at 38.0 and the AIME math exams at 86.7. Pipeline effects are real too: JetBrains measured Qwen3.5-9B at 75.4 on LiveCodeBench v6, well above the 65.6 on Qwen's own card.
The pitch is efficiency rather than frontier quality. With 64 experts and 8 active per token, grouped-query attention and a 131,072-token context, the model does the work of roughly 2.5 billion parameters per word - dense 9B competitors run every parameter on every token. Weights ship in bfloat16 under Apache 2.0 on Hugging Face, GGUF builds start at 7.0GB for llama.cpp, Ollama and LM Studio, and JetBrains says the model runs on a single NVIDIA H200 through vLLM or SGLang. Multi-token prediction support is due within days.
Safety moved too: HarmBench fell from 21.5 to 8.5, where lower is better. Agentic evaluations used the open-source Pi v0.73.1 harness with a 114K-token context and up to 16,000 tokens per turn, which matters when comparing the numbers to other leaderboards - JetBrains re-measured its own Mellum2 with the same pipeline, which is why its published scores differ slightly from the June technical report.
The strategic read is that JetBrains is building for a specific lane: small, self-hostable models that can explore a repository, edit files and check their own work without sending proprietary code to an API. Mellum2.1 will not beat Qwen3.5-9B or a frontier model at fixing hard real-world bugs - the company's own table says so. But for IDE vendors and enterprises that want a coding agent running on their own GPUs under a permissive license, a 47.0 on SWE-bench Verified from a model this small, four months after the same architecture scored 2.0, is a strong argument that post-training - not scale - is where open coding models are being won.
Comments (0)
Log in to join the discussion
Log InNo comments yet