ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data AI Agents Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests Claude Anthropic Makes Cruelty Toward Claude a Policy Violation in First Usage-Policy Rewrite in Over a Year Coding Assistants Harness Buys Augment Code's Cosmos to Complete Its Autonomous Software Factory, From Ticket to Merge-Ready PR Claude Anthropic Turns Claude Into a BI Tool: Dashboards and Motion Enter Beta as Docs, Slides and Design Go GA ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data AI Agents Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests Claude Anthropic Makes Cruelty Toward Claude a Policy Violation in First Usage-Policy Rewrite in Over a Year Coding Assistants Harness Buys Augment Code's Cosmos to Complete Its Autonomous Software Factory, From Ticket to Merge-Ready PR Claude Anthropic Turns Claude Into a BI Tool: Dashboards and Motion Enter Beta as Docs, Slides and Design Go GA

JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning

JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning

JetBrains' Mellum2.1 keeps the 12B-total/2.5B-active MoE shape of Mellum2 but puts reinforcement learning at the center of training, lifting SWE-bench Verified from 2.0 to 47.0 and Terminal-Bench 2.1 from 0.6 to 17.4 - all self-reported. Weights ship under Apache 2.0 with GGUF builds from 7GB.

JetBrains released Mellum2.1 on Wednesday, the second major version of its open coding model in four months, and the headline is a single number: SWE-bench Verified went from 2.0 to 47.0. The 12B-parameter mixture-of-experts model keeps the exact architecture of Mellum2 - 12 billion total parameters with 2.5 billion active per token - and the company attributes nearly all of the gain to one change: reinforcement learning moved from a short final stage to the center of post-training.

The training recipe is the interesting part. JetBrains put the model to work in real software repositories inside sandboxes, giving it a shell and file-editing tools and rewarding it when tests pass, with math, competitive programming, science and tool-use tasks mixed in. The company says it filtered open RL datasets for broken tests and unverifiable answers, and launched millions of sandboxes across thousands of environments. The released checkpoint, Mellum2.1-12B-A2.5B-Thinking, emits its chain of thought before answering and is aimed at three uses: agent worker, general reasoning assistant, and private self-hosted deployment.

The numbers - all JetBrains self-reported, measured with one pipeline in thinking mode - show both the strength and the limits of that recipe. SWE-bench Pro rose from 0.0 to 28.0 and Terminal-Bench 2.1 from 0.6 to 17.4. The model leads its comparison group on LiveCodeBench v6 at 82.0, HumanEval+ at 91.5 and MBPP+ at 79.4, and beats its own predecessor on 15 of 17 listed benchmarks. But Qwen3.5-9B still leads on the hardest agentic software tasks, taking SWE-bench Verified at 50.0, SWE-bench Pro at 38.0 and the AIME math exams at 86.7. Pipeline effects are real too: JetBrains measured Qwen3.5-9B at 75.4 on LiveCodeBench v6, well above the 65.6 on Qwen's own card.

The pitch is efficiency rather than frontier quality. With 64 experts and 8 active per token, grouped-query attention and a 131,072-token context, the model does the work of roughly 2.5 billion parameters per word - dense 9B competitors run every parameter on every token. Weights ship in bfloat16 under Apache 2.0 on Hugging Face, GGUF builds start at 7.0GB for llama.cpp, Ollama and LM Studio, and JetBrains says the model runs on a single NVIDIA H200 through vLLM or SGLang. Multi-token prediction support is due within days.

Safety moved too: HarmBench fell from 21.5 to 8.5, where lower is better. Agentic evaluations used the open-source Pi v0.73.1 harness with a 114K-token context and up to 16,000 tokens per turn, which matters when comparing the numbers to other leaderboards - JetBrains re-measured its own Mellum2 with the same pipeline, which is why its published scores differ slightly from the June technical report.

The strategic read is that JetBrains is building for a specific lane: small, self-hostable models that can explore a repository, edit files and check their own work without sending proprietary code to an API. Mellum2.1 will not beat Qwen3.5-9B or a frontier model at fixing hard real-world bugs - the company's own table says so. But for IDE vendors and enterprises that want a coding agent running on their own GPUs under a permissive license, a 47.0 on SWE-bench Verified from a model this small, four months after the same architecture scored 2.0, is a strong argument that post-training - not scale - is where open coding models are being won.

Comments (0)

Log in to join the discussion

Log In

No comments yet