Tech

IBM Unveils Granite 4.1: Dense Architecture Models Challenge MoE Rivals

The 8B model matches or outperforms a 32B predecessor on key benchmarks, while the suite supports 512K context windows under an Apache 2.0 licence.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
New open-source family from IBM prioritises efficiency and enterprise reliability over parameter bloat.

IBM has launched Granite 4.1, a new family of open-source language models specifically engineered for enterprise applications. The release comprises three dense transformer sizes with 3B, 8B, and 30B parameters, all trained on a dataset of 15 trillion tokens. Unlike previous iterations that relied on Mixture of Experts (MoE) techniques, this generation utilises a dense architecture designed to offer predictable performance and latency without the complexity of sparse routing.

A central highlight of the announcement is the 8B model, which matches or outperforms the previous generation's 32B MoE model across various benchmarks. On the ArenaHard benchmark, the 8B instruct model scored 69.0, surpassing the prior Granite 4.0-H-Small. Similarly, on the BFCL V3 tool-calling standard, the 8B achieved a score of 68.3 compared to 64.7 for the 32B model. IBM attributes this efficiency to a rigorous data curation process that filtered 4.1 million fine-tuning samples using an LLM-as-Judge system to ensure high-quality instruction following and grounding.

The 30B variant also demonstrates significant capability, leading IBM's internal BFCL V3 chart with a score of 73.7, outperforming the Gemma-4-31B model which scored 72.7. The suite is designed to handle extensive context windows of up to 512K tokens. To maintain short-context performance while extending to this capacity, IBM employed a staged extension approach moving from 32K to 128K and finally to 512K, combined with model merging to preserve earlier weights.

Development of the models involved a unique four-stage Reinforcement Learning (RL) pipeline to address performance regressions observed during earlier phases. After initial joint domain training and general chat reinforcement learning, a dedicated math recovery stage was implemented to fix performance drops caused by the chat RLHF phase. This iterative approach resulted in the 8B model achieving a score of 92.5 on GSM8K and 60.8 on BFCL V3, making it competitive for edge deployment despite its smaller parameter count.

Granite 4.1 is distributed under an Apache 2.0 licence, facilitating commercial use without legal restrictions. The models are available for immediate deployment via Ollama, Hugging Face, and IBM's API. The 3B model offers a lightweight option for edge use cases, while the 30B serves high-compute requirements, providing a flexible range for various enterprise infrastructure needs.

Continue reading

More from Tech

Read next: The Walrus warns of collapsing digital memory as AI erodes search reliability
Read next: Open-source tool claims 97 per cent token savings for AI agents
Read next: Valvoline Unveils August 2026 Promotional Offers for Service and Retail Buyers