FAST TAKE · 2026-08-25 · IBM GRANITE 4.2 (3B / 8B / 30B)
Granite 4.2: same 3B / 8B / 30B sizes, now with thinking and 512K context, Apache 2.0
IBM released Granite 4.2 on August 25, 2026 — the model card's own date; the HuggingFace repos were created August 7 and the official GGUF builds on August 12. Same three dense sizes as Granite 4.1, same Apache 2.0 licence, built on the 4.1 base models, with two real changes: native step-by-step reasoning ("thinking") before the final answer, and a 512K context window as the default rather than a late-training variant. The GGUF repos were at 149K–179K downloads each on September 7, ahead of the BF16 originals, which says most people are running these locally.
Verdict: IBM's dense Apache 2.0 family gains native reasoning and 512K context at the same three sizes — and we were two weeks late noticing
The take
The facts, from the `ibm-granite/granite-4.2-8b` card read September 7: "Decoder-only Dense Transformer (Reasoning)", base model Granite-4.1-8B-Base, 8B parameters, Apache 2.0, release date August 25, 2026, 512K context, "reasoning-augmented tool calling" as the headline feature, trained with a final asynchronous reinforcement-learning stage. The 3B and 30B repos carry the same date and licence; safetensors totals are 3.66B, 8.79B and 29.3B. Ollama already serves `granite4.2:8b` (the registry manifest resolves), alongside the 4.1 tags.
What changes for local buyers: the sizes and memory footprints are unchanged, so every hardware fit on our Granite page still holds — roughly 2 GB, 5 GB and 17 GB at Q4. What is new is that the 8B and 30B now compete on the axis where Granite 4.1 lost to Qwen: reasoning and tool-calling. IBM's evaluation tables are the vendor's own and we have not reproduced them, so we are not moving any planner pick on the strength of a card. The 512K default context is the concrete, checkable upgrade: it removes the 4.1 caveat that long context needed the extended-context variants.
Why we were late: Granite 4.2 shipped in the shadow of the Qwen3.8-Flash and GLM-5.3 drops the same week, and our August 28 sweep listed new repos by announcement rather than by download count — the exact failure our "sort by adoption" rule was written for after Qwen3.8-27B. It is the second time the rule has been needed in a month; this sweep applied it, which is how the GGUF download numbers surfaced.
Our call: the `/models/granite-4-1/` page now describes 4.2 as the current generation and keeps its URL; no planner-pick change until independent numbers exist. What would change our mind: third-party reasoning and tool-calling evaluations that put the 8B ahead of Qwen3.5-9B in its memory band, and a clean read of the thinking mode's token overhead at Q4 on a 12 GB card.
Where this fits
Models: IBM Granite 4.1 → 4.2 · Qwen3.8-27B · Gemma 4 (31B dense + 26B A4B MoE + 12B multimodal)
Hardware: NVIDIA RTX 3060 12 GB · RTX 5060 Ti 16 GB · Mac Mini M4 16 GB
Sources
Next step
Try this in the planner→