the AI bench
VERIFIED SEPTEMBER 2026
All models

MODEL · IBM RESEARCH · 3B / 8B / 30B DENSE (INSTRUCT + BASE EACH)

IBM Granite 4.1 → 4.2

IBM's refreshed open-weights enterprise family — three dense decoder-only sizes, Apache 2.0, trained on ~15T tokens with progressive annealing toward technical/scientific/mathematical data plus instruction-following. The 8B instruct claims to match the prior Granite 4.0 32B-A9B MoE flagship on IBM's own benchmarks; cross-vendor comparison (vs Qwen/Gemma/Mistral) is unverified at time of publication.

License: Apache 2.0 · Context: 128K default; 512K via late-training context-extension stage · Released: April 29, 2026 (4.1); August 25, 2026 (4.2)

The decision in five lines

The call
Consider — runnable locally, family reference
Best for
Local evaluation and family reference
Runs on
23 hardware picks fit (cheapest: Intel Arc B580 12 GB · $310)
Watch out
Frontier reasoning or coding-agent workflows where Qwen3-Coder-30B-A3B or GLM-5.1 are the established daily drivers.
Evidence
Estimated · last verified September 2026

3B
PARAMETERS
DENSE
TYPE
128K
CONTEXT
~2.1 GB (3B) / ~5.3 GB (8B) / ~17 GB (30B) — Ollama Q4_K_M
VRAM AT Q4

Where we recommend this

This model isn’t currently in an active planner slot. See the runner notes below if you’re running it anyway.

The call

IBM's refreshed open-weights enterprise family — three dense decoder-only sizes, Apache 2.0, trained on ~15T tokens with progressive annealing toward technical/scientific/mathematical data plus instruction-following. The 8B instruct claims to match the prior Granite 4.0 32B-A9B MoE flagship on IBM's own benchmarks; cross-vendor comparison (vs Qwen/Gemma/Mistral) is unverified at time of publication.

When not to use: Frontier reasoning or coding-agent workflows where Qwen3-Coder-30B-A3B or GLM-5.1 are the established daily drivers. Granite 4.1 is positioned as enterprise-deploy-friendly (clean Apache 2.0, IBM support story, traceable training data) rather than benchmark-chasing.

Runner notes

**Update — September 7, 2026: Granite 4.2 is the current generation.** IBM released 4.2 on August 25, 2026 (model-card date; HF repos created Aug 7, official GGUFs Aug 12) at the same three dense sizes and the same Apache 2.0 licence, built on the 4.1 base models. Two real changes: native reasoning ("thinking") with reasoning-augmented tool calling, and a 512K context window by default rather than via extended-context variants. The official GGUF repos (`ibm-granite/granite-4.2-{3b,8b,30b}-GGUF`) were at 149K–179K downloads each on Sep 7. `granite4.2:8b` resolves on the Ollama registry. Memory footprints are unchanged, so every fit below still holds; we have not reproduced IBM's evaluation tables and no planner pick moves on them. This page keeps its 4.1 URL. Ollama tags live same-day: `granite4.1:3b`, `granite4.1:8b`, `granite4.1:30b`. Default tags ship at 128K context — 512K requires the extended-context training-stage variants from `huggingface.co/ibm-granite`. Q4_K_M is the Ollama default; Q8_0 available via `:8b-q8_0` etc. HuggingFace hosting at `ibm-granite/granite-4.1-*-instruct`.

License
Apache 2.0
Released
April 29, 2026 (4.1); August 25, 2026 (4.2)
Maker
IBM Research

Hardware that fits

Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.

Next step

Find-by-model — see what hardware runs this