FAST TAKE · 2026-07-27 · THINKING MACHINES INKLING-SMALL (276B-A12B)
Inkling-Small — a new lab ships open weights, and they are good
Thinking Machines released Inkling-Small on July 27 under Apache 2.0: a 276B-total / 12B-active multimodal MoE that takes text, image and audio in. The headline is not the size, it is the ratio — on the lab's own comparison table it posts an Artificial Analysis Index of 40.0 against Qwen3.5 397B-A17B's 34.0 and Nemotron 3 Ultra 550B-A55B's 38.0, while activating 12B parameters per token. It is still too big to run at home, so no planner pick, but a genuinely permissive licence on a model this capable is the kind of thing that reshapes what "open" means a year from now.
Verdict: Thinking Machines' first open weights, Apache 2.0 — and it beats every larger open model in its comparison table
The take
Verified from the model card: 276B total / 12B active, a 42-layer decoder routing each token to 6 of 256 experts plus 2 shared experts, hybrid local/global attention, natively multimodal (images via a hierarchical patch encoder, audio via discrete tokens, all projected into one shared space). BF16 and NVFP4 numerics, with an official NVFP4 sibling. The HF repo carries 33 real safetensors shards and the `apache-2.0` tag — not a custom licence with "open" in the name, the actual Apache 2.0.
The evaluation table is unusually forthcoming, and worth reading carefully rather than taking the top line. SWE-Bench Verified 80.2%, SWE-Bench Pro 55.9%, Terminal-Bench 2.1 64.7%, AIME 2026 95.5%, GPQA Diamond 89.5%. Against closed models it lands between Claude Haiku 4.5 and GPT-5.6 Luna on the AA Index (30.0 / 40.0 / 49.0). But the lab flags its own methodology honestly: its numbers use an internal harness while external models are self-reported, and it discloses that some Terminal-Bench solutions were found contaminated from web search and scored zero. It also publishes where it loses — SimpleQA Verified 20.6% against DeepSeek V4 Flash's 34.1%, and an AA Omniscience index of −9.0, meaning it hallucinates more than several models it beats on reasoning. Labs that publish their losses are easier to believe about their wins.
Our call: fast take, no model entry and no planner pick. At 276B total this is multi-GPU-server or nothing — roughly 138 GB even at Q4, past every box on our hardware list, and there are no official GGUF or MLX builds. What makes it worth a dated verdict is the licence and the shape. A 12B-active model performing at this level under Apache 2.0 is the argument that sparse MoE plus permissive licensing is where open weights get genuinely useful, and if Thinking Machines ships a smaller sibling on the same recipe, that one will be a planner candidate. We will look again when it does.
Where this fits
Models: Qwen 3.5 122B-A10B · DeepSeek V4-Flash · gpt-oss-120b · Kimi K3 (2.8T-A50B)
Hardware: NVIDIA DGX Spark · Dual RTX 5090 · Mac Studio M3 Ultra 96 GB
Sources
Next step
Try this in the planner→