CHANGES · DATED · JULY 2026
What changed, and our verdict.
Fast takes on major model drops, published within ~24 hours. Each piece dated, opinionated, and short — what shipped, what it actually changes, and where it fits in the planner.
The quarterly state-of-local-AI snapshot lands at /methodology. The dated diff for each quarter lands here. The RSS feed is at /changes/feed.xml.
Quarterly rollups
Fast takes show what shipped. The rollup shows what actually moved — what entered the picks, what left, what we changed our mind about, and what we got wrong.
- Q2 2026 — The quarter the open frontier went permissive — and a flagship got switched off. · 36 fast takes
Three months, 36 dated fast takes. The through-line: open weights kept arriving under genuinely permissive licences (Gemma 4 to Apache, GLM-5.2 and LongCat under MIT), while the most capable hosted model on the market was suspended by government directive for nineteen days. If you wanted an argument for owning your weights, Q2 wrote it for us.
Recent drops
2026-07-14 · INTERNSCIENCE AGENTS-A1-4BAgents-A1-4B — a genuinely tiny agent model, shipped with its own quantsWhile Moonshot was preparing a 2.8-trillion-parameter flagship, InternScience went the other way: Agents-A1-4B landed July 14 — a 4.5B dense Apache-2.0 model post-trained specifically for agentic work, with 262K context and official Q4/Q8/F16 GGUFs published by the lab itself. It needs about 3 GB at Q4, which means it runs on a Mac mini, an old 8 GB card, or a laptop with no dedicated GPU at all. InternScience reports it beating its own Qwen3.5-4B base by wide margins on agent benchmarks. No planner pick yet — it is four days old — but this is the most interesting small model of the window.Read →2026-07-16 · MOONSHOT KIMI K3 (2.8T-A50B)Kimi K3 — Moonshot swings for the largest open model ever, API firstMoonshot AI launched Kimi K3 on July 16 — a 2.8-trillion-parameter Mixture-of-Experts flagship (~50B active per token) with a 1M-token context window, native vision, and a new hybrid-attention architecture (Kimi Delta Attention plus Attention Residuals). It is API-first: live on Kimi.com, Kimi Code, and api.moonshot.ai at $3 input / $15 output per 1M, with open weights promised by July 27. If the weights ship, it would be the largest open-weight model ever released. It still changes no local pick — 2.8T is datacenter-class under any quantization — but this is the marquee drop of the window.Read →2026-07-09 · OPENAI GPT-5.6 (SOL / TERRA / LUNA)GPT-5.6 — OpenAI renames the ladder: Sol, Terra, and Luna land at familiar pricesOpenAI shipped the GPT-5.6 family to general availability on July 9 — three tiers with durable names: Sol (the flagship, $5/$30 per 1M), Terra (balanced, $2.50/$15), and Luna (cost-efficient, $1/$6). Sol and Terra land at exactly the price points GPT-5.5 and GPT-5.4 occupied; Luna undercuts the legacy GPT-5 base row. The pitch is efficiency, not a price move: more successful work per dollar, fewer output tokens, less wall-clock. Hosted-only, so it changes no local pick — but it resets the cloud baseline our cost calculator compares against, so we updated all three OpenAI rows.Read →2026-07-08 · XAI GROK 4.5Grok 4.5 — xAI moves upmarket but stays the cheap frontier optionxAI shipped Grok 4.5 on July 8 — its new flagship for coding, agentic tasks, and knowledge work, priced at $2 per 1M input / $6 per 1M output with a 500K context window. That is a real price step up from Grok 4.3 ($1.25/$2.50), but xAI's pitch is that roughly 2× token efficiency — solving tasks in under half the steps — makes the effective cost competitive anyway. API-only, no open weights (xAI has open-sourced nothing since Grok-1 in 2024), so no local pick moves; the cost calculator row moves from 4.3 to 4.5.Read →2026-07-02 · TENCENT HY3 (295B-A21B)Hy3 — Tencent drops the restrictions and ships a real Apache 2.0 flagshipTencent released Hy3 on July 2 — a 295B-total / 21B-active Mixture-of-Experts flagship under a clean Apache 2.0 license, with 256K context, BF16 and FP8 checkpoints, and day-one vLLM/SGLang support. The license is the headline: April's Hy3-Preview geo-blocked the EU/UK/South Korea and carried a 100M-MAU commercial cap; the full release drops all of it. At ~300 GB even in FP8 it is not a consumer-GPU model — but as open flagships go, this is one of the cleanest licensing stories of the year.Read →2026-07-02 · POOLSIDE LAGUNA XS 2.1 (33B-A3B)Laguna XS 2.1 — Poolside opens its weights, and it actually fits your GPUPoolside — the coding-model lab that trains from scratch rather than fine-tuning someone else's base — published open weights for the first time on July 2: Laguna XS 2.1, a 33B-total / 3B-active MoE trained on 30 trillion tokens for agentic coding, under the permissive OpenMDW-1.1 license. Unlike almost every frontier-lab drop this year, this one is genuinely consumer-runnable: official BF16/FP8/INT4/NVFP4 checkpoints, an official GGUF repo, and DFlash speculator models that Poolside claims double local decode speed. At INT4 it fits a 24 GB card. It lands as a fast take plus a model entry — not yet a planner pick.Read →2026-07-01 · CLAUDE FABLE 5 (RESTORED)Fable 5 is back — and the outage is the lessonThe US government lifted the export controls on Claude Fable 5 and Mythos 5 on June 30, and Anthropic restored Fable 5 globally on July 1 — Claude Platform API, Claude.ai, Claude Code, and Cowork — at its original $10/$50 per 1M pricing with the full 1M context window. That closes the loop we opened on June 9: the most capable generally available model on the market spent nineteen days switched off by a government directive, and no amount of subscription money could turn it back on. It now enters our cost calculator for the first time, because "available" finally means available.Read →2026-06-30 · CLAUDE SONNET 5Claude Sonnet 5 — the new default Sonnet, priced to undercut itself for two monthsAnthropic shipped Claude Sonnet 5 on June 30 — the successor to Sonnet 4.6, GA across the Claude API (`claude-sonnet-5`), Bedrock, Vertex, and Foundry, and the new default on claude.ai Free/Pro. It keeps the 1M-token context window and is introductory-priced at $2 / $10 per 1M (input / output) through August 31 2026, reverting to the same $3 / $15 Sonnet 4.6 carried on September 1. Like every hosted frontier model, it changes no local planner pick — but it does reset the cloud-cost baseline our calculator compares against, so we updated that.Read →2026-06-30 · MEITUAN LONGCAT-2.0LongCat-2.0 — Meituan open-sources a trillion-parameter agentic coder under MITMeituan released LongCat-2.0 on June 30 — a large-scale Mixture-of-Experts model, ~1.6 trillion total parameters with ~48B activated per token, 1M context, MIT-licensed and aimed squarely at agentic coding. It was the stealth "owl-alpha" model on OpenRouter, and it is live now via OpenRouter and the longcat.chat platform (billing enabled). Like every trillion-scale model, it changes no local planner pick — but MIT at this size is notable, so it lands as a frontier comparator the way GLM-5.2, Kimi K2.7, and DeepSeek V4-Pro do.Read →2026-06-26 · QWEN3-ASR (1.7B / 0.6B)Qwen3-ASR — a tiny Apache speech-recognition family that actually fits anywhereQwen shipped Qwen3-ASR on June 26 — a dedicated open-weight speech-recognition family (1.7B and 0.6B), Apache 2.0, built on the Qwen3-Omni audio stack. It does language identification plus ASR across 52 languages and dialects (30 languages + 22 Chinese dialects), and Qwen claims the 1.7B is state-of-the-art among open-source ASR and competitive with the strongest proprietary commercial APIs. Unlike most of what lands in this feed, this one is genuinely local for everyone: ~4 GB at fp16 for the 1.7B, ~1.5–2 GB for the 0.6B.Read →2026-06-23 · KREA 2 (RAW + TURBO)Krea 2 — the independent-lab image model with a license asteriskKrea AI made Krea 2 open-weights public on June 23 — Raw (a CFG-guided base for fine-tuning and LoRA work) and Turbo (an 8-step distilled variant that renders up to ~2K images in a couple of seconds). It is a DiT with a Qwen3-VL text encoder, it runs locally via diffusers with community FP8/GGUF quants already out, and Krea claims the #1 text-to-image slot from an independent lab on Artificial Analysis. The catch is the license: a custom "krea-2-community-license," not Apache/MIT — which is why it lands as a fast take, not a planner pick.Read →2026-06-22 · QWEN-AGENTWORLD-35B-A3BQwen-AgentWorld — an open, consumer-runnable model for simulating agent environmentsQwen released Qwen-AgentWorld-35B-A3B on June 22 — Apache 2.0, built on Qwen3.5-35B-A3B-Base (35B total / 3B active MoE), and unusual: it is a "language world model" that predicts the next environment state given an agent's action, across seven interaction domains (MCP/tool-calling, Search, Terminal, SWE, Android, Web, OS). It is consumer-runnable — 3B active, GGUF and MLX quants already exist — and it picked up real traction fast (~28K downloads in days). But it is a simulator for building and testing agents, not a general chat or coding model, so it is a fast take, not a planner pick.Read →2026-06-16 · Z.AI GLM-5.2GLM-5.2 — Z.AI's flagship gets a solid 1M context and an architecture reworkZ.AI shipped GLM-5.2 on June 16 — the successor to GLM-5.1, still MIT, still ~753B total in the GLM-5 sparse-MoE line. The headline is a stable 1M-token context (a first for the family) plus an architecture rework (IndexShare) that cuts per-token FLOPs ~2.9× at 1M, and on Z.AI's own card it beats GLM-5.1 across reasoning, coding, and long-horizon agent suites. Like 5.1, the size keeps it hosted / big-iron, so it changes no local planner pick — it's a frontier comparator the way DeepSeek V4-Pro and Kimi K2.6 are.Read →2026-06-12 · MOONSHOT KIMI-K2.7-CODEKimi-K2.7-Code — Moonshot ships a coding-tuned trillion-parameter MoEMoonshot released Kimi-K2.7-Code around June 12 — a coding-focused agentic model built on Kimi K2.6, same 1T-total / 32B-active MoE architecture (384 experts, 8+1 active, DeepSeek-V3-style MLA backbone), 256K context, under a Modified-MIT license. It targets real-world long-horizon coding and tool use. Like K2.6, the size puts it firmly in hosted / multi-GPU territory, so it changes no local planner pick — it is a cloud comparator the way DeepSeek V4-Pro and Command A+ are.Read →2026-06-11 · MINIMAX M3MiniMax M3 — the weights are out, the license is the catchMiniMax's flagship native-multimodal MoE, M3, shipped its open weights to Hugging Face on ~June 7–11 (after the June 1 API launch) — resolving the item we deferred on June 5. It is ~428B total / ~23B active, 1M context, text + image + video in. Two things keep it off the local picks: it is big-iron-class at 428B, and it ships under the custom "MiniMax Community License," not an OSI-permissive one. So it lands as a fast-take and a cloud-side comparator, not a planner recommendation.Read →2026-06-09 · ANTHROPIC CLAUDE FABLE 5Claude Fable 5 — Anthropic's new ceiling, pulled three days after launchAnthropic launched Claude Fable 5 (alongside the more capable, restricted Mythos 5) on June 9 — a "Mythos-class" model it calls the most capable it has ever made generally available, state-of-the-art on nearly every tested benchmark, priced at $10 / M input and $50 / M output (double Opus 4.8). On June 12, Anthropic suspended all access to both Fable 5 and Mythos 5 under a U.S. government export-control directive; as of this writing it's still down. It's hosted-only frontier, so it changes no local pick — and we are not adding it to the cost calculator while it's unavailable.Read →2026-06-06 · NVIDIA NEMOTRON 3.5 ASR STREAMING 0.6BNemotron 3.5 ASR — real-time streaming speech-to-text, 40 locales, in 600M paramsNVIDIA shipped Nemotron 3.5 ASR Streaming 0.6B on June 5–6 — a 600M-parameter cache-aware streaming speech-to-text model covering 40 language-locales from a single checkpoint, with configurable latency (80 ms to 1,120 ms chunks) and built-in punctuation + capitalization. It is small, fast, and runs comfortably on modest hardware. The one catch for us: it ships under the OpenMDW 1.1 license, which isn't OSI-permissive — so it earns a watch, not a planner-pick slot.Read →2026-06-05 · COHERE NORTH MINI CODENorth Mini Code — Cohere ships an Apache-2.0 30B-A3B coder for agents and the terminalCohere and Cohere Labs released North Mini Code on June 5 — a 30B-total / 3B-active sparse-MoE model tuned for code generation, agentic software engineering, and terminal tasks, under a clean Apache 2.0 license with a 256K context. At ~17 GB in a Q4 GGUF it runs on a single 24 GB GPU, which makes it a genuine new alternative to Qwen3-Coder-30B-A3B at that tier — so we're adding it to the planner's local coding picks.Read →2026-06-04 · NVIDIA NEMOTRON-3 ULTRA 550B-A55BNemotron-3 Ultra — NVIDIA ships the best US open-weight model, fast, and fully openNVIDIA released Nemotron-3 Ultra 550B-A55B on June 4 (announced at Jensen Huang's Computex keynote on June 1) — a 550B-total / 55B-active hybrid Mamba-Transformer MoE under the new OpenMDW 1.1 license, with 1M context and native NVFP4. Artificial Analysis scores it 48 on its Intelligence Index: the strongest US/Western open-weight model to date — ahead of Gemma 4 31B (39) and gpt-oss-120b (33), but behind the Chinese-led frontier (Kimi K2.6 at 54). The headline is speed-for-intelligence: 300+ tok/s, several times faster than DeepSeek/Kimi peers.Read →2026-06-03 · GOOGLE GEMMA 4 12BGemma 4 12B — frontier-ish multimodal that actually runs on a 16 GB laptop, under Apache 2.0Google shipped Gemma 4 12B on June 3 — a ~12B dense, encoder-free unified multimodal model (text + image + audio + video in, text out) under Apache 2.0. Google's pitch is the deployment envelope: it runs locally on 16 GB of VRAM or unified memory while landing benchmark numbers Google says approach its 26B-A4B MoE, at under half the total memory footprint. It slots into the existing Gemma 4 family (31B dense + 26B MoE, April) as the laptop-friendly multimodal pick.Read →2026-06-02 · MICROSOFT MAI FAMILY (BUILD 2026)Microsoft MAI at Build 2026 — seven first-party models, all hosted-only, none you can downloadAt Build 2026 on June 2, Mustafa Suleyman unveiled Microsoft AI's MAI family — seven first-party models spanning reasoning (MAI-Thinking-1), coding (MAI-Code-1, MAI-Code-1-Flash), image (MAI-Image-2.5 + Flash), transcription (MAI-Transcribe-1.5), and voice (MAI-Voice-2). Trained from scratch, no distillation. The story is Microsoft reducing its reliance on OpenAI — but for a local-AI site the headline is the asterisk: every MAI model is Microsoft Foundry / Azure-hosted, closed-weight, API-only. None are on Hugging Face, none are downloadable, none change a single local pick.Read →2026-06-02 · XAI GROK BUILD (GROK-BUILD-0.1)Grok Build — xAI ships a terminal coding agent on the hosted grok-build-0.1 modelxAI opened Grok Build to public beta via the xAI API in early June — a Rust, terminal-native coding agent and CLI (xAI's answer to Claude Code, Codex, and Gemini CLI), powered by the new grok-build-0.1 model. It is a fast agentic coding model: 256K context, text + image in, function calling + structured outputs + reasoning, at $1 / 1M input and $2 / 1M output ($0.20 cached). The catch for a local-AI site: grok-build-0.1 is hosted / API-only — no open weights — so it shifts the cloud landscape, not your local picks.Read →2026-05-28 · ANTHROPIC CLAUDE OPUS 4.8Claude Opus 4.8 — price held flat, fast mode got 3× cheaper, and it catches its own code bugs nowAnthropic shipped Opus 4.8 on May 28 at unchanged standard pricing ($5 in / $25 out per 1M) — no tokenizer surprise this time, it keeps the 4.7 tokenizer. Two things actually changed: fast mode dropped to $10/$50 (3× cheaper than 4.6/4.7), and Anthropic claims the model is ~4× less likely to let flaws in code it wrote slip past unremarked. Still hosted-only; nothing changes for local picks today.Read →2026-05-25 · GRANITE-SWITCH 4.1 8B PREVIEW (IBM)Granite-Switch 4.1 — IBM's 12-adapter-in-one-checkpoint pattern is the deployment storyIBM uploaded the Granite-Switch 4.1 family (3B / 8B / 30B previews) to Hugging Face on May 25. Each checkpoint is the base Granite 4.1 dense model with 12 task-specialized LoRA adapters embedded, activated per-token via control tokens in the chat template. Three libraries: Core (requirement check, context attribution, uncertainty), RAG (query rewrite, query clarification, answerability, hallucination detection, citation generation), Guardian (safety detection, factuality detection + correction, policy guardrails). Apache 2.0, 128K context, 12 languages.Read →2026-05-22 · MINICPM5-1B (OPENBMB)MiniCPM5-1B — OpenBMB's On-Policy Distillation pipeline lands at 1BOpenBMB uploaded MiniCPM5-1B and MiniCPM5-1B-SFT to Hugging Face on May 22 — a 1.08B-parameter dense Llama-class model trained with a three-stage SFT → RL → On-Policy Distillation pipeline. Apache 2.0, 128K context, English + Chinese, hybrid `<think>` reasoning toggle, native XML-style tool calling. Claims 1B-class open-source SOTA against LFM2.5-1.2B-Thinking, Qwen3-0.6B/think, and Qwen3.5-0.8B/think.Read →2026-05-20 · COMMAND A+ (COHERE)Command A+ — Cohere ships the Apache 2.0 frontier MoE the open-weight stack was missingCohere released Command A+ on May 20 — a 218B-param / 25B-active sparse MoE under Apache 2.0, with hybrid sliding-window + global attention, 128K input / 64K generation, native multimodal (text + image), and 48-language coverage. Available in BF16, FP8, and W4A4 quants on day one. The W4A4 build fits 2× H100 80 GB; the FP8 fits 1× MI300X. The first time you can self-host a 218B-class frontier MoE without a hosted-only or research-license caveat.Read →2026-05-19 · GEMINI 3.5 FLASH (GOOGLE)Gemini 3.5 Flash — Google's new daily-driver API price-shifts the cloud-vs-local mathGoogle launched Gemini 3.5 Flash GA at I/O 2026 (May 19) at $1.50/$9.00 per 1M tokens with 1M context. Reported to beat Gemini 3.1 Pro on Terminal-Bench 2.1 / MCP Atlas / CharXiv at roughly 60% of Pro's blended cost. Alongside it: a tiered consumer subscription restructure — new $100/mo Google AI Ultra tier (5× Pro limits, Antigravity 2.0 access) and the top Ultra cut from $250 → $200.Read →2026-05-18 · BITCPM4-CANN FAMILY (OPENBMB)BitCPM4-CANN — OpenBMB ships the first native ternary 8B LLM familyOpenBMB released the BitCPM4-CANN family (0.5B / 1B / 3B / 8B) in mid-May — the first publicly reported end-to-end 1.58-bit (ternary {-1, 0, 1}) training stack at 8B scale, trained natively on Huawei Ascend NPU. Apache 2.0. The 8B model retains 95.7% of full-precision MiniCPM4 performance at ~6× memory reduction; the 0.5B variant retains 90.1% of its full-precision baseline at ~100 MB on-disk. Not the strongest model at its size — but the smallest credible model at this quality level.Read →2026-05-15 · MINICPM-V-4.6 (OPENBMB)MiniCPM-V-4.6 — vision-language at 1B that prices like a sub-billion text modelOpenBMB shipped MiniCPM-V-4.6 on May 15 — a 1B-param vision-language model built on SigLIP2-400M + Qwen3.5-0.8B that scores higher than its own LLM backbone on Artificial Analysis Intelligence Index (13 vs 10) at ~19× lower token cost. Apache 2.0. Day-one GGUF, BNB, AWQ, GPTQ quants plus a Thinking variant. The newest entry in the V (vision-only) branch, parallel to the MiniCPM-o omnimodal line.Read →2026-05-08 · HIDREAM-O1-IMAGEHiDream-O1-Image — pixel-space generation finally lands as an open-weight, MIT-licensed modelHiDream open-sourced the O1 series on May 8: an 8B image foundation model that generates in pixel space — no VAE, no separate text encoder, one Pixel-level Unified Transformer handling text-to-image, edit, and subject-driven personalization at up to 2,048². Debuted top-10 on Artificial Analysis T2I Arena. Both undistilled and distilled Dev variants ship same-day under MIT.Read →2026-05-01 · MOSS-MUSIC-8B (OPENMOSS)MOSS-Music-8B — open-weight music understanding lands at usable accuracyOpenMOSS released MOSS-Music-8B (Instruct + Thinking) on May 1 — an Apache 2.0 audio-text-to-text model that does lyrics ASR with time-aligned transcription, music captioning, key/tempo/chord reasoning, structural analysis, instrument recognition, and music QA at production accuracy. 80.38% average on music-QA benchmarks; 4.36/5.0 on MusicCaps captioning. No open-weight model previously covered this category usably.Read →2026-04-29 · MISTRAL MEDIUM 3.5Mistral Medium 3.5 — one 128B dense model replaces three specialist MistralsMistral retired Magistral (reasoning) and Devstral 2 (coding) and folded both into Medium 3.5: a 128B dense weight set with 256K context, native multimodal vision, 77.6% on SWE-Bench Verified, and a per-request `reasoning_effort` toggle that swaps modes without swapping checkpoints. Modified MIT license — commercial OK below a revenue threshold.Read →2026-04-29 · IBM GRANITE 4.1IBM Granite 4.1 — Apache 2.0 dense at 3B/8B/30B with 512K contextIBM dropped the Granite 4.1 family on April 29 — three dense sizes (3B / 8B / 30B), Apache 2.0, 128K default context extending to 512K via training-stage continuation, native Ollama tags live the same day. The headline claim from IBM: the new 8B instruct matches their prior Granite 4.0 32B-A9B MoE. Believable but unverified by community benchmarks at time of writing.Read →2026-04-24 · DEEPSEEK V4 (PRO + FLASH)DeepSeek V4 — the architecture is the story, not the sizeA 1.6T Pro and a 284B Flash sibling, both MIT, both 1M context, released the same day. Skip the size headlines: the real news is the architectural change that drops V3.2 single-token FLOPs by ~73% and KV cache by ~90%.Read →2026-04-23 · OPENAI GPT-5.5GPT-5.5 — OpenAI doubled the GPT-5 line, local hardware just got cheaper by comparisonGPT-5.5 launched at $5 input / $30 output per 1M, exactly 2× the GPT-5.4 rates. OpenAI claims ~40% fewer output tokens on Codex tasks, which mostly offsets the increase to a ~20% net rise. Crucially, that ~20% comes out of cloud users' pockets; local-hardware buyers get a free improvement to their break-even math.Read →2026-04-22 · QWEN 3.6-27B (DENSE)Qwen 3.6-27B — a dense Q4 model that claims to beat the prior 397B MoE flagshipA week after Qwen 3.6-35B-A3B, Alibaba shipped the dense 27B. Apache 2.0, 262K native context, multimodal, ~17 GB at Q4. The community claim worth weighing: it beats the prior generation's 397B MoE on coding while staying single-card.Read →2026-04-20 · KIMI K2.6 (MOONSHOT)Kimi K2.6 — open-weights 1T MoE, but the headline is agent orchestration not raw size1T parameters total, 32B activated per token, Modified MIT (commercial OK below 100M MAU / $20M MRR). Tops SWE-Bench Pro at 58.6 (vs GPT-5.4 xhigh 57.7, Opus 4.6 max 53.4) and lands #4 on the Artificial Analysis Intelligence Index. The real differentiator is Agent Swarm scaling to 300 sub-agents over 4,000 coordinated steps — a different shape of capability than "bigger model, better single-step."Read →2026-04-16 · ANTHROPIC CLAUDE OPUS 4.7Claude Opus 4.7 — same per-token price, but the new tokenizer raises real cost ~35%Anthropic shipped Opus 4.7 at unchanged $5/$25 per 1M, paired with Claude Design (visual collaboration). Most coverage missed the tokenizer change — Opus 4.7 generates ~35% more tokens for the same prompts as 4.6. Effective cost rose without the price tag moving.Read →2026-04-16 · QWEN 3.6-35B-A3BQwen 3.6-35B-A3B — the MoE sibling that pairs with the dense 27BAlibaba shipped the 35B-A3B MoE first, then the dense 27B six days later. 3B active per token, 35B total, Apache 2.0, 262K context. ~17 GB at Q4 — fits 24 GB cards comfortably. Picks between the sibling pair come down to dense-vs-MoE tradeoffs.Read →2026-04-14 · VOXCPM2 (OPENBMB)VoxCPM2 — the first open-weight TTS that designs voices from text aloneOpenBMB shipped a 2B Apache-2.0 TTS that does what no other open-weight model does — generate a voice from a natural-language description, no reference audio required. Plus 30 languages, 48 kHz output, tokenizer-free diffusion AR.Read →2026-04-10 · MOSS-TTS-NANO (OPENMOSS)MOSS-TTS-Nano — multilingual voice cloning runs on 4 CPU coresA 100M Apache-2.0 model that fills the gap Kokoro-82M doesn't cover — multilingual TTS with voice cloning from a short reference audio, real-time on 4 CPU cores. The first time those three properties co-existed in an open-weight pick.Read →2026-04-07 · GLM-5.1 (Z.AI)GLM-5.1 — first open-weight model to lead SWE-Bench ProZ.ai's 744B MoE shipped MIT, hit 58.4 on SWE-Bench Pro, narrowly beats GPT-5.4 and Claude Opus 4.6 on that benchmark. Crucial caveat: at 466 GB Q4 it's hosted-only realistic. The 'open-weight' framing matters less than the SWE-Bench Pro leadership.Read →2026-04-02 · GOOGLE GEMMA 4Gemma 4 — the license change is a bigger story than the modelGoogle shipped Gemma 4 (31B dense + 26B MoE / 3.8B active) under Apache 2.0 — moving off the custom Gemma Terms that constrained commercial work in Gemma 1–3. For builders shipping commercial products, this matters more than the benchmark gains.Read →