the AI bench
VERIFIED JULY 2026

CHANGES · DATED · JULY 2026

What changed, and our verdict.

Fast takes on major model drops, published within ~24 hours. Each piece dated, opinionated, and short — what shipped, what it actually changes, and where it fits in the planner.

The quarterly state-of-local-AI snapshot lands at /methodology. The dated diff for each quarter lands here. The RSS feed is at /changes/feed.xml.


Quarterly rollups

Fast takes show what shipped. The rollup shows what actually moved — what entered the picks, what left, what we changed our mind about, and what we got wrong.


Recent drops

2026-07-14 · INTERNSCIENCE AGENTS-A1-4BAgents-A1-4B — a genuinely tiny agent model, shipped with its own quantsWhile Moonshot was preparing a 2.8-trillion-parameter flagship, InternScience went the other way: Agents-A1-4B landed July 14 — a 4.5B dense Apache-2.0 model post-trained specifically for agentic work, with 262K context and official Q4/Q8/F16 GGUFs published by the lab itself. It needs about 3 GB at Q4, which means it runs on a Mac mini, an old 8 GB card, or a laptop with no dedicated GPU at all. InternScience reports it beating its own Qwen3.5-4B base by wide margins on agent benchmarks. No planner pick yet — it is four days old — but this is the most interesting small model of the window.Read →2026-07-16 · MOONSHOT KIMI K3 (2.8T-A50B)Kimi K3 — Moonshot swings for the largest open model ever, API firstMoonshot AI launched Kimi K3 on July 16 — a 2.8-trillion-parameter Mixture-of-Experts flagship (~50B active per token) with a 1M-token context window, native vision, and a new hybrid-attention architecture (Kimi Delta Attention plus Attention Residuals). It is API-first: live on Kimi.com, Kimi Code, and api.moonshot.ai at $3 input / $15 output per 1M, with open weights promised by July 27. If the weights ship, it would be the largest open-weight model ever released. It still changes no local pick — 2.8T is datacenter-class under any quantization — but this is the marquee drop of the window.Read →2026-07-09 · OPENAI GPT-5.6 (SOL / TERRA / LUNA)GPT-5.6 — OpenAI renames the ladder: Sol, Terra, and Luna land at familiar pricesOpenAI shipped the GPT-5.6 family to general availability on July 9 — three tiers with durable names: Sol (the flagship, $5/$30 per 1M), Terra (balanced, $2.50/$15), and Luna (cost-efficient, $1/$6). Sol and Terra land at exactly the price points GPT-5.5 and GPT-5.4 occupied; Luna undercuts the legacy GPT-5 base row. The pitch is efficiency, not a price move: more successful work per dollar, fewer output tokens, less wall-clock. Hosted-only, so it changes no local pick — but it resets the cloud baseline our cost calculator compares against, so we updated all three OpenAI rows.Read →2026-07-08 · XAI GROK 4.5Grok 4.5 — xAI moves upmarket but stays the cheap frontier optionxAI shipped Grok 4.5 on July 8 — its new flagship for coding, agentic tasks, and knowledge work, priced at $2 per 1M input / $6 per 1M output with a 500K context window. That is a real price step up from Grok 4.3 ($1.25/$2.50), but xAI's pitch is that roughly 2× token efficiency — solving tasks in under half the steps — makes the effective cost competitive anyway. API-only, no open weights (xAI has open-sourced nothing since Grok-1 in 2024), so no local pick moves; the cost calculator row moves from 4.3 to 4.5.Read →2026-07-02 · TENCENT HY3 (295B-A21B)Hy3 — Tencent drops the restrictions and ships a real Apache 2.0 flagshipTencent released Hy3 on July 2 — a 295B-total / 21B-active Mixture-of-Experts flagship under a clean Apache 2.0 license, with 256K context, BF16 and FP8 checkpoints, and day-one vLLM/SGLang support. The license is the headline: April's Hy3-Preview geo-blocked the EU/UK/South Korea and carried a 100M-MAU commercial cap; the full release drops all of it. At ~300 GB even in FP8 it is not a consumer-GPU model — but as open flagships go, this is one of the cleanest licensing stories of the year.Read →2026-07-02 · POOLSIDE LAGUNA XS 2.1 (33B-A3B)Laguna XS 2.1 — Poolside opens its weights, and it actually fits your GPUPoolside — the coding-model lab that trains from scratch rather than fine-tuning someone else's base — published open weights for the first time on July 2: Laguna XS 2.1, a 33B-total / 3B-active MoE trained on 30 trillion tokens for agentic coding, under the permissive OpenMDW-1.1 license. Unlike almost every frontier-lab drop this year, this one is genuinely consumer-runnable: official BF16/FP8/INT4/NVFP4 checkpoints, an official GGUF repo, and DFlash speculator models that Poolside claims double local decode speed. At INT4 it fits a 24 GB card. It lands as a fast take plus a model entry — not yet a planner pick.Read →2026-07-01 · CLAUDE FABLE 5 (RESTORED)Fable 5 is back — and the outage is the lessonThe US government lifted the export controls on Claude Fable 5 and Mythos 5 on June 30, and Anthropic restored Fable 5 globally on July 1 — Claude Platform API, Claude.ai, Claude Code, and Cowork — at its original $10/$50 per 1M pricing with the full 1M context window. That closes the loop we opened on June 9: the most capable generally available model on the market spent nineteen days switched off by a government directive, and no amount of subscription money could turn it back on. It now enters our cost calculator for the first time, because "available" finally means available.Read →2026-06-30 · CLAUDE SONNET 5Claude Sonnet 5 — the new default Sonnet, priced to undercut itself for two monthsAnthropic shipped Claude Sonnet 5 on June 30 — the successor to Sonnet 4.6, GA across the Claude API (`claude-sonnet-5`), Bedrock, Vertex, and Foundry, and the new default on claude.ai Free/Pro. It keeps the 1M-token context window and is introductory-priced at $2 / $10 per 1M (input / output) through August 31 2026, reverting to the same $3 / $15 Sonnet 4.6 carried on September 1. Like every hosted frontier model, it changes no local planner pick — but it does reset the cloud-cost baseline our calculator compares against, so we updated that.Read →2026-06-30 · MEITUAN LONGCAT-2.0LongCat-2.0 — Meituan open-sources a trillion-parameter agentic coder under MITMeituan released LongCat-2.0 on June 30 — a large-scale Mixture-of-Experts model, ~1.6 trillion total parameters with ~48B activated per token, 1M context, MIT-licensed and aimed squarely at agentic coding. It was the stealth "owl-alpha" model on OpenRouter, and it is live now via OpenRouter and the longcat.chat platform (billing enabled). Like every trillion-scale model, it changes no local planner pick — but MIT at this size is notable, so it lands as a frontier comparator the way GLM-5.2, Kimi K2.7, and DeepSeek V4-Pro do.Read →2026-06-26 · QWEN3-ASR (1.7B / 0.6B)Qwen3-ASR — a tiny Apache speech-recognition family that actually fits anywhereQwen shipped Qwen3-ASR on June 26 — a dedicated open-weight speech-recognition family (1.7B and 0.6B), Apache 2.0, built on the Qwen3-Omni audio stack. It does language identification plus ASR across 52 languages and dialects (30 languages + 22 Chinese dialects), and Qwen claims the 1.7B is state-of-the-art among open-source ASR and competitive with the strongest proprietary commercial APIs. Unlike most of what lands in this feed, this one is genuinely local for everyone: ~4 GB at fp16 for the 1.7B, ~1.5–2 GB for the 0.6B.Read →2026-06-23 · KREA 2 (RAW + TURBO)Krea 2 — the independent-lab image model with a license asteriskKrea AI made Krea 2 open-weights public on June 23 — Raw (a CFG-guided base for fine-tuning and LoRA work) and Turbo (an 8-step distilled variant that renders up to ~2K images in a couple of seconds). It is a DiT with a Qwen3-VL text encoder, it runs locally via diffusers with community FP8/GGUF quants already out, and Krea claims the #1 text-to-image slot from an independent lab on Artificial Analysis. The catch is the license: a custom "krea-2-community-license," not Apache/MIT — which is why it lands as a fast take, not a planner pick.Read →2026-06-22 · QWEN-AGENTWORLD-35B-A3BQwen-AgentWorld — an open, consumer-runnable model for simulating agent environmentsQwen released Qwen-AgentWorld-35B-A3B on June 22 — Apache 2.0, built on Qwen3.5-35B-A3B-Base (35B total / 3B active MoE), and unusual: it is a "language world model" that predicts the next environment state given an agent's action, across seven interaction domains (MCP/tool-calling, Search, Terminal, SWE, Android, Web, OS). It is consumer-runnable — 3B active, GGUF and MLX quants already exist — and it picked up real traction fast (~28K downloads in days). But it is a simulator for building and testing agents, not a general chat or coding model, so it is a fast take, not a planner pick.Read →2026-06-16 · Z.AI GLM-5.2GLM-5.2 — Z.AI's flagship gets a solid 1M context and an architecture reworkZ.AI shipped GLM-5.2 on June 16 — the successor to GLM-5.1, still MIT, still ~753B total in the GLM-5 sparse-MoE line. The headline is a stable 1M-token context (a first for the family) plus an architecture rework (IndexShare) that cuts per-token FLOPs ~2.9× at 1M, and on Z.AI's own card it beats GLM-5.1 across reasoning, coding, and long-horizon agent suites. Like 5.1, the size keeps it hosted / big-iron, so it changes no local planner pick — it's a frontier comparator the way DeepSeek V4-Pro and Kimi K2.6 are.Read →2026-06-12 · MOONSHOT KIMI-K2.7-CODEKimi-K2.7-Code — Moonshot ships a coding-tuned trillion-parameter MoEMoonshot released Kimi-K2.7-Code around June 12 — a coding-focused agentic model built on Kimi K2.6, same 1T-total / 32B-active MoE architecture (384 experts, 8+1 active, DeepSeek-V3-style MLA backbone), 256K context, under a Modified-MIT license. It targets real-world long-horizon coding and tool use. Like K2.6, the size puts it firmly in hosted / multi-GPU territory, so it changes no local planner pick — it is a cloud comparator the way DeepSeek V4-Pro and Command A+ are.Read →2026-06-11 · MINIMAX M3MiniMax M3 — the weights are out, the license is the catchMiniMax's flagship native-multimodal MoE, M3, shipped its open weights to Hugging Face on ~June 7–11 (after the June 1 API launch) — resolving the item we deferred on June 5. It is ~428B total / ~23B active, 1M context, text + image + video in. Two things keep it off the local picks: it is big-iron-class at 428B, and it ships under the custom "MiniMax Community License," not an OSI-permissive one. So it lands as a fast-take and a cloud-side comparator, not a planner recommendation.Read →2026-06-09 · ANTHROPIC CLAUDE FABLE 5Claude Fable 5 — Anthropic's new ceiling, pulled three days after launchAnthropic launched Claude Fable 5 (alongside the more capable, restricted Mythos 5) on June 9 — a "Mythos-class" model it calls the most capable it has ever made generally available, state-of-the-art on nearly every tested benchmark, priced at $10 / M input and $50 / M output (double Opus 4.8). On June 12, Anthropic suspended all access to both Fable 5 and Mythos 5 under a U.S. government export-control directive; as of this writing it's still down. It's hosted-only frontier, so it changes no local pick — and we are not adding it to the cost calculator while it's unavailable.Read →2026-06-06 · NVIDIA NEMOTRON 3.5 ASR STREAMING 0.6BNemotron 3.5 ASR — real-time streaming speech-to-text, 40 locales, in 600M paramsNVIDIA shipped Nemotron 3.5 ASR Streaming 0.6B on June 5–6 — a 600M-parameter cache-aware streaming speech-to-text model covering 40 language-locales from a single checkpoint, with configurable latency (80 ms to 1,120 ms chunks) and built-in punctuation + capitalization. It is small, fast, and runs comfortably on modest hardware. The one catch for us: it ships under the OpenMDW 1.1 license, which isn't OSI-permissive — so it earns a watch, not a planner-pick slot.Read →2026-06-05 · COHERE NORTH MINI CODENorth Mini Code — Cohere ships an Apache-2.0 30B-A3B coder for agents and the terminalCohere and Cohere Labs released North Mini Code on June 5 — a 30B-total / 3B-active sparse-MoE model tuned for code generation, agentic software engineering, and terminal tasks, under a clean Apache 2.0 license with a 256K context. At ~17 GB in a Q4 GGUF it runs on a single 24 GB GPU, which makes it a genuine new alternative to Qwen3-Coder-30B-A3B at that tier — so we're adding it to the planner's local coding picks.Read →2026-06-04 · NVIDIA NEMOTRON-3 ULTRA 550B-A55BNemotron-3 Ultra — NVIDIA ships the best US open-weight model, fast, and fully openNVIDIA released Nemotron-3 Ultra 550B-A55B on June 4 (announced at Jensen Huang's Computex keynote on June 1) — a 550B-total / 55B-active hybrid Mamba-Transformer MoE under the new OpenMDW 1.1 license, with 1M context and native NVFP4. Artificial Analysis scores it 48 on its Intelligence Index: the strongest US/Western open-weight model to date — ahead of Gemma 4 31B (39) and gpt-oss-120b (33), but behind the Chinese-led frontier (Kimi K2.6 at 54). The headline is speed-for-intelligence: 300+ tok/s, several times faster than DeepSeek/Kimi peers.Read →2026-06-03 · GOOGLE GEMMA 4 12BGemma 4 12B — frontier-ish multimodal that actually runs on a 16 GB laptop, under Apache 2.0Google shipped Gemma 4 12B on June 3 — a ~12B dense, encoder-free unified multimodal model (text + image + audio + video in, text out) under Apache 2.0. Google's pitch is the deployment envelope: it runs locally on 16 GB of VRAM or unified memory while landing benchmark numbers Google says approach its 26B-A4B MoE, at under half the total memory footprint. It slots into the existing Gemma 4 family (31B dense + 26B MoE, April) as the laptop-friendly multimodal pick.Read →2026-06-02 · MICROSOFT MAI FAMILY (BUILD 2026)Microsoft MAI at Build 2026 — seven first-party models, all hosted-only, none you can downloadAt Build 2026 on June 2, Mustafa Suleyman unveiled Microsoft AI's MAI family — seven first-party models spanning reasoning (MAI-Thinking-1), coding (MAI-Code-1, MAI-Code-1-Flash), image (MAI-Image-2.5 + Flash), transcription (MAI-Transcribe-1.5), and voice (MAI-Voice-2). Trained from scratch, no distillation. The story is Microsoft reducing its reliance on OpenAI — but for a local-AI site the headline is the asterisk: every MAI model is Microsoft Foundry / Azure-hosted, closed-weight, API-only. None are on Hugging Face, none are downloadable, none change a single local pick.Read →2026-06-02 · XAI GROK BUILD (GROK-BUILD-0.1)Grok Build — xAI ships a terminal coding agent on the hosted grok-build-0.1 modelxAI opened Grok Build to public beta via the xAI API in early June — a Rust, terminal-native coding agent and CLI (xAI's answer to Claude Code, Codex, and Gemini CLI), powered by the new grok-build-0.1 model. It is a fast agentic coding model: 256K context, text + image in, function calling + structured outputs + reasoning, at $1 / 1M input and $2 / 1M output ($0.20 cached). The catch for a local-AI site: grok-build-0.1 is hosted / API-only — no open weights — so it shifts the cloud landscape, not your local picks.Read →2026-05-28 · ANTHROPIC CLAUDE OPUS 4.8Claude Opus 4.8 — price held flat, fast mode got 3× cheaper, and it catches its own code bugs nowAnthropic shipped Opus 4.8 on May 28 at unchanged standard pricing ($5 in / $25 out per 1M) — no tokenizer surprise this time, it keeps the 4.7 tokenizer. Two things actually changed: fast mode dropped to $10/$50 (3× cheaper than 4.6/4.7), and Anthropic claims the model is ~4× less likely to let flaws in code it wrote slip past unremarked. Still hosted-only; nothing changes for local picks today.Read →2026-05-25 · GRANITE-SWITCH 4.1 8B PREVIEW (IBM)Granite-Switch 4.1 — IBM's 12-adapter-in-one-checkpoint pattern is the deployment storyIBM uploaded the Granite-Switch 4.1 family (3B / 8B / 30B previews) to Hugging Face on May 25. Each checkpoint is the base Granite 4.1 dense model with 12 task-specialized LoRA adapters embedded, activated per-token via control tokens in the chat template. Three libraries: Core (requirement check, context attribution, uncertainty), RAG (query rewrite, query clarification, answerability, hallucination detection, citation generation), Guardian (safety detection, factuality detection + correction, policy guardrails). Apache 2.0, 128K context, 12 languages.Read →2026-05-22 · MINICPM5-1B (OPENBMB)MiniCPM5-1B — OpenBMB's On-Policy Distillation pipeline lands at 1BOpenBMB uploaded MiniCPM5-1B and MiniCPM5-1B-SFT to Hugging Face on May 22 — a 1.08B-parameter dense Llama-class model trained with a three-stage SFT → RL → On-Policy Distillation pipeline. Apache 2.0, 128K context, English + Chinese, hybrid `<think>` reasoning toggle, native XML-style tool calling. Claims 1B-class open-source SOTA against LFM2.5-1.2B-Thinking, Qwen3-0.6B/think, and Qwen3.5-0.8B/think.Read →2026-05-20 · COMMAND A+ (COHERE)Command A+ — Cohere ships the Apache 2.0 frontier MoE the open-weight stack was missingCohere released Command A+ on May 20 — a 218B-param / 25B-active sparse MoE under Apache 2.0, with hybrid sliding-window + global attention, 128K input / 64K generation, native multimodal (text + image), and 48-language coverage. Available in BF16, FP8, and W4A4 quants on day one. The W4A4 build fits 2× H100 80 GB; the FP8 fits 1× MI300X. The first time you can self-host a 218B-class frontier MoE without a hosted-only or research-license caveat.Read →2026-05-19 · GEMINI 3.5 FLASH (GOOGLE)Gemini 3.5 Flash — Google's new daily-driver API price-shifts the cloud-vs-local mathGoogle launched Gemini 3.5 Flash GA at I/O 2026 (May 19) at $1.50/$9.00 per 1M tokens with 1M context. Reported to beat Gemini 3.1 Pro on Terminal-Bench 2.1 / MCP Atlas / CharXiv at roughly 60% of Pro's blended cost. Alongside it: a tiered consumer subscription restructure — new $100/mo Google AI Ultra tier (5× Pro limits, Antigravity 2.0 access) and the top Ultra cut from $250 → $200.Read →2026-05-18 · BITCPM4-CANN FAMILY (OPENBMB)BitCPM4-CANN — OpenBMB ships the first native ternary 8B LLM familyOpenBMB released the BitCPM4-CANN family (0.5B / 1B / 3B / 8B) in mid-May — the first publicly reported end-to-end 1.58-bit (ternary {-1, 0, 1}) training stack at 8B scale, trained natively on Huawei Ascend NPU. Apache 2.0. The 8B model retains 95.7% of full-precision MiniCPM4 performance at ~6× memory reduction; the 0.5B variant retains 90.1% of its full-precision baseline at ~100 MB on-disk. Not the strongest model at its size — but the smallest credible model at this quality level.Read →2026-05-15 · MINICPM-V-4.6 (OPENBMB)MiniCPM-V-4.6 — vision-language at 1B that prices like a sub-billion text modelOpenBMB shipped MiniCPM-V-4.6 on May 15 — a 1B-param vision-language model built on SigLIP2-400M + Qwen3.5-0.8B that scores higher than its own LLM backbone on Artificial Analysis Intelligence Index (13 vs 10) at ~19× lower token cost. Apache 2.0. Day-one GGUF, BNB, AWQ, GPTQ quants plus a Thinking variant. The newest entry in the V (vision-only) branch, parallel to the MiniCPM-o omnimodal line.Read →2026-05-08 · HIDREAM-O1-IMAGEHiDream-O1-Image — pixel-space generation finally lands as an open-weight, MIT-licensed modelHiDream open-sourced the O1 series on May 8: an 8B image foundation model that generates in pixel space — no VAE, no separate text encoder, one Pixel-level Unified Transformer handling text-to-image, edit, and subject-driven personalization at up to 2,048². Debuted top-10 on Artificial Analysis T2I Arena. Both undistilled and distilled Dev variants ship same-day under MIT.Read →2026-05-01 · MOSS-MUSIC-8B (OPENMOSS)MOSS-Music-8B — open-weight music understanding lands at usable accuracyOpenMOSS released MOSS-Music-8B (Instruct + Thinking) on May 1 — an Apache 2.0 audio-text-to-text model that does lyrics ASR with time-aligned transcription, music captioning, key/tempo/chord reasoning, structural analysis, instrument recognition, and music QA at production accuracy. 80.38% average on music-QA benchmarks; 4.36/5.0 on MusicCaps captioning. No open-weight model previously covered this category usably.Read →2026-04-29 · MISTRAL MEDIUM 3.5Mistral Medium 3.5 — one 128B dense model replaces three specialist MistralsMistral retired Magistral (reasoning) and Devstral 2 (coding) and folded both into Medium 3.5: a 128B dense weight set with 256K context, native multimodal vision, 77.6% on SWE-Bench Verified, and a per-request `reasoning_effort` toggle that swaps modes without swapping checkpoints. Modified MIT license — commercial OK below a revenue threshold.Read →2026-04-29 · IBM GRANITE 4.1IBM Granite 4.1 — Apache 2.0 dense at 3B/8B/30B with 512K contextIBM dropped the Granite 4.1 family on April 29 — three dense sizes (3B / 8B / 30B), Apache 2.0, 128K default context extending to 512K via training-stage continuation, native Ollama tags live the same day. The headline claim from IBM: the new 8B instruct matches their prior Granite 4.0 32B-A9B MoE. Believable but unverified by community benchmarks at time of writing.Read →2026-04-24 · DEEPSEEK V4 (PRO + FLASH)DeepSeek V4 — the architecture is the story, not the sizeA 1.6T Pro and a 284B Flash sibling, both MIT, both 1M context, released the same day. Skip the size headlines: the real news is the architectural change that drops V3.2 single-token FLOPs by ~73% and KV cache by ~90%.Read →2026-04-23 · OPENAI GPT-5.5GPT-5.5 — OpenAI doubled the GPT-5 line, local hardware just got cheaper by comparisonGPT-5.5 launched at $5 input / $30 output per 1M, exactly 2× the GPT-5.4 rates. OpenAI claims ~40% fewer output tokens on Codex tasks, which mostly offsets the increase to a ~20% net rise. Crucially, that ~20% comes out of cloud users' pockets; local-hardware buyers get a free improvement to their break-even math.Read →2026-04-22 · QWEN 3.6-27B (DENSE)Qwen 3.6-27B — a dense Q4 model that claims to beat the prior 397B MoE flagshipA week after Qwen 3.6-35B-A3B, Alibaba shipped the dense 27B. Apache 2.0, 262K native context, multimodal, ~17 GB at Q4. The community claim worth weighing: it beats the prior generation's 397B MoE on coding while staying single-card.Read →2026-04-20 · KIMI K2.6 (MOONSHOT)Kimi K2.6 — open-weights 1T MoE, but the headline is agent orchestration not raw size1T parameters total, 32B activated per token, Modified MIT (commercial OK below 100M MAU / $20M MRR). Tops SWE-Bench Pro at 58.6 (vs GPT-5.4 xhigh 57.7, Opus 4.6 max 53.4) and lands #4 on the Artificial Analysis Intelligence Index. The real differentiator is Agent Swarm scaling to 300 sub-agents over 4,000 coordinated steps — a different shape of capability than "bigger model, better single-step."Read →2026-04-16 · ANTHROPIC CLAUDE OPUS 4.7Claude Opus 4.7 — same per-token price, but the new tokenizer raises real cost ~35%Anthropic shipped Opus 4.7 at unchanged $5/$25 per 1M, paired with Claude Design (visual collaboration). Most coverage missed the tokenizer change — Opus 4.7 generates ~35% more tokens for the same prompts as 4.6. Effective cost rose without the price tag moving.Read →2026-04-16 · QWEN 3.6-35B-A3BQwen 3.6-35B-A3B — the MoE sibling that pairs with the dense 27BAlibaba shipped the 35B-A3B MoE first, then the dense 27B six days later. 3B active per token, 35B total, Apache 2.0, 262K context. ~17 GB at Q4 — fits 24 GB cards comfortably. Picks between the sibling pair come down to dense-vs-MoE tradeoffs.Read →2026-04-14 · VOXCPM2 (OPENBMB)VoxCPM2 — the first open-weight TTS that designs voices from text aloneOpenBMB shipped a 2B Apache-2.0 TTS that does what no other open-weight model does — generate a voice from a natural-language description, no reference audio required. Plus 30 languages, 48 kHz output, tokenizer-free diffusion AR.Read →2026-04-10 · MOSS-TTS-NANO (OPENMOSS)MOSS-TTS-Nano — multilingual voice cloning runs on 4 CPU coresA 100M Apache-2.0 model that fills the gap Kokoro-82M doesn't cover — multilingual TTS with voice cloning from a short reference audio, real-time on 4 CPU cores. The first time those three properties co-existed in an open-weight pick.Read →2026-04-07 · GLM-5.1 (Z.AI)GLM-5.1 — first open-weight model to lead SWE-Bench ProZ.ai's 744B MoE shipped MIT, hit 58.4 on SWE-Bench Pro, narrowly beats GPT-5.4 and Claude Opus 4.6 on that benchmark. Crucial caveat: at 466 GB Q4 it's hosted-only realistic. The 'open-weight' framing matters less than the SWE-Bench Pro leadership.Read →2026-04-02 · GOOGLE GEMMA 4Gemma 4 — the license change is a bigger story than the modelGoogle shipped Gemma 4 (31B dense + 26B MoE / 3.8B active) under Apache 2.0 — moving off the custom Gemma Terms that constrained commercial work in Gemma 1–3. For builders shipping commercial products, this matters more than the benchmark gains.Read →