the AI bench
VERIFIED SEPTEMBER 2026
All hardware

HARDWARE · SILENT WORKSTATION · 96 GB UNIFIED

Mac Studio M3 Ultra 96 GB

The 70B-dense workstation that runs silently on 70 watts.

819 GB/s unified memory bandwidth — the highest in any shipping Mac — plus 96 GB capacity puts Llama 3.3 70B Q4 at 12–18 tok/s on a box that runs at ~70 W idle and fits on a bookshelf. Dual M3 Max dies under one heatsink, no GPU tower, no fan noise.

The decision in five lines

The call
Buy — The 70B-dense workstation that runs silently on 70 watts.
Best for
Silent workstation
Runs well
Qwen3-Coder-30B-A3B (MoE, fits 24GB) · Qwen3.8-27B · Z-Image-Turbo (Apache 2.0)
Watch out
Update — September 7, 2026: the M3 Ultra Mac Studio is discontinued. Apple announced M5 Ultra on August 25, 2026 (pre-order from that day, arriving from September 22); the 96 GB / 1 TB base configuration is $5,499 on apple.com — $200 more than the M3 Ultra price on this page. Apple rates the top 36-core CPU / 80-core GPU M5 Ultra at 1.2 TB/s; the base 30-core / 64-core bin that carries the $5,499 price has no bandwidth figure on the spec page yet, so we do not quote one. The 70B Q4 fit and idle-power story on this page carry over; the tok/s numbers are M3 Ultra measurements.
Evidence
Estimated · last verified September 2026

96
GB UNIFIED
819
GB/S BANDWIDTH
28/60
CPU / GPU CORES
$5,499
M5 ULTRA 96GB/1TB (PRE-ORDER)

What fits at this tier

96 GB unified × 0.67 default = ~64 GB effective LLM budget. Llama 3.3 70B Q4 (~40 GB) runs at 12–18 tok/s, 35B-A3B MoE at 90–110 tok/s, Qwen 3.5 27B 6-bit MLX at 28–35 tok/s. 122B-A10B MoE fits at 4-bit MLX with sysctl tweak at ~55–70 tok/s. This is the Mac-native 70B-comfortable tier.

CODING
Qwen3-Coder-30B-A3B (MoE, fits 24GB) Community daily driver for local coding; 3B-active MoE delivers 30B quality at 3B-dense speed.
CHAT / GENERAL
Qwen3.8-27B August 5 2026 dense refresh, Apache 2.0, ~17 GB at Q4. Native vision, and a Gated DeltaNet hybrid that keeps a growing KV cache on only 16 of 64 layers — roughly half the cache of the 3.6-27B it replaces.
DOCS & RETRIEVAL
Qwen3.8-27B 262K native context extensible to 1M, native vision, single-GPU at Q4. The linear-attention hybrid matters most here: three quarters of its layers hold a constant-size state, so long-doc work costs far less memory than a conventional dense 27B.
IMAGE
Z-Image-Turbo (Apache 2.0) Community daily driver for realism; 6B, 8-step inference, Apache 2.0 — commercial OK.
AGENTS
Muse Glimmer 30B (Apache 2.0, fits 24GB) Meta's Aug 10 2026 open release — a 30B dense model trained around the agent loop itself (plan, call, read, recover) rather than around chat, with a vision encoder so it can read a screenshot mid-task. Quantized under 20 GB to leave the KV cache and drafter inside a 24 GB card, and it ships its own speculative-decoding drafter. Apache 2.0 with no user-count clause. Benchmarks are vendor-run and roughly in line with its size class, not ahead of it — the reason it is here is the agent-loop training and the licence, not a claimed score.
VOICE
Qwen3-Omni-30B-A3B-Instruct Apache 2.0 MoE; audio+video+image+text in, speech+text out; 17GB at Q4. Frontier unified voice.

The call

Buy it if you want a desktop local-AI workstation that stays silent at full inference load and you can live with Mac-native software (Ollama, LM Studio, MLX). The 819 GB/s bandwidth is what justifies the premium over a 48 GB M5 Pro MBP.

Skip it if you need the 128 GB or 256 GB upgrade tier — Apple pulled those Mac Studio configs April 11 2026, and they're either "currently unavailable" or 4–5 months out. The 96 GB base Ultra is the realistically-buyable top config.

Watchouts

  • Update — September 7, 2026: the M3 Ultra Mac Studio is discontinued. Apple announced M5 Ultra on August 25, 2026 (pre-order from that day, arriving from September 22); the 96 GB / 1 TB base configuration is $5,499 on apple.com — $200 more than the M3 Ultra price on this page. Apple rates the top 36-core CPU / 80-core GPU M5 Ultra at 1.2 TB/s; the base 30-core / 64-core bin that carries the $5,499 price has no bandwidth figure on the spec page yet, so we do not quote one. The 70B Q4 fit and idle-power story on this page carry over; the tok/s numbers are M3 Ultra measurements.
  • 128 GB and 256 GB Mac Studio RAM upgrades were pulled April 11 2026 per MacRumors. The 96 GB base Ultra is the only currently-shipping Ultra Studio — and as of mid-2026 it carries a ~13–14 week lead time (a July order lands around October), with resellers (B&H, Best Buy) now showing it discontinued/sold out ahead of the expected fall M5 Studio.
  • Prefill tax is real on long contexts. 70B + 32K prompt can cost 30–60 seconds before the first token — Mac unified memory's honest weakness vs NVIDIA.
  • `sudo sysctl iogpu.wired_limit_mb=80000` is recommended for 70B Q4 work. Without it, macOS's 33% reservation clips the addressable memory below the model's needs.
  • M5 Ultra refresh now expected October 2026 (Bloomberg/Gurman, April 19) — Mac Studio slipped from the rumored June WWDC window due to supply-chain delays. Apple also discontinued the 512 GB option in March 2026, and the 128 GB / 256 GB Ultra configs were pulled April 11. The 96 GB tier remains the buyable Ultra ceiling.

Local vs cloud at this tier

● LOCAL WINS

70B dense at interactive speed, silent, at 70 W idle. MoE 35B-A3B + 122B-A10B unbounded. The quietest 70B-capable rig in the lineup.

● CLOUD WINS

Frontier reasoning (Opus 4.8, GPT-5.4) and first-day model access. Cloud also wins on electricity — this box runs at 70 W idle but pulls 180+ W under sustained load, and the total-cost-of-ownership over 3 years tilts cloud-ward for light users.

The right 70B-dense Mac for anyone who values silence + desk footprint over raw NVIDIA throughput — but Apple's June 25 2026 hike took this from $3,999 to $5,299 (+$1,300), pushing break-even vs a $100/mo ChatGPT Pro plan from ~36 to ~56 months. The buy case now leans much harder on privacy + unlimited use than on cost math; if you don't need 96 GB unified specifically, a 5090 or dual-3090 build is far better value.

Next step

Load this setup into the planner