the AI bench
VERIFIED AUGUST 2026
All hardware

HARDWARE · TOP TIER · 32 GB

NVIDIA RTX 5090

The single-GPU answer for people who refuse to wait.

A 32 GB Blackwell card that runs every modern coding, chat, and agent model at Q4 with headroom, at speeds a used dual-3090 rig can match only with a power bill and a compromise. The brief late-May AIB thaw that pulled the floor to ~$2,910 is long gone, and August moved the whole band up again: every RTX 5090 card in stock at Newegg on August 13 sits at $4,699–$4,899, with marketplace "more options" starting at $3,999. Treat ~$4,000 as the realistic floor now, not $3,500.

The decision in five lines

The call
Buy — The single-GPU answer for people who refuse to wait.
Best for
Top tier
Runs well
Qwen3-Coder-30B-A3B (MoE, fits 24GB) · Qwen3.8-27B · Z-Image-Turbo (Apache 2.0)
Watch out
Street price is now more than double the $1,999 MSRP, and the direction is still up. Floor moved $3,000 (April) → $3,799 (early May) → $2,910–$4,300 (May 28, ASUS TUF AIBs at $2,909.99) → $3,500–$4,300 (June) → $4,000–$4,900 (August 13: verified on Newegg — every in-stock card $4,699–$4,899, marketplace options from $3,999). The $2,909 card of late May is now a ~$4,700 card. DRAM/GDDR7 costs are the driver; IDC pegs the shortage through Q4 2027, so plan for this to get worse rather than better.
Evidence
Measured · last verified August 2026

32
GB GDDR7
1,792
GB/S BANDWIDTH
575
W TDP
~$4,000
FROM (STREET)

What fits at this tier

Unlocks the top band across every use case. MoE 30B-A3B and 35B-A3B at Q4 with room for 32K+ context; 32B dense Q4 fits with headroom; FLUX.2 dev fits without quant pain. 70B dense Q4 (~40 GB) does NOT fit 32 GB — needs 48 GB+ (dual-3090 or dual-5090).

CODING
Qwen3-Coder-30B-A3B (MoE, fits 24GB) Community daily driver for local coding; 3B-active MoE delivers 30B quality at 3B-dense speed.
CHAT / GENERAL
Qwen3.8-27B August 5 2026 dense refresh, Apache 2.0, ~17 GB at Q4. Native vision, and a Gated DeltaNet hybrid that keeps a growing KV cache on only 16 of 64 layers — roughly half the cache of the 3.6-27B it replaces.
DOCS & RETRIEVAL
Qwen3.8-27B 262K native context extensible to 1M, native vision, single-GPU at Q4. The linear-attention hybrid matters most here: three quarters of its layers hold a constant-size state, so long-doc work costs far less memory than a conventional dense 27B.
IMAGE
Z-Image-Turbo (Apache 2.0) Community daily driver for realism; 6B, 8-step inference, Apache 2.0 — commercial OK.
AGENTS
Muse Glimmer 30B (Apache 2.0, fits 24GB) Meta's Aug 10 2026 open release — a 30B dense model trained around the agent loop itself (plan, call, read, recover) rather than around chat, with a vision encoder so it can read a screenshot mid-task. Quantized under 20 GB to leave the KV cache and drafter inside a 24 GB card, and it ships its own speculative-decoding drafter. Apache 2.0 with no user-count clause. Benchmarks are vendor-run and roughly in line with its size class, not ahead of it — the reason it is here is the agent-loop training and the licence, not a claimed score.
VOICE
Qwen3-Omni-30B-A3B-Instruct Apache 2.0 MoE; audio+video+image+text in, speech+text out; 17GB at Q4. Frontier unified voice.

The call

Buy it if you want one card that runs every modern MoE at Q4 without thinking about quant or context.

Skip it if you can stomach used hardware — a pair of RTX 3090s gives you 48 GB for roughly $1,600 all-in, at the cost of noise, heat, and PCIe splitting.

Watchouts

  • Street price is now more than double the $1,999 MSRP, and the direction is still up. Floor moved $3,000 (April) → $3,799 (early May) → $2,910–$4,300 (May 28, ASUS TUF AIBs at $2,909.99) → $3,500–$4,300 (June) → $4,000–$4,900 (August 13: verified on Newegg — every in-stock card $4,699–$4,899, marketplace options from $3,999). The $2,909 card of late May is now a ~$4,700 card. DRAM/GDDR7 costs are the driver; IDC pegs the shortage through Q4 2027, so plan for this to get worse rather than better.
  • 575 W TDP needs a 1,000 W+ PSU and honest cooling. Check case airflow and PSU rail headroom before buying.
  • 12VHPWR connector — follow Nvidia's re-seat guidance. Melted-cable stories are rare but real at this power draw.
  • FE cards run hot under sustained load. AIB partner cards (ASUS, MSI, Gigabyte) with larger heatsinks are worth the $100–$200 premium if you'll keep it loaded for hours.

Local vs cloud at this tier

● LOCAL WINS

Unrestricted context, fully offline, no per-token cost for heavy daily use (>50M tok/mo), and you own the hardware for 3–5 years.

● CLOUD WINS

Frontier reasoning (Claude Opus 4.8, GPT-5.4), first-day access to new models, zero upfront cost, and no electricity or cooling to think about.

At this tier, the break-even vs a $100/mo ChatGPT Pro plan is ~24–30 months at regular usage. Heavy API users break even in under a year.

Next step

Load this setup into the planner