the AI bench
VERIFIED SEPTEMBER 2026
All hardware

HARDWARE · TEAM RED · 24 GB

AMD Radeon RX 7900 XTX

The AMD answer, with an honest asterisk on driver friction.

24 GB at roughly 85–90% of a 4090's throughput under ROCm. The hardware is fine; the software ecosystem is the tax. Plan 5–10 hours on first-time ROCm setup, plus the ongoing friction of Ollama being patchy on AMD. Pricing, verified on Newegg on September 15, 2026 (in stock, sorted by lowest price): new cards start at $1,299.00 (Sapphire Pulse), with most between $1,327 and $1,499. The ~$850 refurbished units we listed in August were no longer in stock.

The decision in five lines

The call
Buy — The AMD answer, with an honest asterisk on driver friction.
Best for
Team red
Runs well
Qwen3-Coder-30B-A3B (MoE, fits 24GB) · Qwen 3.5 35B-A3B (MoE, fits 24GB) · Gemma 4 31B (256K context)
Watch out
AMD post-EOL'd the 7900 XTX, so it is a used-first buy. New cards in stock on Newegg start at $1,299 (September 15, 2026); refurbished stock had gone. Used prices are not verified (no sold-listing data). For fresh inventory, RDNA4 (RX 9070 XT, new from $759.99 on the same date) is the cheaper AMD step.
Evidence
Estimated · last verified September 2026

24
GB GDDR6
960
GB/S BANDWIDTH
355
W TDP
~$1,300
NEW (IN STOCK, SEP 15)

What fits at this tier

Runs 24 GB-tier models at Q4 cleanly — MoE 30B-A3B, 14B dense, gpt-oss-20b — via vLLM or llama.cpp (HIP). Some MoE picks had HIP kernel issues through 2025; check llama.cpp release notes before pulling the latest.

CODING
Qwen3-Coder-30B-A3B (MoE, fits 24GB) 3B-active MoE — benchmark champion for local coding at this tier.
CHAT / GENERAL
Qwen 3.5 35B-A3B (MoE, fits 24GB) 3B active MoE — 30B quality at 3B inference speed.
DOCS & RETRIEVAL
Gemma 4 31B (256K context) 31B dense with 256K context; Gemma commercial-permissive terms; Arena top 5.
IMAGE
HiDream-O1-Image (8B, MIT) May 8, 2026 release. Pixel-space (no VAE, no disjoint text encoder) — debuted top-10 on Artificial Analysis T2I Arena. MIT-licensed 8B; one model handles T2I + edit + subject-driven personalization at up to 2,048².
AGENTS
Qwen 3.5 35B-A3B (MoE, fits 24GB) MoE with native tool use; fits 24GB at Q4; Apache 2.0.
VOICE
VoxCPM2 (2B, Apache 2.0) 30 languages, 48 kHz, tokenizer-free diffusion AR; voice design from text. April 2026 release.

The call

Buy it if you already run Linux, want 24 GB without paying NVIDIA's supply-crunch tax, and treat driver tinkering as acceptable friction rather than pain.

Skip it if your time is worth more than ~$50/hr — you will spend a day on ROCm setup and another day later the first time a new model doesn't JIT cleanly. Also skip if scarcity premiums make the RX 9070 XT (RDNA4, 16 GB at $750–$870) the cleaner AMD entry — the 9070 XT trades 8 GB of headroom for fresh architecture support and ROCm 7+ first-class WMMA.

Watchouts

  • AMD post-EOL'd the 7900 XTX, so it is a used-first buy. New cards in stock on Newegg start at $1,299 (September 15, 2026); refurbished stock had gone. Used prices are not verified (no sold-listing data). For fresh inventory, RDNA4 (RX 9070 XT, new from $759.99 on the same date) is the cheaper AMD step.
  • Budget 5–10 hours for first-time ROCm setup. Official AMD packages are the reliable path; distro repos lag by 6–12 months.
  • Ollama on AMD is still patchy. Use vLLM or llama.cpp with HIP for the reliable runner story.
  • MoE on ROCm is no longer the crash story it was: llama.cpp #19880 (Feb 28), #20024 — the Qwen 3.5 35B-A3B memory fault, fixed on both ROCm and Vulkan (Mar 13) — and #20545 (Jun 2, pinned to the ROCm stack, not llama.cpp) are all closed, and ROCm is now at 7.14 vs the 7.2 those were filed against. Vulkan still usually wins MoE token generation on this card; ROCm tends to win prompt processing. Pick by workload, not to dodge a bug.

Local vs cloud at this tier

● LOCAL WINS

24 GB for less than a 4090, fully offline, no per-token cost. Hardware longevity on AMD is better than NVIDIA historically — less planned obsolescence.

● CLOUD WINS

No driver maintenance, no kernel compatibility issues, immediate model access the day they're released. For anyone whose time is billable, cloud looks compelling compared to the AMD software story.

At ~$1,300 new with regular usage, break-even vs ChatGPT Plus is over five years (this line used to say ~30–36 months at an $850 refurb price that is no longer in stock). The real question is whether your hourly rate makes ROCm debugging a net positive.

Next step

Load this setup into the planner