the AI bench
VERIFIED AUGUST 2026
All fast takes

FAST TAKE · 2026-07-30 · OPENAI GPT-5.6 LUNA / TERRA PRICE CUT

OpenAI cuts Luna 80% and Terra 20% — recompute your break-even

On July 30 OpenAI cut API pricing on two of the three GPT-5.6 tiers: Luna dropped from $1/$6 per million to $0.20/$1.20, an 80% cut, and Terra from $2.50/$15 to $2/$12. Sol is unchanged at $5/$30. This is the kind of change that quietly invalidates a buying decision made two months ago, so we have updated the cost calculator and we are saying plainly what it does to the argument for buying a GPU: at the cheap end, it weakens it.

Verdict: Luna just got 80% cheaper — the cheapest credible hosted tier we track, and the first real dent in the local-vs-cloud payback math this year


The take

The numbers, verbatim from OpenAI's own announcement: "Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna. Sol pricing remains unchanged." The developers.openai.com rate table agrees, and adds the detail that long-context requests bill at double input and 1.5× output. Also new: Fast mode replaces Priority Processing, at twice the standard price for roughly 2.5× the speed.

What it does to our cost calculator. On a blended 50/50 input/output basis Luna moves from $3.50 to $0.70 per million tokens — it is now the cheapest row we track, undercutting Gemini 3.1 Flash-Lite's $0.875, which had held that spot since May. Terra moves from $8.75 to $7.00. Run the calculator honestly and the consequence is uncomfortable for the local case: a heavy user on Luna now needs a very long payback window before a $3,449 Framework Desktop or a $3,639 RTX 5090 earns itself back on token costs alone. If your only argument for local was arithmetic, that argument is weaker this week than last.

Which is worth saying because it is not the argument we actually make. The case for local weights has never rested on beating hosted inference on price per token — hosted providers can and do cut prices whenever it suits them, and a price cut is not a commitment. The case rests on the things a price cut cannot touch: the model cannot be deprecated out from under you, the terms cannot be rewritten, your data does not leave the machine, and it keeps working when the vendor decides your country or your use case is no longer served. We watched exactly that happen this year when Fable 5 was switched off for 19 days under an export-control directive. Buy local for control and privacy, and treat cheap hosted inference as what it is: a genuinely good deal, on someone else's terms.

Where this fits

Models: gpt-oss-20b · Qwen3-Coder-30B-A3B · Qwen 3.6-35B-A3B

Hardware: NVIDIA RTX 5090 · Framework Desktop (Ryzen AI Max+ 395) · RTX 5060 Ti 16 GB

Sources

Next step

Try this in the planner