The number that matters is 80. On July 30 OpenAI cut the API price of GPT-5.6 Luna by 80%, to $0.20 per million input tokens and $1.20 per million output, and cut GPT-5.6 Terra by 20%, to $2 and $12. Sol, the top tier, keeps its price but gains a Fast mode: up to 2.5 times faster at twice the cost, replacing Priority Processing. The cuts were announced in a post titled 'Advancing the price-performance frontier with GPT-5.6,' one day after a companion post explained where the company says the margin came from.
That explanation is the unusual part. OpenAI says GPT-5.6 Sol, running inside Codex, autonomously rewrote and optimized production kernels written in Triton and Gluon, its own open-source GPU languages, cutting end-to-end serving cost by 20%. The model also designed and ran hundreds of experiments on the speculative-decoding draft model that accelerates its own token generation, lifting that efficiency by more than 15%, and intervened on hardware failures and training instability along the way. OpenAI is careful to frame this as 'a human-led process,' and says the AI-written kernels were validated with FpSan, an open-source floating-point sanitizer.
The distribution moved at the same speed as the price. All three GPT-5.6 models are now generally available on Amazon Bedrock, with a new explicit prompt-caching mode: cache reads billed at a 90% discount, writes at 1.25 times the uncached rate, a 30-minute TTL and up to four cache breakpoints per request. OpenAI's benchmark claims for the week: Sol at maximum reasoning beats Anthropic's Claude Fable 5 on Artificial Analysis's Coding Agent Index 'at less than half of the cost,' Terra matches GPT-5.5 on intelligence benchmarks at half the price, and Luna outperforms Fable 5 on Agents' Last Exam at an estimated 99% lower cost per task.
For builders, the practical part is simple: any cost model written in Q2 is stale today, and the new caching economics reward stable prompt prefixes. The fine print takes longer. The widely quoted figure that March's flagship intelligence now sells for one-thirteenth the price is third-party arithmetic, comparing GPT-5.4's score on one public index with Luna's, not an OpenAI number. The self-optimization gains are OpenAI measuring OpenAI, inside a process it defines. And 'the model made itself cheaper' is a claim about infrastructure the buyer cannot inspect. None of that makes the invoices less real. It means the efficiency story now setting prices deserves the same line-by-line reading as the benchmarks before it.
