On July 30, OpenAI announced dramatic price cuts to its GPT-5.6 model family just three weeks after launch. Luna dropped from $1 to $0.20 per million input tokens and from $6 to $1.20 output, while Terra fell from $2.50 to $2 input and from $15 to $12 output. The flagship Sol model remains at $5/$30 but added a Fast mode running 2.5x faster at double the price.
What makes this unusual is how OpenAI achieved the efficiency gains. With Codex, GPT-5.6 Sol autonomously rewrote and optimized OpenAI's production kernels, the core code that executes the mathematical operations making up the model. Combined with broader kernel improvements that Sol identified through the same process, the effort cut end-to-end serving costs by 20%. The model also rewriting its own GPU code to make the 5.6 models 15% more efficient.
The timing reveals competitive pressure. A CNBC investigation published on July 7, 2026, revealed that Chinese models have captured 46% of US enterprise token usage on OpenRouter, with models like DeepSeek V4 Pro priced far below OpenAI's offerings. Sam Altman acknowledged OpenAI wants to "offer the best price/intelligence tradeoff at every level."