Latest AI News

July 31, 2026 · Daily brief

OpenAI's GPT-5.6 Luna Drops 80%—Model Optimized Itself

Sovereignty angle
When AI models optimize their own infrastructure code, the line between tool and toolmaker blurs. If you're building on Luna at $0.20/M tokens, you're now dependent on a model that rewrites itself—a stack you audit less and control even less.

OpenAI slashed GPT-5.6 Luna pricing by 80% to $0.20/$1.20 per million tokens after its Sol model autonomously rewrote GPU kernels, cutting serving costs 20%.

On July 30, OpenAI announced dramatic price cuts to its GPT-5.6 model family just three weeks after launch. Luna dropped from $1 to $0.20 per million input tokens and from $6 to $1.20 output, while Terra fell from $2.50 to $2 input and from $15 to $12 output. The flagship Sol model remains at $5/$30 but added a Fast mode running 2.5x faster at double the price.

What makes this unusual is how OpenAI achieved the efficiency gains. With Codex, GPT-5.6 Sol autonomously rewrote and optimized OpenAI's production kernels, the core code that executes the mathematical operations making up the model. Combined with broader kernel improvements that Sol identified through the same process, the effort cut end-to-end serving costs by 20%. The model also rewriting its own GPU code to make the 5.6 models 15% more efficient.

The timing reveals competitive pressure. A CNBC investigation published on July 7, 2026, revealed that Chinese models have captured 46% of US enterprise token usage on OpenRouter, with models like DeepSeek V4 Pro priced far below OpenAI's offerings. Sam Altman acknowledged OpenAI wants to "offer the best price/intelligence tradeoff at every level."