"In all other use cases, you really don't need GPT-5.6 Sol Fast mode." — OpenAI
GPT-5.6 Luna: dramatic price cuts, maintained performance
OpenAI announced steep reductions to the API price for its GPT-5.6 Luna model, cutting Luna’s cost to $0.20 per million input tokens and $1.20 per million output tokens. Those prices represent an 80% reduction from Luna’s prior rates of $1 per million input tokens and $6 per million output tokens, according to the company’s update.
At the same time OpenAI reported that, in its own test results, it places Luna at the top of its intelligence index among the compared models — a claim the company highlights alongside the substantially lower cost-per-task implied by the new pricing.
GPT-5.6 Terra: modest reductions, continued premium positioning
OpenAI also lowered pricing for GPT-5.6 Terra. Input-token pricing moved from $2.50 to $2.00 per million, and output-token pricing fell from $15 to $12 per million. The cuts for Terra are smaller in percentage terms than Luna’s, but follow the same pattern of reducing per-token costs across the 5.6 family.
GPT-5.6 Sol Fast: speed at twice the price
OpenAI introduced a Fast mode for GPT-5.6 Sol aimed at API customers who need quicker responses. The company said Sol Fast is up to 2.5 times faster than standard processing without reducing the model’s intelligence. That extra performance, however, comes at twice the standard API price for Sol.
OpenAI’s guidance is explicit: the Fast option is tailored for time-sensitive workflows — the company singled out coding, research, and "agentic workloads" — and said that for all other use cases users do not need Sol Fast mode. The announcement also made clear that Sol’s standard pricing remains unchanged for now.
Auto-review, Codex, and ChatGPT Work: accounting and operational impacts
OpenAI said the new prices affect how it counts usage in Codex and ChatGPT Work. Concretely, the company noted that if new tasks use these GPT-5.6 models, they will deduct less from customers’ allowances, "so you can complete more work under the same quota." The firm is also upgrading Auto-review in the ChatGPT app and the Codex CLI from GPT-5.4 to GPT-5.6 Luna; OpenAI estimates this change should reduce the cost of Auto-review by approximately ten times.
What this means for technologists, procurement leaders, and end users
- Technologists and developers: Lower per-token pricing for Luna and Terra means experiments and continuous integration tasks that are token-intensive may become materially cheaper. Developers running Auto-review in Codex or ChatGPT Work will see a concrete cost reduction if their workflows migrate to GPT-5.6 Luna, given OpenAI’s estimate of roughly a tenfold cost saving for Auto-review.
- Procurement and enterprise buyers: Because OpenAI says these changes cause fewer allowance deductions for the same tasks, organizations that meter usage by quotas can expect existing allowances to stretch further when they move workloads to GPT-5.6 models. The availability of a higher-cost, faster Sol Fast option creates a clear procurement trade-off between latency-sensitive workloads and routine tasks.
- End users and product managers: The company’s placement of Luna at the top of its intelligence index alongside much lower per-task cost suggests product teams will weigh Luna as both a capability and a cost lever when choosing models for features that require high throughput or intensive output token generation.
OpenAI’s announcement pairs two themes: making higher-performing models less expensive per token, and introducing a premium fast lane for situations where latency outweighs unit cost. The largest percentage price cut targeted Luna, which OpenAI positions as both the most cost-efficient and, in its tests, the most intelligent of the compared models. Meanwhile, Sol Fast offers a clear but costly option for time-critical API usage — and OpenAI’s own language signals that most users will not need that premium tier.
Source: Bleeping Computer — OpenAI says its new GPT 5.6 models are becoming more cost-efficient




