GPT-6 Sol and Luna: the price cut is the upgrade
GPT-6 Sol and Luna cut API prices while independent benchmarks show smaller gains. I compare the scores, ChatGPT access, and mixed first reports on Reddit.

GPT-6 Sol and Luna launched on 22 September 2026 with their strongest immediate argument in the price table: Sol’s standard API input and output rates are half those of GPT-5.6 Sol. Independent benchmarks show a coding improvement for Sol, but only small or flat intelligence-index changes across the pair. I would treat this as a value upgrade first. [1] [2] [3]
What changed in the API prices?
The price cuts are large enough to matter even if your tasks see little quality improvement. OpenAI halves Sol’s input and output rates, while Luna’s input rate falls by half and its output rate drops from $1.20 to $0.50 per million tokens. [2]
OpenAI standard API rates for short-context requests, retrieved 22 September 2026. Uncached input; not subscription usage rates. [2]
| Model | Input, USD / 1M tokens | Output, USD / 1M tokens |
|---|---|---|
| GPT-5.6 Sol | 4.00 | 20.00 |
| GPT-6 Sol | 2.00 | 10.00 |
| GPT-5.6 Luna | 0.20 | 1.20 |
| GPT-6 Luna | 0.10 | 0.50 |
At the new rates, a hypothetical request with 100,000 uncached input tokens and 10,000 billed output tokens costs $0.30 on Sol, compared with $0.60 on its predecessor. The same token counts cost $0.015 on Luna instead of $0.032. That is a calculation from the price table, not a measurement of how many tokens either model needs to solve a task.
The distinction matters for agents because a cheaper model can spend the saving on extra attempts. Long-context pricing also changes the calculation: OpenAI applies higher rates when input exceeds 272,000 tokens. I would compare bills at the request sizes I actually use rather than assume the headline rates apply everywhere. [2]
This is the economic reason to revisit the over-engineering I saw with GPT-5.6 Sol. Cheaper tokens reduce the cost of an unnecessary detour, but they do not make the detour useful. The new release still has to follow the repository’s conventions and stop when the requested work is done.
Do the GPT-6 benchmarks show a big improvement?
The independent results show a modest aggregate change and a clearer gain on Sol’s terminal work. Artificial Analysis’ Intelligence Index v4.3.2 displays Sol at 48, up from 47, while Luna stays at 37 after rounding. All four entries use maximum effort. These are rounded index points, not percentages of tasks solved; the underlying changes are about +0.55 for Sol and -0.07 for Luna. [4] [5] [6] [7]
| Model family | GPT-5.6 | GPT-6 | Change in displayed score |
|---|---|---|---|
| Sol | 47 | 48 | +1 point |
| Luna | 37 | 37 | 0 points |
That does not mean every task is unchanged. In Artificial Analysis’ Terminal-Bench 4.0 evaluation, Sol rises from 39.9% to 43.9%. The evaluator uses mini-swe-agent across 66 terminal tasks, averaging single-attempt success over three repetitions. It is a four-percentage-point gain under that setup. [3]
Show the data as a table
| Benchmark | GPT-6 Sol | GPT-5.6 Sol |
|---|---|---|
| Terminal-Bench 4.0 | 43.9% | 39.9% |
The improvements are uneven. On Artificial Analysis’ GDPval-AA v2.1, which compares the quality of work products, Sol falls from 1,588 to 1,487 Elo and Luna from 1,443 to 1,367. Elo is a relative rating from comparisons between outputs, so these are not percentage-point losses. The results are a reason to test the documents and analytical work you actually need instead of assuming every GPT-6 output improves. [13]
Sol’s terminal gain is therefore more interesting than its small index change, but the release is not a clean win across evaluations. I would not assign Luna a terminal score that is absent from this comparison. OpenAI’s model guide describes better coding, factual reliability, and communication; the official sources checked for this article did not provide a numerical comparison with both predecessors. [8]
There is a stronger independent coding jump in the separate Opus 5.5 release. For Sol and Luna, the combination of these results and lower rates makes the more persuasive case: useful capability at a lower token price.
Can you use Sol and Luna in ChatGPT?
The new models are rolling out in ChatGPT Work and Codex, not ordinary Chat. OpenAI’s changelog names those surfaces explicitly. If you are looking in the Chat model picker, the absence of GPT-6 Sol does not by itself mean your rollout failed. [1]
That product distinction is easy to miss when a release gets shortened to “ChatGPT 6.” Sol and Luna are model names, while ChatGPT contains different places to use models. They also follow the earlier GPT-6 Astra launch; Astra is not another model newly launched on 22 September.
This distinction matters more than a token discount when you are in the middle of a long job. A plan can still impose a shorter usage window even when the weekly allowance has room, as I described in the Codex Plus five-hour limit. I would keep API calculations and subscription observations separate when judging the release.
What are Reddit users noticing?
The first Reddit threads contain price enthusiasm and conflicting reports about practical quality. These are launch-day anecdotes from 22 September, with different tasks and effort settings. I have not run a controlled comparison of my own for this article.
In r/codex, one person reports a promising generated plan that consumed 1% of their five-hour allowance and explicitly calls it a sample of one. In the same thread, another prefers the previous Luna, while a Sol user complains about excessive links and document paths in answers. Those are useful things to test, but they do not establish a typical usage rate. [9]
The release thread shows the same split: one user likes Luna at maximum effort and reports low usage, while another calls it a poor experience. Some comments describe access errors during rollout. On r/OpenAI, enthusiasm about lower prices sits alongside complaints that the models are unavailable in ordinary Chat. [10] [11]
My next step would be to separate the jobs. OpenAI recommends starting Sol at medium effort and Luna at high, and positions Luna for bounded work such as extraction, classification, and structured summaries. That gives me a sensible starting point for a comparison, rather than setting everything to maximum and waiting to see what the usage meter does. [8]
I would give Sol a known repository change and Luna a well-defined transformation, then compare accepted results, corrections, and actual cost against the previous models. The price cut is already documented. Whether it buys more completed work is the part worth measuring next.





