The price card is an input, not the result
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026. Sol is positioned for complex coding and agentic workflows. Luna is positioned for focused, high-volume tasks. Both expose a 1,050,000-token context window and up to 128,000 output tokens. Sol costs $2 per million input tokens and $10 per million output tokens at the standard rate; Luna costs $0.10 and $0.50. Those numbers are useful, but they do not answer the migration question.
A model migration succeeds when the new route produces more accepted work for the same engineering budget without breaking task boundaries. Token price alone misses retries, longer outputs, human review time, failed tool calls, regressions, and escalation to a stronger model. A cheap call that produces an unusable patch is not a bargain. A more expensive call that closes a high-risk issue with a small verified diff may be.
The first week of community discussion makes that distinction visible. OpenAI describes improved price-performance and lower API prices. Large Reddit threads report a mixture of successful routine work, perceived coding regressions, and uncertainty about how API economics map to Codex subscriptions. A GitHub issue also reports reasoning-effort behavior on the ChatGPT-backed Codex path. These are signals to test, not proof that every Sol or Luna deployment has the same behavior.
Treat the model name as a candidate configuration. The unit of promotion is a task class with evidence.
The release changed after launch
On September 25, OpenAI fixed an image-encoding bug that had degraded image understanding in Sol and Luna across the API and Codex, including computer-use tasks. The changelog explicitly tells affected users to rerun evaluations and retry impacted workflows. That is a compact lesson in why evaluation fixtures need version and date metadata. A result recorded on September 23 for an image-heavy workflow is not equivalent to a result recorded after the fix.
This does not make Sol or Luna unreliable by definition. It means the deployed model is part of a moving service. Evaluation evidence has a freshness boundary. If a provider announces a behavioral fix, changes a model alias, or updates a supported control, the relevant fixtures return to the queue.
Define each tier as an operational contract
Sol and Luna share headline features, but their intended work differs. Sol is the default candidate for ambiguous repository work, multi-step repair, and agent workflows where reasoning quality is worth paying for. Luna is the default candidate for narrow transformations, classification, extraction, and high-volume steps with a deterministic verifier. Neither should be a universal default.
| Decision | GPT-6 Sol | GPT-6 Luna | Migration implication |
| Provider positioning | Complex coding and agentic workflows | Focused, high-volume tasks | Start with separate fixture sets |
| Standard text price | $2 input / $10 output per 1M | $0.10 input / $0.50 output per 1M | Calculate full task cost, not rate alone |
| Reasoning effort | none through max | none through max | Pin and record the requested effort |
| Built-in tools | Responses API | Responses API | Do not assume endpoint parity |
| Chat Completions tools | Function calling only at none | Function calling only at none | Test endpoint and effort together |
| Long input pricing | Above 272K input tokens, full-request input/cache rates double and output is 1.5x | Long context can reverse a cheap-looking route |
The endpoint constraint is easy to miss. OpenAI's model pages say to use the Responses API for built-in tools and function calling. Chat Completions supports function calling only when reasoning effort is set to none. A migration that preserves the model ID but changes endpoint behavior is not equivalent. Freeze the tuple of model, endpoint, effort, tool schema, system contract, and processing mode.
version: 1
routes:
routine_transform:
model: gpt-6-luna
endpoint: responses
reasoning_effort: low
max_output_tokens: 6000
verifier: schema_and_snapshot
fallback: gpt-5.6-luna
repository_repair:
model: gpt-6-sol
endpoint: responses
reasoning_effort: high
max_output_tokens: 24000
verifier: tests_diff_scope_human
fallback: gpt-5.6-sol
promotion:
minimum_pass_rate_delta: 0
maximum_review_minutes_delta: 0
maximum_scope_violation_rate: 0.01
minimum_fixture_count: 30
This routing file is a policy artifact, not a performance claim. Its value is that it makes the expected route and promotion rule reviewable. A release manager can see when the fallback changes, which verifier applies, and whether a high-volume task was accidentally pointed at Sol max.
Build acceptance fixtures around work your team already understands
Public benchmarks answer broad questions. They do not include your repository conventions, internal APIs, flaky tests, review expectations, or tool permissions. Use completed tasks where the correct outcome and failure patterns are known. Keep the task request, starting commit, relevant environment, expected checks, and accepted patch. Remove secrets and unstable external dependencies.
Separate fixtures by task class. A classification result should not offset a failed migration. A good suite might include routine structured extraction, small bug fixes, dependency upgrades, multi-file features, visual tasks, and tool-heavy incident triage. Sol and Luna can then be promoted independently instead of through one blended score.
for fixture in acceptance_suite:
candidate = run(route=fixture.candidate_route, input=fixture.input)
baseline = run(route=fixture.baseline_route, input=fixture.input)
for result in (candidate, baseline):
result.schema_ok = validate_schema(result.output)
result.tests_ok = run_frozen_tests(result.patch)
result.scope_ok = changed_files(result.patch) <= fixture.allowed_files
result.review_minutes = blinded_human_review(result)
result.total_cost = tokens + tools + retries + review_minutes_cost
record(compare(candidate, baseline), service_version, run_date)
promote_only_if(all_required_gates_pass_by_task_class)
Record requested and observed controls
Do not assume a reasoning-effort parameter was honored because the request succeeded. One current Codex issue reports that changing effort mid-thread appeared not to change behavior for Sol and Luna on the ChatGPT backend. The reporter explicitly did not test the public API, so the issue cannot support a general API claim. It does support a test design: log the requested setting, response metadata, endpoint, client version, and whether the task changed in a measurable way.
For tool use, record each call, argument summary, result status, retry, and permission denial. For code changes, store the patch, tests, build output, and reviewer decision. For image and computer-use tasks, include the exact image fixture and provider-change date because the September 25 fix altered the relevant path.
Use a blinded review when quality is subjective
Developer preference is noisy when the reviewer sees the model name. Strip model labels from code explanations, design proposals, or UI drafts and ask reviewers to score correctness, completeness, scope, and revision effort. Preserve a free-text rejection reason. The rejection taxonomy often teaches more than the average score: missing constraint, wrong tool, incomplete edge case, excess diff, fabricated API, or unreadable explanation.
Freeze the verifier before comparing routes
A migration experiment becomes circular when the candidate model can change the test that decides whether it succeeded. Freeze expected outputs, deterministic checks, allowed-file boundaries, and reviewer instructions before the first candidate run. If a fixture genuinely contains a bad assertion, repair it in a separate versioned change and rerun every route against the new fixture version. Never let one model receive the corrected test while its baseline keeps the old one.
For code tasks, require the original regression test to fail against the starting revision and pass against the candidate patch. Run the unchanged surrounding suite as a second gate. A patch that deletes an assertion, widens a type to silence an error, or replaces a precise failure with a catch-all should be rejected even if the command exits successfully. For extraction or classification, keep a hidden answer key and score field-level errors separately; a perfect average can conceal a dangerous miss in the one field that authorizes an action.
Sample the negative space as well. Include tasks the model should refuse, tasks with insufficient evidence, tool calls outside the granted scope, and requests whose correct result is a clarifying question. Sol's extra capability is not valuable if it makes confident progress past a missing requirement. Luna's throughput is not valuable if a queue of cheap calls creates more exceptions than reviewers can resolve.
Finally, set the decision rule before seeing the results. Define which metrics are hard gates, which may trade off, and how much regression is acceptable for each workload. Require confidence intervals or repeated runs when outcomes vary. A useful promotion record names the fixture set, model snapshot, endpoint, reasoning effort, client and harness versions, run count, evaluator, exceptions, and rollback owner. That record lets the team distinguish a real model change from a prompt edit, client upgrade, transient service issue, or reviewer preference.
Measure cost per accepted task
The useful denominator is accepted output, not call count. Start with provider token charges, then add tool fees, retries, fallback calls, and human review. Divide by the number of outputs that pass the required gate. Keep latency and queue pressure as separate service metrics so a cheap but slow route does not hide an operational problem.
cost_per_accepted_task =
(input_tokens_cost
+ cache_write_cost
+ output_tokens_cost
+ tool_call_cost
+ retry_and_fallback_cost
+ reviewer_minutes * loaded_reviewer_rate)
/ accepted_tasks
Consider two routes for 1,000 narrow extraction tasks. Luna might cost very little in tokens, but if 12 percent require a Sol fallback and reviewers spend two extra minutes resolving schema drift, the final cost changes materially. Conversely, Sol may be cheaper for a difficult repair if it avoids three failed iterations and 20 minutes of senior review. This is why the same model can be the right choice for one lane and the wrong choice for another.
Context length needs its own budget. Both model pages state that prompts above 272K input tokens move the entire request to higher rates. A 1.05M context window is a capability ceiling, not permission to send a repository indiscriminately. Retrieval quality, prompt caching, and context compaction remain architectural work.
Migration failure modes
| Failure | Why it happens | Detection | Response |
| Alias-only rollout | Team assumes family continuity | No frozen before/after fixtures | Restore baseline and run acceptance suite |
| Cheap-token illusion | Retries and review are omitted | Token cost falls while accepted throughput does not | Use full task-cost equation |
| Endpoint mismatch | Tool behavior differs by API path and effort | Missing or rejected function calls | Pin endpoint and effort in route contract |
| Long-context surprise | Requests cross 272K input threshold | Step change in per-request cost | Compact context and add budget alerts |
| Stale visual eval | Fixture predates September 25 fix | Run metadata shows earlier service date | Rerun image and computer-use tests |
| Sentiment as benchmark | Forum reports are generalized | No reproducible local evidence | Use reports to select tests, not conclusions |
| No rollback lane | Old model removed on day one | Regressions block production queue | Keep versioned routes and fallback capacity |
Community reports deserve neither dismissal nor promotion to universal fact. A high-engagement complaint can expose a task class worth testing. It cannot establish that every account, endpoint, prompt, or repository behaves the same way. The right response is a fixture that reproduces the claimed failure under controlled conditions.
FAQ
Should every GPT-5.6 Sol workload move to GPT-6 Sol?
No. Similar naming does not establish task equivalence. Replay representative fixtures and promote only task classes where verified quality and total cost meet the gate.
When is GPT-6 Luna the better default?
Luna is the natural candidate for focused, high-volume work with tight inputs and deterministic verification. It is not automatically suitable for ambiguous, high-impact decisions simply because it supports high reasoning effort.
Can I compare Codex subscription experience with API pricing?
Only with care. Subscription quotas, product routing, client behavior, and API billing are different systems. Keep observations attached to the surface where they were measured.
What should be rerun after the September 25 fix?
Any acceptance fixture that depends on image understanding, screenshots, visual interfaces, or computer use and was executed before the fix should be rerun.
What is the minimum evidence for promotion?
Use enough representative fixtures to cover the task's failure modes, then require correctness, scope, verification, review-time, and cost gates. Thirty fixtures can be a reasonable starting floor for a narrow lane, but risk and diversity should determine the final number.