Three model releases on the same day invite a shallow comparison: benchmark rank and price per million tokens. The cost that matters is whether work is completed, how much human repair it needs, whether caching hits, and how long recovery takes after a failure.
On September 22, OpenAI released GPT-6 Sol and GPT-6 Luna in its developer changelog, while Anthropic released Claude Opus 5.5. This guide turns the announcements into a practical routing plan: separate work by consequence and repetition, run a controlled pilot, then decide on completed-task cost.

The graphic does not declare a universal winner. High-volume work with deterministic checks has a different selection rule from ambiguous work, long code chains, or work with serious quality consequences.
Confirm model IDs, pricing and boundaries first
Official API list prices are per million tokens. The table uses the OpenAI September 22 changelog and Anthropic’s release post. Check the console for long-context, Fast mode, batch and regional terms before deployment.
| Model | Good starting workload | Input / cached input / output | Verify first |
|---|---|---|---|
| GPT-6 Luna | Classification, extraction, structured drafts, batch jobs | $0.10 / $0.01 / $0.50 | Volume, cache hit rate, structured output |
| GPT-6 Sol | Supervised debugging, analysis, tool use, multi-step execution | $2 / $0.20 / $10 | Success rate, tool failures, output length |
| Claude Opus 5.5 | High-value deliverables, complex code, long audits, quality-sensitive work | $4 / $0.20 / $20 | Repair reduction, permissions and safeguards |
OpenAI says Sol and Luna are available through Responses and Chat Completions, with interface-specific details in its model guidance. Anthropic lists Opus 5.5 at $4 input, $20 output and $0.20 cache-read per million tokens, and documents safeguards and routing for some sensitive work. Those boundaries are part of the selection decision.
Do not substitute the lowest unit price for cost accounting
Start with this API estimate:
task API cost =
uncached input tokens × input price
+ cached input tokens × cache-read price
+ output tokens × output price
+ cache writes, tool calls and long-context fees
task API cost =
uncached input tokens × input price
+ cached input tokens × cache-read price
+ output tokens × output price
+ cache writes, tool calls and long-context fees
Then add the costs that appear in an actual team:
completed-task cost =
API cost
+ human review and repair time × hourly cost
+ failed reruns × cost per rerun
completed-task cost =
API cost
+ human review and repair time × hourly cost
+ failed reruns × cost per rerun
Luna can be the right choice for 100,000 structured extraction jobs when output remains valid and retries are rare. A payment migration or repository-wide permission audit can justify a more expensive model when it prevents one incorrect change and a day of rework.
Cache pricing is not a footnote. Keep durable system instructions, tool descriptions and knowledge references in a stable prefix. Put the daily task later. Changing every part of the prefix on every request destroys the cache benefit.
Divide work into three tiers
Do not migrate all traffic at once. Sample current requests and sort them by consequence and repetition.
Tier 1: repetitive, low-risk, easy to validate
Classification, field extraction, format normalization, first-pass summaries, batch translation and ticket routing fit here. Start with Luna only when there is a schema, a sample set and a failure fallback.
Take 100 de-identified requests that humans already resolved. Write mechanical acceptance rules: fields are present, amounts are unchanged, dates use the right format. Run two passes, record valid-result rate, latency, retry count and repair minutes. Expand traffic only after the business threshold is reached. External writes should require validation, a confidence condition and approval.
Tier 2: reasoning and tool-assisted execution
Debugging, data analysis, workflows with tool calls and longer document synthesis are good Sol candidates. Fix the task set, toolset, maximum turns and evaluator before comparing models. If you change model, prompt, tools and context together, the result is not comparable.
Record tool calls as well as tokens. A model that retries the same broken endpoint ten times may look inexpensive by token count and still cost you latency and incident response.
Tier 3: expensive failures, ambiguity or deep review
Repository migration, permissions audit, complex repair, a consequential report and final publication review belong here. Anthropic’s release material gives Opus 5.5 as a candidate for long and complex work. Treat vendor performance claims as a reason to test, not a promise that your workload will reproduce them.
Compare with a fixed acceptance gate: tests passing, change scope, fact traceability and number of human review edits. When a lower-priced route fails the same quality gate twice, escalate rather than adding more uncontrolled reasoning turns.
Run a seven-day pilot
| Day | Action | Output |
|---|---|---|
| 1 | Choose 20–50 tasks per tier; set evaluators and data boundaries | Task set and score sheet |
| 2 | Run the incumbent route; save input version, output, time and cost | Baseline |
| 3 | Run Luna on Tier 1 only; do not auto-write external systems | Batch results |
| 4 | Run Sol on Tier 2 with fixed tools and turn cap | Execution results |
| 5 | Run Opus 5.5 on Tier 3 with human code or fact review | High-value results |
| 6 | Classify failures; repair prompt, tool contract or data | Failure ledger |
| 7 | Route by completed-task cost and prepare rollback | Routing policy |
Track task ID, model and parameters, first-pass success, human repair minutes, tokens and tool fees, and final deliverability. Without the final deliverability column, a team can be fooled by attractive isolated scores.
Start with an explicit routing rule
if work is high-volume and has a deterministic validator:
route to GPT-6 Luna
elif work needs tools or multi-step execution with human review:
route to GPT-6 Sol
elif failure has legal, financial, security, production, or publication impact:
route to Claude Opus 5.5 or the team's highest approved model
else:
start with Sol, cap attempts, and escalate after two failed validations
if work is high-volume and has a deterministic validator:
route to GPT-6 Luna
elif work needs tools or multi-step execution with human review:
route to GPT-6 Sol
elif failure has legal, financial, security, production, or publication impact:
route to Claude Opus 5.5 or the team's highest approved model
else:
start with Sol, cap attempts, and escalate after two failed validations
Review failures weekly. If Luna fails because the schema is unclear, repair the validator. If Sol loops on tools, improve tool outputs and retry policy. If Opus 5.5 costs more without reducing rework, narrow its route. With ClaudeAPI, keep task tiers, cost records and human review in the same workflow so a model switch remains auditable.
Create a task card before you create routing rules
“Write an analysis” and “fix a bug” are not routing inputs. Give each work type a task card: type, stable prefix and context size, tool use, success definition, consequence of failure, human minutes and fallback. Similar tasks can carry very different risks.
| Field | What to record | Why it matters |
|---|---|---|
| success definition | JSON validation, tests, human score, cited facts | replaces “looks good” with a finish line |
| failure consequence | rerun, customer impact, production impact | sets escalation and approval needs |
| cost fields | input, cache, output, tools, human minutes | makes total cost auditable |
| fallback | retry, upgrade, human handoff, stop write | prevents endless loops |
Use one sample to calculate what cheap means
Do not use different tasks for each model and compare invoices. Use the same completed tasks with the same prompt and validator. Track first-pass rate, human repair minutes, median latency and total cost per delivered result.
task_id,task_tier,model,validator_pass,review_minutes,input_tokens,cached_tokens,output_tokens,tool_calls,retries,final_status
T-001,extract,Luna,1,0,4200,3000,220,0,0,delivered
T-002,debug,Sol,1,12,18000,14000,1600,3,1,delivered
T-003,release_review,Opus,1,28,25000,20000,2400,4,0,delivered
task_id,task_tier,model,validator_pass,review_minutes,input_tokens,cached_tokens,output_tokens,tool_calls,retries,final_status
T-001,extract,Luna,1,0,4200,3000,220,0,0,delivered
T-002,debug,Sol,1,12,18000,14000,1600,3,1,delivered
T-003,release_review,Opus,1,28,25000,20000,2400,4,0,delivered
A low-priced request that needs two reruns and fifteen minutes of repair does not automatically belong in the default route.
Write escalation rules instead of treating retries as a strategy
Passes a deterministic validator → eligible for the low-cost batch pool
Fails once → inspect input and schema before changing models
Fails the same task twice → escalate one tier and keep the failure sample
Writes externally, changes permissions, touches production, or publishes → require human approval
Makes a numeric, policy, or factual claim → require a source field or withhold delivery
Passes a deterministic validator → eligible for the low-cost batch pool
Fails once → inspect input and schema before changing models
Fails the same task twice → escalate one tier and keep the failure sample
Writes externally, changes permissions, touches production, or publishes → require human approval
Makes a numeric, policy, or factual claim → require a source field or withhold delivery
Review retries, escalations and human takeovers weekly. If most escalations originate in one ambiguous field, fix the form. If a tool endpoint is repeatedly called, fix its return values and errors. If a higher-cost model reduces rework only for one task class, keep it there. Let routing move with this evidence, not launch-day excitement.



