Skip to main content
This site is an independent third-party technical service provider. Claude™ and Anthropic® are trademarks of Anthropic, PBC. This site has no affiliation, endorsement, or partnership with Anthropic.

Claude Sonnet 5.5 migration guide: route work by task and effort, not benchmark scores

Use Claude Sonnet 5.5 with a task-routing, effort-testing, cost-tracking, and rollback workflow for coding and knowledge work.

Dev GuidesClaude Sonnet 5.5model routingClaude APIagentscost optimizationEst. read6 min
2026.09.29 published
Claude Sonnet 5.5 routing dashboard shows effort, cost, and task acceptance

Moving from Sonnet 5 to Sonnet 5.5 is a delivery decision, not a leaderboard decision. A team needs to know which jobs finish faster, what each usable result costs, and how a failed rollout returns to the prior path. This guide uses task segmentation, effort selection, controlled trials, and rollback rules.

Anthropic’s Sonnet 5.5 announcement says the model keeps Sonnet 5 pricing at $2 per million input tokens and $10 per million output tokens, while improving output speed by more than 30% and reducing per-task cost by up to 30% for most work in Anthropic’s testing. The same announcement keeps an important boundary: Opus 5.5 remains stronger for complex, open-ended work that requires sustained judgment.

Claude Sonnet 5.5 effort routing workflow

Workflow: choose model and effort by task class, then use acceptance and rollback data to refine the route.

Segment work before changing a default model

Classify work by acceptance criteria and error cost.

Work type First trial Keep the stronger review path when
Bounded work: tested bug fixes, template documents, structured tables Sonnet 5.5 at Low or Medium output can take a production action without approval
Stable multi-step work: repository changes, weekly reports, sourced drafts Sonnet 5.5 at Medium with review scope changes or the conclusion needs cross-functional ownership
Open-ended decisions: architecture, incidents, compliance, risky migrations Opus 5.5 or dual review a benchmark result is the only reason to downgrade

A bounded task has explicit tests, prohibited changes, and a named owner. “Fix the login timeout” becomes useful only after you add a reproduction command, test command, affected directory, and interface constraints.

Treat effort as a budget control

Anthropic states that Claude apps and Claude Code default to Medium effort, while Claude Platform defaults to High. Higher effort can make the model reason and check longer; it can also raise latency and cost. Start with routing rather than a single global setting.

routing:
  routine_fix:
    model: claude-sonnet-5-5
    effort: low
    acceptance: "tests_pass && diff_within_scope"
  documented_change:
    model: claude-sonnet-5-5
    effort: medium
    acceptance: "tests_pass && reviewer_approves"
  ambiguous_decision:
    model: claude-opus-5-5
    effort: high
    acceptance: "evidence_cited && decision_owner_approves"
routing:
  routine_fix:
    model: claude-sonnet-5-5
    effort: low
    acceptance: "tests_pass && diff_within_scope"
  documented_change:
    model: claude-sonnet-5-5
    effort: medium
    acceptance: "tests_pass && reviewer_approves"
  ambiguous_decision:
    model: claude-opus-5-5
    effort: high
    acceptance: "evidence_cited && decision_owner_approves"

Confirm current model IDs, parameters, and SDK syntax in the Claude Platform documentation before deploying.

Run a seven-day controlled trial

Choose 20–30 completed tasks with reviewable outcomes. Run the old path and Sonnet 5.5 with the same inputs, permissions, tools, and reviewer. Log task class, model, effort, input/output tokens, latency, first-pass acceptance, rework minutes, and failure reason.

Use usable-delivery cost, not token cost alone:

usable delivery cost = model-call cost + rework minutes × team cost per minute
usable delivery cost = model-call cost + rework minutes × team cost per minute

If first-pass acceptance falls, rework concentrates in one task class, or a permission/data-boundary issue occurs, route that class back before broadening the trial.

Put coding changes between tests and review

Anthropic reports a 70.6% result for Sonnet 5.5 on Terminal-Bench 4.0, but a benchmark is not repository acceptance. Require a plan, a minimal diff, isolated test execution, and reviewer sign-off. For existing workflows that disable upfront thinking, Anthropic’s migration guide calls out the new between_tools setting; test it in a non-production repository first.

Check API compatibility before routing real traffic

Migration is more than changing model to claude-sonnet-5-5. Anthropic’s migration guidance identifies changes that can return HTTP 400 or quietly hide content in your UI. Replay an existing request for each case:

Old dependency Check on Sonnet 5.5 Likely failure
thinking: disabled Use adaptive thinking, or between_tools for no upfront thinking at Low/Medium/High effort HTTP 400
Forced tool_choice: any/tool Use auto, a strict tool schema, and an explicit instruction for when to call it HTTP 400 or changed tool behavior
Reading only content[0].text Process text, thinking, and tool-use blocks by type Empty UI or missing progress
Legacy computer-use tool Move to the new toolset on Claude API/Google Cloud; check Bedrock separately HTTP 400 or broken agent loop
Edited conversation history Keep conversations append-only and test cross-model thinking blocks Inconsistent context
Existing advisor model Confirm that executor/advisor pairing is supported Failed request

Do not treat these field names as universal gateway configuration: platforms and SDK versions differ. Smoke-test a text-only call, a tool call, a multi-turn conversation, and a failed/retried call. Save request settings, response-block types, and error codes before changing the route for the 20–30 real trial tasks.

Make knowledge-work output auditable

For reports, slides, and spreadsheets, define the audience, source files, allowed conclusions, mandatory citations, and prohibited extrapolations. Ask the model to list unanswered questions before drafting. Good layout does not verify numbers.

Calculate the cost of an accepted result

Anthropic lists Sonnet 5.5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens. Opus 5.5 is listed at $4, $20, and $0.20 respectively. These are billing rates, not the full cost of a completed task. For an illustration, not a measured result, a job using one million input tokens and 200,000 output tokens costs about $4 on Sonnet and $8 on Opus before cache, tools, and retries. If the Sonnet output requires 15 extra minutes of human repair, charge that time back to the task.

Metric Record for every trial Why it matters
First-pass acceptance accepted tasks / total tasks Detects cheaper calls that create rework
Accepted-delivery cost model + tools + human repair Aligns comparison with actual delivery
P50 and P95 elapsed time request to accepted result Captures slow outliers
High-severity failure unauthorized action, fabricated evidence, data exposure Can stop a route regardless of average cost

Anthropic’s “30%+ faster” and “up to 30% lower cost per task” are results under its testing conditions, not a discount every team will realize. Use the same task set, tool permissions, acceptance tests, and reviewers for both paths. Changing the prompt and tooling at the same time should be recorded as a separate experiment.

Define rollback before enabling the route

Set thresholds by task class rather than one global switch. Example rules: revert a class if first-pass acceptance drops more than five percentage points against the old route, if P95 completion time grows by more than 20%, or if an unauthorized tool action occurs. Those numbers are illustrative and should reflect your team’s risk tolerance. Keep the prior model configuration, prompt version, tool grants, and test cases so rollback is a routing change, not an emergency rebuild.

At the end of the week, report each class in one row: sample size, effort setting, acceptance rate, mean repair minutes, total accepted-delivery cost, and failure modes that still need Opus or human review. That record remains useful at the next model release.

Release checklist

  • [ ] Work is segmented by clarity and error cost.
  • [ ] Every task class has acceptance criteria.
  • [ ] Low, Medium, and High effort each have real samples.
  • [ ] Logs include tokens, cache use, latency, tool calls, retries, and rework.
  • [ ] Production changes have tests, review, and a rollback path.
  • [ ] Model IDs and data-retention settings are confirmed in the console.

Sonnet 5.5 is most useful when it moves well-bounded work into an accepted state faster. Record task, effort, and acceptance outcome together before deciding that a lower token bill is a lower delivery cost.

Primary sources

Related Articles