Migrating to Claude Sonnet 5.5: Thinking, Tools and API Errors
Migrate an integration to Claude Sonnet 5.5: check thinking, forced tools, HTTP 400 errors, streaming, trial costs and a practical rollback plan.
Syntalith
Changing a model identifier can produce an HTTP 400 error or an interface that goes quiet between tool calls. Sonnet 5.5 changes request configuration and response handling. Check the API client, conversation history and streaming before moving production traffic.
The model was released on 28 September 2026. Documentation checked on 1 October 2026 lists claude-sonnet-5-5 for the Claude API and describes incompatibilities with earlier integrations. Sonnet 5.5 overview.
Inspect the affected parts
| Area | Application check |
|---|---|
| Sampling parameters | Non-default temperature, top_p or top_k returns HTTP 400 |
| Thinking | Disabling up-front thinking uses between_tools |
| Forced tools | Earlier forced-tool requests can be incompatible |
| History | Thinking blocks are tied to their model and conversation |
| Computer use | computer_20251124 is not accepted on Claude API and Google Cloud |
| Streaming | Text between tools can arrive as thinking blocks the UI does not display |
The overview also documents restrictions on older advisor models. Check these if your integration uses that tool. Platform-specific identifiers and options must be verified for Claude API, Bedrock or the other platform you use.
Inspect the actual request
In a test environment, inspect fields sent by the application, including defaults added by a library or intermediary. Keep keys and full customer data out of the diagnostic report. Parameter names, version, HTTP status and a sanitised error are often sufficient.
If a shared client sets temperature for every model, changing the identifier will not remove that field. Make configuration depend on the selected model's capabilities, then verify it with a real test request.
A quiet interface may still be working
Text between tool calls now has different response handling. An interface listening only for ordinary text blocks may stop displaying progress. The documentation points to display configuration or between_tools, depending on the intended behaviour.
Decide which progress messages the application should expose rather than forwarding raw events without inspection. Test interruption and conversation resumption too. History from another model needs compatible handling; thinking blocks cannot simply be copied between arbitrary models and conversations.
Test before changing production traffic
Include ordinary requests and exceptions: an empty tool result, malformed output, a long history and an interruption after an operation. Measure correctness, duration and usage together. Define acceptance criteria before running the comparison.
Keep the prior configuration and a rollback path. Do not run both variants against the same production write tools, which could duplicate operations. Use copied data and test tools. Reverting the model does not undo a payment or sent message.
Budget for the trial
Standard Claude API prices are $2 per million input tokens and $10 per million output tokens. One million input tokens plus 200,000 output tokens would cost 2 + 0.2 × 10 = $4, before tools, retries and other charges. Cache and Batch have separate pricing rules.
Implementation cost also includes client changes, evaluation and monitoring. A vendor's task-cost improvement is not a savings guarantee for your application. See the OpenAI and Claude pricing comparison for broader calculations.
Syntalith helps evaluate and develop AI integrations. A process audit starts at €600 excluding VAT; migration work is quoted from the application and required checks. Describe your integration and issue so we can scope an evaluation against your system's data and rules.
Syntalith is an OpenAI Select Partner in the OpenAI Partner Network.
We help your team respond to customers faster and find information in company documents. We choose and set up the right AI tools, then teach your team how to use them.
How to introduce ChatGPT and OpenAI at workFree process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary: a possible direction, missing information and the next step.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form