Claude Sonnet 5.5 warrants a trial in a business workflow, with a separate budget for max effort. In its September 28, 2026 report, Artificial Analysis records 56 points and second place on the Intelligence Index, roughly 193,000 output tokens and $7.60 per task at that setting. Some evaluations used a prerelease deployment with a structured-output bug; the authors announced reruns. These are the reported results, not our measurements. Source: Artificial Analysis.

Token prices and the cost of an agent’s work

Anthropic lists prices of $2 per million input tokens and $10 per million output tokens. It also reports task costs up to 30% below Sonnet 5 in its own testing. Source: the Claude Sonnet 5.5 launch.

Comparing these reports requires knowing the tasks and settings. A company deploying an agent needs a narrower answer: what does an accepted result cost in its process? We propose counting retries, tool calls and the time spent reviewing the output. An abandoned task still consumes a budget.

Consider an agent fixing an application bug. Its patch might pass the tests while also changing a function’s public interface. Reviewing that patch then includes inspecting the extra changes and rerunning integration tests. A completed session alone is insufficient evidence that the job is done.

Lazarus Systems commentary: match effort to the task

Our recommendation is to compare several configurations of the same model first. We would test simple data extraction, a code fix and document analysis requiring comparisons across sources as separate task groups. Each group needs an agreed error tolerance, time limit and method for checking answers before the trial begins.

For that pilot, we would select bounded tasks with reference answers. Reviewers should not know which setting produced each result. This helps reduce the temptation to rate an answer more highly because it came from the more expensive configuration.

We would retain higher effort where it improves the accepted-result rate enough to justify the additional cost and waiting time. When configurations perform similarly, repeat the runs and inspect the actual failures. One successful run is insufficient grounds for changing an entire workflow. This is our proposed evaluation approach; we have not run our own Sonnet 5.5 benchmark.

Blue particles radiating across a dark digital field.
Editorial illustration.

Migrating to Sonnet 5.5 includes checking the integration

Anthropic documents breaking changes. between_tools replaces thinking: disabled, while forcing a tool call with tool_choice: any or tool returns an error. The default effort in the Claude API is high. Source: Sonnet 5.5 migration documentation.

A migration test should therefore include cases where the model replies with text instead of performing the expected operation. In our proposed workflow, the application checks the result type and required fields before writing to the destination system. Incomplete data goes back for review. Permission to write remains a separate application decision.

Dividing work between Sonnet and local models

For a process involving internal documents, we would consider keeping retrieval and context preparation inside the company’s environment. Only approved data would reach the external model. That design requires tracing the entire flow, including prompts, logs, attachments and tool responses.

For narrowly defined tasks, include local Qwen models in the comparison. Our Gemini 4 Argon article develops the task-cost calculation further. Use the same acceptance criteria across these comparisons so that the results help select a model for a particular job.

To discuss an AI agent deployment with Lazarus Systems, prepare example tasks, expected answers and a list of systems the agent would access. Those inputs give us a basis for defining a pilot and assessing whether Sonnet 5.5 fits the process.

LAZARUS SYSTEMSLet’s talk about your project ↗