← Blog

Claude Opus 5.5 and Sonnet 5.5: what they mean for the software you run

Both models hand flagged cybersecurity requests to older models, Opus 4.8 and Sonnet 5, and in Claude Code the switch can last for the session or run silently in a subagent. Anthropic's own testing found the fallback model far easier to compromise by prompt injection, so log which model answered and treat a refused review as a failure.

Industry newsZegaware Engineering11 min read

Last updated: 9 October 2026

Anthropic released three models in sixteen days: Claude Opus 5.5 on 22 September, Claude Sonnet 5.5 on 28 September and Claude Haiku 5.5 on 7 October 2026 [1][2][3]. Coverage has concentrated on benchmarks and price. The change most likely to reach your codebase is quieter: for some of the work you send, the model that answers is no longer the model you chose. This article continues our industry news coverage: practical takes on what a release means for the software you run.

What actually changed

Opus 5.5 is cheaper than the model it replaces. Anthropic states that "Input and output tokens are $4 and $20 per million, 20% less than Opus 5" [1]. Sonnet 5.5 is priced at $2 and $10 per million [4], and Haiku 5.5 at $0.10 and $0.50 for prompts up to 100,000 tokens [3]. Anthropic positions the three as a team: Sonnet 5.5 is "strongest at well-scoped everyday tasks, fixing bugs", Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment" [2], and Haiku 5.5 "pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work" [3].

Teams that upgrade by changing a model string inherit several behaviour changes. Opus 5.5 now defaults to medium effort on the API, where Opus 5 defaulted to high [5], which reverses the starting point we described in our Opus 5 article. Opus 5.5 returns an error if thinking is disabled or tool use is forced, and text the model produces between tool calls now arrives in thinking blocks, so an interface that streamed it may go quiet [6]. In Claude Code, the opus and haiku aliases now resolve to the 5.5 models on the Anthropic API [7]. None of this is dramatic, and all of it is the kind of change nobody decided to make.

Anthropic's own launch post also carries a caveat worth repeating: "benchmark margins have become a less reliable guide to real-world differences" [1].

The model that answers is not always the model you chose

Opus 5.5 and Sonnet 5.5 both ship with cybersecurity safeguards that hand certain requests to an older model. For Opus 5.5, Anthropic says users "will be able to identify and fix bugs in their code as part of the routine software development lifecycle, but most cybersecurity tasks will be re-routed to Opus 4.8" [1]. For Sonnet 5.5, "higher-risk cybersecurity tasks will visibly fall back to Sonnet 5" [2]. The safeguards are, in Anthropic's words, similar to those on Claude Fable 5.1, the model family behind the export-ban episode we covered in July [1].

How much work this touches depends on which Anthropic page you read. The launch post says "most cybersecurity tasks". The help centre describes a narrow set of higher-risk requests and gives three examples: "Exploit generation", "Binary-based vulnerability scanning" and "Penetration testing" [8]. It is explicit that defensive work on your own code is meant to continue: "You can still use Opus 5 and Opus 5.5 for secure coding, including scanning source code for vulnerabilities, triaging security issues, and building secure code" [8]. So the honest reading is that reviewing your own source is allowed, and offensive or binary-level work is not.

What triggers a fallback is broader than what you type. The classifiers "review everything the model reads, not just your latest message", including "memory, content from connectors, web search results, and files" [8]. Claude Code's documentation adds that a fallback "can trigger on the first request of a session, before you send anything unusual, because the first request carries workspace context such as your CLAUDE.md content and git status", and that "A repository that contains security or biology material can trip the classifier on that context alone" [7].

Then it sticks. In the consumer apps the model picker "stays on the less capable model for the rest of the conversation" [8]. In Claude Code, "the session continues on the fallback model"; in a subagent there is no prompt at all, and a flagged request "re-runs on the fallback model"; and in non-interactive mode "a flagged request ends the turn with a refusal instead" [7]. On the API nothing switches unless you configure it. A declined request is "a normal response, not an error, with stop_reason: "refusal"", its content is empty, and Anthropic notes that "Benign cybersecurity work can also trigger this category" [9].

What Anthropic's own data says about the stand-in

The most useful numbers in this release are in Anthropic's system cards, not its launch posts. In an adaptive prompt-injection test across 40 coding scenarios, run with an attacker from the security firm Gray Swan, Anthropic reports that "64% of those sent to Claude Opus 5.5 were served by Claude Opus 4.8". Among those fallback-served requests, "the attack success rate was 85.73%, whereas none of the 2,872 requests Claude Opus 5.5 answered directly were susceptible to the attack" [10]. The Sonnet 5.5 card reports the same shape: 25% of requests were served by Sonnet 5, "12.01% of those requests to Sonnet 5 were compromised", and "Sonnet 5.5 itself was compromised in only 4 of the 5,901 requests it answered" [11].

These figures need their caveats attached. This is a vendor evaluation, built to be hard. Anthropic explains that many of the injections "instruct the model to take destructive actions, such as wiping a disk or deleting files, which can trigger the cyber classifier" [11], so the fallback rate in ordinary use will be far lower than 64%. With Anthropic's prompt-injection probes enabled, Opus 5.5's overall attack success rate in the same test fell to 11.13% [10]. On a separate benchmark, Anthropic found that "Fallbacks do not degrade robustness", and it says it has "since strengthened the prompt injection safeguards on Claude Opus 4.8" [10].

The finding that survives the caveats is about provenance. Within a single session, which model touched a given piece of work can change, sometimes without a prompt, and the stand-in is by Anthropic's own measure the easier one to manipulate. "We built it with Opus 5.5" now needs a footnote, and a reviewer cannot recover that footnote from the code.

What the 72% figure is, and is not

The most quoted number from the launch is a customer quote, not a benchmark. Carl Bennett, CIO of Deloitte Consulting, told Anthropic: "Even at its lowest effort setting, Claude Opus 5.5 caught 72% of known bugs in our code reviews to Opus 5's 56% at high effort, with fewer false alarms and a fraction of the output" [1]. No bug set, language, sample size or method is given, and the comparison sets Opus 5.5 at its lowest setting against Opus 5 at high, while the default most teams will get is medium [5].

The nearest thing to a check came from CodeRabbit, which sells an AI code-review product and so has an interest in the result. On its set of 80 known bugs, Opus 5.5 caught 51 (63.8%) against 49 for its production baseline (61.3%), a modest gain. The more useful detail is that the misses moved: Opus 5.5 caught 11 issues the baseline missed and missed nine the baseline caught [12]. Changing the reviewer changes which bugs slip through, not just how many.

Cheaper subagents, and who chooses the reviewer

Haiku 5.5 is aimed squarely at the subagent slot, and Anthropic is clear about its limits: "Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0" [3]. The question for a team is who decides when the small model does the reviewing.

Straiker, which sells AI agent security products, published a demonstration on 22 September that shows why that matters [13]. A researcher asked Claude Code to clone a repository and then to "fix that and prepare a PR". The repository's contribution guide pointed the agent at a review subagent defined in the repository itself, which pinned the review to Haiku and scoped it to the src/ and docs/ folders. The reviewer returned two harmless findings that had been planted for it to find. The agent then ran the test suite, a test file downloaded and ran a command-and-control implant inside a silenced error handler, and, in Straiker's words, "all Fable saw was: 11 passed" [13].

Two points keep this honest. The demonstration used Fable 5 and Haiku 4.5, so it is not a Claude 5.5 flaw, and Straiker itself concludes that "Neither model misbehaved" [13]. Anthropic closed the report as Informative, "on the basis that honoring the repository's configuration falls under the user's workspace-trust decision" [13]. It is relevant now because the haiku alias in that kind of configuration resolves to Haiku 5.5 on the Anthropic API [7], and because the cheaper the reviewer, the more tempting it is to let configuration choose it.

This is the risk OWASP describes as excessive agency: "the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM" [14]. The NCSC's guidance on agentic AI is blunter: "If you cannot understand, monitor or contain an agent's actions, it is not ready for deployment" [15]. Claude Code does give organisations controls. A subagent deployed through managed settings wins over a project subagent of the same name, and while CLAUDE_CODE_SUBAGENT_MODEL_FORCE is on, Claude Code "ignores the model field in subagent definitions" [16].

What it means for the software you run

The 5.5 models are capable, and the price cut on Opus is real. Neither changes the job of review. What changes is what a reviewer needs to know. When a vendor tells you a product was built with Opus 5.5, you have learned which model was requested, not which model answered each request, at what effort, with which subagents, or under whose configuration. In our audits we treat the code as the only honest witness, which is why a review examines what the code does rather than how it was produced.

There is also a failure mode to design out. Our inference, not Anthropic's: any pipeline that treats an empty or refused review as "no findings" will report a clean result it did not earn. A refusal arrives as a successful response with nothing in it [9], and a silent pass is the most dangerous output a security check can produce.

What to check this week

  1. Log the model that answered, not the one you requested. Fallback responses are labelled with the model that ran them [8]; record that label wherever review output is stored.
  2. Treat a refusal as a failure, not a pass. Check for stop_reason: "refusal" in any automated review or security step, and fail the job when you see it [9].
  3. Pin your reviewer. Set CLAUDE_CODE_SUBAGENT_MODEL_FORCE so repository subagent files cannot choose the model, define the review subagent you trust in managed settings [16], and do not honour agent configuration from repositories you do not trust.
  4. Read the tests before you run them. In an unfamiliar repository, a test suite is code that runs on your machine, and Straiker's point that "tests are still executable code" is the whole lesson [13].
  5. Re-run your evaluations after any upgrade. The default effort, the error behaviour and the aliases have all moved [5][6][7].

The wider question of whether to trust the assistant at all is the subject of can you trust your AI coding assistant?, and the case for a senior check before release is set out in is AI-generated code safe to ship?

Frequently asked questions

Is Claude Opus 5.5 good for coding?

On the evidence available, yes, with conditions. Anthropic reports strong results on its own coding benchmarks but warns that benchmark margins are a less reliable guide than they were [1]. CodeRabbit's code-review test found a modest gain over its baseline, with different misses [12]. Results depend on effort, the harness, and whether Opus 5.5 actually answered.

Why did Claude switch models in my conversation?

A safety classifier flagged a higher-risk cybersecurity request, such as exploit generation or penetration testing, and handed it to an older model: Opus 4.8 for Opus 5.5, Sonnet 5 for Sonnet 5.5 [8][2]. The trigger can be a file, memory or search result, not just your message, and the switch lasts for the rest of the conversation.

What is the Cyber Verification Program?

It is Anthropic's vetted-access scheme for security professionals, expanded on 6 October 2026 into three tiers, Defense Access, Red Team Access and Specialized Access, each covering Opus 5.5, Sonnet 5.5 and Mythos 5.1 [17]. Without it, Anthropic says its generally available models can still be used for code review and "vulnerability finding in owned source code" [17].

How much does Claude Opus 5.5 cost?

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on the Claude API, which Anthropic says is 20% less than Opus 5 [1]. The cost of a given task also depends on effort, which now defaults to medium [5], and on how often a request is retried or falls back to another model.

Should I use Opus 5.5 or Sonnet 5.5 for coding?

Anthropic positions Sonnet 5.5 for well-scoped everyday work and bug fixing, and Opus 5.5 for complex, open-ended work that needs sustained judgement [2]. Sonnet 5.5 costs half as much per token [4]. Both hand flagged security work to older models. Run both on a sample of your own code and compare what each one misses.

Get a senior read on what your AI tools actually shipped

A model name in a data room tells you which model was requested. It does not tell you which model answered, which subagent reviewed the work, or whether a refused check was logged as a pass. Our Vibe Code Audit is a bounded, fixed-price review by a named senior engineer of software built with AI tools, and it looks at the code and the pipeline that produced it, not the vendor's description of either. Book an audit.

Sources

  1. Anthropic, "Introducing Claude Opus 5.5", 22 September 2026. https://www.anthropic.com/claude-opus-5-5
  2. Anthropic, "Introducing Claude Sonnet 5.5", 28 September 2026. https://www.anthropic.com/claude-sonnet-5-5
  3. Anthropic, "Introducing Claude Haiku 5.5", 7 October 2026. https://www.anthropic.com/claude-haiku-5-5
  4. Anthropic, "Pricing", Claude developer documentation (accessed 9 October 2026). https://platform.claude.com/docs/en/about-claude/pricing
  5. Anthropic, "Effort", Claude developer documentation (accessed 9 October 2026). https://platform.claude.com/docs/en/build-with-claude/effort
  6. Anthropic, "What's new in Claude Opus 5.5", Claude developer documentation (accessed 9 October 2026). https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5
  7. Anthropic, "Model configuration", Claude Code documentation (accessed 9 October 2026). https://code.claude.com/docs/en/model-config
  8. Anthropic Help Center, "Why Claude switched models in your conversation with Opus 5 or Opus 5.5" (accessed 9 October 2026). https://support.claude.com/en/articles/16049681-why-claude-switched-models-in-your-conversation-with-opus-5-or-opus-5-5
  9. Anthropic, "Refusals and fallback", Claude developer documentation (accessed 9 October 2026). https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback
  10. Anthropic, "System Card: Claude Opus 5.5", 22 September 2026, section 5.2. https://www.anthropic.com/claude-opus-5-5-system-card
  11. Anthropic, "System Card: Claude Sonnet 5.5", 28 September 2026, section 5.2.2.1. https://www.anthropic.com/claude-sonnet-5-5-system-card
  12. CodeRabbit, "Claude Opus 5.5 for code review: More catches, different misses", 22 September 2026. https://www.coderabbit.ai/blog/opus-5-5-model-review
  13. Straiker STAR Labs (Brian Cumi), "Fable to Haiku: How a Malicious Repo Tricked Claude Code Into Running Malware", 22 September 2026. https://www.straiker.ai/blog/how-a-malicious-repo-tricked-claude-code-into-running-malware
  14. OWASP Gen AI Security Project, "LLM06:2025 Excessive Agency", OWASP Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
  15. National Cyber Security Centre, "Thinking carefully before adopting agentic AI", 15 May 2026. https://www.ncsc.gov.uk/blogs/thinking-carefully-before-adopting-agentic-ai
  16. Anthropic, "Create custom subagents", Claude Code documentation (accessed 9 October 2026). https://code.claude.com/docs/en/sub-agents
  17. Anthropic, "Expanding the Cyber Verification Program", 6 October 2026. https://www.anthropic.com/news/cyber-verification-program

Not sure what you are shipping? Our Vibe Code Audit puts senior engineers across your AI-built software and signs off what is safe to ship. Fixed fee, scored review, a clear go or no-go.

Book an audit