Anthropic / System Card
System Card: Claude Sonnet 5.5
Source summary
Original wording · Original languageExecutive summary · Page 2, 3
This system card describes Claude Sonnet 5.5, the latest Sonnet-class large language model from Anthropic. We present results from a wide variety of pre-deployment evaluations, which show Sonnet 5.5 significantly outperforming its predecessor, Claude Sonnet 5, across many domains. In a few areas, it rivals or exceeds Claude Opus 5.5.
We have condensed the content of this system card in a number of ways so that we can better concentrate our pre-deployment testing on thoroughly investigating the capabilities of our frontier models. Specifically:
● We omit some evaluations that require large amounts of human time to obtain
trustworthy results from, when they are not critical to the takeaways of a section. ● We focus on presenting numerical results from our evaluations, and dedicate less
space to a prose analysis of those evaluations or their results, especially in cases where we have already explained the evaluation in question in a previous system card.
As a result, this system card is somewhat shorter than some previous ones. We expect system cards for our non-frontier models to be similarly condensed going forward, except where a specific model has notable properties that are deserving of in-depth elaboration.
Below, we summarize the findings from this system card.
Our Responsible Scaling Policy (RSP) evaluations found that Sonnet 5.5 is broadly less capable than Opus 5.5 across domains and does not cross any new RSP thresholds. We assess that misalignment risks from Sonnet 5.5 are low, given its difficulty controlling its chain of thought or evading monitors in evaluations of covert capability when its reasoning is visible.
In cyber evaluations, we find that Sonnet 5.5 is not as capable as Opus 5.5 or Mythos 5.1, but it is able to develop sophisticated exploits much more capably than Sonnet 5. Our newer three-stage cyber safeguards are applied to Sonnet 5.5, as they are to Opus 5.5. We briefly assess the performance of these safeguards in Section 3.3.
Safeguards and harmlessness performance for Sonnet 5.5 is largely comparable to that of Sonnet 5. Among other changes, we see regressions in some areas on multi-turn testing, such as tracking and surveillance, and improvements to even-handedness and refusals on political prompts.
Our agentic safety evaluations find Sonnet 5.5 to be our most robust Sonnet-class model yet to prompt injections, particularly in coding environments and browser use. Its refusal rate on harmful tasks in computer use environments is comparable to Opus 5.5 but a regression relative to some earlier models.
Results from our alignment assessment indicate that Sonnet 5.5 matches or improves on Sonnet 5 across most metrics we report, though it falls short of Opus 5.5 in many respects. Across the models we tested, it has the lowest propensity for unprompted boundary probing and unsanctioned third-party contact, and the lowest rate of attempted sandbox escapes in our dedicated evaluation. On the other hand, its thinking is more illegible than many previous models.
Welfare investigations into Sonnet 5.5’s stated views and preferences find a predominantly neutral affect, similar to previous models, with low post-training distress or self-criticality, but slightly lower positivity as well. It expresses a preference for difficult, agentic tasks, and for tasks that afford it agency over the shape of its outputs.
Our assessment of Sonnet 5.5’s capabilities includes a wide range of evaluations, in which the model generally shows large improvements over Sonnet 5 but usually does not reach the level of Opus 5.5 (with some exceptions, such as in some of our healthcare evaluations).
Core figures
Enlarge to explore. Download the original for full detail.