Anthropic brought its lightest and most powerful member of the family to the forefront on October 8th - Claude Haiku 5.5 was officially launched. The official positioning is clear: this is the fastest and most capable small model from the company, specifically designed to tackle high-frequency, cost-sensitive scenarios, such as article summaries, content classification, database queries, customer service responses, browser automation operations, and sub-tasks that large AI agents break down when doing heavy work.

This small model's approach is very aggressive. The average operating cost of Haiku 5.5 is reduced by about 75% compared to the previous generation, Haiku 4.5. For requests within 100,000 token units, the input and output costs are only $0.1 and $0.5 per million tokens, respectively - meaning that the money previously spent on a batch of lightweight tasks now almost just a fraction. More importantly, for the first time, the Haiku series has included adjustable reasoning intensity: users no longer need to pay full computing power for every sentence, but can adjust the depth of thinking according to the weight of the task, allowing the affordable small model to think a few more steps when needed.

The price lever of its neighbor Sonnet 5.5 was also adjusted. Anthropic halved the cache reading price of Sonnet 5.5, reducing it to $0.1 per million tokens. According to its statement, this will reduce the operational cost of most of its agent tasks by about 20%. Cache reading is the most frequently called part when the agent system repeatedly reviews long contexts. This cut effectively frees up those autonomous agents running in the background all day long.