Anthropic recently officially launched the latest lightweight model in the Claude 5.5 series, Claude Haiku 5.5. For regular requests within 100,000 tokens, the input and output prices have been reduced to $0.10 and $0.50 per million tokens, respectively, representing a 90% drop in listed prices compared to the previous generation.

As the fastest responding model currently available, Haiku 5.5 focuses on high-concurrency, low-latency, and cost-effective scenarios, making it highly suitable for frequent tasks such as real-time customer service, summary extraction, and data classification. Additionally, the system has introduced adaptive thinking and dynamic adjustment features for the first time, which can automatically balance reasoning depth and computing costs based on task complexity.

image.png

New Tokenization Mechanism Attracts Attention, Actual Cost May Increase

However, the official also warned that due to Haiku 5.5 using the same new tokenizer as advanced models, the number of tokens generated from the same text input has increased. Testing shows that long-text tasks may consume 25% to 30% more tokens, so the overall cost reduction after considering actual workloads is approximately 75%.

It should be noted that once a single request exceeds the 100,000-token threshold, both the input and output unit prices will surge to five times the base price. This means developers must accurately assess token consumption for specific workflows while enjoying the low-unit-price advantage, to avoid falling into hidden cost traps.