In short
The price per million tokens doesn’t show the full cost of Claude Haiku 5.5: the model counts tokens differently and becomes sharply more expensive with long context. I break down where its pricing looks attractive and why you should check the size of your requests before choosing it.
For short prompts, Claude Haiku 5.5 costs the same as GPT-6 Luna. But if the context exceeds 100,000 tokens, Haiku’s rate increases fivefold, and the tokenizer can quietly increase the bill even before that threshold.
As long as prompts stay under 100,000 tokens, both models cost the same: $0.10 per million input tokens and $0.50 per million output tokens. Simon Willison writes that Haiku 5.5 delivers higher benchmark results. However, his tokenizer comparison test showed that the same long prompt uses approximately 1.25 times more tokens than it does with Haiku 4.5.
After 100,000 tokens, Haiku’s price rises to $0.50 per million input tokens and $2.50 per million output tokens. Luna’s threshold is higher, at 272,000 tokens, after which the prices are $0.20 and $0.75. In other words, the “same as the competitor” pricing only applies to prompts that fit within the lower range and don’t expand because of tokenization differences.
In Willison’s test, Haiku with a low reasoning level generated an SVG of a pelican on a bicycle in 7 seconds and cost 0.0936 cents. The maximum level took 5 minutes 9 seconds and cost 3.3826 cents. Anthropic also added API credits for Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 for Team, shared among all team users.
There are limitations as well. Reasoning cannot be disabled in Haiku 5.5, and the medium level is selected by default. For long prompts, the pricing quickly becomes less attractive, and additional tokens can increase spending. API credits are available to the specified subscribers, but the author separately notes that for active API use, OpenAI’s Codex subscription may still be a better deal.
Do your prompts usually stay within 100,000 tokens, or do costs more often increase because of long contexts?
Source: Simon Willison's Weblog