Claude Haiku 5.5 API

Claude Haiku 5.5 launched on 2026-10-07 as Anthropic’s cheapest and fastest model and the first Haiku with effort levels. It is priced by prompt length: 90% cheaper than Haiku 4.5 up to 100K tokens, and 5x per token for the whole request once the prompt passes 100K. This page covers how the price works, the official benchmarks, how to choose between it, GPT-6 Luna and Sonnet 5.5, community reviews, the 400 errors you will hit migrating from Haiku 4.5, and how to turn it on in Claude Code.

Price check: official, OpenRouter and Wokey

USD per million tokens. Official and OpenRouter rates were checked by hand on 2026-10-08, excluding tax; the Wokey rate is read from the live price list.

ChannelInputCache readOutput
Official API$0.1$0.01$0.5
OpenRouter$0.1$0.01$0.5
Wokey$0.02$0.002$0.1

The table shows prompts up to 100,000 tokens. When the prompt (input + cache reads + cache writes) exceeds 100,000, the official API and OpenRouter bill the whole request at $0.5 input / $2.5 output; Wokey also prices by prompt length, with the long-prompt rates in the pricing table below.

· Official pricing page · OpenRouter endpoint list

Use it in Claude Code

This model supports /v1/messages on Wokey. Set these variables, then start claude.

export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="$WOKEY_API_KEY"
export ANTHROPIC_MODEL="claude-haiku-5-5"

Claude Haiku 5.5 at a glance

Release date
2026-10-07, the last Claude 5.5 model (Opus 5.5 shipped 09-22 and Sonnet 5.5 09-28). From Claude Code 2.1.293 the haiku alias points to it on the Anthropic API
Positioning
High-volume, latency-sensitive work (classification, extraction, routing, summaries, compaction) and sub-agents under Opus 5.5 or Sonnet 5.5. Anthropic calls it its fastest model, about 75% cheaper to run than Haiku 4.5 on average
Official price (tiered)
Prompts up to 100K tokens: $0.10 input / $0.50 output, cache read $0.01, 5-minute cache write $0.125, 1-hour $0.20. Prompts over 100K: every meter 5x ($0.50 / $2.50). Prompt = input + cache reads + cache writes; Batch is a further 50% off
Context and output
1M context with no beta header; 128K max output, or 300K on the Batch API with the output-300k-2026-03-24 beta header
Thinking and effort
Adaptive thinking on by default, default effort medium on the API and in Claude Code, five levels low / medium / high / xhigh / max; thinking text is omitted by default
Tokenizer
Same as Claude 4.7 and later: the same text counts about 30% more tokens than on Haiku 4.5, so recount token-based costs and max_tokens
Model ID
claude-haiku-5-5 (no date suffix and no separate alias; anthropic.claude-haiku-5-5 on Bedrock)

· Sources: Anthropic · Migration guide

Official benchmarks: Haiku 5.5 vs Haiku 4.5 vs GPT-6 Luna vs Sonnet 5.5

From Anthropic’s Haiku 5.5 launch page; these are vendor-reported results. “—” means no score was published. The Sonnet 5.5 column is Anthropic’s reference; its FrontierCode score is at xhigh.

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1162073514371840
AA-Briefcase v1.1157861413361824
OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%
Humanity's Last Exam (no tools)45.9%10.2%—56.9%
Humanity's Last Exam (tools)57.4%18.7%—64.5%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
FrontierCode 1.1 Main46.4%—42.4%52.1%
Chartography (no tools)46.4%6.4%29.1%61.6%

On Terminal-Bench 4.0 Haiku 5.5 ran in Claude Code and GPT-6 Luna in Codex (public leaderboard), so the two are not a like-for-like comparison. OSWorld here is the offline subset and does not match the Sonnet 5.5 launch figure. In Artificial Analysis’ independent Intelligence Index (v4.3.2) Haiku 5.5 scores 43 at max (xhigh 41, high 38, medium 34, low 29), versus 17 for Haiku 4.5, 56 for Sonnet 5.5 and 58 for Opus 5.5; several reports cite GPT-6 Luna (max) at 38 on the same index.

Which to use: Haiku 5.5, GPT-6 Luna or Sonnet 5.5

vs Haiku 4.5

Same role, drop-in: 90% cheaper per token up to 100K and 50% cheaper above, and ahead on every published benchmark (Terminal-Bench 4.0 from 0% to 39.2%, OSWorld from 15.7% to 72.4%). Check the migration list below first, and recount tokens: the same text is about 30% more tokens with the new tokenizer.

vs GPT-6 Luna

Up to 100K tokens both list at $0.10 / $0.50, and Haiku 5.5 scores higher on the official benchmarks and the Artificial Analysis index (43 vs 38). But Luna’s long-context surcharge starts at 272K and is only 2x input / 1.5x output, so Luna is clearly cheaper on long prompts, and community tests saw Haiku 5.5 use more tokens at high effort. Pick Haiku 5.5 when most prompts stay under 100K and quality matters; pick Luna for long context or the lowest unit cost.

vs Sonnet 5.5

Use Sonnet 5.5 for complex agentic coding (Terminal-Bench 70.6% vs 39.2%); hand well-scoped lookups, summaries, classification, codebase exploration and sub-agent work to Haiku 5.5 at about 1/20 of Sonnet 5.5’s list price. A common setup runs Opus 5.5 or Sonnet 5.5 in the main session with Haiku 5.5 sub-agents.

Choosing effort

The default medium fits most work; use low for simple classification or extraction, where Artificial Analysis measured 6.3 seconds to the first answer token and about $0.02 per index task. Max averages about $0.21 per task with very high output, and community runs saw max take tens of times longer than low. If you need xhigh or max, try Sonnet 5.5 at low or medium first.

Effort documentation

Third-party and community reviews

  • Artificial Analysis: Intelligence Index (v4.3.2) 43 at max, 26 points above Haiku 4.5; about 243 output tokens/s at max. An index task averaged $0.21 at max and $0.02 at low at official prices; at max the whole index produced 440M output tokens, which is verbose. Every Wokey meter is 20% of the official rate, so the same max-effort task comes to about $0.042.
  • Hacker News (638 points, 321 comments): Positive: finally priced like GPT-6 Luna (Simon Willison noted Haiku 4.5 cost 10x Luna), noticeably smarter than Luna, and good enough to run as a sub-agent. Critical: the 100K cutoff is low for agentic sessions; the new tokenizer adds about 30% tokens; by Artificial Analysis’ cost per task Haiku is still roughly 3x Luna; quality drops at low, and xhigh to max is a large jump in time and tokens.

Migrating from Haiku 4.5: requests that return 400

These changes come from Anthropic’s migration guide. Check your requests before switching the model to claude-haiku-5-5. Through Wokey the gateway handles some of them; each item says which.

  • thinking: {"type": "enabled", "budget_tokens": N} returns 400; use {"type": "adaptive"} with effort. thinking: {"type": "disabled"} only works at high effort or below. Wokey drops manual budget_tokens and drops disabled at xhigh / max, falling back to default adaptive thinking.
  • Non-default temperature, top_p or top_k returns 400 (temperature must be 1, top_p must be 0.99, any top_k fails). Wokey removes all three automatically.
  • An assistant prefill (messages ending with an assistant turn) returns 400, even with thinking off. End with a user turn and use structured outputs or the system prompt instead.
  • The computer_20250124 computer use tool returns 400; use the computer_toolset_20260801 toolset.
  • Thinking blocks replay only in the account that produced them and are silently dropped elsewhere; sending them back after editing system, tools or earlier messages returns 400, so keep conversations append-only.
  • Responses can start with a thinking block whose text is omitted by default, so select blocks by type; thinking counts toward max_tokens, so raise small limits.
  • Forced tool_choice (any or a named tool) works but the response starts with the tool call and no thinking; refusals return stop_reason: "refusal"; Priority Tier is not supported.

Run haiku and sub-agents on Haiku 5.5 in Claude Code (Wokey)

export ANTHROPIC_BASE_URL="https://api.wokey.ai"
export ANTHROPIC_AUTH_TOKEN="<your Wokey API key>"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="claude-haiku-5-5"
# Optional: run sub-agents that do not name a model on Haiku 5.5
export CLAUDE_CODE_SUBAGENT_MODEL="haiku"
claude
Model ID
claude-haiku-5-5
Vendors
anthropic
Supply status
Currently callable

Catalog pricing snapshot

Prices are per 1M tokens; request-time pricing controls settlement.

MeterWokeyOfficial referenceSavings
Input$0.02$0.180%
Output$0.1$0.580%
Cache read$0.002$0.0180%
Cache write$0.04$0.280%

Prompts over 100K tokens

Prompt = input + cache reads + cache writes. When the prompt exceeds 100,000 tokens, the whole request, output and cache included, bills at this table; the table above covers prompts up to 100K.

MeterWokeyOfficial referenceSavings
Input$0.1$0.580%
Output$0.5$2.580%
Cache read$0.01$0.0580%
Cache write$0.2$180%

Official price source: https://platform.claude.com/docs/en/models/haiku-5-5/overview

Model capabilities

Context
1,000,000 tokens
Max output
128,000 tokens
Streaming
Supported
Tools
Supported
Vision
Supported
Supported API forms
/v1/messages

Send a request

/v1/messages

curl https://api.wokey.ai/v1/messages \
  -H "x-api-key: $WOKEY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-haiku-5-5","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'

Use Claude Haiku 5.5 in your tools

The gateway is https://api.wokey.ai (OpenAI-compatible clients usually take https://api.wokey.ai/v1), and the model ID is claude-haiku-5-5. Open the guide for the client you use.

Frequently asked questions

How does the Claude Haiku 5.5 100K pricing tier work? Does output cost more above 100K tokens too?

The official price has two tiers by prompt length. When the prompt (input + cache reads + cache writes) is up to 100,000 tokens: $0.10 input, $0.50 output, $0.01 cache read and $0.125 5-minute cache write per million tokens. Above 100,000 tokens every meter of the whole request is 5x ($0.50 input, $2.50 output), so a 120K-token prompt also pays $2.50 for output. Anthropic says about 90% of Haiku 4.5 requests stayed under 100K. On Wokey it costs $0.02 / $0.1 input / output per million tokens for prompts up to 100K and $0.1 / $0.5 above 100K (official $0.10 / $0.50, and $0.50 / $2.50 above 100K), with the same prompt-length tiers.

Claude Haiku 5.5 vs GPT-6 Luna: which is better and which is cheaper?

Up to 100K tokens both list at $0.10 / $0.50. Haiku 5.5 leads Anthropic’s published benchmarks (GDPval-AA 1620 vs 1437, OSWorld 72.4% vs 48.9%) and the Artificial Analysis index (43 vs 38). But GPT-6 Luna’s long-context surcharge starts at 272K and is only 2x input / 1.5x output, while Haiku 5.5 multiplies every meter by 5 from 100K, and Haiku’s tokenizer counts about 30% more tokens for the same text. Choose Haiku 5.5 for short prompts where quality matters and Luna for long context or the lowest unit cost; test both on your own workload.

How good are the Claude Haiku 5.5 benchmarks compared with Haiku 4.5?

In Anthropic’s benchmarks Terminal-Bench 4.0 rose from 0% to 39.2%, OSWorld 2.1 (offline subset) from 15.7% to 72.4%, GDPval-AA from 735 to 1620, and Humanity’s Last Exam with tools from 18.7% to 57.4%. Artificial Analysis’ Intelligence Index went from 17 to 43 at max, second in its price class. Sonnet 5.5 is still well ahead on complex coding.

Should I use Claude Haiku 5.5 or Sonnet 5.5?

Use Sonnet 5.5 for complex agentic coding and tasks that need sustained judgment (Terminal-Bench 70.6% vs 39.2%). Use Haiku 5.5 for classification, extraction, summaries, compaction, codebase search and sub-agents at about 1/20 of Sonnet 5.5’s list price. Anthropic itself recommends Haiku 5.5 as a sub-agent for Opus 5.5 and Sonnet 5.5.

How do I use Claude Haiku 5.5 in Claude Code?

From Claude Code 2.1.293 the haiku alias points to Haiku 5.5 on the Anthropic API with medium effort by default; run claude update first. Behind a gateway such as Wokey, set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key and ANTHROPIC_DEFAULT_HAIKU_MODEL="claude-haiku-5-5", so background work and agents declared as haiku run on Haiku 5.5; add CLAUDE_CODE_SUBAGENT_MODEL="haiku" to send sub-agents without a model there too. Switch the main session with /model claude-haiku-5-5. Claude Code cannot turn Haiku 5.5’s thinking off.

Is there a cheaper Claude Haiku 5.5 API than the official price?

Wokey serves claude-haiku-5-5 over the native Anthropic /v1/messages API. On Wokey it costs $0.02 / $0.1 input / output per million tokens for prompts up to 100K and $0.1 / $0.5 above 100K (official $0.10 / $0.50, and $0.50 / $2.50 above 100K). With "Official verification" turned on for the API key, passthrough Messages responses carry a tee.proof signature you can check locally at /tools/verify to confirm the response is the official upstream’s original output.

How much does the Claude Haiku 5.5 API cost, and how much does it save versus the official API?

For prompts (input + cache reads + cache writes) up to 100,000 tokens, Wokey lists input at $0.02 per 1M tokens and output at $0.1 per 1M tokens; the official references are $0.1 and $0.5. That is 80% lower for input and 80% lower for output. Above 100,000 prompt tokens the whole request bills at $0.1 input and $0.5 output per 1M tokens (official $0.5 and $2.5), cache prices included. Request-time pricing controls settlement, and the vendor pricing source is linked on this page.

What is the Claude Haiku 5.5 API model ID?

The model ID for Claude Haiku 5.5 on Wokey is claude-haiku-5-5. Use it exactly as written in the request body's model field or in your client's model setting (for example ANTHROPIC_MODEL in Claude Code, or model in the Codex CLI config.toml); GET https://api.wokey.ai/v1/models also lists it.

Which API forms does Claude Haiku 5.5 support, and how do I call it through Wokey?

Claude Haiku 5.5 supports /v1/messages (Messages requests). Set the API base URL to https://api.wokey.ai, authenticate with a Wokey API key, and use the canonical model ID claude-haiku-5-5; you can start with /v1/messages. The Send a request section includes a curl example for every supported API form.

How do I use Claude Haiku 5.5 in Claude Code?

Claude Code calls Claude Haiku 5.5 over /v1/messages. Set ANTHROPIC_BASE_URL="https://api.wokey.ai", ANTHROPIC_AUTH_TOKEN to your Wokey API key, and ANTHROPIC_MODEL="claude-haiku-5-5", then start claude. Claude Desktop also uses /v1/messages; the setup guides cover both step by step.

What are the context window and maximum output for Claude Haiku 5.5?

The catalog context window is 1,000,000 tokens and the maximum output is 128,000 tokens. These are separate limits, and a request must also satisfy the upstream constraints in effect when it is processed.

How are cache reads and cache writes priced for Claude Haiku 5.5?

Cache read: Wokey lists $0.002 per 1M tokens and the official reference is $0.01 per 1M tokens. Cache write: Wokey lists $0.04 per 1M tokens and the official reference is $0.2 per 1M tokens.

How does calling Claude Haiku 5.5 through Wokey differ from OpenRouter?

Wokey lists Claude Haiku 5.5 at $0.02 input / $0.1 output per 1M tokens, against an official reference of $0.1 / $0.5. OpenRouter states it adds no markup on inference and charges its fee when you buy credits (5.5% by card). Wokey carries a curated set of models and has no per-provider routing parameters, so OpenRouter fits better if you need a wider catalog. The OpenRouter alternative page has the full comparison.

How do Claude Haiku 5.5 and Claude Haiku 4.5 differ in price and context?

In the published price comparison, Claude Haiku 5.5 lists input/output at $0.02 / $0.1 with a 1,000,000-token context window; Claude Haiku 4.5 lists $0.2 / $1 with a 200,000-token context window. This is a factual price-and-capacity comparison, not a quality ranking; also compare their supported API forms and capabilities.

Related models and documentation

Get API key · claude-haiku-5-5