Compare models
Put models side by side on price, context, capabilities, and benchmarks.
A July 2025 refreshed release of Qwen3 235B-A22B, Alibaba's large MoE model, with improved instruction following and general capability over the original version. Suited for demanding general-purpose, coding, and agentic workloads.
Medicine & Healthcare#86Business, Management, & Finance#94Software & IT Services#101Life, Physical, & Social Science#102
Cost rate
0.1x
Context
262K
Input
Output
Reasoning
Tool calling
Structured outputs
Pricing/M tokens
Input$0.09
Cached input$0.00
Output$0.55
Context
Context length262,144
Max output tokens16,384
Knowledge cutoff2025-06-30
Performance
Throughput59 tok/s
Latency2.65s
Claude Haiku 4.5 is Anthropic's fast, affordable small model in the Claude 4 generation, balancing strong instruction-following and coding ability with low latency. It is designed for high-throughput agentic and chat applications where speed and cost efficiency are priorities.
Document analysis#31Software & IT Services#100Mathematical#101Writing, Literature, & Language#109
Cost rate
1x
Context
200K
Input
Output
Reasoning
Tool calling
Structured outputs
Pricing/M tokens
Input$1.00
Cached input$0.10
Output$5.00
Context
Context length200,000
Max output tokens64,000
Knowledge cutoff-
Performance
Throughput90 tok/s
Latency1.04s