Compare models
Put models side by side on price, context, capabilities, and benchmarks.
Qwen3.7 Max is Alibaba's flagship API-only model (no open weights) in the Qwen3.7 generation, positioned as the text flagship for advanced reasoning and agentic tasks.
Cost rate
1x
Context
1M
Input
Output
Reasoning
Tool calling
Structured outputs
Pricing/M tokens
Input$1.48
Cached input$0.29
Output$4.42
Context
Context length1,000,000
Max output tokens131,072
Knowledge cutoff-
Performance
Throughput201 tok/s
Latency2.26s
Claude Haiku 4.5 is Anthropic's fast, affordable small model in the Claude 4 generation, balancing strong instruction-following and coding ability with low latency. It is designed for high-throughput agentic and chat applications where speed and cost efficiency are priorities.
Document analysis#31Software & IT Services#100Mathematical#101Writing, Literature, & Language#109
Cost rate
1x
Context
200K
Input
Output
Reasoning
Tool calling
Structured outputs
Pricing/M tokens
Input$1.00
Cached input$0.10
Output$5.00
Context
Context length200,000
Max output tokens64,000
Knowledge cutoff-
Performance
Throughput90 tok/s
Latency1.04s
Benchmarks
Head-to-head scores across reasoning, coding, and agentic indices, plus per-domain rankings.
Intelligence
62.5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.1
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
60.9