Anthropic · all models ›

AClaude Opus 4API

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows.

200KContext window · tokens
$15Input price · per M tokens
$75Output price · per M tokens
AnthropicProvider

Prices update automatically — checked daily against provider list prices.

See open-source alternatives → Compare all model prices

Benchmarks & performance

Independent benchmark scores for Claude Opus 4, measured by Artificial Analysis. Higher is better. Measurement mode: Reasoning.

Intelligence index31.7
Math index73.3
GPQA79.6%
MMLU-Pro87.3%
Humanity's Last Exam12.3%
Long Context Reasoning36.3%
LiveCodeBench63.6%
SciCode39.8%
MATH-50098.2%
AIME75.7%
AIME 202573.3%
IFBench53.7%
τ²-Bench73.4%
Terminal-Bench Hard31.1%
💰 Blended price$30 / 1M tokens
📈 Value1.1 intelligence points per $
vs. models measured heretop 45%
Scores higher than 55% of the 273 models measured by Artificial Analysis and tracked here.
Models at this level cost $1 per 1M tokens (median of 20) — this one costs $30.
Cheaper and better on this index: Claude Opus 5 (Fast) · Claude Opus 5 · Claude Opus 5 (batch) and 110 more
Benchmark data by Artificial Analysis

About this model

Claude Opus 4 is a commercial AI model by Anthropic. The specifications below are tracked automatically: pricing is refreshed daily from public list prices, so the numbers on this page reflect the current cost of using the model through its API.

Frequently asked questions

What is Claude Opus 4?

Claude Opus 4 is an AI language model from Anthropic. It is a proprietary model, available through an API. It scores 31.7 on the Artificial Analysis intelligence index.

Is Claude Opus 4 free?

Claude Opus 4 is not free: it costs $15 per million input tokens and $75 per million output tokens. Open-weight alternatives can be self-hosted at no per-token cost.

What is Claude Opus 4 good at?

Independent benchmarks from Artificial Analysis give it GPQA 79.6%, MMLU-Pro 87.3%, Humanity's Last Exam 12.3%, Long Context Reasoning 36.3%, LiveCodeBench 63.6%, SciCode 39.8%, MATH-500 98.2%, AIME 75.7%, AIME 2025 73.3%, IFBench 53.7%, τ²-Bench 73.4%, Terminal-Bench Hard 31.1%. It is particularly used for mathematical reasoning.

Related models

Claude Opus 4.7 (Fast)AnthropicClaude Opus 4.1AnthropicClaude Fable 5AnthropicClaude Opus 5 (Fast)AnthropicClaude Opus 4.8 (Fast)AnthropicClaude Opus 4.1 (batch)Anthropic