AI Tools Projects Directory
🏆 Leaderboard 📡 News ✨ Prompts ⚖️ Compare Models 🔧 Tools 📦 Projects 🚀 Spaces
Blog
Price data

The intelligence gap is 5%. The price gap is 3×.

The best open-weight model now scores 59.7 on the Artificial Analysis intelligence index against 63.1 for the best closed one — 3.4 points apart, or 5.4%. Per point of intelligence, the median open model still costs 3.1× less. Here is where the closed premium makes sense, and where it does not.

Updated 19 August 2026·7 min read·No paywall

TL;DR — the short version

The capability race is close to a tie. The best closed model in our index, Claude Opus 5, scores 63.1 on the Artificial Analysis intelligence index. The best open one, Kimi K3, scores 59.7 — a gap of 3.4 points, or 5.4%.

The price race is not. The median closed model charges $5.00 per million output tokens; the median open-weight model charges $1.00.

Normalise by capability and the gap holds: $0.1731 per intelligence point for the median closed model against $0.0552 for the median open one — open weights are 3.1× cheaper per point.

The sharpest single trade: GLM 5.3 scores 59.5 for $4.40/M where Claude Opus 5 scores 63.1 for $25/M — 5.7× less money for 3.6 points less intelligence.

For most of the modern LLM era, the argument for paying closed-lab prices was simple: nothing else came close. Our catalogue no longer supports that argument in its old form. Across the 270 models in our index that carry both an Artificial Analysis intelligence score and a published API price — 129 closed, 141 open-weight — the distance between the best of each camp has narrowed to 3.4 points. The distance between their price lists has not followed.

A 3.4-point gap at the top

The top of the closed shelf is Claude Opus 5 at 63.1. The top of the open shelf is Moonshot AI’s Kimi K3 at 59.7. That is a 5.4% difference on the index — close enough that on many individual tasks the ordering will flip, and far enough that on the hardest ones it will not.

What has changed is the depth behind the leader. GLM 5.3 sits at 59.5, Qwen3.8 Max at 58.1, Qwen3.8 2.4T A95B at 57.7. The open column is no longer a single outlier chasing the frontier; it is a cluster parked just below it. Our leaderboard tracks the full ranking as it moves.

What a point of intelligence costs

Sticker prices tell the first half of the story. The median closed model with a published price charges $5.00 per million output tokens; the median open-weight model charges $1.00. But a cheap weak model is not a bargain, so the fairer measure divides the price by the score. On that measure, the median closed model costs $0.1731 per million output tokens for each point it earns on the intelligence index. The median open model costs $0.0552. Open weights deliver a point of intelligence for 3.1× less money.

The price of a point of intelligence Median output price — $ per million tokens Closed $5.00 Open $1.00 Median price per intelligence point — $ per million tokens, per index point Closed $0.1731 Open $0.0552 Open weights: 3.1× cheaper per intelligence point 270 models with both a benchmark score and a published price · measured 19 August 2026
Median output price: $5.00 per million tokens for closed models against $1.00 for open ones. Median price per point on the Artificial Analysis intelligence index: $0.1731 for closed models against $0.0552 for open ones — open weights are 3.1× cheaper per point. 270 models with both a benchmark score and a published price: 129 closed, 141 open.

Both medians matter, because they answer different questions. The raw median describes the shelf as it is stocked. The per-point median describes the deal you actually get — and it shows that the open discount is not an artefact of open labs shipping weaker models. Even adjusted for what the models can do, the open side charges 3.1× less.

The top of each shelf, side by side

ModelWeightsIntelligence indexOutput price ($/M tokens)
Claude Opus 5Closed63.1$25
Claude Fable 5Closed62.1$50
Kimi K3Open59.7$15
GLM 5.3Open59.5$4.40
Qwen3.8 MaxOpen58.1$6
Qwen3.8 2.4T A95BOpen57.7$6
DeepSeek V4 Pro 0423Open53.2$1.98

Closed pricing is also more layered than a single sticker. Claude Opus 5 lists at $25/M, with a Fast tier at $50 and batch processing at $12.50; Claude Fable 5 lists at $50/M with batch at $25. Batch discounts narrow the gap for offline workloads. They do not close it: even Opus 5’s batch price sits above the sticker price of every open model in the table except Kimi K3.

The open column has its own spread. Kimi K3’s $15/M is closed-shelf pricing for open weights — the price of being the open leader. A step below it, GLM 5.3 at $4.40/M and the Qwen3.8 pair at $6/M are where the open value argument actually lives, and DeepSeek V4 Pro 0423 at $1.98/M anchors the budget end at 53.2 points.

The 5.7× question

The sharpest comparison in the table is not between the leaders. It is GLM 5.3 against Claude Opus 5: 59.5 points for $4.40/M against 63.1 points for $25/M. That is 5.7× less money for 3.6 points less intelligence. Every closed-versus-open procurement decision this year is some version of that line item.

The question is no longer “which model is better?” The closed one is, by 3.4 points at the very top. The question is whether the marginal capability is worth the multiple on your workload — and that depends entirely on what the workload is. You can run this exact head-to-head for any pair of models on our model comparison pages, with the current prices from the pricing table.

When the premium is rational

The 3.4 points at the top of the index are not evenly distributed across tasks. Composite scores compress a model’s behaviour into a single number, and the places where frontier models separate from the pack — long multi-step reasoning, agentic work where an early mistake poisons everything downstream — are exactly the places a small index gap understates the difference in outcomes.

That points to a clear rule of thumb. Pay the closed premium when errors compound: agent pipelines that chain dozens of tool calls, the kind of workloads we mapped in our study of the MCP ecosystem, or high-stakes single answers where a retry is expensive. Pay it, too, when model spend is a rounding error next to the salary of the person reviewing the output. In those settings the last few points are the product, and $25/M is cheap insurance.

When it is not

Most token volume is not frontier work. Summarisation, extraction, classification, routine code, first-draft anything — these are tasks where the open cluster is already past the capability bar, and paying a multiple for points you will never notice is a subsidy, not a strategy. At $1.98/M, DeepSeek V4 Pro 0423 exists precisely for this tier of work, and the open alternatives to most commercial defaults are no longer a compromise.

Open weights also carry an option closed models structurally cannot offer: the price floor is your own hardware. Our companion piece on what actually runs on your GPU works through which of these models fit on local machines — and our local AI guide covers the tooling. The open discount quoted here is the hosted discount; self-hosting is a different economics again.

The honest counterweight is longevity. An API contract comes with a company attached; an open repository comes with a maintainer who may walk away, a pattern we measured in our study of open-source AI abandonment. Before building on an open model’s ecosystem, it is worth a glance at the project health signals we track for the repositories in the catalogue.

What this measures, and what it does not

This is a study of list prices and index scores: 270 models in our catalogue that carry both an Artificial Analysis intelligence score and a published output price, 129 closed and 141 open-weight. Medians are computed within each camp; per-point figures divide a model’s output price by its index score. It does not capture throughput, latency, rate limits, context length or negotiated enterprise pricing — all of which move real bills. It also compares hosted APIs to hosted APIs; the case for running open weights yourself is a separate, stronger one.

The conclusion survives the caveats. The capability gap between the best closed and the best open model has closed to 5.4%. The per-point price gap stands at 3.1×. When those numbers were far apart in the other direction, defaulting to closed was the safe choice. Today the safe choice is to ask, per workload, which side of the 3.4 points you are actually on.

Keep reading

Data: olud.ai catalogue, refreshed nightly — method on /methodology.html.