Live per-token pricing for every major model — proprietary and open-weight — in one place. See costs at a glance in the chart, work out your monthly bill, compare models side by side, and find out what the same workload would cost to run on your own hardware: for open-weight models, nothing but electricity.
Pricing pulled from live data · tracked over time · olud.ai
Bars scale to the most expensive model shown. Lower is cheaper.
Every price above is what a provider charges to run a model for you. But open-weight models — DeepSeek, Llama, Qwen, Mistral, Gemma — are free to download and run on your own GPU. After the hardware, the marginal cost per token is basically electricity. For steady, high-volume workloads that flips the maths entirely.
For each popular paid model, here's the open-weight model people switch to, and where to read up on it:
Method. Prices are per million tokens, pulled from the OpenRouter API and refreshed daily. “Blended” cost weights input and output 3:1, a common single-number proxy. Self-host cost is marginal (electricity) and excludes hardware — actual figures depend on your GPU, utilisation and power price. Price-move history is logged from launch, so the moves section fills in over the coming weeks.
Pricing tells you what a model costs. The leaderboard tells you whether it's any good — open-weight models ranked by real-world usage.
Open the leaderboard →