Mistral Large 2407 vs GPT-5.6 Luna compared — price per token, context window, multimodality, openness and which to choose. Full 2026 breakdown.
Mistral Large 2407 — full profile › · GPT-5.6 Luna — full profile ›
Prices & specs refreshed from live data · olud.ai
Open-model prices = cheapest provider via OpenRouter; official maker rates may be higher.
| Spec | Mistral Large 2407 | GPT-5.6 Luna | Winner |
|---|---|---|---|
| Maker | Mistral AI | OpenAI | – |
| Type | Proprietary | Proprietary | – |
| Context window | 131K tokens | 1.1M tokens | |
| Input price | $2/M | $0.2/M | |
| Output price | $6/M | $1.2/M | |
| Vision / multimodal | No | Yes | |
| Tool / function calling | Yes | Yes | = Tie |
| Self-hostable | No (API only) | No (API only) | – |
| License | Proprietary | Proprietary | – |
GPT-5.6 Luna is ~5.0× cheaper than Mistral Large 2407 on output tokens ($1.2 vs $6 per M tokens).
| Capability | Mistral Large 2407 | GPT-5.6 Luna |
|---|---|---|
| Open weights (downloadable) | ✗ | ✗ |
| Self-hostable | ✗ | ✗ |
| Runs fully offline | ✗ | ✗ |
| Vision / multimodal | ✗ | ✓ |
| Tool / function calling | ✓ | ✓ |
| 1M+ context window | ✗ | ✓ |
Independent benchmark scores measured by Artificial Analysis. Higher is better (except latency).
Benchmark data by Artificial Analysis.
| Criterion | Mistral Large 2407 | GPT-5.6 Luna |
|---|---|---|
| Cost-efficiency | 4.0 | 4.5 |
| Context window | 3.5 | 5.0 |
| Openness | 1.5 | 1.5 |
| Self-hosting | 1.0 | 1.0 |
| Multimodality | 3.5 | 5.0 |
Scores come from live data — output price (cost), context length, open vs closed weights (openness & self-hosting) and vision/tool support (multimodality). Raw task quality isn't scored here; it depends on your benchmark — see the verdict.
This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/m
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.6 Luna is cheaper on output ($1.2/M vs $6/M).
GPT-5.6 Luna offers the larger context window (1.1M tokens).
Choose GPT-5.6 Luna for the lower output price ($1.2/M vs $6/M). It also gives you the larger 1.1M context window.
Browse the full open-source model leaderboard, thousands of tools and live benchmarks — all in one place.
Open the leaderboard →