The best open-weight, self-hostable alternatives to Mistral Large 3 2512 (batch) in 2026 — compared on price, context window and capabilities. Run them locally and cut API costs.
Refreshed from live data · olud.ai
Mistral Large 3 2512 (batch) is a proprietary, API-only model. These open-weight models can be self-hosted, run offline and used at a fraction of the cost — here's how the top ones stack up.
29.8 pts ABOVE Mistral Large 3 2512 (batch) on the Artificial Analysis intelligence index · 1.3× cheaper per million output tokens
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
DeepSeek V4.1 Flash vs Mistral Large 3 2512 (batch) →7.0 pts ABOVE Mistral Large 3 2512 (batch) on the Artificial Analysis intelligence index · 3.4× cheaper per million output tokens
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemma 4 26B A4B vs Mistral Large 3 2512 (batch) →Within 0.4 pts of Mistral Large 3 2512 (batch) on the Artificial Analysis intelligence index · 1.1× cheaper per million output tokens
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Llama 4 Maverick vs Mistral Large 3 2512 (batch) →Live ranking of open-weight models with pricing, context windows and capabilities.
Open the leaderboard →