The best open-weight, self-hostable alternatives to Mistral Medium 3.5 (batch) in 2026 — compared on price, context window and capabilities. Run them locally and cut API costs.
Refreshed from live data · olud.ai
Mistral Medium 3.5 (batch) is a proprietary, API-only model. These open-weight models can be self-hosted, run offline and used at a fraction of the cost — here's how the top ones stack up.
24.6 pts ABOVE Mistral Medium 3.5 (batch) on the Artificial Analysis intelligence index · 6.3× cheaper per million output tokens
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
DeepSeek V4.1 Flash vs Mistral Medium 3.5 (batch) →14.7 pts ABOVE Mistral Medium 3.5 (batch) on the Artificial Analysis intelligence index · 3.1× cheaper per million output tokens
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax M3 vs Mistral Medium 3.5 (batch) →11.2 pts ABOVE Mistral Medium 3.5 (batch) on the Artificial Analysis intelligence index · 3.1× cheaper per million output tokens
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small vs Mistral Medium 3.5 (batch) →4.6 pts ABOVE Mistral Medium 3.5 (batch) on the Artificial Analysis intelligence index · 3.3× cheaper per million output tokens
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Step 3.7 Flash vs Mistral Medium 3.5 (batch) →1.8 pts ABOVE Mistral Medium 3.5 (batch) on the Artificial Analysis intelligence index · 17.0× cheaper per million output tokens
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemma 4 26B A4B vs Mistral Medium 3.5 (batch) →Live ranking of open-weight models with pricing, context windows and capabilities.
Open the leaderboard →