The best open-weight, self-hostable alternatives to GPT-3.5 Turbo Instruct in 2026 — compared on price, context window and capabilities. Run them locally and cut API costs.
Refreshed from live data · olud.ai
GPT-3.5 Turbo Instruct is a proprietary, API-only model. These open-weight models can be self-hosted, run offline and used at a fraction of the cost — here's how the top ones stack up.
34.0 pts ABOVE GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 3.3× cheaper per million output tokens
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
DeepSeek V4.1 Flash vs GPT-3.5 Turbo Instruct →24.1 pts ABOVE GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 1.7× cheaper per million output tokens
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax M3 vs GPT-3.5 Turbo Instruct →20.9 pts ABOVE GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 2.3× cheaper per million output tokens
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
MiMo-V2.5-Pro vs GPT-3.5 Turbo Instruct →20.6 pts ABOVE GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 1.7× cheaper per million output tokens
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small vs GPT-3.5 Turbo Instruct →20.3 pts ABOVE GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 6.1× cheaper per million output tokens
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
Hy3 vs GPT-3.5 Turbo Instruct →19.4 pts ABOVE GPT-3.5 Turbo Instruct on the Artificial Analysis intelligence index · 33.3× cheaper per million output tokens
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enablin
Ling 3.0 Flash vs GPT-3.5 Turbo Instruct →Live ranking of open-weight models with pricing, context windows and capabilities.
Open the leaderboard →