The best open-weight, self-hostable alternatives to Gemini 3.5 Flash Lite in 2026 — compared on price, context window and capabilities. Run them locally and cut API costs.
Refreshed from live data · olud.ai
Gemini 3.5 Flash Lite is a proprietary, API-only model. These open-weight models can be self-hosted, run offline and used at a fraction of the cost — here's how the top ones stack up.
16.8 pts ABOVE Gemini 3.5 Flash Lite on the Artificial Analysis intelligence index · 4.2× cheaper per million output tokens
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
DeepSeek V4.1 Flash vs Gemini 3.5 Flash Lite →6.9 pts ABOVE Gemini 3.5 Flash Lite on the Artificial Analysis intelligence index · 2.1× cheaper per million output tokens
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
MiniMax M3 vs Gemini 3.5 Flash Lite →3.4 pts ABOVE Gemini 3.5 Flash Lite on the Artificial Analysis intelligence index · 2.1× cheaper per million output tokens
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
Inkling Small vs Gemini 3.5 Flash Lite →Live ranking of open-weight models with pricing, context windows and capabilities.
Open the leaderboard →