AI Models · Paid vs Paid

Gemini 3.1 Flash Lite Paid vs Gemini 2.5 Flash Lite Paid

Gemini 3.1 Flash Lite vs Gemini 2.5 Flash Lite compared — price per token, context window, multimodality, openness and which to choose. Full 2026 breakdown.

Prices & specs refreshed from live data · OpenSourceAI.tech

Open-model prices = cheapest provider via OpenRouter; official maker rates may be higher.

Choose Gemini 2.5 Flash Lite for the lower output price ($0.4/M vs $1.5/M). Choose Gemini 3.1 Flash Lite if you need the larger 1M context window.

Gemini 3.1 Flash Lite vs Gemini 2.5 Flash Lite specs

SpecGemini 3.1 Flash LiteGemini 2.5 Flash Lite
MakerGoogleGoogle
TypeProprietaryProprietary
Context window1M tokens1M tokens
Input price$0.25/M$0.1/M
Output price$1.5/M$0.4/M
Vision / multimodalYesYes
Tool / function callingYesYes
Self-hostableNo (API only)No (API only)
LicenseProprietaryProprietary

Feature comparison

CapabilityGemini 3.1 Flash LiteGemini 2.5 Flash Lite
Open weights (downloadable)
Self-hostable
Runs fully offline
Vision / multimodal
Tool / function calling
1M+ context window

Benchmarks: Gemini 3.1 Flash Lite vs Gemini 2.5 Flash Lite

Independent benchmark scores measured by Artificial Analysis. Higher is better (except latency).

Gemini 3.1 Flash Lite
Gemini 2.5 Flash Lite
Intelligence index
25
6.9
Coding index
34.7
GPQA
82.2%
47.4%
Humanity's Last Exam
16.2%
3.7%
Long Context Reasoning
65.3%
31.3%
SciCode
41.9%
17.7%
IFBench
77.2%
31.5%
τ²-Bench
31.3%
19%
τ-Bench Banking
8.7%
Terminal-Bench
31.1%
Terminal-Bench Hard
24.2%
2.3%
Math index
35.3
MMLU-Pro
72.4%
LiveCodeBench
40%
MATH-500
92.6%
AIME
50%
AIME 2025
35.3%
Speed
326.9 tok/s
231.4 tok/s
Latency
4.94s
0.31s
Intelligence per $
44.4
39.4

Benchmark data by Artificial Analysis.

How Gemini 3.1 Flash Lite and Gemini 2.5 Flash Lite score

🤝 Neck and neck on these criteria (3.4 vs 3.5 / 5).
CriterionGemini 3.1 Flash LiteGemini 2.5 Flash Lite
Cost-efficiency4.55.0
Context window5.05.0
Openness1.51.5
Self-hosting1.01.0
Multimodality5.05.0

Scores come from live data — output price (cost), context length, open vs closed weights (openness & self-hosting) and vision/tool support (multimodality). Raw task quality isn't scored here; it depends on your benchmark — see the verdict.

What each model is

Gemini 3.1 Flash Lite Paid

Google · Proprietary

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Gemini 2.5 Flash Lite Paid

Google · Proprietary

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

Other models in these families

These variants are tracked but not compared here — one page per family keeps the comparison readable.

Other variants tracked
Gemini 2.5 Flash Lite
Other variants tracked
Gemini 3.1 Flash Lite

Frequently asked questions

Gemini 3.1 Flash Lite vs Gemini 2.5 Flash Lite — which is cheaper?

Gemini 2.5 Flash Lite is cheaper on output ($0.4/M vs $1.5/M).

Which has the larger context window?

Gemini 3.1 Flash Lite offers the larger context window (1M tokens).

Gemini 3.1 Flash Lite vs Gemini 2.5 Flash Lite — which should I pick in 2026?

Choose Gemini 2.5 Flash Lite for the lower output price ($0.4/M vs $1.5/M). Choose Gemini 3.1 Flash Lite if you need the larger 1M context window.

People also compare

Explore more open-source AI

Browse the full open-source model leaderboard, thousands of tools and live benchmarks — all in one place.

Open the leaderboard →