Model Terminal

Gemini 3.1 Flash-Lite

Gemini 3.1 Flash-Lite is Google DeepMind's lowest-cost, lowest-latency tier of the Gemini 3.1 model family, designed for high-volume, latency-sensitive API workloads such as extraction, translation, moderation, and agentic pipelines. It offers a significant quality step-up over Gemini 2.5 Flash-Lite while remaining the most cost-efficient model in the Gemini lineup. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Google AI
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed