Google has promoted the Gemini 3.1 Flash-Lite from preview to general availability, making its cheapest and fastest model officially ready for high-concurrency production workloads. The model comes with four levels of reasoning intensity (minimal, low, medium, high), letting users trade speed for quality depending on the use case.
Pricing holds at preview levels: $0.25 per million input tokens and $1.50 per million output tokens. Against its closest competitor, that's one-quarter the input cost of Claude 4.5 Haiku ($0.25 vs. $1.00) and less than one-third the output cost ($1.50 vs. $5.00). It's also cheaper than Google's own previous-gen 2.5 Flash ($0.30 input, $2.50 output). The context window remains 1 million tokens.
Performance punches above its weight class. On GPQA Diamond (graduate-level science reasoning), Flash-Lite scores 86.9% — beating Claude 4.5 Haiku's 73.0% and GPT-5 mini's 82.3%. On MMMU-Pro (multimodal understanding), it hits 76.8%, again leading its segment. Output speed is 363 tokens per second, 45% faster than 2.5 Flash, with 2.5x faster time to first token. On Arena.ai, it holds an Elo rating of 1432.
Several companies are already running Flash-Lite in production. Customer service platform Gladly uses it to power AI agents for text channels, handling millions of weekly interactions at roughly 60% lower cost than similarly capable models, with p95 latency around 1.8 seconds and a 99.6% success rate. JetBrains deploys it for both its IDE AI assistant and the Junie agent. Financial operations platform Ramp relies on it for high-frequency, latency-sensitive scenarios.
Coding is a relative weak spot: Flash-Lite scores 72.0% on LiveCodeBench, trailing GPT-5 mini's 80.4%.