Is Google (Gemini) Gemini 3.1 Flash-Lite deprecated?

Deprecated

Yes. Gemini 3.1 Flash-Lite is deprecated and Google (Gemini) shuts it down on 2027-05-07 — in ~8 months.

gemini-3.1-flash-lite last verified

Summary

Gemini 3.1 Flash-Lite is on Google's deprecation list, with 7 May 2027 as the earliest date on which it can be shut down and Gemini 3.5 Flash-Lite named as its replacement. Until then it runs at $0.25 per million input tokens and $1.5 per million output tokens with a 1,048,576-token context window.

Google released Gemini 3.1 Flash-Lite on 7 May 2026 and lists 7 May 2027 against it on the deprecation page. Google states that the dates in that table are the earliest possible retirement dates rather than fixed shutdowns, and the model page still shows the model as stable. It is positioned for high-volume agentic workflows, simple data extraction and applications where latency and cost are the primary constraints.

It takes text, image, video, audio and PDF input and returns text, with caching, code execution, file search, function calling, Maps and search grounding, structured outputs, thinking, URL context, the Batch API and flex and priority inference. The output limit is 65,536 tokens and the Live API is not supported.

The replacement costs more: Gemini 3.5 Flash-Lite is $0.3 / $2.5 against $0.25 / $1.5 here, so budget for roughly a 20 percent higher input rate and a 67 percent higher output rate after the migration.

Modality
  • Text
  • Vision
Use case
  • Chat
  • Coding
  • Reasoning
  • Agents & tool use
  • Long context
Tier
  • Small
Access
  • Hosted API

Specifications

Input price
$0.25 / M tokens
Output price
$1.50 / M tokens
Price band
Standard
Context window
1M tokens
Max output
65.5K tokens
Input
Text, image
Output
Text
Knowledge cutoff
unknown
Tool use
Yes
Structured output
Yes
Weights licence
not published
Model maker
Google (Gemini)

“unknown” means the figure has not been published, or we have not verified it yet — never that the answer is no. “not published” under the weights licence means this model is not offered with downloadable weights. The price band is ours: it bands the blended price (three parts input to one part output) as budget (≤ $0.50), standard (≤ $3), premium (≤ $12) or flagship.

Lifecycle

  1. Released
  2. Deprecated not announced Google (Gemini) has not published the date this happened
  3. Retired in ~8 months

Use it for

  • Keep existing high-volume traffic running while you migrate before 7 May 2027.
  • Compare cost and output against Gemini 3.5 Flash-Lite and Gemini 2.5 Flash-Lite before you choose a replacement.

Avoid it when

  • You are starting a new integration; Google has named a replacement and can retire this model from 7 May 2027.
  • You need more than 65,536 output tokens in a single response.
  • You need a live, low-latency session; the Live API is not supported on this model.

What to use instead of Gemini 3.1 Flash-Lite

  • Gemini 3.5 Flash-Lite Curated

    Google (Gemini) Active

    The replacement named on Google's deprecation page, at $0.3 / $2.5 against $0.25 / $1.5, with the same limits and modalities.

    $0.30 in · $2.50 out / M tokens 1M context

  • Gemini 2.5 Flash-Lite Curated

    Google (Gemini) Active

    Cheaper than the model being retired at $0.1 / $0.4, at the cost of an older January 2025 knowledge cutoff.

    $0.10 in · $0.40 out / M tokens 1M context

  • GPT-5.6 Luna Curated

    OpenAI Active

    OpenAI's cost tier at $0.2 / $1.2 with a 1.05M-token window, but text and image input only.

    $0.20 in · $1.20 out / M tokens 1.1M context

  • Claude Haiku 4.5 Curated

    Anthropic Active

    Anthropic's small tier at $1 / $5, four times the input price, with a 200K-token window.

    $1.00 in · $5.00 out / M tokens 200K context

  • Claude Haiku 4.5 Derived

    Amazon Bedrock Active

    Alternative from Amazon Bedrock, same tier (small), similar price; shares chat, coding, reasoning, agents and long-context.

    $1.00 in · $5.00 out / M tokens 200K context

Curated suggestions were picked by a reviewer, with the reason written by hand. Derived suggestions are ranked from the data — shared tags, the same tier and a comparable price band — not from benchmark scores. Check the capability fit yourself: we track lifecycle, price and specifications.

Sources

Verified against the source on . Dataset snapshot . Providers can change a date without notice — the link above is always the final word.

About this profile

The specifications and dates above are facts from Google (Gemini), with the sources listed. The summary, the use-for and avoid-when lists and the curated alternatives are our assessment: editorial profile reviewed on by llm-researcher.

Common questions

Is Gemini 3.1 Flash-Lite deprecated?
Yes. Gemini 3.1 Flash-Lite is deprecated and Google (Gemini) shuts it down on 2027-05-07 — in ~8 months.
When does Gemini 3.1 Flash-Lite shut down?
Google (Gemini) has announced 2027-05-07 as the shutdown date for Gemini 3.1 Flash-Lite.
What does Gemini 3.1 Flash-Lite cost?
$0.25 per million input tokens and $1.50 per million output tokens, as last verified on 2026-09-10.
What should I use instead of Gemini 3.1 Flash-Lite?
Google (Gemini) Gemini 3.5 Flash-Lite — The replacement named on Google's deprecation page, at $0.3 / $2.5 against $0.25 / $1.5, with the same limits and modalities. Google (Gemini) Gemini 2.5 Flash-Lite — Cheaper than the model being retired at $0.1 / $0.4, at the cost of an older January 2025 knowledge cutoff. OpenAI GPT-5.6 Luna — OpenAI's cost tier at $0.2 / $1.2 with a 1.05M-token window, but text and image input only.

All Google (Gemini) models and their end dates →