Google (Gemini) Gemini 3.8 Flash — what it is, when to use it and what to use instead
No. Gemini 3.8 Flash is active at Google (Gemini) and has no announced end date.
Summary
Gemini 3.8 Flash is the newest model in Google's Flash line, at $0.75 per million input tokens and $3.75 per million output tokens. It takes text, image, video, audio and PDF input within a 1,048,576-token context window and returns up to 65,536 output tokens.
Google positions Gemini 3.8 Flash for long-horizon software engineering, autonomous agents and complex enterprise workflows at Flash-level speed and cost. It is the most recent Flash release, listed with a September 2026 update, and it carries the same $0.75 / $3.75 price as Gemini 3.7 Flash and Gemini 3.6 Flash.
It supports structured outputs, context caching, function calling, code execution, search grounding, URL context, file search, the Batch API and thinking at low, medium or high. Computer use is available in preview. There is no Live API support and no image or audio generation; the output is text.
No deprecation has been announced.
- Modality
- Use case
- Tier
- Access
Specifications
- Input price
- $0.75 / M tokens
- Output price
- $3.75 / M tokens
- Price band
- Standard
- Context window
- 1M tokens
- Max output
- 65.5K tokens
- Input
- Text, image
- Output
- Text
- Knowledge cutoff
- unknown
- Tool use
- Yes
- Structured output
- Yes
- Weights licence
- not published
- Model maker
- Google (Gemini)
“unknown” means the figure has not been published, or we have not verified it yet — never that the answer is no. “not published” under the weights licence means this model is not offered with downloadable weights. The price band is ours: it bands the blended price (three parts input to one part output) as budget (≤ $0.50), standard (≤ $3), premium (≤ $12) or flagship.
Lifecycle
Google (Gemini) has not announced a deprecation or shutdown date for this model. That is a statement about what has been published, not a guarantee — providers typically announce a shutdown months ahead, and this page updates when they do.
- Released
- Deprecated not announced
- Retired not announced
Use it for
- Run agentic coding and long-running tool loops where a Flash-class price makes many calls affordable.
- Process mixed input in one request, combining text with images, video, audio or PDFs.
- Ground answers in live search results or URL content without a separate retrieval stack.
- Handle inputs up to roughly 1M tokens without chunking.
Avoid it when
- You need more than 65,536 output tokens in a single response; that limit is half of what the Claude and GPT flagships allow.
- You need a live, low-latency voice or streaming session; the Live API is not supported on this model.
- You need image or audio generation, which this model does not do.
- Cost per call is the binding constraint; Gemini 3.5 Flash-Lite costs $0.3 / $2.5 and Gemini 2.5 Flash-Lite $0.1 / $0.4.
Alternatives to consider
-
Gemini 3.7 Flash Curated
The previous Flash release at the same $0.75 / $3.75 and the same limits, if you want to change one generation at a time.
$0.75 in · $3.75 out / M tokens 1M context
-
Gemini 3.5 Flash-Lite Curated
Less than half the input price at $0.3 / $2.5 with the same modalities, for high-throughput or sub-agent work.
$0.30 in · $2.50 out / M tokens 1M context
-
Claude Sonnet 5 Curated
Anthropic's mid tier at $2 / $10 with a 1M-token window and 128K max output, twice this model's output limit.
$2.00 in · $10.00 out / M tokens 1M context
-
GPT-5.6 Terra Curated
OpenAI's balanced tier at $2 / $12 with a 1.05M-token window and 128K max output.
$2.00 in · $12.00 out / M tokens 1.1M context
-
Gemini 3.7 Flash Derived
Alternative from DeepInfra, same tier (mid-tier), similar price; shares chat, coding, reasoning, agents and long-context.
$0.75 in · $3.75 out / M tokens 1M context
Curated suggestions were picked by a reviewer, with the reason written by hand. Derived suggestions are ranked from the data — shared tags, the same tier and a comparable price band — not from benchmark scores. Check the capability fit yourself: we track lifecycle, price and specifications.
Sources
- ai.google.dev https://ai.google.dev/gemini-api/docs/models
- ai.google.dev https://ai.google.dev/gemini-api/docs/pricing
- ai.google.dev https://ai.google.dev/gemini-api/docs/deprecations
- ai.google.dev https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash
Verified against the source on . Dataset snapshot . Providers can change a date without notice — the link above is always the final word.
About this profile
The specifications and dates above are facts from Google (Gemini), with the sources listed. The summary, the use-for and avoid-when lists and the curated alternatives are our assessment: editorial profile reviewed on by llm-researcher.
Common questions
- What is Gemini 3.8 Flash for?
- Gemini 3.8 Flash is the newest model in Google's Flash line, at $0.75 per million input tokens and $3.75 per million output tokens. It takes text, image, video, audio and PDF input within a 1,048,576-token context window and returns up to 65,536 output tokens.
- What does Gemini 3.8 Flash cost?
- $0.75 per million input tokens and $3.75 per million output tokens, as last verified on 2026-09-10.
- What is the context window of Gemini 3.8 Flash?
- 1M tokens.
- Is Gemini 3.8 Flash deprecated?
- No. Gemini 3.8 Flash is active at Google (Gemini) and has no announced end date.
- What should I use instead of Gemini 3.8 Flash?
- Google (Gemini) Gemini 3.7 Flash — The previous Flash release at the same $0.75 / $3.75 and the same limits, if you want to change one generation at a time. Google (Gemini) Gemini 3.5 Flash-Lite — Less than half the input price at $0.3 / $2.5 with the same modalities, for high-throughput or sub-agent work. Anthropic Claude Sonnet 5 — Anthropic's mid tier at $2 / $10 with a 1M-token window and 128K max output, twice this model's output limit.