Gemini 3.8 Flash
AvailableGoogle's stable Flash model for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Access details: Gemini API
Versioned identifier · Identity and lifecycle apply to this exact identifier and serving surface.
gemini-3.8-flashCapabilities & limits
- Context window
- 1,048,576
- Maximum output
- 65,536
- Image input
- Supported
- Tool calling
- Supported
- Structured output
- Supported
Documented interfaces and limits. These are not measures of model quality.
Pricing
USD / 1M tokens- Input
- $0.75USD per million input tokens, standard tier introductory rate through 2026-12-31.
- Output
- $3.75USD per million output tokens including thinking, standard tier introductory rate through 2026-12-31.
USD per million standard text tokens. Detailed cache, batch and other pricing conditions appear in the rate card.
Full rate card 8 rates
Amounts in USD. Check the unit and conditions on each rate.
Standard input introductory
$0.75 USD per 1M tokens- effective through 2026-12-31
Standard output introductory
$3.75 USD per 1M tokens- includes thinking tokens
- effective through 2026-12-31
Standard cache read introductory
$0.075 USD per 1M tokens- effective through 2026-12-31
Cache storage introductory
$0.50 USD per 1M tokens per hour- effective through 2026-12-31
Standard input
$1.50 USD per 1M tokens- effective from 2027-01-01
Standard output
$7.50 USD per 1M tokens- includes thinking tokens
- effective from 2027-01-01
Standard cache read
$0.15 USD per 1M tokens- effective from 2027-01-01
Cache storage
$1.00 USD per 1M tokens per hour- effective from 2027-01-01
Free tier is free of charge subject to availability and limits; Batch, Flex, and Priority tiers are separately priced.
Price history
Input $0.75 per million tokens
USD per million input tokens, standard tier introductory rate through 2026-12-31.
Input $1.50 per million tokens
USD per million input tokens, standard tier rate effective 2027-01-01.
Output $3.75 per million tokens
USD per million output tokens including thinking, standard tier introductory rate through 2026-12-31.
Output $7.50 per million tokens
USD per million output tokens including thinking, standard tier rate effective 2027-01-01.
Technical details
- Training / knowledge cutoff
- January 2025
- Input modalities
- text, image, video, audio, PDF
- Output modalities
- text
3 details not verified
- Parameters
- Active parameters
- Architecture
1,048,576 input-token limit; 65,536 output-token limit; thinking levels low, medium, high.
From Google
Official announcement ↗How they describe it
Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
What they say is different
delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains
Verbatim provider claims · not an independent assessment
Release & lifecycle
Availability applies to this exact identifier on Gemini API.
- First API availability
- Sep 2, 2026
- available
- Observed Sep 13, 2026