Gemini 3.8 Flash API Pricing: $0.75/$3.75 Intro, Cheapest Frontier Coding Model of 2026
Gemini 3.8 Live pricing (row group added September 15, 2026)
| Rate | Published unit | Paid tier | Free tier |
|---|---|---|---|
| Audio input | Per minute of audio | $0.005/min ($3.00 per 1M audio tokens) | Free of charge |
| Audio output (includes thinking tokens) | Per minute of audio | $0.018/min ($12.00 per 1M audio tokens) | Free of charge |
| Text input | Per 1M tokens | $0.75 | Free of charge |
| Text output | Per 1M tokens | $4.50 | Free of charge |
| Image / video input | Per minute, or per 1M tokens | $0.002/min ($1.00 per 1M tokens) | Free of charge |
| Grounding with Google Search | Per search request | 5,000 free requests per month (shared across all Gemini 3.x models), then $14 per 1,000 requests | Supported |
| Extended Thinking premium | Per minute | Not published as of September 15, 2026 — Google publishes no separate Extended Thinking rate; its one row covers gemini-3.8-live and gemini-3.8-live-extended-thinking | — |
| Per-hour or per-session rate | Per hour / per session | Not published as of September 15, 2026 — Google publishes per-minute and per-token rates only | — |
One published row covers both models, and it is not marked as an estimate. Google publishes per-minute audio pricing for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on its Gemini API pricing page: $0.005 per minute of audio input and $0.018 per minute of audio output (output includes thinking tokens). The same published row covers both models, so Google does not publish a separate rate for Extended Thinking. Source: Google, Gemini API pricing, ai.google.dev/gemini-api/docs/pricing (retrieved 15 September 2026). The row group is shared with a third model, gemini-3.1-flash-live-preview, and carries no "Estimated pricing" footnote — unlike the Live Translate and Transcribe rows on the same page. Nothing in this table is inferred from the Gemini 3.8 Flash tiers below: those are different models on the same pricing page.
This page is the site's Gemini API pricing surface, so it now carries the live, audio-to-audio tier as well. Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026 — two live-dialogue models Google describes as "the building blocks for reliable, production-ready voice agents" — and, unlike the Flash tiers, they are published with a per-minute audio rate. Both reach developers in the Gemini API and Google AI Studio; enterprise access is in private preview in Gemini Enterprise, and Extended Thinking additionally rolls out in Gemini Live, with Docs in Google Workspace for Google AI Pro and Ultra subscribers.
What a minute of Gemini 3.8 Live voice costs (worked example)
The published mechanics are simple enough to quote from: audio metered per minute, with reasoning billed inside the output line. The assumptions below are ours — Google publishes no per-call or per-session figure — and every dollar total is derived arithmetic, not Google's number.
- Assumption — 10-minute calls: each call is 10 minutes of audio.
- Assumption — full-duplex audio: one minute of audio in and one minute of audio out per minute of conversation.
- Assumption — no separate reasoning tier: Google's row label is "Output price (including thinking tokens)", so a reasoning-heavy call on Extended Thinking bills at the same published rate as a standard one; there is no per-minute premium to add.
- Assumption — voice only: backend model calls, tool traffic and grounding search requests are excluded (grounding is priced per request, not per minute).
- Derived check on the rate: the same page states audio token billing at 25 tokens per second (footnote to the Live Translate row), i.e. 1,500 audio tokens per minute — $3.00/1M × 1,500 = $0.0045/min, which Google states as $0.005/min, and $12.00/1M × 1,500 = $0.018/min exactly. That is our arithmetic, and it is why the two published figures are consistent with each other.
| Calls per day | Audio minutes per day | Audio input | Audio output | Voice cost per day | Per 30-day month |
|---|---|---|---|---|---|
| 10 | 100 | $0.50 | $1.80 | $2.30 | $69.00 |
| 100 | 1,000 | $5.00 | $18.00 | $23.00 | $690.00 |
| 1,000 | 10,000 | $50.00 | $180.00 | $230.00 | $6,900.00 |
Derived from Google's published $0.005/min audio input and $0.018/min audio output: one full-duplex minute is $0.023 ($0.005 in + $0.018 out), a 10-minute call is $0.23 ($0.05 in, $0.18 out), and an hour of continuous full-duplex audio is $1.38. For scale, the site's GPT-Live-1 voice agent cost per minute page documents the incumbent voice stack at a flat $0.05 per minute of voice, billed per second with the backend model billed separately — so the question for a quote is no longer only the per-minute voice rate, but how many minutes a session actually runs and what reasoning rides along inside them.
What changed for voice-agent quotes
Two capability changes from the launch move the commercial needle. First, background tool execution: "It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background" — and more broadly, "These models handle complex reasoning, real-time visual context, and background task execution without interrupting your conversation." Second, real-time visual context: "Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses."
That is what shifts a voice agent's cost model from tokens per call to minutes plus reasoning: the meter runs on minutes of audio while thinking and tool work ride inside the session, so a quote assembled from per-call token math misses both the session length and the reasoning billed in the output line. Price the minutes, state the reasoning assumption, and check the arithmetic against GPT-Live-1 voice agent cost per minute for the incumbent stack. On the agency side the same change lands in scope and SLA language — see The $0.05 Voice Agent: What GPT-Live-1 Changes for Agencies.
Gemini 3.8 Flash price, in one paragraph
Gemini 3.8 Flash API pricing is an introductory $0.75 per 1M input tokens and $3.75 per 1M output tokens (output includes thinking tokens), announced by Google on Sept 2, 2026 — the same intro rate Gemini 3.7 Flash launched with three weeks earlier. The rate holds through 2026-12-31; from Jan 1, 2027 it reverts to $1.50/$7.50 per 1M. During the intro window it is the cheapest big-lab frontier coding model of 2026 on a per-token basis — the "frontier-level performance at a fraction of the price" story Google and Neowin are pointing at — and it is now a selectable strategy in the AI Agency Pricing Calculator.
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google's best reasoning & coding model yet and the third Flash-family release in six weeks (3.6 Flash → 3.7 Flash → 3.8 Flash). Google's recommended model for software engineering, autonomous agents, and complex multi-step reasoning, it is described as a "workhorse" that often approaches the performance of higher-cost frontier models. Specs: 1M-token context window, 64K max output, multimodal input (text, images, audio, video), and support for customizable effort levels — Google notes 3.8 Flash can use more tokens at higher effort levels, so per-task cost is workload-dependent.
Availability on day one: Google AI Pro/Ultra subscribers (Gemini app, AI Mode in Search, Gemini in Google Sheets), plus developers via Gemini API, AI Studio, Google Antigravity, and Android Studio. Gemini 3.8 Flash Cyber — a cybersecurity variant with frontier-level vulnerability discovery — is gated through the new Fairwind Program for governments and trusted defenders; it is not generally available, so it does not belong in ordinary client pricing.
Gemini 3.8 Flash API pricing table (vs. previous Gemini Flash versions)
| Model | Input $/1M | Output $/1M | Status | Notes |
|---|---|---|---|---|
| Gemini 3.8 Flash (Google) | $0.75 | $3.75 | CURRENT — released Sept 2, 2026 | Intro through 2026-12-31, then $1.50/$7.50; 1M context, 64K max output |
| Gemini 3.7 Flash (Google) | $0.75 | $3.75 | Previous version (Aug 13 – Sept 1, 2026) | Same intro structure; superseded by 3.8 Flash |
| Gemini 3.6 Flash (Google) | $1.50 | $7.50 | Previous version (late July 2026) | Launch rate equals 3.7/3.8 post-intro rate |
Rates verified Sept 2, 2026 against Google's announcement, the DeepMind model card, and third-party listings (OpenRouter: google/gemini-3.8-flash, $0.75/$3.75 per 1M). All USD per 1M tokens. The intro structure for 3.8 Flash is byte-for-byte the same deal Google offered on 3.7 Flash — the page below keeps 3.7/3.6 rows because agencies still run them and need the comparison for routing math.
How Gemini 3.8 Flash compares on price in 2026
The relevant comparison for agencies is list price per 1M tokens during the 3.8 intro window:
| Model | Input $/1M | Output $/1M | Position |
|---|---|---|---|
| Gemini 3.8 Flash (intro) | $0.75 | $3.75 | Cheapest frontier-class vendor API of 2026 (intro window) |
| Grok 4.6 (SpaceXAI) | $2.00 | $6.00 | 2.7x the 3.8 intro input price |
| Qwen 3.8 Max (Alibaba) | $2.00 | $6.00 | Open-weight leader at hosted list price |
| GPT-5.6 Sol (OpenAI promo) | $4.00 | $20.00 | 5.3x the 3.8 intro input price |
| Claude Fable 5.1 (Anthropic) | $10.00 | $50.00 | 13x the 3.8 intro input price |
The honest caveat: open-weight APIs are still cheaper in absolute dollars — GLM-5.3-Flash ($0.15/$0.50, promo $0.075/$0.25 through Sept 9) and Qwen3.8-Flash ($0.15/$0.47) undercut Google's intro input price by ~5x. "Cheapest frontier coding model 2026" is therefore a frontier-vendor claim: among the big labs' frontier-class coding APIs, nothing beats Gemini 3.8 Flash's intro $0.75/$3.75 during the window. After Dec 31, 2026 the rate doubles to $1.50/$7.50 and the crown passes back to whatever Google ships next.
What the benchmarks say (why the price matters)
Google positions 3.8 Flash as its best reasoning & coding model: it tops DeepSWE v1.1 (long-horizon software engineering) at a fraction of larger frontier models' cost, and leads its class on finance-agent and legal-agent benchmarks (Vals Finance Agent V2, Harvey's Legal Agent Benchmark) plus HLE-Verified 54.9% for multi-step reasoning. That is the agency-relevant combo: frontier-class coding at an intro price below the open-weight leader's list rate. Treat vendor benchmarks as directional until third-party replication, and budget for the token-burn at higher effort levels.
What agencies should do with the Gemini 3.8 Flash price
- Re-run routing math on the intro rate. At $0.75/$3.75 the case for routing simple traffic to Flash-Lite shrinks vs. the 3.7-era gap — verify on your own mix. The calculator's Gemini routing estimator now lists 3.8 Flash as the current model.
- Quote the post-intro rate for anything past Dec 31. $1.50/$7.50 from Jan 1, 2027 changes long-term retainer math; label the assumption in fixed-fee quotes.
- Keep 3.7 Flash rows for existing stacks. If a client stack pins
gemini-3-7-flash, the old rate table still applies — that model's intro is identical, so the line item doesn't move. - Do not price in 3.8 Flash Cyber. Gated via Fairwind; not on the public API.
Sources
- Google announcement (X): x.com/Google — "Introducing Gemini 3.8, our best reasoning & coding model yet" (Sept 2, 2026)
- Google blog: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- DeepMind model card: Gemini 3.8 Flash
- Google Gemini API pricing: ai.google.dev/gemini-api/docs/pricing
- 9to5Google: Gemini 3.8 Flash rolling out three weeks after last release
- Ars Technica: Google releases Gemini 3.8 Flash, its third Flash model in six weeks
- Thurrott: Google Releases Gemini 3.8 Flash and Cyber Variant
- Neowin: Google launches Gemini 3.8 Flash with frontier-level performance at a fraction of the price
- OpenRouter model page: google/gemini-3.8-flash (live listing, verified Sept 2, 2026)
- Google blog: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (Sept 15, 2026)
- Google model docs: gemini-3.8-live and gemini-3.8-live-extended-thinking (September 2026)
- DeepMind model card: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking (published 15 September 2026)
Accuracy note: Intro pricing and the Dec 31, 2026 / $1.50/$7.50 reversion verified against Google's own announcement and blog (Sept 2, 2026). 1M context / 64K output per DeepMind model card. Third-party rates (OpenRouter, AA) verified live on publish date; 2026 pricing moves weekly — re-verify before quoting. Per-model intro structure identical to the Gemini 3.7 Flash deal; 3.7 and 3.6 rows preserved as previous versions. Gemini 3.8 Live row group added September 15, 2026 and verified the same day against Google's Gemini API pricing page; no rate in that group is inferred from the 3.8 Flash tiers, and the per-minute voice-agent example is derived arithmetic with its assumptions stated. Benchmarks vendor-reported unless noted.
Model Gemini 3.8 Flash in your next quote
Try the AI Agency Pricing Calculator →Estimate setup fees, retainers, and margin — then compare per-task costs across frontier and open-weight models with the AI Model Cost per Task 2026 page.
Frequently asked questions
What is the Gemini 3.8 Flash price?
Introductory $0.75 per 1M input tokens and $3.75 per 1M output tokens (output includes thinking tokens), valid through 2026-12-31; $1.50/$7.50 from Jan 1, 2027. Same intro structure as Gemini 3.7 Flash.
Is Gemini 3.8 Flash the cheapest frontier coding model in 2026?
Among frontier-class vendor APIs during its intro window, yes — $0.75/$3.75 undercuts Grok 4.6, Qwen 3.8 Max, GPT-5.6 Sol, and Claude Fable 5.1. Open-weight APIs (GLM-5.3-Flash, Qwen3.8-Flash) are still cheaper per token in absolute dollars.
What is Gemini 3.8 Flash and when was it released?
Google's best reasoning & coding model yet, released Sept 2, 2026 — the third Flash release in six weeks. Recommended for software engineering, autonomous agents, and complex multi-step reasoning; 1M context, 64K max output.
Is Gemini 3.8 Flash Cyber available through the API?
No — the Cyber variant is limited-access through Google's Fairwind Program for trusted defenders and is not generally available. Do not price it into client stacks.
Where does Gemini 3.8 Flash show up in the pricing calculator?
It is a selectable Model Strategy on the homepage calculator (factor 0.93 / margin +4 during intro) and the current model in the Gemini API Cost & Model Routing Savings estimator. Gemini 3.7 Flash rows remain selectable, labeled as the previous version.
What is Gemini 3.8 Live Extended Thinking?
Google's higher-intelligence live-dialogue model for high-complexity, multi-step work, alongside the cost-efficiency tier Gemini 3.8 Live. For voice agents, reliability trades against latency: Extended Thinking reasons while speaking, so narration replaces instant answers. Availability: the Gemini API, AI Studio, Gemini Live, and private preview in Gemini Enterprise; the published rates sit in one Gemini 3.8 Live pricing row.
Is voice still more expensive than text for agents?
Not on one meter: Google publishes $0.005 per minute of audio input and $0.018 per minute of audio output, with thinking tokens billed inside that output line, so the voice comparison is minutes plus reasoning rather than tokens per call. The Gemini Live API cost per minute is the unit real-time voice agent pricing is built from.