Gemini 3.8 Flash API Pricing: $0.75/$3.75 Intro, Cheapest Frontier Coding Model of 2026

Published September 2, 2026Updated September 15, 2026By ABD Legacy LLC
gemini api-pricing model-cost frontier-coding agent-stacks

Gemini 3.8 Live pricing (row group added September 15, 2026)

3.8 Live / 3.8 Live Extended Thinking — row group added September 15, 2026
RatePublished unitPaid tierFree tier
Audio inputPer minute of audio$0.005/min
($3.00 per 1M audio tokens)
Free of charge
Audio output (includes thinking tokens)Per minute of audio$0.018/min
($12.00 per 1M audio tokens)
Free of charge
Text inputPer 1M tokens$0.75Free of charge
Text outputPer 1M tokens$4.50Free of charge
Image / video inputPer minute, or per 1M tokens$0.002/min ($1.00 per 1M tokens)Free of charge
Grounding with Google SearchPer search request5,000 free requests per month (shared across all Gemini 3.x models), then $14 per 1,000 requestsSupported
Extended Thinking premiumPer minuteNot published as of September 15, 2026 — Google publishes no separate Extended Thinking rate; its one row covers gemini-3.8-live and gemini-3.8-live-extended-thinking—
Per-hour or per-session ratePer hour / per sessionNot published as of September 15, 2026 — Google publishes per-minute and per-token rates only—

One published row covers both models, and it is not marked as an estimate. Google publishes per-minute audio pricing for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on its Gemini API pricing page: $0.005 per minute of audio input and $0.018 per minute of audio output (output includes thinking tokens). The same published row covers both models, so Google does not publish a separate rate for Extended Thinking. Source: Google, Gemini API pricing, ai.google.dev/gemini-api/docs/pricing (retrieved 15 September 2026). The row group is shared with a third model, gemini-3.1-flash-live-preview, and carries no "Estimated pricing" footnote — unlike the Live Translate and Transcribe rows on the same page. Nothing in this table is inferred from the Gemini 3.8 Flash tiers below: those are different models on the same pricing page.

This page is the site's Gemini API pricing surface, so it now carries the live, audio-to-audio tier as well. Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026 — two live-dialogue models Google describes as "the building blocks for reliable, production-ready voice agents" — and, unlike the Flash tiers, they are published with a per-minute audio rate. Both reach developers in the Gemini API and Google AI Studio; enterprise access is in private preview in Gemini Enterprise, and Extended Thinking additionally rolls out in Gemini Live, with Docs in Google Workspace for Google AI Pro and Ultra subscribers.

What a minute of Gemini 3.8 Live voice costs (worked example)

The published mechanics are simple enough to quote from: audio metered per minute, with reasoning billed inside the output line. The assumptions below are ours — Google publishes no per-call or per-session figure — and every dollar total is derived arithmetic, not Google's number.

Calls per dayAudio minutes per dayAudio inputAudio outputVoice cost per dayPer 30-day month
10100$0.50$1.80$2.30$69.00
1001,000$5.00$18.00$23.00$690.00
1,00010,000$50.00$180.00$230.00$6,900.00

Derived from Google's published $0.005/min audio input and $0.018/min audio output: one full-duplex minute is $0.023 ($0.005 in + $0.018 out), a 10-minute call is $0.23 ($0.05 in, $0.18 out), and an hour of continuous full-duplex audio is $1.38. For scale, the site's GPT-Live-1 voice agent cost per minute page documents the incumbent voice stack at a flat $0.05 per minute of voice, billed per second with the backend model billed separately — so the question for a quote is no longer only the per-minute voice rate, but how many minutes a session actually runs and what reasoning rides along inside them.

What changed for voice-agent quotes

Two capability changes from the launch move the commercial needle. First, background tool execution: "It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background" — and more broadly, "These models handle complex reasoning, real-time visual context, and background task execution without interrupting your conversation." Second, real-time visual context: "Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses."

That is what shifts a voice agent's cost model from tokens per call to minutes plus reasoning: the meter runs on minutes of audio while thinking and tool work ride inside the session, so a quote assembled from per-call token math misses both the session length and the reasoning billed in the output line. Price the minutes, state the reasoning assumption, and check the arithmetic against GPT-Live-1 voice agent cost per minute for the incumbent stack. On the agency side the same change lands in scope and SLA language — see The $0.05 Voice Agent: What GPT-Live-1 Changes for Agencies.

Gemini 3.8 Flash price, in one paragraph

Gemini 3.8 Flash API pricing is an introductory $0.75 per 1M input tokens and $3.75 per 1M output tokens (output includes thinking tokens), announced by Google on Sept 2, 2026 — the same intro rate Gemini 3.7 Flash launched with three weeks earlier. The rate holds through 2026-12-31; from Jan 1, 2027 it reverts to $1.50/$7.50 per 1M. During the intro window it is the cheapest big-lab frontier coding model of 2026 on a per-token basis — the "frontier-level performance at a fraction of the price" story Google and Neowin are pointing at — and it is now a selectable strategy in the AI Agency Pricing Calculator.

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google's best reasoning & coding model yet and the third Flash-family release in six weeks (3.6 Flash → 3.7 Flash → 3.8 Flash). Google's recommended model for software engineering, autonomous agents, and complex multi-step reasoning, it is described as a "workhorse" that often approaches the performance of higher-cost frontier models. Specs: 1M-token context window, 64K max output, multimodal input (text, images, audio, video), and support for customizable effort levels — Google notes 3.8 Flash can use more tokens at higher effort levels, so per-task cost is workload-dependent.

Availability on day one: Google AI Pro/Ultra subscribers (Gemini app, AI Mode in Search, Gemini in Google Sheets), plus developers via Gemini API, AI Studio, Google Antigravity, and Android Studio. Gemini 3.8 Flash Cyber — a cybersecurity variant with frontier-level vulnerability discovery — is gated through the new Fairwind Program for governments and trusted defenders; it is not generally available, so it does not belong in ordinary client pricing.

Gemini 3.8 Flash API pricing table (vs. previous Gemini Flash versions)

ModelInput $/1MOutput $/1MStatusNotes
Gemini 3.8 Flash (Google)$0.75$3.75CURRENT — released Sept 2, 2026Intro through 2026-12-31, then $1.50/$7.50; 1M context, 64K max output
Gemini 3.7 Flash (Google)$0.75$3.75Previous version (Aug 13 – Sept 1, 2026)Same intro structure; superseded by 3.8 Flash
Gemini 3.6 Flash (Google)$1.50$7.50Previous version (late July 2026)Launch rate equals 3.7/3.8 post-intro rate

Rates verified Sept 2, 2026 against Google's announcement, the DeepMind model card, and third-party listings (OpenRouter: google/gemini-3.8-flash, $0.75/$3.75 per 1M). All USD per 1M tokens. The intro structure for 3.8 Flash is byte-for-byte the same deal Google offered on 3.7 Flash — the page below keeps 3.7/3.6 rows because agencies still run them and need the comparison for routing math.

How Gemini 3.8 Flash compares on price in 2026

The relevant comparison for agencies is list price per 1M tokens during the 3.8 intro window:

ModelInput $/1MOutput $/1MPosition
Gemini 3.8 Flash (intro)$0.75$3.75Cheapest frontier-class vendor API of 2026 (intro window)
Grok 4.6 (SpaceXAI)$2.00$6.002.7x the 3.8 intro input price
Qwen 3.8 Max (Alibaba)$2.00$6.00Open-weight leader at hosted list price
GPT-5.6 Sol (OpenAI promo)$4.00$20.005.3x the 3.8 intro input price
Claude Fable 5.1 (Anthropic)$10.00$50.0013x the 3.8 intro input price

The honest caveat: open-weight APIs are still cheaper in absolute dollars — GLM-5.3-Flash ($0.15/$0.50, promo $0.075/$0.25 through Sept 9) and Qwen3.8-Flash ($0.15/$0.47) undercut Google's intro input price by ~5x. "Cheapest frontier coding model 2026" is therefore a frontier-vendor claim: among the big labs' frontier-class coding APIs, nothing beats Gemini 3.8 Flash's intro $0.75/$3.75 during the window. After Dec 31, 2026 the rate doubles to $1.50/$7.50 and the crown passes back to whatever Google ships next.

What the benchmarks say (why the price matters)

Google positions 3.8 Flash as its best reasoning & coding model: it tops DeepSWE v1.1 (long-horizon software engineering) at a fraction of larger frontier models' cost, and leads its class on finance-agent and legal-agent benchmarks (Vals Finance Agent V2, Harvey's Legal Agent Benchmark) plus HLE-Verified 54.9% for multi-step reasoning. That is the agency-relevant combo: frontier-class coding at an intro price below the open-weight leader's list rate. Treat vendor benchmarks as directional until third-party replication, and budget for the token-burn at higher effort levels.

What agencies should do with the Gemini 3.8 Flash price

Sources

Accuracy note: Intro pricing and the Dec 31, 2026 / $1.50/$7.50 reversion verified against Google's own announcement and blog (Sept 2, 2026). 1M context / 64K output per DeepMind model card. Third-party rates (OpenRouter, AA) verified live on publish date; 2026 pricing moves weekly — re-verify before quoting. Per-model intro structure identical to the Gemini 3.7 Flash deal; 3.7 and 3.6 rows preserved as previous versions. Gemini 3.8 Live row group added September 15, 2026 and verified the same day against Google's Gemini API pricing page; no rate in that group is inferred from the 3.8 Flash tiers, and the per-minute voice-agent example is derived arithmetic with its assumptions stated. Benchmarks vendor-reported unless noted.

Model Gemini 3.8 Flash in your next quote

Try the AI Agency Pricing Calculator →

Estimate setup fees, retainers, and margin — then compare per-task costs across frontier and open-weight models with the AI Model Cost per Task 2026 page.

Frequently asked questions

What is the Gemini 3.8 Flash price?

Introductory $0.75 per 1M input tokens and $3.75 per 1M output tokens (output includes thinking tokens), valid through 2026-12-31; $1.50/$7.50 from Jan 1, 2027. Same intro structure as Gemini 3.7 Flash.

Is Gemini 3.8 Flash the cheapest frontier coding model in 2026?

Among frontier-class vendor APIs during its intro window, yes — $0.75/$3.75 undercuts Grok 4.6, Qwen 3.8 Max, GPT-5.6 Sol, and Claude Fable 5.1. Open-weight APIs (GLM-5.3-Flash, Qwen3.8-Flash) are still cheaper per token in absolute dollars.

What is Gemini 3.8 Flash and when was it released?

Google's best reasoning & coding model yet, released Sept 2, 2026 — the third Flash release in six weeks. Recommended for software engineering, autonomous agents, and complex multi-step reasoning; 1M context, 64K max output.

Is Gemini 3.8 Flash Cyber available through the API?

No — the Cyber variant is limited-access through Google's Fairwind Program for trusted defenders and is not generally available. Do not price it into client stacks.

Where does Gemini 3.8 Flash show up in the pricing calculator?

It is a selectable Model Strategy on the homepage calculator (factor 0.93 / margin +4 during intro) and the current model in the Gemini API Cost & Model Routing Savings estimator. Gemini 3.7 Flash rows remain selectable, labeled as the previous version.

What is Gemini 3.8 Live Extended Thinking?

Google's higher-intelligence live-dialogue model for high-complexity, multi-step work, alongside the cost-efficiency tier Gemini 3.8 Live. For voice agents, reliability trades against latency: Extended Thinking reasons while speaking, so narration replaces instant answers. Availability: the Gemini API, AI Studio, Gemini Live, and private preview in Gemini Enterprise; the published rates sit in one Gemini 3.8 Live pricing row.

Is voice still more expensive than text for agents?

Not on one meter: Google publishes $0.005 per minute of audio input and $0.018 per minute of audio output, with thinking tokens billed inside that output line, so the voice comparison is minutes plus reasoning rather than tokens per call. The Gemini Live API cost per minute is the unit real-time voice agent pricing is built from.