GPT-Live-1 Cost Calculator: What a Voice Agent Really Costs Per Call

Published September 10, 2026 Updated September 18, 2026 By ABD Legacy LLC
Voice: $0.05/min Billed by the second Backend billed separately Concurrency-planned In the API since Sep 10, 2026

The short answer: GPT-Live-1 is billed on two meters, and the $0.05 per minute you have read about is only the first one. Voice time is $0.05 per minute, billed by the second, with no round-up — but the agent's backend reasoning, tool calls and web search are billed separately at the backend model's normal token rates. On the default case on this page (4-minute calls, 1,000 calls a day, 6 reasoning turns per call on GPT-5.6 Terra) voice alone is $200.00 per day and the backend adds $39.60 per day on top. Change the backend from Terra to GPT-6 Astra and the backend line becomes $180.00 per day — within 10% of the voice bill, for a model that is doing the thinking, not the talking.

Developer access is live. GPT-Live-1 has been callable from the API since September 10, 2026 — the full-duplex voice model behind ChatGPT's voice mode, now at a published $0.05 per minute of session time, billed by the second. There is no ChatGPT-only gate and no waitlist, so what changes for a developer is the integration surface (what the API exposes) and the fact that the voice layer can now be dropped onto a managed Agents API session, where container, tool and storage lines bill alongside it (managed-session overhead).

That failure mode is the reason this calculator exists. Nearly every published breakdown of the GPT-Live-1 API stops at the $0.05 per minute figure, which is a voice number. A production phone agent is a voice session plus a reasoning workload, and the second half scales with how hard the caller's request is, not with how long the call lasts. It also has a capacity dimension nobody models: GPT-Live session limits are measured in concurrent sessions, not dollars, so a deployment can run out of session headroom before it runs out of budget.

Everything below is computed live from editable inputs. Published OpenAI rates are labelled as such; every modelling assumption (call mix, token shapes, peak factor, month definition) is labelled as ours and is editable. The machine-readable export at the bottom of the calculator is the exact parameter set and result set the article and this page cite.

GPT-Live-1 Voice + Backend Cost Estimator

All rates shown as of 2026-09-10. Rate inputs marked OpenAI rate are OpenAI-published; inputs marked your assumption are modelling choices and are editable. Nothing here is billed by this site.

:
Backend reasoning (the second meter)
35%
GPT-Realtime-2.1 comparison (what token metering would have cost)
Total cost per call: — voice — + backend —

1. Voice meter ($0.05/min, billed by the second)

Session duration used—
Billed voice seconds (init rule applied)—
Init rule—
Voice cost per call—
Of which is billed silence / backend work—
Voice cost per day—
Voice cost per month —
Same day rate at other month lengths—

2. Backend meter (separate bill, separate scaling law)

Backend model tokens per call—
Tool / web-search spend per call—
Backend total per call—
Backend as a share of total cost—
Backend cost per day—
Backend cost per month —
Effective backend cost per voice minute—

3. Total (the number to quote)

Total per call—
Total per day—
Total per month —
Same day rate at other month lengths—
Voice share of total—

4. Concurrency planner (capacity, not budget)

Calls per hour in the busy window—
Average concurrent sessions needed—
Peak concurrent sessions needed—
Derivation—
Tier checked against—
Headroom at the peak—
Verdict—
Tier qualification floor—
TierCeilingUtilisation at peakFits?
Tier 125——
Tier 250——
Tier 3200——
Tier 4300——
Tier 5500——
Free——Not supported

5. Voice-only quick reference (WebRTC, $0.05/min)

4:00 call—
1:30 call (OpenAI's own worked example)—
0:40 call—
0:05 call, WebRTC (15 s floor)—
0:05 call, WebSocket (no documented floor)—

GPT-Realtime-2.1 token metering, for comparison

Audio cost per minute at your token rates ($32/M in, $64/M out)—
Audio cost per call—
Flat-rate break-even at your in:out audio token mix—
Verdict—

$32/M and $64/M are GPT-Realtime-2.1's audio rows. The same model also has text ($4/$24) and image ($5 in) rows, so an audio-only comparison is a modelling floor, not a full Realtime bill.

billed_voice_seconds = max(session_seconds, 15) # WebRTC session creation only voice_cost = billed_voice_seconds / 60 * $0.05 backend_cost = sum over turns of (in_tokens * in_rate + cached_in * cached_rate + out_tokens * out_rate) / 1e6 + tool_turns * tool_cost_per_invocation total_cost = voice_cost + backend_cost required_concurrent = (calls_per_day / operating_hours) * avg_call_seconds / 3600 * peak_factor

Machine-readable export: the parameters above, the computed results, which numbers are OpenAI-published and which are our assumptions, and the month convention. This is the exact model the page and the article cite. The default parameter set is also served as a file at /gpt-live-1-cost-model.json. Copy shareable link encodes the whole parameter set in the URL hash, and the backend model travels with it — a shared link reopens on the same backend and the same cost, with any rate you edited yourself carried as an explicit override.

—

Nothing on this page is financial advice or a quote from OpenAI. The $0.05 per minute rate, the backend token rates, the 15-second init rule (WebRTC) and the 25/50/200/300/500 concurrent-session ceilings are OpenAI-published as of 2026-09-10; every volume, token-shape, cache, tool, idle-share, peak-factor and month input is a modelling assumption you can change. The Agents API managed lines quoted on this page (no per-session fee, web search $10.00 per 1,000 calls, file search $2.50 per 1,000 calls plus $0.10 per GB per day, and containers at $0.03 / $0.12 / $0.48 / $1.92 per 20-minute session at 1 / 4 / 16 / 64 GB with a 5-minute minimum) were read from OpenAI's pricing page on 2026-09-18; the container per-minute figures are derived from the printed session price.

How the two meters work

Meter one is duration. A GPT-Live-1 session is billed on active session time at $0.05 per minute. OpenAI's cost guide defines active session time as including periods when the user speaks, the assistant speaks, both are silent, and the backend is working — and it notes that muting the microphone does not close the session. Two consequences follow: silence is billable, and delegating a slow reasoning task to the backend does not pause the voice meter. Billing is by the second with no round-up, so a 40-second call is $0.0333 rather than a full minute.

Meter two is reasoning. Backend Responses calls use the normal pricing for the configured model and tools, and OpenAI's cost guide says to include model input and output tokens, cached input where supported, and any applicable image or tool charges — plus the cost of any other services your application calls. The published WebRTC example delegates to GPT-5.6 Terra with hosted web search, so the default on this page uses Terra at its list rates: $2.00 per 1M input tokens, $0.20 cached input, $12.00 per 1M output tokens (short-context column).

The 15-second session-init rule, precisely. Creating a WebRTC session with POST /v1/live/sessions bills 15 seconds of voice duration while the session initializes, and that amount is credited against duration charges once the session starts running. OpenAI's own worked example bills a 90-second session as 90 seconds, "not 105 seconds". The net effect is therefore a 15-second floor, max(session_seconds, 15), not a 15-second adder — and both of the primary statements scope it to WebRTC session creation. The WebSocket guide documents no init charge, which is why transport is a toggle here.

Usage accounting. For reconciliation, session.usage.updated replaces a cumulative duration snapshot rather than adding to it; record the final usage.seconds from session.closed once. Backend work can also outlive the voice session, so a deployment can accrue backend cost with zero voice cost and the reverse.

Why the backend line decides whether this product works

The voice meter is linear, predictable and cheap: 4 minutes is 20 cents, no matter what the caller asks. The backend meter is neither. It scales with turns, with context, with how many turns trigger a tool or a search, and with the price of the model you delegated to. On this page's defaults the backend is a modest $0.0396 per call against $0.2000 of voice time. Switch to Astra, OpenAI's suggested model for complex customer issues, and the same token shape costs $0.1800 per call — roughly 5× the Terra backend, or 90% of the entire voice bill.

That is the "cheap voice, expensive reasoning" failure mode in one number. A deployment priced on the $0.05 per minute rate, then handed to an expensive reasoner with a high tool-call share, can double or triple its unit economics without a single line of its voice code changing. The calculator's backend-share line exists to make that visible before it shows up on an invoice.

Backend (6 turns × 1,500 in / 300 out)Backend $/callVoice $/callBackend share of total
GPT-5.6 Luna — cost-sensitive$0.00396$0.20001.9%
GPT-5.6 Terra — OpenAI's example config$0.0396$0.200016.5%
GPT-5.6 Sol — promo through 2026-11-21$0.0720$0.200026.5%
GPT-6 Astra — complex reasoning$0.1800$0.200047.4%

Derived from OpenAI's published per-1M token rates (Standard, short-context column) applied to the same token shape; reproduce any row in the calculator above. Cache-write tokens are not modelled because they are paid once per cached prefix rather than per turn.

Sizing the line: concurrency, not dollars

GPT-Live rate limits are measured in concurrent sessions, and the ceilings are 25 / 50 / 200 / 300 / 500 for Tiers 1–5. The Free tier is not supported. That makes capacity a first-class cost question: a deployment can be comfortably affordable and still fail at 9 a.m. because too many callers are talking at once.

The sizing formula is Little's Law applied to a call centre line:

concurrent_sessions = (calls_per_day / operating_hours) × average_call_seconds / 3600 × peak_factor

At 1,000 calls a day in an 8-hour window with 4-minute calls, that is 125 calls per hour × 240 seconds ÷ 3600 = 8.33 concurrent sessions on average. A 2.5× busy-hour peak takes it to 20.8 — still inside Tier 1's ceiling of 25, with roughly 17% headroom left. Push the same call volume into a 4-hour window, or let average calls run to 6 minutes, and the peak requirement crosses 25 and you are capacity-limited before you are budget-limited: the constraint is session headroom, not spend, and the fix is a tier change rather than a cheaper model.

It is worth noticing how far apart the two constraints sit. Tier 1 qualification starts at $5 paid and $100 per month of API spend, and the default case on this page burns $200.00 of voice time per day. Spend is not what limits a voice deployment of this size; concurrent sessions are. That is exactly why the planner is on the page.

This page is the estimator. If you would rather read the same model written out — both meters in one place, the flat-rate vs GPT-Realtime-2.1 audio-token comparison with its break-even, and the capacity ladder from 25 up to 500 sessions including the call volume at which Tier 1 runs out — the long-form GPT-Live-1 voice agent cost breakdown prints the worked examples, the sourcing label on every figure and the concurrency table behind these numbers.

What is not included in the $0.05 per minute

Developer access: what the API exposes since September 10, 2026

GPT-Live-1 shipped to the API on September 10, 2026 at a published rate, so the developer question is no longer whether the model can be called but how. Three facts from OpenAI's own model card decide the integration shape:

Nothing on this page assumes a ChatGPT-only model and nothing here is gated on one. The rate, the init rule, the tier ceilings and the delegation rules are reproduced from the developer-facing pages listed under Sources, and the same numbers are what the calculator above is built on.

Worked examples: a 5-minute call and a 1,000-minute month

Both figures use the default call shape on this page — 6 reasoning turns of 1,500 input / 300 output tokens per call on a GPT-5.6 Terra backend, WebRTC transport, with 30% of the session billed to silence and backend wait. The voice half is the published $0.05 per minute with no round-up; the backend half is our arithmetic on Terra's published rates ($2.00 per 1M input, $12.00 per 1M output); the managed rows are OpenAI's published Agents API rates listed under Sources.

One 5-minute call.

One 5:00 callCostHow it is derived
Voice meter$0.2500300 billed seconds ÷ 60 × $0.05/min
Backend meter — GPT-5.6 Terra$0.03966 turns × (1,500 in × $2.00 + 300 out × $12.00) per 1M
Two-meter total$0.2896voice + backend
Managed — per-session fee$0.0000OpenAI publishes no Agents API session fee
Managed — 4 GB container, 5 min$0.03005-minute minimum × the derived $0.0060/min at the $0.12 per 20-minute tier
Managed — one web search$0.0100$10.00 per 1,000 calls
Managed total$0.3296two-meter total + $0.0400 of managed overhead

The managed rows price the same session on OpenAI's infrastructure instead of yours. They do not change the voice rate: the container and search lines are extra meters sitting alongside the $0.05, which is exactly why a “$0.05-a-minute voice agent” is a front-end price and not a total.

A 1,000-minute month is 200 of those 5-minute calls, so the whole table above scales by 200:

1,000 voice minutes (200 × 5:00 calls)CostHow it is derived
Voice meter$50.001,000 minutes × $0.05/min — the figure to quote for the voice layer alone
Backend meter — GPT-5.6 Terra$7.92200 calls × $0.0396
Two-meter total$57.92voice + backend
Managed overhead — container + search$8.00200 sessions × ($0.0300 + $0.0100); per-session fee $0.0000
Managed total$65.92$57.92 + $8.00

A month figure needs no month convention because the input is already a total of billed minutes; the per-day and per-month rows in the calculator above do, which is why every one of them prints its day count. The example also shows which half of a voice-agent bill the voice layer is: 86.3% of the two-meter total at this call shape, falling to 75.8% once the managed container and search lines are on it. The model's reasoning is the smaller half here — change the backend to GPT-6 Astra and it stops being smaller, which is the failure mode the calculator's backend-share row exists to surface.

The same two worked examples, the managed overlay and the estimator that produces them are on the long-form GPT-Live-1 voice agent cost breakdown; both pages run one cost model, so a figure quoted from either reconciles with the other.

Managed-session and tool overhead when the agent runs on OpenAI's infrastructure

Since September 10, 2026 a GPT-Live-1 front end can sit on an Agents API session, where OpenAI hosts the orchestration, the container and the tool loop. That adds meters but not a platform fee: OpenAI states there are no additional fees for using the Agents API, so the published per-session price is $0.00 — a verified zero, not an unpublished estimate. The lines that do bill are these:

Managed linePublished rateUnit, minimum and note
GPT-Live-1 voice front end$0.05 / minBilled per second with no round-up, so $3.00 per open hour; silence, backend wait and a muted microphone are all inside the clock
Per-session fee$0.00No Agents API session fee is published — this row is a verified zero
Web search$10.00 / 1,000 callsPer call, plus the search-content tokens at the backend model's input rate
File search$2.50 / 1,000 callsPer call, plus $0.10 per GB per day of retrieval storage after the first free GB
Hosted container$0.03 / $0.12 / $0.48 / $1.92Per 20-minute session at 1 / 4 / 16 / 64 GB, billed by the minute with a 5-minute minimum
Backend model tokensThat model's ratesInput, cached input and output tokens for the delegated reasoner — the second meter above

The container-per-minute figures used in the worked examples above ($0.0060 a minute at 4 GB) are derived from OpenAI's printed session price — OpenAI publishes the 20-minute session tier and the 5-minute minimum, not a per-minute dollar figure. Session duration caps, session concurrency caps and the free tier are likewise not published for the Agents API. The full managed-session meter set, including file-search storage and three scenario presets, is priced in the managed-session rows on the AI agent API cost calculator, and the commercial read of the same shift is on findaiagency's Agents API economy page.

Methodology and provenance

Every number on this page is either (a) an OpenAI-published rate reproduced from the pages listed under Sources, (b) arithmetic derived from those rates, or (c) a modelling assumption labelled as such in the calculator and in the JSON export. Default parameters and their provenance: voice rate, init seconds and scope, tier ceilings and backend rates are OpenAI-published; calls per day (1,000), call length (4:00), operating window (8 h), peak factor (1×), turns per call (6), tokens per turn (1,500 in / 300 out), tool-call share (35%), cached-input share (0%), idle share (30%) and the 30-day month are our modelling choices, chosen to match a mid-size phone-support deployment and fully editable.

The month convention is printed on every monthly row because OpenAI defines no month. This page's default is a 30-day month ($200.00 per day = $6,000.00); the same day rate is $6,200.00 over 31 days and $6,083.33 at 365÷12, and both alternatives are shown next to the primary figure.

Known limits of this model. The published short-context / long-context rate boundary is not stated by OpenAI, so the long-context toggle is manual. Cache-write tokens are not modelled. The tokens-per-minute rates in the Realtime comparison are ours, because OpenAI publishes no conversion from audio minutes to audio tokens — that is why the comparison prints a break-even token rate as well as a cost. Tool and search spend is a per-invocation input, not an estimate of search-result token volume. And no benchmark figure on this page is presented as independently reproduced: OpenAI's published evaluations are OpenAI's.

Frequently asked questions

What does GPT-Live-1 cost per minute?

A GPT-Live-1 voice session is billed at $0.05 per minute of active session time (OpenAI-published rate, as of September 10, 2026). Billing is by the second with no round-up to the next whole minute, so a 40-second call bills $0.0333 and a 4-minute call bills $0.2000. Backend reasoning is billed separately and is not included in the $0.05.

Is GPT-Live-1 billed per second or per minute?

Per second. OpenAI states the $0.05 per minute rate is not rounded up to the next whole minute, so 40 seconds of session time bills $0.0333 rather than a full minute. The one exception is session creation: a WebRTC session bills 15 seconds while it initializes, and that amount is credited against duration charges, which makes 15 seconds a floor rather than an extra charge.

Is the backend model included in the $0.05 per minute?

No. GPT-Live-1 voice time and backend reasoning are two separate meters. Backend Responses calls use the normal pricing for the configured model and tools, so the backend bill scales with how hard the agent thinks, not with how long the caller talks. On the calculator's default — 6 reasoning turns of 1,500 input and 300 output tokens per call on GPT-5.6 Terra — the backend adds $0.0396 per 4-minute call on top of $0.2000 of voice time.

How many concurrent sessions does Tier 1 allow?

Tier 1 allows 25 concurrent GPT-Live sessions. Tier 2 allows 50, Tier 3 allows 200, Tier 4 allows 300 and Tier 5 allows 500; the Free tier is not supported for GPT-Live. Rate limits are measured in concurrent sessions, and Tier 1 qualification starts at $5 paid and $100 per month of API spend.

How do I size concurrency for 1,000 calls a day?

Use Little's Law on your busy window: concurrent sessions = (calls per day ÷ operating hours) × average call seconds ÷ 3600 × peak factor. At 1,000 calls a day over 8 hours with 4-minute calls that is 125 calls per hour × 240 seconds ÷ 3600 = 8.33 concurrent sessions on average, or 20.8 at a 2.5× peak — which still fits Tier 1's ceiling of 25.

What does the 15-second session-init charge actually cost?

15 seconds of voice time is $0.0125 at the $0.05 per minute rate. It is not an adder: OpenAI's cost guide says the amount billed at session creation is credited against duration charges, and its own worked example bills a 90-second session as 90 seconds, not 105. So the effective rule is a 15-second floor, and OpenAI documents it for WebRTC session creation only; the WebSocket guide states no init charge.

Can I use a non-OpenAI backend model?

Yes. OpenAI's announcement says a GPT-Live-1 session can delegate reasoning and tool calls to a backend text model like GPT-6 Astra or a third-party model, and the backend is billed at that model's own rates. That is why this calculator lets you type any input and output token rate instead of picking only from OpenAI's price list.

Is $0.05 per minute the whole bill for a phone agent?

No. Billed voice time includes periods when the user speaks, the assistant speaks, both are silent, or the backend is working, so idle time and thinking time are billed at $0.05 per minute; muting the microphone does not close the session. On top of that, backend model calls, tool calls and web search are billed separately. The two-meter total in this calculator is the number to quote, not the voice rate alone.

Is GPT-Live-1 available to developers through the API, or only inside ChatGPT?

It is available to developers. GPT-Live-1 shipped to the API on September 10, 2026 with a published rate of $0.05 per minute of session time billed by the second, so there is no ChatGPT-only gate and no waitlist. The limit is surface, not access: the model card lists Live Sessions (v1/live/sessions) as its supported endpoint and marks Chat Completions, Responses, Realtime, Batch and the transcription endpoints as not supported for this model ID.

What does a 5-minute GPT-Live-1 call cost, and what does 1,000 voice minutes a month cost?

A 5-minute call is 300 billed seconds = $0.2500 of voice, plus the backend: 6 reasoning turns of 1,500 input / 300 output tokens on GPT-5.6 Terra add $0.0396, for $0.2896 two-meter. Managed infrastructure adds $0.0400 to the same call - a $0.0000 per-session fee, 5 minutes of 4 GB container at the derived $0.0060 a minute, and one web search at $0.0100 - so $0.3296 all-in. A 1,000-voice-minute month is 200 of those calls: $50.00 of voice, $7.92 of Terra backend, $57.92 two-meter and $65.92 with the managed lines on top.

Sources