AI Agent API Cost Calculator: Estimate Monthly Agent Workload Costs
The short answer: a moderately active AI agent costs roughly $50–$150/month in raw API spend on mid-tier 2026 models, and a full agent deployment is easy to model once you separate the four cost drivers: number of agents, calls per agent per day, tokens (or price) per call, and agency overhead. This calculator walks an agency owner through each driver and turns it into a defensible monthly number — plus a suggested client price using agent-first pricing.
This page is the pricing companion to "AI Agents Are the API Economy's Biggest Customers: What Agencies Should Charge" — the full analysis of the PYMNTS report is there; the math is here. AI agents are the fastest-growing class of API consumers (PYMNTS, Aug 24, 2026), OpenRouter passed 1 trillion tokens/day in late 2025, and Cloudflare's Matthew Prince cites a single agent task querying ~5,000 sites versus ~5 for a human. That machine-scale consumption is exactly what an agency is paying for — and what it should bill for, deliberately.
The short answer for a managed session: since September 10, 2026 OpenAI also sells the agent as a service — the Agents API, a managed Codex harness — and that changes the shape of the bill without changing what you quote on. A managed session is still tokens at your chosen model's rates, but it adds three meters the application code used to be: tool calls (web search $10.00 per 1,000 calls; file search $2.50 per 1,000 calls, plus $0.10 per GB per day of retrieval storage after the first free GB), container-minutes ($0.03 / $0.12 / $0.48 / $1.92 per 20-minute session at 1, 4, 16 and 64 GB, billed by the minute with a 5-minute minimum), and — only if you put a voice front end on it — GPT-Live-1 at $0.05 per minute of open session, which is $3.00 per open hour. There is no per-session fee: OpenAI states there are no additional fees for using the Agents API, and its pricing page carries no Agents API section at all. The managed block at the bottom of the calculator models all of it, with three scenario presets — a prototype agent, a production cloud agent on managed infrastructure, and an agency-managed deployment.
Capacity context: Anthropic's reported $45 billion Nscale commitment and its SpaceX Colossus 1 compute deal signal real Claude capacity growth — but also pricing pressure ahead of the record IPO. For the full breakdown, see Anthropic's $45B Nscale compute deal: what it means for Claude capacity and pricing.
Agent API Cost Estimator
Agent-first client pricing (suggested)
Agent reliability / durable execution
New Sept 18, 2026Every mode above prices work that succeeds. This row prices the layer that keeps an agent going when something fails — the compute a failed session re-runs while its state is recovered and the retry finishes. It is additive: the reliability layer is added to whichever headline total your calculator mode is showing. It is not a per-session charge (the Agents API per-session fee is published as $0.00 above, and stays that way), and it is not a vendor quote: no durable-execution platform publishes a list price, which is why the flat fee below defaults to $0.00. All five inputs are this page's model — replace them with your own logs.
Reliability / durable execution — line items
model defaults — tune to your own logsMarket anchor — Temporal's $550M Series E at a $12.55B valuation, 14 September 2026 (Temporal · The Next Web). The round funds the layer Temporal calls Durable Execution, which keeps applications, “including AI agents”, “running when something fails” — the argument for pricing reliability as its own line, not for buying one vendor's stack.
How the math works
Four steps, each with one formula. The calculator runs these live; the same formulas are what you'd put in a spreadsheet or a quote.
- Monthly calls.
agents × calls/agent/day × working days × complexity. The complexity multiplier is where agents differ from humans: retries re-pay full context, subagent fan-out re-reads context per turn, and loops still bill. 1.5× is a defensible default for real agentic work; heavy fan-out runs closer to 3×. The retry and fan-out multipliers, worked on a fixed 10,000-task workload, decide whether a metered routing layer pays for itself — a layer that adds a pass of its own makes the bill worse, not better. - Monthly tokens.
monthly calls × tokens per call, split into input and output. A typical agent turn sends a large system prompt + conversation context (input-heavy) and returns a smaller completion (output). - Raw API cost.
(input tokens ÷ 1M × input price × (1 − cache hit rate)) + (input tokens ÷ 1M × cache price × cache hit rate) + (output tokens ÷ 1M × output price). Caching matters: a 40% input cache hit rate at $0.20/1M vs a $2.00/1M miss price cuts the input leg nearly in half. For per-call pricing, raw cost is simplymonthly calls × price per call. - Fully-loaded cost.
raw × (1 + overhead). Overhead covers the human and tooling layer — monitoring, oversight, integration, prompt maintenance, re-baselining — which doesn't scale with tokens but is real cost. Then cost per agent =loaded ÷ agents.
Outcome-based mode (new). Switch the mode selector to Outcome-based to compare the same workload against paying per completed task instead of per token. Enter your expected completed tasks per month and the price per completed outcome (market references cluster around $0.50–$2.00 per resolution — Intercom Fin $0.99, Zendesk Verified Resolutions ~$1.20–$1.50, HubSpot Customer Agent $0.50, Salesforce Agentforce $2). The calculator shows monthly and yearly cost for both models side by side, plus the delta. Outcome pricing transfers retry/failure risk to the vendor: it tends to win when your success rate is low or your workload burns many calls per completion, and loses at high volume when the per-outcome price exceeds your effective per-token cost.
One dimension the token math above does not carry is capacity. An API budget scales with spend, but the voice layer of the same stack is limited in concurrent sessions — GPT-Live-1 runs 25 / 50 / 200 / 300 / 500 across Tiers 1–5, and the Free tier cannot call the model at all — so a phone deployment can be comfortably affordable and still fail at the busy hour. The concurrent-session sizing for a voice line (sessions, not requests per minute) applies Little's Law to a call centre line and shows where 1,000 and 1,500 calls a day land against the Tier 1 ceiling. Size calls, not RPM, before you quote the retainer.
Reference: current per-1M rates (Aug 2026) — the inputs for your own math
Same reference table as our per-task cost benchmarks, current as of Aug 28, 2026 except the DeepSeek rows, which are on DeepSeek's Sept 10, 2026 sheet. Rates move fast — re-verify before quoting a client.
| Model | Input ($/1M) | Output ($/1M) | Notes |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | −80% Jul 30, 2026 (from $1/$6); 13.8x usage in the Jul 27 – Aug 14 discount window |
| DeepSeek V4.1-Flash (off-peak) | $0.15 | $0.60 | API name deepseek-flash (DeepSeek sheet, Sept 10, 2026); peak $0.30/$1.20; cache-hit input $0.003 off-peak / $0.006 peak. The retired deepseek-v4-flash IDs are aliases billed at the Flash price |
| DeepSeek V4-Pro (off-peak) | $0.66 | $1.98 | Peak $1.32/$3.96; cache-hit input $0.022 off-peak / $0.044 peak. DeepSeek V4 Pro vs V4.1-Flash: the withdrawn Sept 14 reroute and the 100x cached-input spread |
| Gemini 3.8 Flash (intro) | $0.75 | $3.75 | Released Sept 2, 2026 — current Gemini Flash; intro through 2026-12-31, then $1.50/$7.50; 1M context |
| Gemini 3.7 Flash (intro — previous version) | $0.75 | $3.75 | Intro through 2026-12-31, then $1.50/$7.50; 1M context; superseded by 3.8 Flash Sept 2, 2026 |
| Meta Muse Glimmer (hosted, Together AI) | $0.35 | $1.50 | Local self-host = electricity after hardware amortized |
| GPT-5.6 Terra | $2.00 | $12.00 | −20% Jul 30, 2026 (from $2.50/$15); 5.6x usage in the discount window |
| Grok 4.6 (SpaceXAI) | $2.00 | $6.00 | 500K context; cache hit $0.50 |
| Qwen 3.8 Max (open weights) | $2.00 | $6.00 | Hosted API list price |
| Claude Sonnet 5 | $2.00 | $10.00 | Permanent as of Aug 10, 2026 |
| GPT-5.6 Sol | $4.00 | $20.00 | Official OpenAI promo (Aug 21 – Nov 21, 2026); cached input $0.40 |
OpenAI's discount experiment (Jul 27 – Aug 14, 2026): when Terra and Luna were discounted 50% via OpenRouter, daily Terra token usage rose 5.6x and Luna 13.8x, while Sol at list price rose only 1.1x (control); ~1/3 of users who tried a discounted model kept using it after expiry. Terra/Luna went from 0.7% to 7.8% of OpenRouter tokens. The same demand elasticity applies to any client workload you quote on a discounted model. See OpenAI's discount experiment: what 13.8x usage growth means for agency pricing for the full analysis.
Managed-harness containers: per 20-minute session per container (1, 4, 16, 64 GB)
One runtime line does not belong in the per-1M table above, because its unit is not tokens — the same reason our per-task benchmarks sit in their own table rather than in the per-1M columns. This is the line an agency pays when OpenAI hosts the Codex harness: a per-container session price by memory tier, on top of model tokens and tool calls, with no Agents-API-specific fee and no Agents API row on the pricing page. The full managed-session meter set — tool calls, container minutes, retrieval storage, voice minutes and the token pass-through — is priced in the September 2026 managed-session section below.
| Container memory | OpenAI printed rate (per 20-minute session per container) | Derived per-minute · 5-min floor (derived, not published) | Note |
|---|---|---|---|
| 1 GB | $0.03 | $0.0015 · $0.0075 | developers.openai.com/api/docs/pricing#built-in-tools, read 2026-09-12; the row is labelled "Hosted Shell and Code Interpreter" and does not name the Agents API; the footnote bills container sessions OpenAI calls "Eligible" and they are billed by the minute, with a 5-minute minimum per session; the per-minute dollar figures are derived (tier ÷ 20), OpenAI publishes none; "not a final bill" (usage fields are best-effort, can be null). |
| 4 GB | $0.12 | $0.006 · $0.03 | |
| 16 GB | $0.48 | $0.024 · $0.12 | |
| 64 GB | $1.92 | $0.096 · $0.48 |
Basis: the four printed prices are OpenAI's, read from developers.openai.com/api/docs/pricing#built-in-tools on 2026-09-12 — 1 GB $0.03, 4 GB $0.12, 16 GB $0.48, 64 GB $1.92 per 20-minute session per container. Everything in the third column is this page's arithmetic on two quoted OpenAI statements, and is labelled derived because OpenAI prints no per-minute dollar figure for this row: per-minute = tier ÷ 20 (OpenAI's changelog for Jun 2, 2026 says eligible container sessions are billed per minute instead of at the full 20-minute session rate, and that "the underlying per-minute rate will remain the same"), and the floor = per-minute × 5 for the footnote's 5-minute minimum. At every tier the floor is 25% of the printed session price, the session price is 4.0× the floor, and the two units cross at exactly 20 billed minutes.
Worked illustration (derived, assumptions stated): one 4 GB container running 3 sessions a day across 22 working days is 66 sessions a month. At the printed session unit that is 66 × $0.12 = $7.92. If each of those 66 sessions bills 12 minutes, the derived per-minute figure gives 66 × 12 × $0.006 = $4.752 — cheaper than the session unit, because 12 minutes is short of the 20-minute crossover. Neither figure is a vendor quote. For what that line does to a client quote, see what this does to a client quote.
Inputs this meter needs, and what the estimator above does not take
- Containers in use — distinct containers per workload, not agents.
- Memory tier per container — 1, 4, 16 or 64 GB.
- Sessions per container per day.
- Session length in minutes — this decides which of the two units is cheaper and where the 5-minute floor bites.
- Working days per month — the estimator's own default is 22.
The estimator at the top of this page is a per-month token-and-call model with a retry/fan-out multiplier and an outcome mode. None of the five inputs above exists in it, and this block does not add them — the same unit problem that keeps our per-task benchmark table out of the per-1M columns above.
What the container row does not say
- OpenAI publishes no per-minute dollar figure. The derived column is this page's arithmetic (tier ÷ 20, × 5) and nothing more.
- "Eligible" is undefined. OpenAI never defines the qualifier on any surface read, so the printed rate and the billed rate are not provably the same sessions.
- The meter is not the bill. Session and turn usage is best-effort, can be
null, "Missing usage does not mean zero usage", and OpenAI states in writing that "These counts are not a final bill"; cache-write pricing cannot be reproduced from the usage fields. - Tokens and tools bill separately at model and tool rates — the Agents API examples run on
gpt-6-astraat $10.00/$50.00 per 1M short-context input/output — so the container line must never be presented as the total. - Session length is your assumption, not a vendor rule. Nothing on any OpenAI surface read states that idle, retry or stalled time bills, so the minutes you enter are the agency's own estimate.
What a managed agent session costs under the Agents API (September 2026)
OpenAI shipped the Agents API in public beta on September 10, 2026: the Codex harness as managed infrastructure, with OpenAI owning sessions, orchestration, context compaction and recovery. The pricing question the launch raises is not what does a session cost but which meter does this session hit. OpenAI's own answer is that the platform adds nothing — the company's stated position is that there are no additional fees for using the Agents API — so the six rows below are the whole bill, and five of them have a published rate.
| Meter | Published rate (read Sept 18, 2026) | What triggers it |
|---|---|---|
| Per-session fee | $0.00 — verified, not an estimate | Nothing. There is no agents row and no agents section on OpenAI's pricing page; the platform fee is affirmatively denied. |
| Model token pass-through | Your selected model's own input, cached-input and output rates | Every turn's prompt and completion. Cached input is the only discount — GPT-6 Astra is $10.00/1M in and $50.00/1M out, against $1.00/1M for cached input. |
| Web search | $10.00 per 1,000 calls, plus search-content tokens at model rates | Each search the agent runs. The tokens the search returns bill twice over — once as the call, once as context. |
| File search | $2.50 per 1,000 calls; storage $0.10 per GB per day, first 1 GB free | Calls bill per session. Storage bills per day, whether or not a session runs — it is the only line on this page that accrues while the agent is idle. |
| Hosted container (sandbox) | 1 GB $0.03 · 4 GB $0.12 · 16 GB $0.48 · 64 GB $1.92, per 20-minute session | Billed by the minute with a 5-minute minimum, on hosted sandboxes only — a no-environment session pays nothing here. |
| GPT-Live-1 voice front end | $0.05 per minute of open session, billed per second without rounding up | Wall-clock voice time, silence included, which is $3.00 per open hour. Backend model and tool usage bill separately, so a "$0.05/minute voice agent" is a front-end price, not a total. |
Not published — and therefore not guessed at: per-step or per-turn fee, session duration cap, session concurrency cap, free tier. The only published clock is a one-hour sandbox idle timeout that OpenAI says is not configurable, and the only published caps anywhere in this stack are GPT-Live-1's concurrent sessions (25 / 50 / 200 / 300 / 500 across Tiers 1–5, with no free tier). The calculator prints each of those as a state instead of a default, and labels container-minutes-per-session as your own estimate, because OpenAI publishes the container rate and its 5-minute minimum but no session-length rule for agents.
The three presets, hand-checked
Preset buttons in the calculator set every input below; the arithmetic is reproduced here so you can check it without clicking anything. All rates are OpenAI's, read 2026-09-18; every line is this page's arithmetic on those rates.
The finding that matters for a quote: on preset (b) the voice layer is 31% of the session bill on a frontier backend — and 81% of it if you swap the backend for a cheap one. The same 6-minute session on GPT-5.6 Luna instead of GPT-6 Astra costs $0.3692 instead of $0.9660, of which $0.3000 is still voice. Model routing moves the token line; only closing idle voice sessions, or not using voice, moves the voice line. These two figures are the same ones on the managed-agents section of the API-economy analysis, so the two pages cannot quietly disagree.
Two practical notes on quoting from this block. First, the managed monthly total is an infrastructure number: the overhead and markup sliders in the estimator above apply to it exactly as they do to the self-hosted rows, because monitoring, re-baselining and accountability are still yours. Second, if the deployment takes phone calls, the binding constraint may be concurrency rather than cost — GPT-Live-1 caps at 25 concurrent sessions on Tier 1, so see concurrent-session sizing for a voice line and the GPT-Live-1 voice + backend cost calculator before you promise a number.
What a realistic deployment costs (worked example)
Default calculator inputs — 10 agents, 200 calls/day, 22 days, 1.5× complexity, 8K in / 1.2K out tokens, Claude Sonnet 5 rates ($2/$10), 40% input cache hit, 25% overhead:
- Monthly calls: 10 × 200 × 22 × 1.5 = 66,000
- Tokens: 528M input, 79.2M output
- Raw API cost: ≈ $1,468/month (input ≈ $676 after cache, output ≈ $792)
- Fully-loaded: ≈ $1,835/month → $183/agent/month
- Suggested client price at 30% markup: ≈ $2,385/month (23% gross margin), or roughly $1.39 per resolution at 50 calls/resolution — which is above Intercom's $0.99 benchmark. That is the honest finding: on mid-tier frontier rates with a 50-call outcome, per-resolution math is not automatically cheaper than the SaaS per-outcome price.
Now move the model lever and the picture changes by an order of magnitude:
- DeepSeek V4.1-Flash off-peak ($0.15/$0.60, cache $0.003): same workload ≈ $96/month raw, ≈ $120 loaded, ≈ $0.09 per resolution — dramatically under every per-outcome benchmark. This is the cheapest-model margin story.
- GPT-5.6 Sol ($4/$20, cache $0.40): same workload ≈ $2,936/month raw, ≈ $3,670 loaded, ≈ $2.78 per resolution — above Salesforce's $2/conversation anchor.
Model choice is the single biggest lever in the model — which is why routing small tasks to cheap models is a margin decision, not a footnote.
The same question, managed. If OpenAI hosts the harness instead, the unit changes and the calculator's third mode takes over. Preset (b) — a production voice support line on managed infrastructure — prices at $0.9660 per session and $28,980.00 a month at 1,000 sessions a day, with the voice layer alone taking 31% of that session (81% if you drop to a cheap backend). Preset (a), a prototype agent with no sandbox, comes to $0.0652 a session. The per-line arithmetic for all three presets, and the meters the self-hosted rows above cannot see — tool calls, container minutes, retrieval storage and voice minutes — are in the managed-session section above.
Why the PYMNTS picture makes this math urgent
PYMNTS reported (Aug 24, 2026) that AI agents are the fastest-growing class of API consumers — and that the API economy's pricing, identity, and trust infrastructure "were not designed for this" and "are being rebuilt now." Three consequences for an agency's cost model:
- Volume is machine-scale. One agent task can touch thousands of endpoints. Per-seat assumptions break; per-usage math is the only honest frame.
- Token prices collapsed ~200x in 16 months (GPT-4 $30/1M input, Mar 2023 → GPT-4o mini $0.15/1M, Jul 2024), so the same client deliverable costs a fraction of what it did 18 months ago — and keeps falling 30–50%/year. If you bill hourly or pass through API costs as a flat line item, clients with a calculator will ask why their bill didn't fall.
- The market has already repriced for agents. Salesforce Agentforce lists $2/conversation and $0.10/action; Intercom Fin charges $0.99/resolution. Those are your pricing anchors — per-outcome, not per-seat.
The full argument, with sources, is in the companion post: AI Agents Are the API Economy's Biggest Customers: What Agencies Should Charge.
Outcome-based pricing: when it's cheaper, when it's riskier
What changed (Aug 30–31, 2026): OpenAI has begun letting some of its largest enterprise customers pay only when its AI actually completes the job — the example cited is a customer-support interaction handled end-to-end — instead of per token, per query, or per compute time. For those accounts it replaces conventional usage-based pricing. The arrangement is limited to select major accounts, is not generally available, and OpenAI has not formally announced it — the terms, the customers, and the prices are all unknown (TNW / The Information, Aug 31, 2026).
When outcome pricing wins: it shifts retry and failure risk to the vendor. Under per-token billing you pay for every failed attempt, retry, and fan-out loop; under outcome billing the vendor absorbs failures. So outcome pricing is usually cheaper when your success rate is low or your workload burns many calls per completed task — exactly the workloads where a per-resolution price like Intercom's $0.99 or Zendesk's ~$1.20–$1.50 Verified Resolution undercuts your effective token cost. That is why per-success prices "cluster around a dollar rather than a cent": the vendor is loading the failure rate into the price.
When it's riskier: at high volume on a cheap model, the per-outcome price can exceed your effective per-token cost. The worked example above shows a DeepSeek V4.1-Flash workload at ~$0.09 per resolution — paying $0.99–$2.00 per outcome for the same work is 11–22x more expensive. And "completed" needs a definition: a resolution is countable, but agentic work involves multi-step tasks where completion is a matter of judgment. Stripe's outcome-pricing guidance warns that without explicit attribution rules, customers will dispute whose outcome it was. Price per-outcome only when success is measurable and the contract defines it.
Use the outcome-based mode above to test your own workload: set your tasks/month and per-outcome price, and the calculator shows the monthly and yearly delta against the same work done per-token. Related: AI Agency Pricing 2026: Sell Outcomes, Not Services and OpenAI's Discount Experiment: What 13.8x Usage Growth Means for Agency Pricing.
How to price agent work for clients (agent-first)
- Pass through API costs with a documented margin (10–20%). Put token/model costs in the contract as a pass-through line at a known markup, re-baselined quarterly. Transparency is your defense when prices move — and they move weekly in 2026.
- Or price per outcome. If an agent resolves a ticket or closes a lead, bill per resolution — Intercom's $0.99 and Salesforce's $2 are client-accepted anchors. Your loaded cost per resolution (calculator output above) tells you the floor.
- Keep a retainer floor + usage overage. Flat retainers still work for oversight, governance, and maintenance — the human-in-the-loop work that doesn't scale with tokens. Add a metered bucket for agent consumption on top. Hybrid is the safest structure in 2026.
- Model the loops, not just the tokens. Quote retries, subagent fan-out, and context reloads explicitly, and offer budget rails (hard cap + kill switch) as a sellable feature. The agencies that forecast honestly win the renegotiation.
One blind spot in the math above: it prices tokens per call, and a voice agent does not bill that way. Audio transport is metered by the minute — billed by the second, with the WebRTC init floor charged even on a session that lasts a few seconds — while the reasoning behind each turn is a second, separate meter with its own model, cache, tool and search rates. Before you quote a voice agent, run the same workload through the GPT-Live-1 voice + backend cost calculator, which prints both meters per call, per day and per month and shows which one actually drives the invoice.
Model the full picture: setup, retainers, and margin
Open the AI Agency Pricing Calculator →The full calculator covers setup fees, retainers, model strategy (DeepSeek, Gemini, Grok 4.6, Qwen, Sonnet 5, local), failure/retry risk, and margin — current 2026 rates.
Frequently asked questions
How do I estimate the API cost of an AI agent workload?
Multiply the number of agents by average API calls per agent per day and working days per month to get monthly calls. Then multiply by tokens per call (or price per call) and the provider's per-token rate, and add agency overhead for retries, oversight, and integration. The formula on this page does all of that live.
What is a realistic API cost per AI agent per month in 2026?
It depends on model and volume. On current 2026 rates a moderately active agent making ~200 calls/day with ~8K input / 1.2K output tokens per call runs about $50–$150/month in raw API cost on mid-tier models (Claude Sonnet 5 $2/$10 per 1M, Grok 4.6 $2/$6) before agency overhead — and a fraction of that on cheap models like DeepSeek V4.1-Flash off-peak ($0.15/$0.60 per 1M).
What is agent-first pricing?
Agent-first pricing bills on agent activities, completed tasks, outcomes, or resources used — not human seats. Examples: Salesforce Agentforce at $2 per conversation and $0.10 per action, Intercom Fin at $0.99 per resolution. For agencies it means passing through API costs with a documented margin, pricing per outcome, or charging a retainer floor plus usage-based overage.
Why do AI agent API bills exceed simple token estimates?
Agents don't make one clean call per task. They fan out subagents, retry failed steps, reload context, and bill every intermediate call. A single agent task can query thousands of endpoints (Cloudflare's Matthew Prince cited ~5,000 sites vs ~5 for a human). Add a workload-complexity multiplier for retries and fan-out or your estimate will be low by 1.5–3x.
Can I pay for AI only when it works?
Yes, but only for select large enterprise customers so far. In late August 2026, OpenAI began letting some of its largest customers pay only when its AI actually completes the job — for example a customer-support interaction handled end-to-end — instead of per token or per call. The arrangement is limited to select major accounts, is not generally available, and OpenAI has not formally announced it; the terms, customers, and prices are not public. Use the outcome-based mode on this calculator to model what that would cost your workload.
When is outcome-based pricing cheaper than per-token, and when is it riskier?
Outcome-based pricing shifts retry and failure risk to the vendor: you pay only when the AI completes the job, so it is cheaper when your success rate is low or your workload burns many calls per completed task. It is riskier at high volume when the per-outcome price exceeds your effective per-token cost — for example a $0.99 per resolution price vs $0.09 per resolution of token cost on a cheap model. Market references for per-resolution pricing: Intercom Fin $0.99 per conversation resolved, Zendesk Verified Resolutions ~$1.20–$1.50, HubSpot Customer Agent $0.50, Salesforce Agentforce $2 per conversation.
How much should I mark up API costs when billing clients?
A transparent 10–20% pass-through margin is the defensible norm — the margin compensates for forecasting risk, monitoring, and re-baselining, not just the tokens. If you price per outcome instead, set the per-resolution price at or below established benchmarks (Intercom Fin $0.99, Salesforce Agentforce $2) while keeping your loaded cost per resolution well under it.
What does an OpenAI Agents API session cost?
There is no per-session fee: OpenAI states there are no additional fees for using the Agents API, so a session bills as model tokens at your chosen model's rates, plus tool calls (web search $10.00 per 1,000 calls, file search $2.50 per 1,000 calls), plus container-minutes when you use a hosted sandbox ($0.03 / $0.12 / $0.48 / $1.92 per 20-minute session at 1, 4, 16 and 64 GB, billed by the minute with a 5-minute minimum), plus $0.05 per minute if you put the GPT-Live-1 voice front end on it. Worked example: a 6-minute voice support session on a GPT-6 Astra backend with a 4 GB container and two web searches comes to $0.9660 per session, of which the voice layer alone is $0.3000. Rates checked September 18, 2026.
Does OpenAI charge a per-session or per-step fee for agents?
No. OpenAI publishes no per-session fee for the Agents API and denies one in writing — there are no additional fees for using the Agents API — and its pricing page carries no Agents API section at all. There is no published per-step or per-turn rate either: what you pay per step is that step's tokens and its tool calls. The session duration cap, session concurrency cap and free tier are also not published; the only published clock is a one-hour sandbox idle timeout that OpenAI says is not configurable.
Sources
- PYMNTS, "AI Agents Become the API Economy's Biggest New Customers" (Aug 24, 2026): pymnts.com
- a16z / OpenRouter, State of AI — 100T-token study, 1T tokens/day (Dec 4, 2025): a16z.com/state-of-ai
- Cloudflare Radar — bots >50% of HTML requests; Prince ~5,000 sites per agent task (Jun 6, 2026): stackfutures.com
- Salesforce Agentforce pricing — $2/conversation, $0.10/action (list, mid-2026): eesel.ai
- Intercom Fin — $0.99/resolution (Mar 3, 2026): myaskai.com
- TokenCost AI Price Index — 200x token price collapse (Mar 20, 2026): tokencost.app
- OpenRouter Blog, "GPT 5.6 Discounts & Jevons Paradox" (Aug 25, 2026): openrouter.ai
- OpenAI API pricing (platform docs, verified Aug 28, 2026 — Terra $2/$12, Luna $0.20/$1.20, Sol $4/$20 promo): platform.openai.com/docs/pricing
- Per-1M model rates as of Aug 28, 2026 (DeepSeek, Google, Together AI, SpaceXAI, Anthropic, OpenAI) — see AI Model Cost per Task 2026 for the full table and dated sources
- DeepSeek API docs, Models & Pricing —
deepseek-flashoff-peak $0.15 input / $0.60 output / $0.003 cached input per 1M, peak exactly 2× (Sept 10, 2026 sheet; the legacydeepseek-v4-flashIDs are aliases billed at the Flash price): api-docs.deepseek.com/quick_start/pricing - OpenAI, “Introducing the Agents API” (Sept 10, 2026) — public beta, the open-source Codex harness as a managed service, and the platform-fee position: openai.com/index/introducing-the-agents-api (this host is 403'd by openai.com; the copy read for this page was the Wayback capture of 2026-09-16, cross-checked against the live developer docs below)
- OpenAI, Agents API overview — what is managed (sessions, orchestration, context compaction, recovery) and that tokens bill at the selected model's rates: developers.openai.com/api/docs/guides/agents-api/overview
- OpenAI, API pricing — read 2026-09-18 for every managed-session rate on this page: web search $10.00 / 1k calls, file search $2.50 / 1k calls with storage $0.10 / GB per day (1 GB free), containers 1/4/16/64 GB $0.03 / $0.12 / $0.48 / $1.92 per 20-minute session billed by the minute with a 5-minute minimum, and GPT-6 Astra $10.00 in / $1.00 cached / $50.00 out per 1M: developers.openai.com/api/docs/pricing
- OpenAI, Voice latency and cost — GPT-Live-1 is $0.05 per minute of session, billed per second without rounding up, and billable time includes silence and backend work: developers.openai.com/api/docs/guides/voice-latency-cost
- OpenAI, GPT-Live-1 model page — concurrency caps 25 / 50 / 200 / 300 / 500 across Tiers 1–5, free tier unsupported: developers.openai.com/api/docs/models/gpt-live-1
- OpenAI, OpenAI-hosted environments — the one-hour sandbox idle timeout, “This timeout isn’t configurable”: developers.openai.com/api/docs/guides/agents-api/environments/openai-hosted
Accuracy note: The calculator uses published 2026 list rates (Aug 28, 2026 snapshot, plus DeepSeek's Sept 10, 2026 V4.1-Flash sheet) and user-supplied inputs; all figures are estimates, not guarantees. GPT-5.6 Terra ($2/$12) and Luna ($0.20/$1.20) are OpenAI's official Jul 30, 2026 list prices; the usage multiples (Luna 13.8x, Terra 5.6x, Sol 1.1x control) and ~1/3 retention are OpenRouter's measured discount-window data (Jul 27 – Aug 14, 2026), attributed as such (research brief t_5c8841ca). The "fastest-growing class of API consumers" framing follows PYMNTS (Aug 24, 2026) — no public dataset measures agents' absolute share of API traffic. Salesforce and Intercom prices are list prices as of mid-2026; real bills can stack additional platform fees. Outcome-based mode: OpenAI's outcome-based pricing (Aug 30–31, 2026 reporting via TNW/The Information) is limited to select enterprise accounts, not generally available, and unannounced; the per-outcome market references (Intercom $0.99, Zendesk ~$1.20–$1.50, HubSpot $0.50, Salesforce $2) are list prices as of mid-2026 and real contract terms vary (research brief t_9974fbbc). Re-verify provider rates before quoting clients — 2026 pricing moves weekly (Gemini 3.8 Flash Sept 2, GPT-5.6 Terra/Luna Jul 30, GPT-5.6 Sol Aug 21, DeepSeek Aug 16, DeepSeek V4.1-Flash Sept 10, Gemini 3.7 Flash Aug 13, Grok 4.6 Aug 12, Claude Sonnet 5 Aug 10).
Managed-session mode (added Sept 18, 2026): every rate in that block is published by OpenAI and was read on 2026-09-18 from developers.openai.com/api/docs/pricing, the Agents API overview and the voice latency & cost guide. The per-session fee is published as zero — OpenAI states there are no additional fees for using the Agents API — and is therefore rendered as a verified state, not an estimate; no placeholder price appears anywhere in the managed block. The session duration cap, the session concurrency cap, the per-step fee and the free tier are not published by OpenAI and are rendered as states rather than defaults, and container-minutes-per-session is labelled as the reader's own estimate because OpenAI publishes the container rate and its 5-minute minimum but no session-length rule for agents. The three preset totals ($0.0652, $0.9660 and $0.6585 per session) are this page's arithmetic on those published rates: they are reproducible line by line in the calculator, and they match the September 2026 research pack (card t_ee08c89d) and the figures published on findaiagency.com. A managed monthly total is infrastructure cost only — it excludes the agency's own overhead, governance and accountability work, which the sliders above still model.