AI Agent API Cost Calculator: Estimate Monthly Agent Workload Costs

Published August 24, 2026 · Updated September 18, 2026By ABD Legacy LLC
AI agent API costs Managed agents Agent-first pricing Agency billing

The short answer: a moderately active AI agent costs roughly $50–$150/month in raw API spend on mid-tier 2026 models, and a full agent deployment is easy to model once you separate the four cost drivers: number of agents, calls per agent per day, tokens (or price) per call, and agency overhead. This calculator walks an agency owner through each driver and turns it into a defensible monthly number — plus a suggested client price using agent-first pricing.

This page is the pricing companion to "AI Agents Are the API Economy's Biggest Customers: What Agencies Should Charge" — the full analysis of the PYMNTS report is there; the math is here. AI agents are the fastest-growing class of API consumers (PYMNTS, Aug 24, 2026), OpenRouter passed 1 trillion tokens/day in late 2025, and Cloudflare's Matthew Prince cites a single agent task querying ~5,000 sites versus ~5 for a human. That machine-scale consumption is exactly what an agency is paying for — and what it should bill for, deliberately.

The short answer for a managed session: since September 10, 2026 OpenAI also sells the agent as a service — the Agents API, a managed Codex harness — and that changes the shape of the bill without changing what you quote on. A managed session is still tokens at your chosen model's rates, but it adds three meters the application code used to be: tool calls (web search $10.00 per 1,000 calls; file search $2.50 per 1,000 calls, plus $0.10 per GB per day of retrieval storage after the first free GB), container-minutes ($0.03 / $0.12 / $0.48 / $1.92 per 20-minute session at 1, 4, 16 and 64 GB, billed by the minute with a 5-minute minimum), and — only if you put a voice front end on it — GPT-Live-1 at $0.05 per minute of open session, which is $3.00 per open hour. There is no per-session fee: OpenAI states there are no additional fees for using the Agents API, and its pricing page carries no Agents API section at all. The managed block at the bottom of the calculator models all of it, with three scenario presets — a prototype agent, a production cloud agent on managed infrastructure, and an agency-managed deployment.

Capacity context: Anthropic's reported $45 billion Nscale commitment and its SpaceX Colossus 1 compute deal signal real Claude capacity growth — but also pricing pressure ahead of the record IPO. For the full breakdown, see Anthropic's $45B Nscale compute deal: what it means for Claude capacity and pricing.

Agent API Cost Estimator

40%
25%
30%
Monthly raw API cost—
Total API calls / month—
Input tokens / month—
Output tokens / month—
Fully-loaded cost (raw + overhead)—
Cost per agent / month (loaded)—
Cost per 1,000 calls (loaded)—

Agent-first client pricing (suggested)

Suggested client price (loaded × (1+markup))—
Equivalent price per resolution*—
Implied gross margin at that price—
*Assuming 50 API calls per client-facing outcome — tune to your actual workflow. Benchmarks for agent-first pricing: Intercom Fin $0.99 per resolution; Salesforce Agentforce $2 per conversation + $0.10 per action. A markup on cost converts to margin as markup ÷ (1 + markup): 30% markup ≈ 23% margin.
All figures are estimates on current published 2026 list rates (see reference table below). Your real bills will vary with model choice, caching, off-peak scheduling, and retry patterns — re-run quarterly.

Agent reliability / durable execution

New Sept 18, 2026

Every mode above prices work that succeeds. This row prices the layer that keeps an agent going when something fails — the compute a failed session re-runs while its state is recovered and the retry finishes. It is additive: the reliability layer is added to whichever headline total your calculator mode is showing. It is not a per-session charge (the Agents API per-session fee is published as $0.00 above, and stays that way), and it is not a vendor quote: no durable-execution platform publishes a list price, which is why the flat fee below defaults to $0.00. All five inputs are this page's model — replace them with your own logs.

Reliability / durable execution — line items

model defaults — tune to your own logs
Failed sessions / month—
Recovery time / month (failed sessions)—
Retry compute (failed sessions repriced)—
Orchestration platform fee, monthly—
Reliability layer — monthly—
All-in monthly (headline total + reliability)—

Market anchor — Temporal's $550M Series E at a $12.55B valuation, 14 September 2026 (Temporal · The Next Web). The round funds the layer Temporal calls Durable Execution, which keeps applications, “including AI agents”, “running when something fails” — the argument for pricing reliability as its own line, not for buying one vendor's stack.

How the math works

Four steps, each with one formula. The calculator runs these live; the same formulas are what you'd put in a spreadsheet or a quote.

  1. Monthly calls. agents × calls/agent/day × working days × complexity. The complexity multiplier is where agents differ from humans: retries re-pay full context, subagent fan-out re-reads context per turn, and loops still bill. 1.5× is a defensible default for real agentic work; heavy fan-out runs closer to 3×. The retry and fan-out multipliers, worked on a fixed 10,000-task workload, decide whether a metered routing layer pays for itself — a layer that adds a pass of its own makes the bill worse, not better.
  2. Monthly tokens. monthly calls × tokens per call, split into input and output. A typical agent turn sends a large system prompt + conversation context (input-heavy) and returns a smaller completion (output).
  3. Raw API cost. (input tokens ÷ 1M × input price × (1 − cache hit rate)) + (input tokens ÷ 1M × cache price × cache hit rate) + (output tokens ÷ 1M × output price). Caching matters: a 40% input cache hit rate at $0.20/1M vs a $2.00/1M miss price cuts the input leg nearly in half. For per-call pricing, raw cost is simply monthly calls × price per call.
  4. Fully-loaded cost. raw × (1 + overhead). Overhead covers the human and tooling layer — monitoring, oversight, integration, prompt maintenance, re-baselining — which doesn't scale with tokens but is real cost. Then cost per agent = loaded ÷ agents.

Outcome-based mode (new). Switch the mode selector to Outcome-based to compare the same workload against paying per completed task instead of per token. Enter your expected completed tasks per month and the price per completed outcome (market references cluster around $0.50–$2.00 per resolution — Intercom Fin $0.99, Zendesk Verified Resolutions ~$1.20–$1.50, HubSpot Customer Agent $0.50, Salesforce Agentforce $2). The calculator shows monthly and yearly cost for both models side by side, plus the delta. Outcome pricing transfers retry/failure risk to the vendor: it tends to win when your success rate is low or your workload burns many calls per completion, and loses at high volume when the per-outcome price exceeds your effective per-token cost.

Monthly raw API cost = (monthlyCalls × inTok ÷ 1M × inPrice × (1 − cacheHit)) + (monthlyCalls × inTok ÷ 1M × cachePrice × cacheHit) + (monthlyCalls × outTok ÷ 1M × outPrice) Fully-loaded = raw × (1 + overhead%) Suggested client price = loaded × (1 + markup%) (30% markup on cost ≈ 23% gross margin) Cost per agent = loaded ÷ agents Cost per resolution = loaded ÷ (monthlyCalls ÷ 50) Outcome-based mode: Outcome monthly = completedTasks × pricePerOutcome Outcome yearly = outcomeMonthly × 12 Per-token monthly = loaded (same workload) Delta = outcomeMonthly − loaded (negative = outcome cheaper)

One dimension the token math above does not carry is capacity. An API budget scales with spend, but the voice layer of the same stack is limited in concurrent sessions — GPT-Live-1 runs 25 / 50 / 200 / 300 / 500 across Tiers 1–5, and the Free tier cannot call the model at all — so a phone deployment can be comfortably affordable and still fail at the busy hour. The concurrent-session sizing for a voice line (sessions, not requests per minute) applies Little's Law to a call centre line and shows where 1,000 and 1,500 calls a day land against the Tier 1 ceiling. Size calls, not RPM, before you quote the retainer.

Reference: current per-1M rates (Aug 2026) — the inputs for your own math

Same reference table as our per-task cost benchmarks, current as of Aug 28, 2026 except the DeepSeek rows, which are on DeepSeek's Sept 10, 2026 sheet. Rates move fast — re-verify before quoting a client.

ModelInput ($/1M)Output ($/1M)Notes
GPT-5.6 Luna$0.20$1.20−80% Jul 30, 2026 (from $1/$6); 13.8x usage in the Jul 27 – Aug 14 discount window
DeepSeek V4.1-Flash (off-peak)$0.15$0.60API name deepseek-flash (DeepSeek sheet, Sept 10, 2026); peak $0.30/$1.20; cache-hit input $0.003 off-peak / $0.006 peak. The retired deepseek-v4-flash IDs are aliases billed at the Flash price
DeepSeek V4-Pro (off-peak)$0.66$1.98Peak $1.32/$3.96; cache-hit input $0.022 off-peak / $0.044 peak. DeepSeek V4 Pro vs V4.1-Flash: the withdrawn Sept 14 reroute and the 100x cached-input spread
Gemini 3.8 Flash (intro)$0.75$3.75Released Sept 2, 2026 — current Gemini Flash; intro through 2026-12-31, then $1.50/$7.50; 1M context
Gemini 3.7 Flash (intro — previous version)$0.75$3.75Intro through 2026-12-31, then $1.50/$7.50; 1M context; superseded by 3.8 Flash Sept 2, 2026
Meta Muse Glimmer (hosted, Together AI)$0.35$1.50Local self-host = electricity after hardware amortized
GPT-5.6 Terra$2.00$12.00−20% Jul 30, 2026 (from $2.50/$15); 5.6x usage in the discount window
Grok 4.6 (SpaceXAI)$2.00$6.00500K context; cache hit $0.50
Qwen 3.8 Max (open weights)$2.00$6.00Hosted API list price
Claude Sonnet 5$2.00$10.00Permanent as of Aug 10, 2026
GPT-5.6 Sol$4.00$20.00Official OpenAI promo (Aug 21 – Nov 21, 2026); cached input $0.40

OpenAI's discount experiment (Jul 27 – Aug 14, 2026): when Terra and Luna were discounted 50% via OpenRouter, daily Terra token usage rose 5.6x and Luna 13.8x, while Sol at list price rose only 1.1x (control); ~1/3 of users who tried a discounted model kept using it after expiry. Terra/Luna went from 0.7% to 7.8% of OpenRouter tokens. The same demand elasticity applies to any client workload you quote on a discounted model. See OpenAI's discount experiment: what 13.8x usage growth means for agency pricing for the full analysis.

Managed-harness containers: per 20-minute session per container (1, 4, 16, 64 GB)

One runtime line does not belong in the per-1M table above, because its unit is not tokens — the same reason our per-task benchmarks sit in their own table rather than in the per-1M columns. This is the line an agency pays when OpenAI hosts the Codex harness: a per-container session price by memory tier, on top of model tokens and tool calls, with no Agents-API-specific fee and no Agents API row on the pricing page. The full managed-session meter set — tool calls, container minutes, retrieval storage, voice minutes and the token pass-through — is priced in the September 2026 managed-session section below.

Container memoryOpenAI printed rate (per 20-minute session per container)Derived per-minute · 5-min floor (derived, not published)Note
1 GB$0.03$0.0015 · $0.0075developers.openai.com/api/docs/pricing#built-in-tools, read 2026-09-12; the row is labelled "Hosted Shell and Code Interpreter" and does not name the Agents API; the footnote bills container sessions OpenAI calls "Eligible" and they are billed by the minute, with a 5-minute minimum per session; the per-minute dollar figures are derived (tier ÷ 20), OpenAI publishes none; "not a final bill" (usage fields are best-effort, can be null).
4 GB$0.12$0.006 · $0.03
16 GB$0.48$0.024 · $0.12
64 GB$1.92$0.096 · $0.48

Basis: the four printed prices are OpenAI's, read from developers.openai.com/api/docs/pricing#built-in-tools on 2026-09-12 — 1 GB $0.03, 4 GB $0.12, 16 GB $0.48, 64 GB $1.92 per 20-minute session per container. Everything in the third column is this page's arithmetic on two quoted OpenAI statements, and is labelled derived because OpenAI prints no per-minute dollar figure for this row: per-minute = tier ÷ 20 (OpenAI's changelog for Jun 2, 2026 says eligible container sessions are billed per minute instead of at the full 20-minute session rate, and that "the underlying per-minute rate will remain the same"), and the floor = per-minute × 5 for the footnote's 5-minute minimum. At every tier the floor is 25% of the printed session price, the session price is 4.0× the floor, and the two units cross at exactly 20 billed minutes.

Worked illustration (derived, assumptions stated): one 4 GB container running 3 sessions a day across 22 working days is 66 sessions a month. At the printed session unit that is 66 × $0.12 = $7.92. If each of those 66 sessions bills 12 minutes, the derived per-minute figure gives 66 × 12 × $0.006 = $4.752 — cheaper than the session unit, because 12 minutes is short of the 20-minute crossover. Neither figure is a vendor quote. For what that line does to a client quote, see what this does to a client quote.

Inputs this meter needs, and what the estimator above does not take

The estimator at the top of this page is a per-month token-and-call model with a retry/fan-out multiplier and an outcome mode. None of the five inputs above exists in it, and this block does not add them — the same unit problem that keeps our per-task benchmark table out of the per-1M columns above.

What the container row does not say

What a managed agent session costs under the Agents API (September 2026)

OpenAI shipped the Agents API in public beta on September 10, 2026: the Codex harness as managed infrastructure, with OpenAI owning sessions, orchestration, context compaction and recovery. The pricing question the launch raises is not what does a session cost but which meter does this session hit. OpenAI's own answer is that the platform adds nothing — the company's stated position is that there are no additional fees for using the Agents API — so the six rows below are the whole bill, and five of them have a published rate.

MeterPublished rate (read Sept 18, 2026)What triggers it
Per-session fee$0.00 — verified, not an estimateNothing. There is no agents row and no agents section on OpenAI's pricing page; the platform fee is affirmatively denied.
Model token pass-throughYour selected model's own input, cached-input and output ratesEvery turn's prompt and completion. Cached input is the only discount — GPT-6 Astra is $10.00/1M in and $50.00/1M out, against $1.00/1M for cached input.
Web search$10.00 per 1,000 calls, plus search-content tokens at model ratesEach search the agent runs. The tokens the search returns bill twice over — once as the call, once as context.
File search$2.50 per 1,000 calls; storage $0.10 per GB per day, first 1 GB freeCalls bill per session. Storage bills per day, whether or not a session runs — it is the only line on this page that accrues while the agent is idle.
Hosted container (sandbox)1 GB $0.03 · 4 GB $0.12 · 16 GB $0.48 · 64 GB $1.92, per 20-minute sessionBilled by the minute with a 5-minute minimum, on hosted sandboxes only — a no-environment session pays nothing here.
GPT-Live-1 voice front end$0.05 per minute of open session, billed per second without rounding upWall-clock voice time, silence included, which is $3.00 per open hour. Backend model and tool usage bill separately, so a "$0.05/minute voice agent" is a front-end price, not a total.

Not published — and therefore not guessed at: per-step or per-turn fee, session duration cap, session concurrency cap, free tier. The only published clock is a one-hour sandbox idle timeout that OpenAI says is not configurable, and the only published caps anywhere in this stack are GPT-Live-1's concurrent sessions (25 / 50 / 200 / 300 / 500 across Tiers 1–5, with no free tier). The calculator prints each of those as a state instead of a default, and labels container-minutes-per-session as your own estimate, because OpenAI publishes the container rate and its 5-minute minimum but no session-length rule for agents.

The three presets, hand-checked

Preset buttons in the calculator set every input below; the arithmetic is reproduced here so you can check it without clicking anything. All rates are OpenAI's, read 2026-09-18; every line is this page's arithmetic on those rates.

(a) Prototype agent — GPT-5.6 Luna backend, no container, no voice tokens (60,000 × $0.20 + 120,000 × $0.02 + 9,000 × $1.20) ÷ 1M = $0.0252 web 4 ÷ 1,000 × $10.00 = $0.0400 per session = $0.0652 440 sessions (20/day × 22 days) = $28.69 / month (b) Production cloud agent on managed infra — GPT-Live-1 voice + GPT-6 Astra backend voice 6.0 min × $0.05/min = $0.3000 tokens (30,000 × $10.00 + 60,000 × $1.00 + 5,000 × $50.00) ÷ 1M = $0.6100 sandbox 6 min at $0.12 per 20-min session (5-min minimum applies) = $0.0360 web 2 ÷ 1,000 × $10.00 = $0.0200 per session = $0.9660 30,000 sessions (1,000/day × 30 days) = $28,980.00 / month (c) Agency-managed deployment — five clients, GPT-5.6 Terra backend, retrieval voice 8.0 min × $0.05/min = $0.4000 tokens (40,000 × $2.00 + 80,000 × $0.20 + 6,000 × $12.00) ÷ 1M = $0.1680 sandbox 8 min at $0.12 per 20-min session = $0.0480 web 3 ÷ 1,000 × $10.00 = $0.0300 file 5 ÷ 1,000 × $2.50 = $0.0125 per session = $0.6585 4,400 sessions (200/day × 22 days) = $2,897.40 storage (2 GB − 1 GB free) × $0.10/GB/day × 22 days = $2.20 managed month total = $2,899.60

The finding that matters for a quote: on preset (b) the voice layer is 31% of the session bill on a frontier backend — and 81% of it if you swap the backend for a cheap one. The same 6-minute session on GPT-5.6 Luna instead of GPT-6 Astra costs $0.3692 instead of $0.9660, of which $0.3000 is still voice. Model routing moves the token line; only closing idle voice sessions, or not using voice, moves the voice line. These two figures are the same ones on the managed-agents section of the API-economy analysis, so the two pages cannot quietly disagree.

Two practical notes on quoting from this block. First, the managed monthly total is an infrastructure number: the overhead and markup sliders in the estimator above apply to it exactly as they do to the self-hosted rows, because monitoring, re-baselining and accountability are still yours. Second, if the deployment takes phone calls, the binding constraint may be concurrency rather than cost — GPT-Live-1 caps at 25 concurrent sessions on Tier 1, so see concurrent-session sizing for a voice line and the GPT-Live-1 voice + backend cost calculator before you promise a number.

What a realistic deployment costs (worked example)

Default calculator inputs — 10 agents, 200 calls/day, 22 days, 1.5× complexity, 8K in / 1.2K out tokens, Claude Sonnet 5 rates ($2/$10), 40% input cache hit, 25% overhead:

Now move the model lever and the picture changes by an order of magnitude:

Model choice is the single biggest lever in the model — which is why routing small tasks to cheap models is a margin decision, not a footnote.

The same question, managed. If OpenAI hosts the harness instead, the unit changes and the calculator's third mode takes over. Preset (b) — a production voice support line on managed infrastructure — prices at $0.9660 per session and $28,980.00 a month at 1,000 sessions a day, with the voice layer alone taking 31% of that session (81% if you drop to a cheap backend). Preset (a), a prototype agent with no sandbox, comes to $0.0652 a session. The per-line arithmetic for all three presets, and the meters the self-hosted rows above cannot see — tool calls, container minutes, retrieval storage and voice minutes — are in the managed-session section above.

Why the PYMNTS picture makes this math urgent

PYMNTS reported (Aug 24, 2026) that AI agents are the fastest-growing class of API consumers — and that the API economy's pricing, identity, and trust infrastructure "were not designed for this" and "are being rebuilt now." Three consequences for an agency's cost model:

The full argument, with sources, is in the companion post: AI Agents Are the API Economy's Biggest Customers: What Agencies Should Charge.

Outcome-based pricing: when it's cheaper, when it's riskier

What changed (Aug 30–31, 2026): OpenAI has begun letting some of its largest enterprise customers pay only when its AI actually completes the job — the example cited is a customer-support interaction handled end-to-end — instead of per token, per query, or per compute time. For those accounts it replaces conventional usage-based pricing. The arrangement is limited to select major accounts, is not generally available, and OpenAI has not formally announced it — the terms, the customers, and the prices are all unknown (TNW / The Information, Aug 31, 2026).

When outcome pricing wins: it shifts retry and failure risk to the vendor. Under per-token billing you pay for every failed attempt, retry, and fan-out loop; under outcome billing the vendor absorbs failures. So outcome pricing is usually cheaper when your success rate is low or your workload burns many calls per completed task — exactly the workloads where a per-resolution price like Intercom's $0.99 or Zendesk's ~$1.20–$1.50 Verified Resolution undercuts your effective token cost. That is why per-success prices "cluster around a dollar rather than a cent": the vendor is loading the failure rate into the price.

When it's riskier: at high volume on a cheap model, the per-outcome price can exceed your effective per-token cost. The worked example above shows a DeepSeek V4.1-Flash workload at ~$0.09 per resolution — paying $0.99–$2.00 per outcome for the same work is 11–22x more expensive. And "completed" needs a definition: a resolution is countable, but agentic work involves multi-step tasks where completion is a matter of judgment. Stripe's outcome-pricing guidance warns that without explicit attribution rules, customers will dispute whose outcome it was. Price per-outcome only when success is measurable and the contract defines it.

Use the outcome-based mode above to test your own workload: set your tasks/month and per-outcome price, and the calculator shows the monthly and yearly delta against the same work done per-token. Related: AI Agency Pricing 2026: Sell Outcomes, Not Services and OpenAI's Discount Experiment: What 13.8x Usage Growth Means for Agency Pricing.

How to price agent work for clients (agent-first)

  1. Pass through API costs with a documented margin (10–20%). Put token/model costs in the contract as a pass-through line at a known markup, re-baselined quarterly. Transparency is your defense when prices move — and they move weekly in 2026.
  2. Or price per outcome. If an agent resolves a ticket or closes a lead, bill per resolution — Intercom's $0.99 and Salesforce's $2 are client-accepted anchors. Your loaded cost per resolution (calculator output above) tells you the floor.
  3. Keep a retainer floor + usage overage. Flat retainers still work for oversight, governance, and maintenance — the human-in-the-loop work that doesn't scale with tokens. Add a metered bucket for agent consumption on top. Hybrid is the safest structure in 2026.
  4. Model the loops, not just the tokens. Quote retries, subagent fan-out, and context reloads explicitly, and offer budget rails (hard cap + kill switch) as a sellable feature. The agencies that forecast honestly win the renegotiation.

One blind spot in the math above: it prices tokens per call, and a voice agent does not bill that way. Audio transport is metered by the minute — billed by the second, with the WebRTC init floor charged even on a session that lasts a few seconds — while the reasoning behind each turn is a second, separate meter with its own model, cache, tool and search rates. Before you quote a voice agent, run the same workload through the GPT-Live-1 voice + backend cost calculator, which prints both meters per call, per day and per month and shows which one actually drives the invoice.

Model the full picture: setup, retainers, and margin

Open the AI Agency Pricing Calculator →

The full calculator covers setup fees, retainers, model strategy (DeepSeek, Gemini, Grok 4.6, Qwen, Sonnet 5, local), failure/retry risk, and margin — current 2026 rates.

Frequently asked questions

How do I estimate the API cost of an AI agent workload?

Multiply the number of agents by average API calls per agent per day and working days per month to get monthly calls. Then multiply by tokens per call (or price per call) and the provider's per-token rate, and add agency overhead for retries, oversight, and integration. The formula on this page does all of that live.

What is a realistic API cost per AI agent per month in 2026?

It depends on model and volume. On current 2026 rates a moderately active agent making ~200 calls/day with ~8K input / 1.2K output tokens per call runs about $50–$150/month in raw API cost on mid-tier models (Claude Sonnet 5 $2/$10 per 1M, Grok 4.6 $2/$6) before agency overhead — and a fraction of that on cheap models like DeepSeek V4.1-Flash off-peak ($0.15/$0.60 per 1M).

What is agent-first pricing?

Agent-first pricing bills on agent activities, completed tasks, outcomes, or resources used — not human seats. Examples: Salesforce Agentforce at $2 per conversation and $0.10 per action, Intercom Fin at $0.99 per resolution. For agencies it means passing through API costs with a documented margin, pricing per outcome, or charging a retainer floor plus usage-based overage.

Why do AI agent API bills exceed simple token estimates?

Agents don't make one clean call per task. They fan out subagents, retry failed steps, reload context, and bill every intermediate call. A single agent task can query thousands of endpoints (Cloudflare's Matthew Prince cited ~5,000 sites vs ~5 for a human). Add a workload-complexity multiplier for retries and fan-out or your estimate will be low by 1.5–3x.

Can I pay for AI only when it works?

Yes, but only for select large enterprise customers so far. In late August 2026, OpenAI began letting some of its largest customers pay only when its AI actually completes the job — for example a customer-support interaction handled end-to-end — instead of per token or per call. The arrangement is limited to select major accounts, is not generally available, and OpenAI has not formally announced it; the terms, customers, and prices are not public. Use the outcome-based mode on this calculator to model what that would cost your workload.

When is outcome-based pricing cheaper than per-token, and when is it riskier?

Outcome-based pricing shifts retry and failure risk to the vendor: you pay only when the AI completes the job, so it is cheaper when your success rate is low or your workload burns many calls per completed task. It is riskier at high volume when the per-outcome price exceeds your effective per-token cost — for example a $0.99 per resolution price vs $0.09 per resolution of token cost on a cheap model. Market references for per-resolution pricing: Intercom Fin $0.99 per conversation resolved, Zendesk Verified Resolutions ~$1.20–$1.50, HubSpot Customer Agent $0.50, Salesforce Agentforce $2 per conversation.

How much should I mark up API costs when billing clients?

A transparent 10–20% pass-through margin is the defensible norm — the margin compensates for forecasting risk, monitoring, and re-baselining, not just the tokens. If you price per outcome instead, set the per-resolution price at or below established benchmarks (Intercom Fin $0.99, Salesforce Agentforce $2) while keeping your loaded cost per resolution well under it.

What does an OpenAI Agents API session cost?

There is no per-session fee: OpenAI states there are no additional fees for using the Agents API, so a session bills as model tokens at your chosen model's rates, plus tool calls (web search $10.00 per 1,000 calls, file search $2.50 per 1,000 calls), plus container-minutes when you use a hosted sandbox ($0.03 / $0.12 / $0.48 / $1.92 per 20-minute session at 1, 4, 16 and 64 GB, billed by the minute with a 5-minute minimum), plus $0.05 per minute if you put the GPT-Live-1 voice front end on it. Worked example: a 6-minute voice support session on a GPT-6 Astra backend with a 4 GB container and two web searches comes to $0.9660 per session, of which the voice layer alone is $0.3000. Rates checked September 18, 2026.

Does OpenAI charge a per-session or per-step fee for agents?

No. OpenAI publishes no per-session fee for the Agents API and denies one in writing — there are no additional fees for using the Agents API — and its pricing page carries no Agents API section at all. There is no published per-step or per-turn rate either: what you pay per step is that step's tokens and its tool calls. The session duration cap, session concurrency cap and free tier are also not published; the only published clock is a one-hour sandbox idle timeout that OpenAI says is not configurable.

Sources

Accuracy note: The calculator uses published 2026 list rates (Aug 28, 2026 snapshot, plus DeepSeek's Sept 10, 2026 V4.1-Flash sheet) and user-supplied inputs; all figures are estimates, not guarantees. GPT-5.6 Terra ($2/$12) and Luna ($0.20/$1.20) are OpenAI's official Jul 30, 2026 list prices; the usage multiples (Luna 13.8x, Terra 5.6x, Sol 1.1x control) and ~1/3 retention are OpenRouter's measured discount-window data (Jul 27 – Aug 14, 2026), attributed as such (research brief t_5c8841ca). The "fastest-growing class of API consumers" framing follows PYMNTS (Aug 24, 2026) — no public dataset measures agents' absolute share of API traffic. Salesforce and Intercom prices are list prices as of mid-2026; real bills can stack additional platform fees. Outcome-based mode: OpenAI's outcome-based pricing (Aug 30–31, 2026 reporting via TNW/The Information) is limited to select enterprise accounts, not generally available, and unannounced; the per-outcome market references (Intercom $0.99, Zendesk ~$1.20–$1.50, HubSpot $0.50, Salesforce $2) are list prices as of mid-2026 and real contract terms vary (research brief t_9974fbbc). Re-verify provider rates before quoting clients — 2026 pricing moves weekly (Gemini 3.8 Flash Sept 2, GPT-5.6 Terra/Luna Jul 30, GPT-5.6 Sol Aug 21, DeepSeek Aug 16, DeepSeek V4.1-Flash Sept 10, Gemini 3.7 Flash Aug 13, Grok 4.6 Aug 12, Claude Sonnet 5 Aug 10).

Managed-session mode (added Sept 18, 2026): every rate in that block is published by OpenAI and was read on 2026-09-18 from developers.openai.com/api/docs/pricing, the Agents API overview and the voice latency & cost guide. The per-session fee is published as zero — OpenAI states there are no additional fees for using the Agents API — and is therefore rendered as a verified state, not an estimate; no placeholder price appears anywhere in the managed block. The session duration cap, the session concurrency cap, the per-step fee and the free tier are not published by OpenAI and are rendered as states rather than defaults, and container-minutes-per-session is labelled as the reader's own estimate because OpenAI publishes the container rate and its 5-minute minimum but no session-length rule for agents. The three preset totals ($0.0652, $0.9660 and $0.6585 per session) are this page's arithmetic on those published rates: they are reproducible line by line in the calculator, and they match the September 2026 research pack (card t_ee08c89d) and the figures published on findaiagency.com. A managed monthly total is infrastructure cost only — it excludes the agency's own overhead, governance and accountability work, which the sliders above still model.