Workload calculator

LLM API Cost Calculator

Estimate monthly API spend from tokens, requests, cache reuse, and batch pricing. Enter your workload first, or use a preset as a shortcut.

Usage preset

Choose a preset or scroll for custom input.

Explore benchmark scores to see which models perform best on specific tasks.

Input

Advanced options

Cost estimates use the generated model database last built on Aug 10, 2026. Pricing, lifecycle, and capability fields can be incomplete or provider-specific, so verify production decisions with the official provider.

How to use this page

Start with a simple preset, then change tokens and request volume to match your product. If you already know the workload type, the use-case pages give a more guided comparison.

How pricing is calculated

Costs are calculated as (tokens ÷ 1,000,000) × price per 1M tokens. Input and output tokens are priced separately — output is typically 3–5× more expensive. The cache hit rate reduces the effective input cost by applying a lower cached price to matched requests. The formula: total = (inputTokens × inputPrice + outputTokens × outputPrice) × requests × (1 − cacheDiscount).

Need concrete examples first?

Read the calculator examples guide for chatbot, RAG, summarization, and coding-agent inputs before changing the fields.

Read examples
Not sure if you need API pricing?

Check when a subscription is enough and when usage-based API pricing matters for product work.

API vs subscription
Which provider is cheapest overall?

Compare OpenAI, Anthropic, Google, Mistral, and DeepSeek across cheapest chat, mid-range, and reasoning model pricing.

Provider comparison
What do benchmark scores mean?

See MMLU, GPQA, HumanEval, and other benchmark scores across providers to understand model quality beyond pricing.

Benchmark scores
Want to compare models side by side?

Put 2-3 models next to each other to compare pricing, context windows, modalities, and capabilities.

Compare models