Compare LLM Models: Pricing, Context, Capabilities & Benchmarks

Pick two or three models and compare the main fields on one page. Compare OpenAI, Anthropic, Google, Mistral, DeepSeek, and other LLM API pricing, context windows, capabilities, and benchmark results side by side to find the best model for your workload.

Start faster

Use a preset if you have not picked models yet

These presets reuse existing article and use-case scenarios. They only preselect models; the comparison table still uses the current database fields.

Compact model article

Start with GPT-5.4 Mini, Gemini 3.1 Flash Lite, and Claude Haiku 4.5.

Low-cost chat shortlist

Open the lowest-cost text-output chat candidates from the current chatbot guide scenario.

RAG context shortlist

Open the context-ready low-cost candidates from the current RAG guide scenario.

Summarization shortlist

Open the low-cost input-heavy candidates from the current summarization guide scenario.

Coding-agent shortlist

Open the context-ready low-cost candidates from the current coding-agent guide scenario.

Budget tier models

Under $0.30 / 1M tokens. Fast and cheap for classification, extraction, and simple Q&A.

Mid-tier models

$0.30 to $2.00 / 1M tokens. Balanced capability for RAG, summarization, and customer chat.

Premium tier models

$2.00+ / 1M tokens. Frontier models for complex reasoning, coding, and analysis.

Reasoning tier models

Chain-of-thought models for math, logic, multi-step planning, and self-critique chains.

Agentic AI models

Function-calling models sorted by combined price. Top picks for tool-use and multi-step agent workflows.

Cache reuse models

Models with cached input pricing sorted by combined price. Best value when repeated context allows cache hits.

Batch processing models

Models with batch API pricing sorted by combined batch price. Cheapest paths for async, high-volume jobs.

Lowest output-cost ratio

Models where output is cheapest relative to input. Best starting point for generation-heavy workloads. See the output-vs-input pricing guide.

Model routing cascade

One budget, mid, premium, and reasoning model — the 60/30/10 rule for agentic routing: 60% budget, 30% mid, 10% premium+reasoning.

Open-weight providers

Cheapest model from each open-weight inference provider — Groq, Fireworks AI, DeepInfra, Together AI, and more. Prices are near 1:1 output-to-input ratio.

Top benchmark models

Highest-scoring models across SWE-bench, GPQA, HumanEval, and other major benchmarks. Start here to compare best-in-class capabilities.