Groq API Pricing and Model Catalog: Use-Case Guide

Groq offers ultra-low latency inference on open-weight models. Here's how to use their API cost-effectively.

|

Pricing data sourced from our catalog. Check data sources for provenance and freshness.

What is Groq?

Groq is an inference provider specializing in:

  • Ultra-low latency: Sub-100ms time-to-first-token
  • Open-weight models: Llama, Mixtral, Gemma, and more
  • Speed-optimized: Built for real-time applications
  • Simple pricing: Predictable per-token costs

Groq model catalog

Chat models

Model Input Output Context
llama-3.1-8b-instant
$0.0500 $0.0800 8K
gemma-7b-it
$0.0500 $0.0800 8K
gpt-oss-20b
$0.0750 $0.3000 33K
gpt-oss-safeguard-20b
$0.0750 $0.3000 66K
llama-guard-4-12b
$0.2000 $0.2000 8K

Cost estimation example

Let's estimate costs for a typical Groq chat workload:

Input tokens per request: 500
Output tokens per request: 200
Requests per month: 100,000
Monthly cost: $4.10

Groq strengths

  • Ultra-low latency: Sub-100ms time-to-first-token
  • Open-weight models: Access to the latest open-weight models
  • Simple pricing: Predictable per-token costs
  • No vendor lock-in: Use the same models elsewhere
  • Real-time applications: Ideal for chatbots and interactive apps

When to use Groq

  • Real-time chat: When latency matters most
  • Interactive applications: When users expect instant responses
  • Open-weight models: When you want model flexibility
  • Cost-sensitive workloads: When you need predictable costs

Cost optimization tips

  • Use smaller models: For less complex tasks
  • Batch processing: Process multiple requests together
  • Caching: Cache responses for repeated queries
  • Right-size your model: Match model size to task complexity
  • Monitor usage: Track token usage to optimize costs

Compare Groq models

Ready to compare Groq models side by side? Use our tools:

Related guides

Frequently asked questions

What is Groq's cheapest model?

Groq's Llama 3.1 8B and Gemma 2 9B models are typically the most cost-effective.

How does Groq compare to direct providers?

Groq offers lower latency but may have higher per-token costs than direct providers like OpenAI or Anthropic.

Does Groq offer batch pricing?

Groq focuses on real-time inference and does not currently offer batch pricing.

Pricing data sourced from official provider documentation. Prices may vary by region and usage tier.