Groq API Pricing and Model Catalog: Use-Case Guide
Groq offers ultra-low latency inference on open-weight models. Here's how to use their API cost-effectively.
Pricing data sourced from our catalog. Check data sources for provenance and freshness.
What is Groq?
Groq is an inference provider specializing in:
- Ultra-low latency: Sub-100ms time-to-first-token
- Open-weight models: Llama, Mixtral, Gemma, and more
- Speed-optimized: Built for real-time applications
- Simple pricing: Predictable per-token costs
Groq model catalog
Chat models
| Model | Input | Output | Context |
|---|---|---|---|
| llama-3.1-8b-instant | $0.0500 | $0.0800 | 8K |
| gemma-7b-it | $0.0500 | $0.0800 | 8K |
| gpt-oss-20b | $0.0750 | $0.3000 | 33K |
| gpt-oss-safeguard-20b | $0.0750 | $0.3000 | 66K |
| llama-guard-4-12b | $0.2000 | $0.2000 | 8K |
Cost estimation example
Let's estimate costs for a typical Groq chat workload:
Groq strengths
- Ultra-low latency: Sub-100ms time-to-first-token
- Open-weight models: Access to the latest open-weight models
- Simple pricing: Predictable per-token costs
- No vendor lock-in: Use the same models elsewhere
- Real-time applications: Ideal for chatbots and interactive apps
When to use Groq
- Real-time chat: When latency matters most
- Interactive applications: When users expect instant responses
- Open-weight models: When you want model flexibility
- Cost-sensitive workloads: When you need predictable costs
Cost optimization tips
- Use smaller models: For less complex tasks
- Batch processing: Process multiple requests together
- Caching: Cache responses for repeated queries
- Right-size your model: Match model size to task complexity
- Monitor usage: Track token usage to optimize costs
Compare Groq models
Ready to compare Groq models side by side? Use our tools:
Related guides
Open-weight inference provider pricing comparison
Compare Groq, Fireworks, DeepInfra, and Together AI for open-weight inference.
Cross-provider pricing comparison
How pricing compares across OpenAI, Anthropic, Google, Mistral, and DeepSeek.
Building a chatbot
How to choose models and optimize costs for chatbot applications.
Hidden costs of LLM APIs
Rate limits, latency, evaluation overhead, and vendor risk beyond per-token pricing.
Frequently asked questions
What is Groq's cheapest model?
Groq's Llama 3.1 8B and Gemma 2 9B models are typically the most cost-effective.
How does Groq compare to direct providers?
Groq offers lower latency but may have higher per-token costs than direct providers like OpenAI or Anthropic.
Does Groq offer batch pricing?
Groq focuses on real-time inference and does not currently offer batch pricing.