LLM Fine-Tuning Cost Comparison: Training and Inference Pricing
Published 2026-08-01 · Updated 2026-08-01
Fine-tuning an LLM on your data adds three cost layers: training tokens, inference markup for the fine-tuned model, and optional hosting fees. This guide compares exact pricing across providers so you can estimate the total cost before starting.
Quick comparison: Training cost per 1M tokens
Training costs vary dramatically by model size and method. LoRA (Low-Rank Adaptation) is cheaper than full fine-tuning but may produce lower quality results for complex tasks.
| Provider | Method | Small (≤16B) | Medium (17-70B) | Large (70B+) |
|---|---|---|---|---|
| Together AI | LoRA SFT | $0.48 | $1.50 | $2.90 |
| Together AI | Full SFT | $1.20 | $3.75 | $7.25 |
| Fireworks AI | LoRA SFT | $1.00 | $6.00 | $12.00 |
| Fireworks AI | Full SFT | $2.00 | $12.00 | $24.00 |
| OpenAI | Supervised | $1.50 (nano) | $5.00 (mini) | $25.00 (4o/4.1) |
| AWS Bedrock | Managed | $1.49 (Llama 13B) | $7.99 (Llama 70B) | Contact sales |
All prices in USD per 1M training tokens. LoRA = Low-Rank Adaptation. SFT = Supervised Fine-Tuning. DPO = Direct Preference Optimization (typically 10-20% more than SFT).
Inference cost: Fine-tuned vs base model
After training, fine-tuned models typically cost more to run than the base model. Some providers charge 2-3x the base rate; others charge the same.
| Provider | Inference markup | Example |
|---|---|---|
| Fireworks AI | Same as base | Llama 3.1 70B: $0.90/1M input (same for fine-tuned) |
| Together AI | Dedicated endpoint | Billed by minute, not per-token |
| OpenAI | 2-3x base | GPT-4o: $3.75 → $15.00/1M output (4x for fine-tuned) |
| AWS Bedrock | Hourly hosting | $23.50/hour for provisioned throughput |
Key insight: Fireworks AI's "same as base" inference pricing makes it the most cost-effective for production fine-tuned workloads. Together AI's dedicated endpoints suit high-volume use cases where predictable pricing matters.
OpenAI fine-tuning: Winding down
OpenAI's fine-tuning platform is being wound down for new users. Existing fine-tuned models remain available, but new fine-tuning jobs are restricted. If you're currently using OpenAI fine-tuning, plan a migration to Together AI or Fireworks AI.
| Model | Training | Inference (Input) | Inference (Output) |
|---|---|---|---|
| GPT-4.1 | $25.00/1M | $3.00/1M | $12.00/1M |
| GPT-4.1 mini | $5.00/1M | $0.80/1M | $3.20/1M |
| GPT-4.1 nano | $1.50/1M | $0.20/1M | $0.80/1M |
| GPT-4o | $25.00/1M | $3.75/1M | $15.00/1M |
| GPT-4o mini | $3.00/1M | $0.30/1M | $1.20/1M |
OpenAI offers a 50% inference discount if you enable data sharing for model improvement.
Together AI: Most flexible pricing
Together AI offers the widest range of fine-tuning options with clear tiered pricing by model size. LoRA SFT starts at $0.48/1M tokens for models up to 16B parameters.
| Model | LoRA SFT | LoRA DPO | Full SFT | Minimum |
|---|---|---|---|---|
| Llama 4 Scout | $3.00/1M | $7.50/1M | Contact | $6.00 |
| Llama 4 Maverick | $8.00/1M | $20.00/1M | Contact | $16.00 |
| DeepSeek-R1 | $10.00/1M | $25.00/1M | Contact | $20.00 |
| Qwen3-235B | $6.00/1M | $15.00/1M | Contact | None |
Minimum charge: $4.00 per job. DPO (Direct Preference Optimization) is typically 10-20% more expensive than SFT but produces models better aligned with human preferences.
Fireworks AI: Best for production
Fireworks AI charges the same inference price for fine-tuned models as the base model. This makes it the most cost-effective choice for production workloads where you need consistent pricing.
| Model Size | LoRA SFT | Full SFT | Inference markup |
|---|---|---|---|
| ≤16B | $1.00/1M | $2.00/1M | None |
| 16-80B | $6.00/1M | $12.00/1M | None |
| 80-300B | $12.00/1M | $24.00/1M | None |
| >300B | $20.00/1M | $40.00/1M | None |
On-demand GPU pricing is also available for self-managed training: H100 at $7.00/hour, B200 at $10.00/hour, B300 at $12.00/hour.
Cost estimation example
Let's estimate the total cost of fine-tuning Llama 3.1 70B on a 500k-token dataset for 3 epochs, then running 100k inference requests/month.
| Component | Together AI (LoRA) | Fireworks AI (LoRA) |
|---|---|---|
| Training tokens | 500k × 3 = 1.5M | 500k × 3 = 1.5M |
| Training cost | 1.5M × $1.50/1M = $2.25 | 1.5M × $6.00/1M = $9.00 |
| Monthly inference (100k req × 2k tokens) | 200M tokens at dedicated rate | 200M tokens at base rate |
| Monthly inference cost | ~$180 (dedicated endpoint) | ~$180 (same as base) |
Bottom line: Training is a one-time cost ($2-9). The real cost is ongoing inference. Choose a provider with low or zero inference markup for production workloads.
When to fine-tune vs prompt engineering
Fine-tuning is not always the right choice. Consider these factors:
| Factor | Prompt engineering | Fine-tuning |
|---|---|---|
| Setup cost | Zero | $2-100+ training |
| Ongoing cost | Same as base model | 2-4x base (OpenAI) or same (Fireworks) |
| Quality improvement | Limited by context window | Significant for domain-specific tasks |
| Iteration speed | Immediate | Hours to days per training run |
| Data requirement | None | 1k-100k+ examples |
Start with prompt engineering and few-shot examples. Only fine-tune when you've measured a clear quality gap that prompting cannot close.
Start here
Use the cost calculator to estimate your monthly inference cost with different models. Then compare providers to find the best combination of training and inference pricing for your workload.