← Back to articles

LLM Fine-Tuning Cost Comparison: Training and Inference Pricing

Published 2026-08-01 · Updated 2026-08-01

Fine-tuning an LLM on your data adds three cost layers: training tokens, inference markup for the fine-tuned model, and optional hosting fees. This guide compares exact pricing across providers so you can estimate the total cost before starting.

Quick comparison: Training cost per 1M tokens

Training costs vary dramatically by model size and method. LoRA (Low-Rank Adaptation) is cheaper than full fine-tuning but may produce lower quality results for complex tasks.

Provider Method Small (≤16B) Medium (17-70B) Large (70B+)
Together AI LoRA SFT $0.48 $1.50 $2.90
Together AI Full SFT $1.20 $3.75 $7.25
Fireworks AI LoRA SFT $1.00 $6.00 $12.00
Fireworks AI Full SFT $2.00 $12.00 $24.00
OpenAI Supervised $1.50 (nano) $5.00 (mini) $25.00 (4o/4.1)
AWS Bedrock Managed $1.49 (Llama 13B) $7.99 (Llama 70B) Contact sales

All prices in USD per 1M training tokens. LoRA = Low-Rank Adaptation. SFT = Supervised Fine-Tuning. DPO = Direct Preference Optimization (typically 10-20% more than SFT).

Inference cost: Fine-tuned vs base model

After training, fine-tuned models typically cost more to run than the base model. Some providers charge 2-3x the base rate; others charge the same.

Provider Inference markup Example
Fireworks AI Same as base Llama 3.1 70B: $0.90/1M input (same for fine-tuned)
Together AI Dedicated endpoint Billed by minute, not per-token
OpenAI 2-3x base GPT-4o: $3.75 → $15.00/1M output (4x for fine-tuned)
AWS Bedrock Hourly hosting $23.50/hour for provisioned throughput

Key insight: Fireworks AI's "same as base" inference pricing makes it the most cost-effective for production fine-tuned workloads. Together AI's dedicated endpoints suit high-volume use cases where predictable pricing matters.

OpenAI fine-tuning: Winding down

OpenAI's fine-tuning platform is being wound down for new users. Existing fine-tuned models remain available, but new fine-tuning jobs are restricted. If you're currently using OpenAI fine-tuning, plan a migration to Together AI or Fireworks AI.

Model Training Inference (Input) Inference (Output)
GPT-4.1 $25.00/1M $3.00/1M $12.00/1M
GPT-4.1 mini $5.00/1M $0.80/1M $3.20/1M
GPT-4.1 nano $1.50/1M $0.20/1M $0.80/1M
GPT-4o $25.00/1M $3.75/1M $15.00/1M
GPT-4o mini $3.00/1M $0.30/1M $1.20/1M

OpenAI offers a 50% inference discount if you enable data sharing for model improvement.

Together AI: Most flexible pricing

Together AI offers the widest range of fine-tuning options with clear tiered pricing by model size. LoRA SFT starts at $0.48/1M tokens for models up to 16B parameters.

Model LoRA SFT LoRA DPO Full SFT Minimum
Llama 4 Scout $3.00/1M $7.50/1M Contact $6.00
Llama 4 Maverick $8.00/1M $20.00/1M Contact $16.00
DeepSeek-R1 $10.00/1M $25.00/1M Contact $20.00
Qwen3-235B $6.00/1M $15.00/1M Contact None

Minimum charge: $4.00 per job. DPO (Direct Preference Optimization) is typically 10-20% more expensive than SFT but produces models better aligned with human preferences.

Fireworks AI: Best for production

Fireworks AI charges the same inference price for fine-tuned models as the base model. This makes it the most cost-effective choice for production workloads where you need consistent pricing.

Model Size LoRA SFT Full SFT Inference markup
≤16B $1.00/1M $2.00/1M None
16-80B $6.00/1M $12.00/1M None
80-300B $12.00/1M $24.00/1M None
>300B $20.00/1M $40.00/1M None

On-demand GPU pricing is also available for self-managed training: H100 at $7.00/hour, B200 at $10.00/hour, B300 at $12.00/hour.

Cost estimation example

Let's estimate the total cost of fine-tuning Llama 3.1 70B on a 500k-token dataset for 3 epochs, then running 100k inference requests/month.

Component Together AI (LoRA) Fireworks AI (LoRA)
Training tokens 500k × 3 = 1.5M 500k × 3 = 1.5M
Training cost 1.5M × $1.50/1M = $2.25 1.5M × $6.00/1M = $9.00
Monthly inference (100k req × 2k tokens) 200M tokens at dedicated rate 200M tokens at base rate
Monthly inference cost ~$180 (dedicated endpoint) ~$180 (same as base)

Bottom line: Training is a one-time cost ($2-9). The real cost is ongoing inference. Choose a provider with low or zero inference markup for production workloads.

When to fine-tune vs prompt engineering

Fine-tuning is not always the right choice. Consider these factors:

Factor Prompt engineering Fine-tuning
Setup cost Zero $2-100+ training
Ongoing cost Same as base model 2-4x base (OpenAI) or same (Fireworks)
Quality improvement Limited by context window Significant for domain-specific tasks
Iteration speed Immediate Hours to days per training run
Data requirement None 1k-100k+ examples

Start with prompt engineering and few-shot examples. Only fine-tune when you've measured a clear quality gap that prompting cannot close.

Start here

Use the cost calculator to estimate your monthly inference cost with different models. Then compare providers to find the best combination of training and inference pricing for your workload.