DeepSeek V4.1 Flash
DeepSeek · Chat model
DeepSeek V4.1 Flash is listed here as a chat model from DeepSeek. This page shows simple API pricing, token limits, and capability flags so you can compare it with similar options.
Provider and model identifiers are kept in their original form for accuracy.
deepseek-deepseek-flash
| Period | Input | Cached input | Output |
|---|---|---|---|
| Off-peak All times outside the Peak windows, including weekends and Chinese public holidays. | $0.1500 / 1M tokens | $0.0030 / 1M tokens | $0.6000 / 1M tokens |
| Peak Monday-Friday 01:00-04:00 and 06:00-10:00 UTC, excluding Chinese public holidays. | $0.3000 / 1M tokens | $0.0060 / 1M tokens | $1.2000 / 1M tokens |
Off-peak rates are half of Peak rates. Peak hours use UTC and exclude Chinese public holidays; all other times are Off-peak.
Source checked: 2026-09-21 · Official pricing
Catalog generated: Sep 22, 2026
Quick read
Best for
Use this page when you need a fast view of cost, context size, and supported features before testing the model in your own workload.
Things to verify
Always check the provider page for discounts, cache pricing, region rules, and any model limits that may not appear in public metadata.
Same base model routes
These route variants are shown only when reviewed source metadata ties exact provider routes to the same base model. Prices remain route-specific. Reviewed identity: deepseek-flash.
| Route | Route ID | Scope | Input | Output | Open |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash DeepSeek Current route | deepseek-flash Base model ID | Base | N/A | N/A | Open route |
| DeepSeek V4 Flash (legacy alias) DeepSeek | deepseek-v4-flash Provider alias | Base | N/A | N/A | Open route |
Pricing
No matching data
Limits
Capabilities
| Capability | Supported |
|---|---|
| Vision | Supported |
| Function calling | Supported |
| Parallel function calling | - |
| Tool choice | - |
| Prompt caching | Supported |
| Reasoning | Supported |
| Response schema | - |
| System messages | - |
| Audio input | - |
| Audio output | - |
| Web search | - |
| PDF input | - |
| Video input | - |
| Native streaming | - |
| Computer use | - |
| Assistant prefill | - |
| Structured output | - |
| Output config | - |
| URL context | - |
Related articles
Articles relevant to this model's provider, capabilities, and use cases.
LLM API Provider Pricing Comparison
Compare OpenAI, Anthropic, Google, Mistral, and DeepSeek across cheapest chat, mid-range, and reasoning model pricing.
Reasoning LLM API Pricing Guide
Which reasoning models are cheapest, which are most expensive, and when to use reasoning vs non-reasoning models.
DeepSeek Provider Guide
DeepSeek-V3-0324 and DeepSeek-R1 use cases, pricing, and when to choose DeepSeek.
Model Routing Cascade
A four-tier Budget, Mid, Premium, and Reasoning model selection framework for routing queries to the lowest-cost suitable tier.
Sources
| Source links | |
| Pricing data | LiteLLM model cost map |
| Synced at | 2026-09-21 |
| Pricing verified | 2026-09-21 |
| Catalog generated | 2026-09-22T00:01:59.924Z |