TRULYSOVEREIGN AI · RADAR

DeepSeek V4.1 Flash pricing dropped, max completion tokens cut by 66%

OpenRouter model catalogue · published 2026-09-24 03:31 UTC
ACTION NEEDED VERIFY_ONLY PRICINGLIMITSMODEL
WHAT CHANGEDEffective now: prompt pricing decreased from $0.15 to $0.14 per million tokens, completion pricing decreased from $0.60 to $0.42 per million tokens (30% reduction). Maximum completion tokens reduced from 393,216 to 131,072.
WHY IT MATTERSIf you generate responses longer than 131,072 tokens, calls will now fail or truncate. The pricing drop reduces costs for existing workloads but the token limit is a breaking change for long-form generation.
WHAT TO DOCheck your application logs for any DeepSeek V4.1 Flash responses exceeding 131,072 tokens in the past 30 days. If found, either switch to a model with higher limits or redesign the prompt to stay under the new ceiling.
CONFIDENCE 95% Primary source ↗ written by anthropic/claude-sonnet-4.5
Show the reasoning — raw diff, triage verdict, confidence
RAW DIFF
 id: deepseek/deepseek-v4.1-flash
 name: DeepSeek: DeepSeek V4.1 Flash
-pricing.prompt: 0.00000015
-pricing.completion: 0.0000006
+pricing.prompt: 0.00000014
+pricing.completion: 0.00000042
 context_length: 1048576
-top_provider.max_completion_tokens: 393216
+top_provider.max_completion_tokens: 131072
HOW IT WAS JUDGED
triage verdictMATERIAL at 0.98, needed 0.5
triage reasoningPricing dropped on both dimensions; output token limit cut significantly.
signalsprompt price decreased · completion price decreased · max_completion_tokens reduced by 66%
suppression
triage modelanthropic/claude-haiku-4.5 ($0.00000)
writer modelanthropic/claude-sonnet-4.5 ($0.00000)