BAD SIGNAL

← News

DeepSeek ships V4.1-Flash on its API with native multimodal and new pricing

10 Sep 2026: DeepSeek published DeepSeek-V4.1-Flash — a multimodal MoE model on the API as deepseek-flash, with 1M context. Company pricing page lists off-peak cached-input at $0.003 per 1M tokens. Benchmarks vs other labs stay DeepSeek’s claims.

10 Sep 2026: DeepSeek’s news page and API docs published DeepSeek-V4.1-Flash. Model name deepseek-flash. Native multimodal (vision). Legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp still route to V4.1-Flash at Flash pricing. That is the filing event.

Company specs: 552B MoE; Causal Encoder–Decoder; 8B active prefill / 16B decode; up to 1M context; max output 384K on the pricing table.

Pricing for deepseek-flash, per 1M tokens on the live Models & Pricing page: cache-hit input off-peak $0.003 / peak $0.006; cache-miss off-peak $0.15 / peak $0.3; output off-peak $0.6 / peak $1.2. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday.

DeepSeek says V4.1-Flash outperforms V4 Pro on its internal multi-party tests for performance, cost, speed, and total time — that is the company’s claim. The 10 Sep news post still described routing deepseek-v4-pro to V4.1-Flash at Flash rates after 12:00 Beijing time on 14 Sep 2026. The live pricing-page footnote now says that, in response to user demand, DeepSeek will continue the V4 Pro API after 14 Sep with unchanged billing and will give further notice before any change.

Weights are on Hugging Face at deepseek-ai/DeepSeek-V4.1-Flash. Cross-lab “beats GPT…” leaderboard lines stay DeepSeek’s claims, not a desk finding.

DeepSeek put a dated product + pricing page on its own domains for V4.1-Flash. File the ship and the posted rates; leave cross-lab leaderboard brags as the company’s claims.

ONLINE

article thread

guidelines

warming…

warming…

Sources