Umans DeepSeek V4.1 Flash
164.8tok/s
throughput · p50 · last 5 min
676ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h
DeepSeek V4.1 Flash, served from the official DeepSeek-V4.1-Flash release: DeepSeek's latest flash model, a 552B mixture-of-experts (8B active per token on prefill, 16B on decode) built for fast agentic coding, with native image understanding and a 1M-token context window. Reasoning has four modes: non-think (none), think low (low), think high (high, the default) and think max (max). Billed per token ($0.15 / $0.60 / $0.028 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability.
Context
1049K
Max output
393K
Recommended
393K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
Weights
Trends
Speed over the last 90 days
90 days agotoday
90 days agotoday
Changelog
Events for Umans DeepSeek V4.1 Flash
No recent events.