umans/status/umans-deepseek-v4.1-flash
Live · refreshes every 30s
← all models
Umans DeepSeek V4.1 Flash
umans-deepseek-v4.1-flash · DeepSeek-V4.1-Flash · DeepSeek
Operational
164.8tok/s
throughput · p50 · last 5 min
676ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4.1 Flash, served from the official DeepSeek-V4.1-Flash release: DeepSeek's latest flash model, a 552B mixture-of-experts (8B active per token on prefill, 16B on decode) built for fast agentic coding, with native image understanding and a 1M-token context window. Reasoning has four modes: non-think (none), think low (low), think high (high, the default) and think max (max). Billed per token ($0.15 / $0.60 / $0.028 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability.

90 days agotoday
Context
1049K
Max output
393K
Recommended
393K
Vision
Yes
Tools
Yes
Reasoning
Toggle · none/low/high/max
Weights
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
90 days agotoday
TTFT p50 · time to first token, lower is better
90 days agotoday
Changelog

Events for Umans DeepSeek V4.1 Flash

incl. gateway-wide announcements
No recent events.