umans/status/umans-deepseek-v4-flash-0731
Live · refreshes every 30s
← all models
Umans DeepSeek V4 Flash Experimental
umans-deepseek-v4-flash-0731 · DeepSeek-V4-Flash · DeepSeek
In testing
314.2tok/s
throughput · p50 · last 5 min
1.66s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent model. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-glm-5.2.

90 days agoin production since Aug 3, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
now 321.2 tok/s
90 days agopre-release before Aug 3, 2026today
TTFT p50 · time to first token, lower is better
best 1.50s · Aug 1now 1.52s
90 days agopre-release before Aug 3, 2026today
Changelog

Events for Umans DeepSeek V4 Flash

incl. gateway-wide announcements
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.