DeepSeek began beta testing V4.1 Flash on Sept. 8, a new model the Hangzhou-based artificial intelligence lab describes as faster, cheaper and natively multimodal — and the first real look at where its flagship Flash line is heading next.
The company announced the test in an official WeChat group notice, according to Gate News. Developers can reach the model through the existing API by leaving the base URL unchanged and setting the model name to deepseek-v4.1-flash-expires-on-0910. Pricing stays the same as deepseek-v4-flash, and each account is capped at 20 concurrent requests. The model name advertises its own deadline: the route is expected to disappear on Sept. 10.
A new architecture, with images built in
DeepSeek's notice says V4.1 Flash uses a new model architecture with native multimodal support, stronger overall capabilities, higher speed and lower cost. Native multimodality would mean text and images are handled by the same model rather than routed through a separate vision system — a notable shift for a lab whose flagship models have been text-first.
The direction has been visible for weeks. On Aug. 21, DeepSeek released an experimental vision variant, V4-Flash-Vision-Exp, which it said matched the official model on text tasks while improving visual-understanding benchmarks. Early testers reported mixed experiences: some described full image input, while one developer's setup saw none. A direct call to the beta API, however, returns an accurate description of an uploaded test image — confirming the capability is live.
Testers report a large speed jump
Speed, however, is not in dispute. Chinese crypto outlet MetaEra reported initial tests hitting 420 tokens per second, generating more than 71,000 tokens in under three minutes. Developers on X and the NVIDIA developer forums logged decode rates between 190 and 350 tokens per second, with one tester calling the model "over 2x faster" than V4 Flash with lower time-to-first-token.
For reference, Artificial Analysis measures the previous V4 Flash 0731 at about 122 tokens per second on DeepSeek's own API. The same testers flagged a familiar trade-off: the model "still overthinks a lot," one wrote, while another said it "seems to be (over)thinking a lot."
A fast year for DeepSeek
The beta caps a rapid 2026. DeepSeek's V4 family arrived in April, and the V4 Flash 0731 refresh — 284 billion total parameters with 13 billion active at inference — scored 50 on the Artificial Analysis Intelligence Index, a 10-point jump over the original V4 Flash and six points ahead of V4 Pro. The company's aggressive cache pricing, including a roughly 98% discount on cache hits, keeps its cost per task well below Western rivals.
DeepSeek has also pushed on tools and open source. It shipped Harness v0.1, a developer preview for agent builders, on Aug. 13 under the MIT license — a release that drew more than 70,000 GitHub stars within a day — and brought V4 Pro to general availability the same day.
Why it matters for the diaspora
For overseas Chinese developers and startups, V4.1 Flash is a reminder of how fast Chinese open-weight models are closing the gap on price and speed. A one-million-token context window at Flash pricing makes long-document, bilingual and agentic workloads cheaper to run, and because DeepSeek's weights are open, diaspora teams can deploy the models on their own infrastructure instead of depending on a single cloud vendor.
What comes next
The Sept. 10 expiry makes this a test rather than a launch: no permanent model name, final price or public release date has been announced. Developers who want to compare results later should record their prompts, latency and token counts now. Given DeepSeek's pace this year, a fuller V4.1 release is likely to follow — but for the next two days, the fastest way to see where the Flash line is going is to switch a model ID and watch the tokens fly.
Cover image: West Lake and the Hangzhou skyline. Photo: Windmemories via Wikimedia Commons, CC BY-SA 4.0.



