All coverage
Milestone note Jul 31, 2026

DeepSeek ships V4-Flash in public beta — same architecture, better agent scores, and a price war underneath

DeepSeek released the public beta of V4-Flash (build 0731): the same 284B-total/13B-active MoE architecture as the April preview, with gains coming entirely from re-post-training. It scores above DeepSeek's own V4-Pro-Preview on all nine published agent and coding benchmarks, speaks OpenAI's Responses API format natively, and works with Codex from day one. OpenAI cut prices days earlier.

DeepSeek released the official public beta of its V4-Flash API on July 31, under the build designation V4-Flash-0731. The architecture is unchanged from the April preview — a 284B-total, 13B-activated mixture-of-experts — with the improvements coming from re-post-training rather than a new design. On the nine agent and coding benchmarks DeepSeek published, the 0731 build scores above the company's own larger V4-Pro-Preview. It natively speaks OpenAI's Responses API format and works with Codex from day one. Reporting around the launch (the-decoder) puts its cost well below comparable US flagship models; OpenAI cut its own prices days earlier, which investors including Michael Burry read as positioning ahead of this release.

Why it matters

The interesting engineering claim is that a smaller, cheaper model beat its own bigger sibling on agent tasks purely through post-training — evidence that the frontier is moving to training technique rather than parameter count, at exactly the moment Chinese labs are competing on cost. The interesting commercial claim is the API compatibility: adopting OpenAI's Responses format and Codex support removes switching friction by design, so the price gap converts into migrations rather than benchmarks. That is the same squeeze Kimi K3 and Qwen3.8 applied from the capability side, now aimed at the developer's integration layer — and it lands as DeepSeek raises at ~$74B ahead of an onshore IPO, giving the discount a balance sheet behind it.

What to watch

Independent benchmark replication now that the beta is public, whether OpenAI's price cuts hold or deepen, and any sign of enterprise migrations rather than developer experimentation.

Who's involved