DeepSeek's Retrained V4-Flash Now Beats Its Own Flagship on Every Agent Benchmark
DeepSeek-V4-Flash-0731 posts 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, outscoring the larger V4-Pro-Preview on all nine published agentic and coding benchmarks despite using only 13B active parameters.
DeepSeek released the official version of DeepSeek-V4-Flash-0731 on July 31, and the numbers invert what “flagship” is supposed to mean. The retrained Flash build now outscores DeepSeek’s own larger V4-Pro-Preview on all nine agentic and coding benchmarks the company has published — a smaller, cheaper model beating the bigger one it was supposed to sit beneath.
The jump from the earlier Flash preview is stark. On Terminal Bench 2.1, a benchmark that measures how well a model handles real command-line and agentic tasks, V4-Flash-0731 scores 82.7 versus 72.1 for V4-Pro-Preview and just 61.8 for the original Flash preview. DeepSWE, which tests software-engineering task completion, jumped from 7.3 to 54.4 — roughly a sevenfold gain in a single retraining pass. On the Artificial Analysis Intelligence Index, the model lands at 50, putting it in range of models several times its active parameter count.
What makes those numbers notable is architectural. V4-Flash-0731 keeps 284 billion total parameters but activates only 13 billion per token, runs a one-million-token context window, and ships under the MIT license — the most permissive open license in common use, with no field-of-use or commercial restrictions at all. DeepSeek didn’t scale up to get these gains; it ran a new round of post-training, the phase that shapes how a model deploys knowledge it already has, rather than adding parameters or data volume. That’s a materially cheaper way to close the gap with larger models, and it’s a repeatable playbook other labs will now have to answer.
Context: this lands the same week LG shipped a 750-billion-parameter K-EXAONE 2.0 and just over a week after Anthropic’s AMD compute deal and Opus 5 launch. Open-weight labs are no longer trailing frontier closed labs by matching scale — DeepSeek just demonstrated that post-training alone, on a mid-sized MoE, can beat a same-family model with dramatically more active compute. For teams choosing between hosted frontier APIs and self-hosted open weights for agentic coding workloads, V4-Flash-0731’s price-to-performance ratio at 13B active parameters is now hard to ignore.
One caveat worth stating plainly: all of these benchmark numbers are vendor-reported by DeepSeek and haven’t yet been independently verified by third-party evaluators. Treat the relative gains as directionally credible given the consistency across nine benchmarks, but wait for outside verification before betting production infrastructure on the absolute scores.