DeepSeek V4 Flash Benchmarks

The official 0731 release was only re-post-trained — same architecture, same size — yet agent benchmarks now far exceed V4 Pro Preview.

Agentic Benchmarks

BenchmarkV4 Flash ScoreNote
Terminal-Bench 2.049.1%Exceeds V4 Pro Preview
MCP Atlas64%Tool-use orchestration
Toolathlon40.7%Multi-tool chaining
Claw-Eval57.8%Agent reliability
Gert Labs54.35%Real-world agent tasks

Coding Benchmarks

BenchmarkV4 Flash ScoreNote
SWE-bench VerifiedTop tierOpen-weight leader
LiveCodeBenchCompetitivevs models 4x the size
Aider PolyglotStrongMulti-language coding
HumanEval+HighCode generation accuracy
Terminal-Bench CodingLeadingCLI-based coding tasks

Reasoning Benchmarks

BenchmarkV4 Flash ScoreNote
GPQA DiamondCompetitiveGraduate-level QA
MATH-500StrongMathematical reasoning

What the Numbers Mean

The headline: a 13B-active model now matches or exceeds the agent performance of its 49B-active sibling. DeepSeek says the architecture is identical to the preview — the improvement comes entirely from post-training. That means V4 Pro's official release will likely close the gap, but for now V4 Flash is the strongest agent model per dollar in the open-weight space.

Caveat: multilingual and multimodal benchmarks are not measured. V4 Flash is a text-only model optimized for agent, coding, and reasoning tasks. If you need vision or audio, look at DeepSeek's separate VL line.

Benchmark FAQ

Why does V4 Flash beat V4 Pro on agent tasks?
The 0731 release applies improved post-training specifically targeting agent capabilities. DeepSeek's changelog states the architecture is unchanged — the gains come entirely from better training, suggesting V4 Pro's official release will likely improve too.
What is Terminal-Bench 2.0?
A benchmark that evaluates AI models on real terminal-based coding and system administration tasks. V4 Flash scores 49.1%, which exceeds V4 Pro Preview — notable because Pro has nearly 4x the active parameters.
What is MCP Atlas?
MCP Atlas tests a model's ability to use tools through the Model Context Protocol. V4 Flash scores 64%, indicating strong capability in multi-tool orchestration — a critical skill for AI agents.
Are these benchmarks verified?
Agentic benchmarks (5 of 5) and coding benchmarks (5 of 5) are verified from published sources. Reasoning has 2 verified scores. Multilingual and multimodal categories are not measured — It is text-only.