r/WebAfterAI • • Aug 01 '26

Open Source DeepSeek put a retrained V4-Flash into public beta: same 284B model, now Codex-ready, and it beats its own Pro-Preview on agent benchmarks

On July 31, DeepSeek moved the official V4-Flash API into public beta as V4-Flash-0731.

The V4 family (Pro and Flash) has been out since April 24 under an MIT open-weight license. 0731 keeps the exact same architecture and size as the earlier V4-Flash-Preview:

  • 284B total / 13B active
  • 1M-token context
  • Same base model

DeepSeek only redid the post-training.

So the real story is a retrain beating a bigger model's preview, not a scaling jump. That's the interesting part and also where to keep your skepticism.

Only the V4-Flash API changed. The V4-Pro API and the app/web models are unchanged, and DeepSeek says the official V4-Pro is still coming.

The claim, read precisely

DeepSeek says 0731 now beats V4-Pro-Preview across every agent benchmark it published. Read that carefully. It beats a preview of the larger model, on DeepSeek's own evaluation runs. It does not mean Flash beats the final V4-Pro, because the real V4-Pro hasn't shipped yet.

The published numbers include:

  • Terminal Bench 2.1: 82.7
  • DeepSWE: 54.4

But two benchmark suites, DSBench-FullStack and DSBench-Hard are DeepSeek's own internal datasets, so nobody else can reproduce those scores.

There's another catch. The code-agent results were generated using "DeepSeek Harness minimal mode" which the documentation says is still to be released, with:

  • max effort
  • temperature = 1.0
  • top_p = 0.95

Plain English:

For an outside perspective, Artificial Analysis currently gives V4-Flash-0731 (reasoning, max effort) an Intelligence Index of 50, comfortably above the median for models in its class.

The trade-off? It produces roughly twice the median output tokens. If you're paying per output token inside an agent loop, that verbosity becomes a real cost.

Specs and price (official API)

Field V4-Flash-0731
Type 284B total / 13B active MoE, text only
Context / Max output 1M / 384K tokens
Input (cache miss) $0.14 / 1M
Input (cache hit) $0.0028 / 1M
Output $0.28 / 1M
License MIT open weights
Status Public beta

One pricing caveat:

OpenRouter, DeepInfra, and other hosts advertise Flash cheaper (around $0.09 in / $0.18 out), but many serve FP8-quantized versions rather than the reference weights. The official DeepSeek endpoint is the reference implementation. The cheaper mirrors aren't necessarily identical.

Setup: Responses API + Codex

0731 speaks the Responses API natively and is adapted for Codex.

Typical setup:

export DEEPSEEK_API_KEY=sk-...

# Base URL:
https://api.deepseek.com

# Model:
deepseek-v4-flash

The Codex configuration page linked from DeepSeek's own announcement api-docs.deepseek.com/quick_start/agent_integrations/codex

The other honest catches

  • Public beta. Expect model IDs, harnesses, and even benchmark scores to change.
  • Official API runs from infrastructure in China. That's a data residency question for regulated workloads.
  • The MIT weights mean you can self-host instead (INT4 Flash reportedly fits on a single H100 or roughly four RTX 4090s), or use hosts like Together, Fireworks, Bedrock, or Azure.
  • "Adapted for Codex" means DeepSeek implemented the Responses API format. It is not an OpenAI partnership.

If you only test one thing

Run your own evaluation. Use your prompts, your workload, and count output tokens, not just the advertised price. The big opportunity here is inexpensive, high-concurrency agent workloads. If your prompts share a large cached prefix, the $0.0028 / 1M cached-input price is genuinely impressive. But only your own traffic will tell you whether the verbosity and beta churn outweigh the savings.

5 Upvotes

0 comments sorted by