On July 31, 2026, Artificial Analysis scored DeepSeek V4 Flash at 50, 10 points above its April preview and level with Gemini 3.6 Flash. Yet in a separate Scale AI and CAIS test, the best agents earned $1,810 of the $143,991 available—about 1.26% of the economic value.
Key takeaways
- Artificial Analysis scored DeepSeek V4 Flash at 50 on July 31, 2026—10 points above its April preview, tied with Gemini 3.6 Flash and one point behind GPT-5.6 Luna.
- V4 Flash’s AA-Omniscience accuracy remained at 37%, while its reported hallucination rate fell 11 percentage points to 84%.
- V4 Flash rose from 1,189 to 1,559 Elo on GDPval-AA v2, a gain of 370 points.
- DeepSeek launched the official V4 Flash API in public beta on July 31, while V4 Pro and its app and web models remained unchanged.
- DeepSeek priced V4 Flash at $0.14 per million input tokens, compared with $1.74 for V4 Pro.
V4 Flash gained 10 composite points while its AA-Omniscience accuracy stayed at 37%; its reported hallucination rate fell 11 percentage points, to 84%. On that benchmark, a buyer watching only the falling hallucination rate would miss flat accuracy.
The two records do not test the same system: the Remote Labor Index result is not a V4 Flash result and cannot rank DeepSeek against Gemini. The cited July 31 reports also identify no enterprise that shortlisted V4 Flash because of its 50. The score supports a trial; it does not document adoption or production reliability.
The score gets DeepSeek through the first gate
Artificial Analysis’s July 31 index placed V4 Flash one point behind GPT-5.6 Luna. Bloomberg reported the same day that DeepSeek had launched the official V4 Flash API in public beta, touting enhanced agent capabilities and benchmark results that surpassed the V4 Pro Preview.
DeepSeek’s public beta gave developers access to the model behind the new rank. A procurement team still has to load its own documents, connect its own tools, enforce its permissions and apply its error tolerances.
Labs, developers and buyers can reuse the same number for different decisions. DeepSeek can claim parity, a developer can sort APIs and a procurement team can choose a shortlist. The score has done its job once the shortlist exists.
DeepSeek’s gain separates accuracy from hallucination
DeepSeek retained the preview model’s architecture and size for V4 Flash. It released the upgrade through its API while V4 Pro, its app and its web model remained unchanged. Buyers should record the exact endpoint and version they test; the Flash API result applies to that surface.
Artificial Analysis introduced AA-Omniscience in November 2025 to test knowledge and hallucination across more than 40 topics. Its separate accuracy and hallucination fields expose a distinction that a composite score can obscure. A calibration claim would require confidence, abstention and evidence that confidence tracks correctness.
Before approving a document-search assistant, a procurement team can seed internal files with answerable and unanswerable questions. It should score correct answers, unsupported answers, abstentions and citations separately, then route high-risk misses to human review.
V4 Flash also rose from 1,189 to 1,559 Elo on GDPval-AA v2, an evaluation of agentic real-world work tasks. To turn that gain into a delegation decision, a buyer should replay representative jobs end to end and count accepted outputs, human corrections, unsafe actions and time to completion.
Each purchase claim needs its own gate
A procurement team buys several claims at once: capability, workflow performance, calibration, control, service quality and economics. Each claim needs its own evidence.
| Claim being purchased | Evidence the buyer should require | Failure question |
|---|---|---|
| General capability | Composite index with component-level results | Which underlying capability actually moved? |
| Workflow performance | Repeated end-to-end completion in a representative environment | Does performance survive task variation and tool use? |
| Calibration | Confidence, error, abstention and escalation rates | Does confidence track correctness, and does the system stop when needed? |
| Security and control | Permission tests, action logs and approval boundaries | Can the agent exceed its authorized scope? |
| Service quality | Latency distribution under expected load | Does responsiveness hold outside a demonstration? |
| Economics | Total cost per accepted outcome, including retries and review | What does completed work cost? |
| Deployment fit | Supported platforms, data boundaries and operating support | Can the organization run and govern the system it tested? |
Agent deployments add connector and approval failures, while changing prices can reverse a benchmark-led shortlist.
Agents expose connector and approval failures
An agent combines a model with tools, external data, orchestration and permissioned actions. The model proposes; the surrounding system decides what the proposal can read, change, spend or send.
The Scale AI and CAIS Remote Labor Index tested agents on economically valuable freelance tasks, including design, video editing, game development and administrative work.
That ratio measures economic value captured, not the share of tasks completed. The index does not report that 98.74% of tasks failed, and it did not test V4 Flash. It warns against translating any composite model score into a completed-work percentage.
A procurement team should run the same task through the model, connector, data source, orchestration policy and human approval path. It should record ambiguous tool returns, denied permissions, retries and escalations to see whether the agent can recover without exceeding its scope.
Reviewers should log corrections, preserve an audit trail and identify who can approve consequential actions. Once an agent can act, buyers must verify control separately from model quality.
A one-point lead can lose on cost
DeepSeek priced V4 Flash at $0.14 per million input tokens and V4 Pro at $1.74. A buyer still has to account for output tokens, retries, tool calls and human review.
On July 30, Axios reported that OpenAI cut the price of GPT-5.6 Luna by about 80% and GPT-5.6 Terra by 20% after improving serving efficiency. The next day, Artificial Analysis placed Luna one index point ahead of V4 Flash. Luna held that narrow lead after a large price change.
A buyer can spend lower serving costs on more attempts, larger contexts or wider deployment. Interactive teams also need latency measurements under expected load because a slow response can stop an otherwise economical workflow.
A procurement team should divide total spend—including output tokens, retries, tool calls and review—by accepted outcomes. That calculation can favor a model that loses the public index if it completes more work within the buyer’s budget and response-time limit.
Frequently asked questions
When will DeepSeek V4 Flash leave public beta?
The cited evidence gives no general-availability date. Buyers would need to ask DeepSeek about the release schedule, support terms and whether endpoints could change during beta.
Which benchmark components produced V4 Flash’s 10-point composite gain?
The piece does not provide that decomposition. The flat 37% AA-Omniscience accuracy shows that the composite increase cannot by itself establish an accuracy improvement.
Does V4 Flash have published calibration or abstention results?
None are reported here. The available figures cover accuracy and hallucination, not whether confidence tracks correctness or whether the model abstains appropriately.
What is V4 Flash’s full cost after output tokens and agent operations?
The cited $0.14 rate covers one million input tokens only. No complete figure is provided for output tokens, retries, tool calls, human review or cost per accepted result.
What benchmark score is high enough for production approval?
No universal cutoff is established in the evidence. A production threshold would depend on the buyer’s required accuracy, risk controls, latency, budget and tolerance for human review.
On July 31, the 50 put V4 Flash beside Gemini in a public index. The Remote Labor Index’s $1,810 result, from different systems on a different test, does not downgrade DeepSeek; it marks the claim the index cannot make. The 50 can carry V4 Flash into a trial. Only accepted work can carry it out.