On this page
These are the two largest models in the current open-ish field. DeepSeek-V4-Pro runs 1.6 trillion parameters with 49 billion active; Qwen3.8-Max runs 2.4 trillion with 95 billion active. Both hold a million tokens of context.
They overlap on exactly one published benchmark, which is a problem for anyone hoping for a leaderboard and a useful lesson about how this season's releases report themselves.
| Metric | DeepSeek-V4-Pro | Qwen3.8-Max |
|---|---|---|
| Vendor | DeepSeek | Alibaba (Qwen) |
| Released | 2026-04-26 | 2026-08-03 |
| Parameters | 1.6T (MoE) | 2.4T (MoE) |
| Active params | 49B | 95B |
| Context window | 1M tokens | 1M tokens |
| Native vision | No | Yes |
| License | Open weights (MIT) | Not published (API-only at launch) |
| Input $/1M | $0.44 | $2 |
| Output $/1M | $0.87 | $6 |
Sources: DeepSeek V4 Pro API pricing, MarkTechPost: Alibaba releases Qwen3.8-Max
The one shared number
GPQA Diamond, a graduate-level science reasoning test: Qwen3.8-Max 92.6, DeepSeek-V4-Pro 90.1.
Qwen takes it by 2.5 points. On a single vendor-reported benchmark, between models trained by different labs on different harnesses, that is a result worth noting and not worth deciding anything on.
Everything else they published lands on separate tests:
| DeepSeek-V4-Pro | Qwen3.8-Max | ||
|---|---|---|---|
| LiveCodeBench | 93.5 | Terminal-Bench 2.1 | 86.6 |
| SWE-bench Multilingual | 76.2 | SWE-bench Pro | 67.7 |
| MMLU-Pro | 87.5 | FrontierSWE | 73.5 |
| DeepSWE v1.1 | 56.6 |
There is no honest way to combine those columns into a ranking. This is the version and harness problem in its purest form: two capable models, almost no common ground to stand them on.
So decide on what is knowable
Three things are unambiguous, and together they settle it for most uses.
Licence. DeepSeek-V4-Pro is MIT. Qwen3.8-Max had not published a licence at launch and was available through QwenCloud only, with a smaller Qwen3.8-27B sibling slated to open. If you need weights you can hold, that is the whole comparison.
Price. DeepSeek-V4-Pro lists at $0.44 in and $0.87 out. Qwen3.8-Max lists at $2.00 and $6.00. DeepSeek is roughly 4.5× cheaper on input and 7× cheaper on output.
| Model | Input $/1M | Output $/1M |
|---|---|---|
| DeepSeek-V4-Pro | $0.44 | $0.87 |
| Qwen3.8-Max | $2 | $6 |
Sources: DeepSeek V4 Pro API pricing, MarkTechPost: Alibaba releases Qwen3.8-Max
Modality. Qwen3.8-Max takes text, images and video. DeepSeek-V4-Pro is text-only. If your pipeline reads screenshots or scanned documents, that is decisive in the other direction.
Where Qwen3.8-Max earns its price
The multimodal results are the strongest part of Alibaba's case, and they are not small: MathVision 95.2, LogicVista 91.9, OSWorld-Verified 86.1. It also reports PaperBench at 93.0, ahead of the proprietary models Alibaba lined up against it.
Read as a whole, Qwen3.8-Max looks strong on agentic and multimodal work and less dominant on core coding, which is where its published SWE-bench Pro of 67.7 sits below the frontier proprietary comparisons in the same table.
If your workload is visual, meaning it reads interfaces, documents or video, Qwen3.8-Max is doing something DeepSeek-V4-Pro cannot do at all, and the price difference stops being the point.
Reading DeepSeek's numbers carefully
Two caveats specific to DeepSeek-V4-Pro, both worth knowing before quoting it.
The card reports a Pro-Max configuration. The benchmark table on the model card is labelled for the Pro-Max instruct variant, so treat those figures as the ceiling of the family rather than a guarantee for whichever endpoint you call.
Pricing has peak and off-peak tiers. DeepSeek moved to time-of-day pricing in August 2026, and output rates rise substantially at peak. The $0.87 above is the standard rate; budget against your actual traffic pattern rather than the headline.
Neither undermines the model. Both are the kind of detail that turns a surprise into a plan.
What we would use
For text-based agentic coding at scale, DeepSeek-V4-Pro is the stronger default of these two: MIT weights, a 1M context, and a price that makes high-volume work affordable. LiveCodeBench 93.5 and SWE-bench Multilingual 76.2 describe a genuinely capable coding model.
For anything that has to look at images or video, Qwen3.8-Max is the one that can, and the multimodal benchmarks say it does it well.
For most everyday engineering work, neither is the first thing to reach for. GLM-5.3-Flash at $0.15 and $0.50 is multimodal, MIT, and close enough on coding benchmarks that these giants are better used as escalation targets than as defaults, the pattern in running open models at frontier level.
Sources
Written by
Cho Yin Yong
Principal AI Solutions Engineer, XY Space
Principal AI Solutions Engineer at XY Space. University of Toronto lecturer for five years, co-author of two patents, winner of two competitive AI awards, and nine years of regulated engineering leadership.
More from Cho Yin YongShare this article
Work with us
We build the systems these posts describe, and we'll tell you in the first call whether yours is worth building.
Start a project