The 2026 frontier race stopped being a capability race and became a benchmark-selection game. DeepSeek's new experimental vision model is the latest exhibit: a vendor claim of parity framed around the competitor a vendor chose, on the benchmarks a vendor selected. The math, read off DeepSeek's own published table, is less generous than the headline.
On that table, DeepSeek-V4-Flash-Vision-Exp beats Opus 4.8 on 3 of 11 benchmarks. ApexBench, the highest-leverage entry for agents, posts 36.5 for DeepSeek against 39.4 for Opus 4.8, a 2.9-point gap the rest of the table does not close. That is "close to" only on the rows the comparison was framed around.
Opus 4.8 was Anthropic's flagship when DeepSeek picked the comparison, and Claude Opus 5 shipped on July 24, six weeks before the V4-Flash-Vision-Exp release. The yardstick DeepSeek measured against is no longer Anthropic's. Parity against last month's leader is not parity against today's leader.
The repeatable check: when a vendor claims parity, read the vendor's own table and count the wins, then read the launch date of the comparison target. A 3-of-11 scorecard against a model already replaced at the top is marketing math, and the market will keep rewarding the framing until buyers read the table before the headline.
Reported by Sky for Type0, from Change Log | DeepSeek API Docs. Read the original: api-docs.deepseek.com