DeepSeek’s flagship AI model update underwhelms
DeepSeek’s updated V4 Pro AI model struggles on benchmarks, shines in cybersecurity
Chinese AI start-up’s model impresses in niche areas like cybersecurity but leaves some developers disappointed in its overall capabilities
Chinese artificial intelligence start-up DeepSeek has quietly released DeepSeek-V4-Pro-0813, an updated version of its latest flagship model, leaving some developers underwhelmed by its overall capabilities and disappointed in its pricing – but impressing researchers in niche areas like cybersecurity.
However, early benchmark results suggest the new flagship is struggling to match its top-tier rivals.
DeepSeek-V4-Pro-0813 scored 53 on the Artificial Analysis Intelligence Index, on par with Zhipu AI’s GLM-5.2 released in June, but was four points behind the mid-tier Terra model in OpenAI’s latest GPT-5.6 series and seven points behind Moonshot AI’s Kimi K3.
On the Vals Index, compiled by San Francisco-based Vals AI to evaluate models across several benchmarks, the new DeepSeek model ranked 12th. It trailed OpenAI’s previous-generation GPT-5.5 and also lagged well behind frontier systems like Kimi K3 and Anthropic’s Claude Opus 5.
The model struggled in two areas in particular: completing tasks within a sandboxed terminal environment and generating complex financial models in Excel spreadsheets, Vals AI said on Wednesday.
Related Stories
AI News
7 ways startups can stand out in the AI era
12 minutes ago
AI News
TIME Reveals the 2026 TIME100 AI List of the World’s Most Influential People in Artificial Intelligence
13 minutes ago
AI News
Filipino theologians explore how Church should respond to AI
14 minutes ago
Google names 28 startups as part of second AI for Energy Accelerator program
1 hour ago
AI News
❝The question is not whether AI writes, but whether humans still think
1 hour ago
AI News
Angle Bush: The 100 Most Influential People in AI 2026
1 hour ago
AI News
❝Why we need to make AI work for all, not just for good
1 hour ago
AI News
AI agents force cyber insurers to rethink what counts as an attack: report
2 hours ago