Monday, Sep 28, 2026
1 Grok 4.7 Tops First Enterprise AI Cyber Defense IndexSPCXNVDAIBM 🤖 AI Sep 28, 8:28 AM EDT 11/7
Artificial Analysis partnered with Nvidia and IBM to release a new set of benchmarks measuring the ability of AI agents to find and patch software vulnerabilities. MiMo-V2.6-Pro and Grok 4.7 (xhigh) lead the results with scores of 56, while GPT-6 Luna and GLM-5.3-Flash follow with 53 and 50, respectively. Several frontier models, including Claude Opus 5.5 and Gemini 3.8 Flash, trail the leaders by 19 to 31 points after declining between 32% and 38% of the tasks on safety grounds.
The evaluation uses three frameworks: CWE-Bench-AA for language-specific patching, DeepsecBench-AA for vulnerability discovery, and CyberGym-E2E-AA for memory safety fixes. In the CyberGym test, models such as GPT-6 Sol and GPT-6 Astra refused every task, while Claude Fable 5.1 declined 99% of assignments. The Cyber Index Alliance, which includes Vercel and CollinearAI, aims to establish an industry standard for defense as AI models gain more advanced cyber offense capabilities.