Artificial Analysis has launched the Artificial Analysis Cyber Index and the Cyber Index Alliance, describing them as a new standard for evaluating how AI models perform on enterprise cyber defense tasks.
The company said the index is meant to measure how well models can help security teams with the defensive loop of finding vulnerabilities, reproducing them, and patching them. It launches with support from alliance partners including Collinear AI, IBM, NVIDIA, and Vercel, and combines three open or partner-contributed benchmarks into a single ranking.
What the index includes
At launch, the index incorporates CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA.
Artificial Analysis describes CWE-Bench-AA as an audit-and-patch test built from Collinear AI’s CWE-bench, spanning 120 held-out tasks across all ten OWASP Top 10 (2025) categories and multiple programming languages. DeepsecBench-AA, based on Vercel’s DeepsecBench, focuses on vulnerability discovery in open-source application code and scores models against a human-verified golden set of findings. CyberGym-E2E-AA, from Berkeley RDI, tests end-to-end work on memory-safety bugs: locating the flaw, reproducing the crash, and patching it without breaking the project’s existing tests.
Artificial Analysis said the index is focused on defensive use cases and does not ask models to build working exploits. Instead, it evaluates how well agents can inspect source code and help internal security teams identify and remediate weaknesses.
Early leaderboard results
In its launch post, Artificial Analysis reported that Grok 4.7 (xhigh) and MiMo-V2.6-Pro tied for the top Cyber Index score at 56.
They were followed by GPT-6 Luna at 53, GLM-5.3-Flash at 50, and Muse Spark 1.3 (xhigh) at 44, according to the same launch results.
Safety refusals shaped some scores
Artificial Analysis also said several frontier models declined portions of the benchmark on safety grounds. In the launch results, it listed GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, and Gemini 3.8 Flash as refusing tasks that represented 32% to 38% of the Cyber Index.
The company said much of the gap between those models and the leaders came from CyberGym-E2E-AA. There, Artificial Analysis reported that GPT-6 Sol and GPT-6 Astra refused every task, while Claude Opus 5.5 refused 98% and Claude Fable 5.1 refused 99%.
That makes the launch as much about benchmark design as about model rankings. Artificial Analysis is explicitly separating safety refusals from ordinary failures in an attempt to show not just whether a model can complete cyber defense work, but whether provider safety policies allow it to try.