• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology
Announcement

Artificial Analysis launches Cyber Index as Grok 4.7 and MiMo-V2.6-Pro top early rankings

Artificial Analysis launched its Cyber Index and Cyber Index Alliance to benchmark AI models on enterprise cyber defense, with Grok 4.7 and MiMo-V2.6-Pro leading the debut rankings at 56.

Artificial AnalysisAA
1 Source, 11d ago, first seen 11d ago

TLDR

Artificial Analysis has launched the Artificial Analysis Cyber Index and the Cyber Index Alliance, framing them as a new standard for measuring how AI models perform on enterprise cyber defense tasks. The debut index combines three benchmarks — CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA — aimed at testing vulnerability discovery and patching. In launch results, Artificial Analysis reported Grok 4.7 (xhigh) and MiMo-V2.6-Pro tied for the lead at 56, ahead of GPT-6 Luna at 53, GLM-5.3-Flash at 50, and Muse Spark 1.3 (xhigh) at 44. Artificial Analysis also reported that several frontier models declined sizable shares of tasks on safety grounds, with the biggest gaps coming from CyberGym-E2E-AA refusals.

Combined views

140K

1 Source, first seen 11d ago

1K likes86 comments223 saves80 reposts

Combined views

140K

1 Source, first seen 11d ago

1K likes86 comments223 saves80 reposts

Artificial Analysis has launched the Artificial Analysis Cyber Index and the Cyber Index Alliance, describing them as a new standard for evaluating how AI models perform on enterprise cyber defense tasks.

Featured Source

The company said the index is meant to measure how well models can help security teams with the defensive loop of finding vulnerabilities, reproducing them, and patching them. It launches with support from alliance partners including Collinear AI, IBM, NVIDIA, and Vercel, and combines three open or partner-contributed benchmarks into a single ranking.

What the index includes

At launch, the index incorporates CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA.

Artificial Analysis describes CWE-Bench-AA as an audit-and-patch test built from Collinear AI’s CWE-bench, spanning 120 held-out tasks across all ten OWASP Top 10 (2025) categories and multiple programming languages. DeepsecBench-AA, based on Vercel’s DeepsecBench, focuses on vulnerability discovery in open-source application code and scores models against a human-verified golden set of findings. CyberGym-E2E-AA, from Berkeley RDI, tests end-to-end work on memory-safety bugs: locating the flaw, reproducing the crash, and patching it without breaking the project’s existing tests.

Artificial Analysis said the index is focused on defensive use cases and does not ask models to build working exploits. Instead, it evaluates how well agents can inspect source code and help internal security teams identify and remediate weaknesses.

Early leaderboard results

In its launch post, Artificial Analysis reported that Grok 4.7 (xhigh) and MiMo-V2.6-Pro tied for the top Cyber Index score at 56.

They were followed by GPT-6 Luna at 53, GLM-5.3-Flash at 50, and Muse Spark 1.3 (xhigh) at 44, according to the same launch results.

Safety refusals shaped some scores

Artificial Analysis also said several frontier models declined portions of the benchmark on safety grounds. In the launch results, it listed GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, and Gemini 3.8 Flash as refusing tasks that represented 32% to 38% of the Cyber Index.

The company said much of the gap between those models and the leaders came from CyberGym-E2E-AA. There, Artificial Analysis reported that GPT-6 Sol and GPT-6 Astra refused every task, while Claude Opus 5.5 refused 98% and Claude Fable 5.1 refused 99%.

That makes the launch as much about benchmark design as about model rankings. Artificial Analysis is explicitly separating safety refusals from ordinary failures in an attempt to show not just whether a model can complete cyber defense work, but whether provider safety policies allow it to try.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Artificial Analysis launches Cyber Index Alliance to benchmark AI agents on cyber defense tasks

Techmeme, citing Artificial Analysis, names Collinear, IBM, Nvidia and Vercel as partners in the alliance.

Aravind Srinivas Claims Perplexity Tops Search at Any Compute

Perplexity AI CEO Aravind Srinivas made the claim on X.

2 Sources

artificialanalysis.aiAnnouncing the Artificial Analysis Cyber Index Alliance: toward better benchmarking of agentic cyber defense15d
Artificial Analysis@ArtificialAnlysAnnouncing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.11d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Artificial AnalysisArtificial Analysis Cyber Index

    2 Sources

    artificialanalysis.aiAnnouncing the Artificial Analysis Cyber Index Alliance: toward better benchmarking of agentic cyber defense15d
    Artificial Analysis@ArtificialAnlysAnnouncing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.11d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet