• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Artificial Analysis launches Cyber Index and alliance for AI cyber defense

Artificial Analysis says Grok 4.7 (xhigh) and MiMo-V2.6-Pro tied for the top index score of 56, while safety refusals held back several other models.

Elon MuskEM
Guillermo RauchGR
Nazneen RajaniNR
16 Sources, 12d ago, first seen 12d ago

TLDR

Artificial Analysis announced a Cyber Index that combines three benchmarks testing how AI agents find and fix vulnerabilities. It named CollinearAI, IBM, Nvidia and Vercel as launch partners in an accompanying alliance. Artificial Analysis says Grok 4.7 (xhigh) and MiMo-V2.6-Pro tied for the top score of 56. Its scoring counts tasks declined on safety grounds as zero and reports those refusals separately from other failures.

Combined views

6.7M

16 Sources, first seen 12d ago

14.2K likes1.7K comments533 saves2.1K reposts

Combined views

6.7M

16 Sources, first seen 12d ago

14.2K likes1.7K comments533 saves2.1K reposts

Sentiment

Positive23.6%76.4%Negative

Summary

Positive replies welcomed the Artificial Analysis Cyber Index for standardizing AI model security evaluations, while negative accounts argued it undervalues open models that did the actual work and unfairly penalizes high refusal rates.

Based on 61 sentiment-bearing replies from 57 accounts across 2 conversations.

Featured Source

Sentiment

Positive23.6%76.4%Negative

Summary

Positive replies welcomed the Artificial Analysis Cyber Index for standardizing AI model security evaluations, while negative accounts argued it undervalues open models that did the actual work and unfairly penalizes high refusal rates.

Based on 61 sentiment-bearing replies from 57 accounts across 2 conversations.

Related

Artificial Analysis held its first Korea event in Seoul

Artificial Analysis says the gathering brought builders and researchers together for talks about AI’s future, including benchmarking and physical AI.

GPT-6.1 Sol scores one point below GPT-6 Astra on the Intelligence Index at under a quarter of its cost per task at max effort

Artificial Analysis says GPT-6.1 Sol replaced GPT-6 Sol after seven days and gained four points on its Intelligence Index. Its input and output token prices remain $2/$10 per million, but the cache-read discount rises from 90% to 95%.

Claude Sonnet 5.5 launches with promised speed gains and up to 30% lower costs for most work

Anthropic says Sonnet 5.5 runs more than 30% faster than Sonnet 5. In pre-release max-effort tests, Artificial Analysis measured per-task costs roughly 50% higher than Sonnet 5’s.

16 Sources

Artificial Analysis@ArtificialAnlysAnnouncing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.12d
Nazneen Rajani@nazneenrajaniToday we’re releasing CWE-Bench v1: 120 held-out audit-and-patch tasks spanning 73 CWEs, all OWASP Top 10 2025 categories, and 8 languages. Highest programmatic Pass@4: 81% (at least one success in four tries). The frontier is climbing fast (about 15% improvements in just one generation of models). Defense is still the harder test; our goal is to test every known vulnerability. Stay tuned for v2. Learn more: https://cwe-bench.com/12d
Guillermo Rauch@rauchg@ArtificialAnlys Awesome to partner with you guys on this. Super important for the world to understand what the cybersecurity model frontier looks like!12d
Florian Brand@xeophonweird… if capabilities in open models are only gained by adversarial distillation, why do open models outperform closed ones with less/no refusals? 🤔🤔🤔12d
Elon Musk@elonmuskNot bad11d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Artificial AnalysisArtificial Analysis Cyber IndexVercel
    Nazneen Rajani

    16 Sources

    Artificial Analysis@ArtificialAnlysAnnouncing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.12d
    Nazneen Rajani@nazneenrajaniToday we’re releasing CWE-Bench v1: 120 held-out audit-and-patch tasks spanning 73 CWEs, all OWASP Top 10 2025 categories, and 8 languages. Highest programmatic Pass@4: 81% (at least one success in four tries). The frontier is climbing fast (about 15% improvements in just one generation of models). Defense is still the harder test; our goal is to test every known vulnerability. Stay tuned for v2. Learn more: https://cwe-bench.com/12d
    Guillermo Rauch@rauchg@ArtificialAnlys Awesome to partner with you guys on this. Super important for the world to understand what the cybersecurity model frontier looks like!12d
    Florian Brand@xeophonweird… if capabilities in open models are only gained by adversarial distillation, why do open models outperform closed ones with less/no refusals? 🤔🤔🤔12d
    Elon Musk@elonmuskNot bad11d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet