• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

Grok 4.7 and Claude Opus 5.5 score zero on CancerBench's cancer-cure metric

CancerBench's creator says the two models have been added and now tie for first and last place. The benchmark's only metric counts how many types of cancer a model has cured.

Tanishq Mathew Abraham, Ph.D.TM
2 Sources, 12d ago, first seen 12d ago

TLDR

CancerBench's creator says they made the benchmark in response to AI lab CEOs talking about curing cancer. After adding Grok 4.7 and Claude Opus 5.5, the creator says both score zero on its sole metric: the number of cancer types a model has cured.

Combined views

14.8K

2 Sources, first seen 12d ago

183 likes8 comments11 saves10 reposts

Combined views

14.8K

2 Sources, first seen 12d ago

183 likes8 comments11 saves10 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Open models and scientists' roles in AI for science

A panelist says a September 2026 session at San Francisco's Open-Source AI Summit discussed how open models can accelerate scientific discovery and how scientists should use AI.

A telescopic language model aims to be valid at every depth

A post shares an excerpt describing a nested-capacity Transformer trained with a randomly truncated capacity prefix alongside a full-capacity pass.

Tests suggest cheap verifiers can work well in LLM post-training

A post quotes a paper reporting that, in Qwen3 tests on HealthBench and PRBench, higher agreement with frontier LLM reference judges did not consistently identify the best training verifier.

2 Sources

Tanishq Mathew Abraham, Ph.D.@iScienceLuvrLast week, two new frontier models were released: 1. Grok 4.7 2. Claude Opus 5.5 Both models have now been added to CancerBench. As expected, they tie for first and last place, which a score of zero. Still waiting for the AI labs to saturate this benchmark 😄12d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Tanishq Mathew Abraham, Ph.D.Claude Opus 5.5Grok 4.7

    2 Sources

    Tanishq Mathew Abraham, Ph.D.@iScienceLuvrLast week, two new frontier models were released: 1. Grok 4.7 2. Claude Opus 5.5 Both models have now been added to CancerBench. As expected, they tie for first and last place, which a score of zero. Still waiting for the AI labs to saturate this benchmark 😄12d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet