• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Opus 5.5 and GPT-6 Luna reportedly shift AI's cost-quality frontier

A post says tests of the latest AI models on 2,400 engineers led its team to name Opus 5.5 its new daily default. It describes GPT-6 Luna as 20 times cheaper than Opus 5.5 for high-volume routing.

Yuchen JinYJ
Patrick WendellPW
5 Sources, 12d ago, first seen 12d ago

TLDR

A team says it tested the latest AI models on 2,400 engineers and found a major shift in the cost-quality frontier. It calls Opus 5.5 its new daily default, claiming higher quality and a 20% lower cost than prior baselines. For high-volume routing, it describes GPT-6 Luna as capable and 20 times cheaper than Opus 5.5.

Combined views

107.3K

5 Sources, first seen 12d ago

213 likes19 comments128 saves48 reposts

Combined views

107.3K

5 Sources, first seen 12d ago

213 likes19 comments128 saves48 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

River API is pitched as a cheap route to task-specific open models

Igor Babuschkin claims fine-tuned open-weight models can beat GPT-6 Astra and Claude Opus 5.5 for under 1% of the cost.

Grok Bot reportedly shifting to a best-model router, with Claude Opus 5.5, MidJourney and Suno in the mix

Elon Musk said Grok Bot will use “the best back end model for any given task,” naming Claude Opus 5.5, MidJourney, Suno and other APIs in a future-facing plan rather than a confirmed live rollout.

Claude Opus 5.5 reportedly scores 94.2% on NerfBench, within normal variance

BridgeMind AI says Opus 5.5 was scoring above its launch level a day earlier but cautions that the dip is not yet evidence of a nerf.

5 Sources

Patrick Wendell@pwendellCrazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier. Results below (online workload analysis of N=2,400 engineers, plus offline evals): 1. Two of the three models released last week clearly expand the cost/quality frontier: Opus 5.5 and GPT-6 Luna. 2. Opus 5.5 is now the highest quality mid-tier model. It is better than all prior Opus models, better than GPT-6 Sol, and better than GPT-5.6 Sol. 3. Opus 5.5 reduces same-task costs consistently by 20% in both offline and online analysis. This is against a baseline of Opus 4.8, the prior least-cost Opus model (Opus 5.0 was a bit of a dud with high costs and barely noticeable quality improvements). 4. Due to best-in-class quality and lower costs, Opus 5.5 is a strong candidate as an “every day default” model for coding, and we are now encouraging it for this purpose at Databricks. 5. GPT-6 Luna is very, very, very cheap. It was at least 20 times cheaper per-task than Opus 5.5 in every offline benchmark we tested and in observed online use. 6. GPT-6 Luna is surprisingly capable given how cheap it is. On one of our most difficult evaluation suites it roughly matches Opus 4.6 performance, while being 99.3% cheaper per-task than Opus 4.6 was at that time. That's a 100X cost reduction in ~9 months! This finding is preliminary and we are still evaluating Luna quality on a broader set of offline and online tests. Our production setup: Unity Gateway to route workloads across models and trace agentic interactions. A mix of end-user harnesses including: Omingent (meta-harness), Claude Code, Codex, and Cursor.12d
Yuchen Jin@Yuchenj_UWRT @pwendell: Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier…12d
Ali Ghodsi@alighodsiWe just tested the latest AI models on 2,400 engineers. The cost/quality frontier just shifted massively. The TL;DR: • Opus 5.5: Our new daily default - higher quality and 20% cheaper than prior baselines. • GPT-6 Luna: Shockingly capable and 20x cheaper than Opus 5.5 for high-volume routing. Must-read for anyone managing production AI workloads 👇12d
Akshay Kothari@akothariView post on X11d
Ion Stoica@istoica05RT @pwendell: Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier…11d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Claude Opus 5.5GPT-6 LunaDatabricks

    5 Sources

    Patrick Wendell@pwendellCrazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier. Results below (online workload analysis of N=2,400 engineers, plus offline evals): 1. Two of the three models released last week clearly expand the cost/quality frontier: Opus 5.5 and GPT-6 Luna. 2. Opus 5.5 is now the highest quality mid-tier model. It is better than all prior Opus models, better than GPT-6 Sol, and better than GPT-5.6 Sol. 3. Opus 5.5 reduces same-task costs consistently by 20% in both offline and online analysis. This is against a baseline of Opus 4.8, the prior least-cost Opus model (Opus 5.0 was a bit of a dud with high costs and barely noticeable quality improvements). 4. Due to best-in-class quality and lower costs, Opus 5.5 is a strong candidate as an “every day default” model for coding, and we are now encouraging it for this purpose at Databricks. 5. GPT-6 Luna is very, very, very cheap. It was at least 20 times cheaper per-task than Opus 5.5 in every offline benchmark we tested and in observed online use. 6. GPT-6 Luna is surprisingly capable given how cheap it is. On one of our most difficult evaluation suites it roughly matches Opus 4.6 performance, while being 99.3% cheaper per-task than Opus 4.6 was at that time. That's a 100X cost reduction in ~9 months! This finding is preliminary and we are still evaluating Luna quality on a broader set of offline and online tests. Our production setup: Unity Gateway to route workloads across models and trace agentic interactions. A mix of end-user harnesses including: Omingent (meta-harness), Claude Code, Codex, and Cursor.12d
    Yuchen Jin@Yuchenj_UWRT @pwendell: Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier…12d
    Ali Ghodsi@alighodsiWe just tested the latest AI models on 2,400 engineers. The cost/quality frontier just shifted massively. The TL;DR: • Opus 5.5: Our new daily default - higher quality and 20% cheaper than prior baselines. • GPT-6 Luna: Shockingly capable and 20x cheaper than Opus 5.5 for high-volume routing. Must-read for anyone managing production AI workloads 👇12d
    Akshay Kothari@akothariView post on X11d
    Ion Stoica@istoica05RT @pwendell: Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier…11d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet