• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Reaction

A proposed AI reward signal for a persistence–trustworthiness tradeoff

A user argues that a tradeoff they say Saachi Jain and OpenAI reported to the press stems from a scalar, totally ordered reward signal. They believe a preference-based alternative could avoid it at a small additional compute cost.

davidad 🎇D🎇
1 Source, 11d ago, first seen 11d ago

TLDR

The user proposes generating concurrent, independent model rollouts, then having the same model checkpoint judge them using detailed rubrics and map its overall preferences in a Hasse diagram. They believe this reward signal could avoid a persistence–trustworthiness tradeoff they attribute to scalar rewards, at a small additional compute cost. They say Saachi Jain and OpenAI reported the tradeoff to the press.

Combined views

3.5K

1 Source, first seen 11d ago

37 likes4 comments17 saves2 reposts

Combined views

3.5K

1 Source, first seen 11d ago

37 likes4 comments17 saves2 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Codex adds beta next-message suggestions for Pro users

OpenAI says the feature uses your conversation and how you talk to Codex to suggest what to say next.

OpenAI Rolls Out "Ultrafast" Mode for GPT-6.1 Sol

OpenAI says the mode offers up to 8x faster speeds than Sol Standard in the API, Codex, and ChatGPT Work.

OpenAI Rolls Out "Ultrafast" Mode for GPT-6.1 Sol
OpenAI’s $50B versus $70B revenue debate and AI infrastructure demand

One post questions how much revenue OpenAI keeps after partners take a cut and argues token growth matters more.

1 Source

davidad 🎇@davidadThe tradeoff between “persistence” and trustworthiness that @saachi_jain_ @openai recently reported to the press is a consequence of a scalar (totally ordered) reward signal. I believe the method below would avoid this, at a small cost of additional compute. cc @jachiam0 @tszzl11d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    OpenAIsaachi_jain_davidad

    1 Source

    davidad 🎇@davidadThe tradeoff between “persistence” and trustworthiness that @saachi_jain_ @openai recently reported to the press is a consequence of a scalar (totally ordered) reward signal. I believe the method below would avoid this, at a small cost of additional compute. cc @jachiam0 @tszzl11d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet