• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Agent Plasticity is proposed to measure how efficiently AI agents improve on held-out environments

Researchers study agents that turn experience into reusable artifacts; they say the top performer need not learn most efficiently.

AKAK
Aran KomatsuzakiAK
Sanjeev AroraSA
18 Sources, 1d ago, first seen 1d ago

TLDR

Researchers propose Agent Plasticity as held-out performance gain per unit of learning cost, meant to measure how efficiently agents learn from experience without weight updates. In tests across Chess, Go and Hex, they report Claude Fable 5 reached the highest fitted performance while GPT-5.6 Sol learned about five times more efficiently. They also found that frequently reusing learned artifacts did not guarantee improvement.

Combined views

16.1K

18 Sources, first seen 1d ago

52 likes12 comments38 saves95 reposts

Combined views

16.1K

18 Sources, first seen 1d ago

52 likes12 comments38 saves95 reposts

Today's AI benchmarks measure what an agent has learned. The researchers behind Agent Plasticity want to measure something different: how efficiently an agent can turn experience into reusable tools or memory, then improve on unseen tasks without changing its model weights.

Featured Source

Researcher Anirudh Goyal described Agent Plasticity as the gain in held-out performance per unit of learning cost, measured until performance saturates. Across Chess, Go and Hex, the team reports Claude Fable 5 reached the highest fitted performance, while GPT-5.6 Sol learned roughly five times more efficiently — evidence, they argue, that raw capability and learning efficiency are not the same.

Reuse alone did not guarantee better results. In the team's account, four models used learned artifacts in 94% to 98% of relevant decisions but improved by sharply different amounts. Among stronger improvers, 83% to 99% of remaining failures occurred while an artifact was already in use.

One chess example shows how the process can alter later behavior. After hitting a 300-ply limit and scoring zero, Claude Fable 5 changed its engine to value draws near that limit; at a later checkpoint, the researchers say it held a draw until Stockfish erred, then delivered mate.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

OpenAI Cuts GPT-5.6 Sol API Prices by 20 Percent

The cut applies for three months to API and credit usage for the model.

Codex Engineer Shares Extended Context Setup for GPT-5.6 Sol

OpenAI Codex engineer Tibo Sottiaux posted configuration steps to unlock the larger window for ChatGPT accounts.

OpenAI Previews Ultrafast Mode for GPT-5.6 Sol

The mode launches first via the OpenAI API to a select group of customers.

OpenAI Previews Ultrafast Mode for GPT-5.6 Sol

18 Sources

Harman Singh @ ICML 🇰🇷🇰🇷@Harman26Singh🤖 How well do agents learn from experience? We study agents that amortize experience into reusable artifacts and introduce Agent Plasticity to measure how efficiently they improve on held-out environments. Top performer ≠ most efficient learner.1d
Rulin Shao@RulinShaoGreat work from @Harman26Singh on agent plasticity! It measures an agent’s ability to improve through self-evolution. We’re seeing a growing trend toward measuring and learning capabilities beyond task performance, such as learning from experience, scaling collaboration, managing context, and self-training. Developing benchmarks and metrics that isolate these capabilities could be very valuable.1d
Anirudh Goyal@anirudhg9119Today's AI benchmarks measure what agents have learned. But how well can they learn from experience? Can agents turn experience into reusable tools and memory without updating model weights and improve on unseen tasks? How efficiently? We investigate these questions with Agent Plasticity. 🧵1d
AK@_akhaliqAgent Plasticity Measuring Self-Improvement Through Experience paper: https://huggingface.co/papers/2610.089021d
Aran Komatsuzaki@arankomatsuzakiRT @Harman26Singh: 🤖 How well do agents learn from experience? We study agents that amortize experience into reusable artifacts and introd…1d
Lukasz Kaiser@lukaszkaiserRT @Harman26Singh: 🤖 How well do agents learn from experience? We study agents that amortize experience into reusable artifacts and introd…1d
Sanjeev Arora@prfsanjeevaroraRT @anirudhg9119: This raises a broader question: Can we benchmark not just what models can do, but how well they learn from experience? W…1d
elvis@omarsar0RT @omarsar0: Banger paper from Meta Superintelligence Labs on self-improving agents. (bookmark it) It's hard to know exactly what drives…8h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    GPT-5.6 SolClaude Fable 5Claude Fable

    18 Sources

    Harman Singh @ ICML 🇰🇷🇰🇷@Harman26Singh🤖 How well do agents learn from experience? We study agents that amortize experience into reusable artifacts and introduce Agent Plasticity to measure how efficiently they improve on held-out environments. Top performer ≠ most efficient learner.1d
    Rulin Shao@RulinShaoGreat work from @Harman26Singh on agent plasticity! It measures an agent’s ability to improve through self-evolution. We’re seeing a growing trend toward measuring and learning capabilities beyond task performance, such as learning from experience, scaling collaboration, managing context, and self-training. Developing benchmarks and metrics that isolate these capabilities could be very valuable.1d
    Anirudh Goyal@anirudhg9119Today's AI benchmarks measure what agents have learned. But how well can they learn from experience? Can agents turn experience into reusable tools and memory without updating model weights and improve on unseen tasks? How efficiently? We investigate these questions with Agent Plasticity. 🧵1d
    AK@_akhaliqAgent Plasticity Measuring Self-Improvement Through Experience paper: https://huggingface.co/papers/2610.089021d
    Aran Komatsuzaki@arankomatsuzakiRT @Harman26Singh: 🤖 How well do agents learn from experience? We study agents that amortize experience into reusable artifacts and introd…1d
    Lukasz Kaiser@lukaszkaiserRT @Harman26Singh: 🤖 How well do agents learn from experience? We study agents that amortize experience into reusable artifacts and introd…1d
    Sanjeev Arora@prfsanjeevaroraRT @anirudhg9119: This raises a broader question: Can we benchmark not just what models can do, but how well they learn from experience? W…1d
    elvis@omarsar0RT @omarsar0: Banger paper from Meta Superintelligence Labs on self-improving agents. (bookmark it) It's hard to know exactly what drives…8h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet