• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

RLM harnesses are claimed to help models generalize to similar unseen tasks 8–32 times longer than training tasks

The author says well-designed harnesses can make structurally similar tasks look nearly identical to individual model calls, even when the tasks come from different domains.

Christopher ManningCM
Gary MarcusGM
Milad Khademi Nori, PhDMK
6 Sources, 12d ago, first seen 12d ago

TLDR

In a July 2026 post, the author says RLMs trained only on short tasks can fully generalize to similar unseen tasks 8–32 times longer because the harness produces near-identical task trajectories. The author also reports that training on grouping essays by author can improve performance on grouping math problems by similar solutions.

Combined views

28.2K

6 Sources, first seen 12d ago

154 likes10 comments152 saves446 reposts

Combined views

28.2K

6 Sources, first seen 12d ago

154 likes10 comments152 saves446 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Prime Intellect Releases Prime Agent Coding Harness

Company presents the harness as token-efficient for autonomous coding work.

Prime Intellect Releases Prime Agent Coding Harness
Harnesses Let Transformers Generalize to Longer Tasks and New Domains
Thore Graepel’s AI reasoning argument sparked a public debate over what LLMs are missing

In posts tied to a new MIT Technology Review piece, former Google DeepMind researcher Thore Graepel argued that advanced AI needs auditable reasoning, then drew responses from Yann LeCun, David Duvenaud and Gary Marcus.

6 Sources

Milad Khademi Nori, PhD@khademinori@TheFalconer2219 @GaryMarcus All sota agentic ai tech use harness engineering and to see how much performance boost (generalization) the harness offers, see the following work from MIT.12d
Gary Marcus@GaryMarcusstunning confirmation of my central work12d
Christopher Manning@chrmanningRT @a1zhang: Read full blog, condensed summary below: https://alexzhang13.github.io/blog/2026/harness/11d
Peter J. Liu@peterjliuTo take a lesson from machine learning applied to harness/prompt engineering, for better generalization it’s important to have an L2 regularization term in your objective function based on the length of the code of the harness (including prompts). And the regularization weight should increase as the model gets better.9d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    alex zhangGary Marcus

    6 Sources

    Milad Khademi Nori, PhD@khademinori@TheFalconer2219 @GaryMarcus All sota agentic ai tech use harness engineering and to see how much performance boost (generalization) the harness offers, see the following work from MIT.12d
    Gary Marcus@GaryMarcusstunning confirmation of my central work12d
    Christopher Manning@chrmanningRT @a1zhang: Read full blog, condensed summary below: https://alexzhang13.github.io/blog/2026/harness/11d
    Peter J. Liu@peterjliuTo take a lesson from machine learning applied to harness/prompt engineering, for better generalization it’s important to have an L2 regularization term in your objective function based on the length of the code of the harness (including prompts). And the regularization weight should increase as the model gets better.9d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet