• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

Free tokens offered to robotics researchers benchmarking open models alongside GPT 6 Astra

A user invites researchers using GPT 6 Astra for robotic control or agentic real2sim to request free tokens for testing Kimi K3 and Qwen 3.8 Flash Next as comparisons.

Eric JangEJ
Aran KomatsuzakiAK
Tanishq Mathew Abraham, Ph.D.TM
10 Sources, 12d ago, first seen 12d ago

TLDR

A user offers free token access to robotics researchers who want to benchmark the agentic capabilities of Kimi K3 and Qwen 3.8 Flash Next alongside their work with GPT 6 Astra. In a follow-up, the user says Astra has recently made new work possible in robotics and inverse graphics, pointing to agentic real2sim as an example.

Combined views

55.7K

10 Sources, first seen 12d ago

306 likes20 comments136 saves19 reposts

Combined views

55.7K

10 Sources, first seen 12d ago

306 likes20 comments136 saves19 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

10 Sources

Eric Jang@ericjang11In the last few weeks, GPT 6 Astra has really made a lot of new things possible in robotics + inverse graphics that were previously extremely surprising. Here's a short list of incredible work that has already come out: agentic real2sim12d
Aran Komatsuzaki@arankomatsuzakiAgreed. I tried comparing various open and closed models on Robobench and adjacent tasks to measure their supervisor / executive capability to help policy model. One most notable thing beside significant score gap is Astra's token efficiency. It consumes only 100 tokens for reasoning, while other models, open or closed, often consume 400 - 1k tokens for reasoning, even when I set reasoning level lower. Among open models, I found Qwen 3.8 27B and GLM 5.3 Flash performing well. At least on this benchmark, they were close to Grok and Muse. hy-embodied-vlm-1.0 and embodied r1.5 were much worse.12d
Tanishq Mathew Abraham, Ph.D.@iScienceLuvrThis is a good example of where i think the future of specialized models is headed. Despite the remarkable capabilities of frontier models, specialized models will still play a role as tools that the frontier models tot enhance their capabilities.12d
Jun Gao@JunGao33210520This is an early and exciting exploration we did! A lot of models/methods we developed in 3D (and many other domains) are basically tools for LLMs, and we can integrate them together in an agentic pipeline to solve challenging problems. For real2sim task, GPT6 can do very well already, but when we really analyze the detailed working traces from GPT, we see a lot of problems. E.g., it's trying to do SfM (from scratch) to estimate camera poses; it's trying to guess object dimensions and scale. These are problems we're solving everyday is 3D vision! The models we developed can be leveraged by GPT through harness to make the whole pipeline more accurate and robust. Maybe the next big question we need to think about is: what kind of tools are suitable for LLMs, and how can we best integrate these tools into LLMs through an agentic workflow (Harness). These tools are not necessarily 3D tools, but can be tools in many other fields. This is an exciting era, and more to come!12d
Bilawal Sidhu@bilawalsidhuGive your agent the right tools and it gets much better at inverse graphics - turning reality into editable, simulation ready 3d worlds12d
Kosta Derpanis@CSProfKGDRT @JunGao33210520: This is an early and exciting exploration we did! A lot of models/methods we developed in 3D (and many other domains) a…11d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    GPT 6 AstraKimi K3

    10 Sources

    Eric Jang@ericjang11In the last few weeks, GPT 6 Astra has really made a lot of new things possible in robotics + inverse graphics that were previously extremely surprising. Here's a short list of incredible work that has already come out: agentic real2sim12d
    Aran Komatsuzaki@arankomatsuzakiAgreed. I tried comparing various open and closed models on Robobench and adjacent tasks to measure their supervisor / executive capability to help policy model. One most notable thing beside significant score gap is Astra's token efficiency. It consumes only 100 tokens for reasoning, while other models, open or closed, often consume 400 - 1k tokens for reasoning, even when I set reasoning level lower. Among open models, I found Qwen 3.8 27B and GLM 5.3 Flash performing well. At least on this benchmark, they were close to Grok and Muse. hy-embodied-vlm-1.0 and embodied r1.5 were much worse.12d
    Tanishq Mathew Abraham, Ph.D.@iScienceLuvrThis is a good example of where i think the future of specialized models is headed. Despite the remarkable capabilities of frontier models, specialized models will still play a role as tools that the frontier models tot enhance their capabilities.12d
    Jun Gao@JunGao33210520This is an early and exciting exploration we did! A lot of models/methods we developed in 3D (and many other domains) are basically tools for LLMs, and we can integrate them together in an agentic pipeline to solve challenging problems. For real2sim task, GPT6 can do very well already, but when we really analyze the detailed working traces from GPT, we see a lot of problems. E.g., it's trying to do SfM (from scratch) to estimate camera poses; it's trying to guess object dimensions and scale. These are problems we're solving everyday is 3D vision! The models we developed can be leveraged by GPT through harness to make the whole pipeline more accurate and robust. Maybe the next big question we need to think about is: what kind of tools are suitable for LLMs, and how can we best integrate these tools into LLMs through an agentic workflow (Harness). These tools are not necessarily 3D tools, but can be tools in many other fields. This is an exciting era, and more to come!12d
    Bilawal Sidhu@bilawalsidhuGive your agent the right tools and it gets much better at inverse graphics - turning reality into editable, simulation ready 3d worlds12d
    Kosta Derpanis@CSProfKGDRT @JunGao33210520: This is an early and exciting exploration we did! A lot of models/methods we developed in 3D (and many other domains) a…11d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet