• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

SAIL proposes testing robot movement plans in simulation before physical runs

Sakana AI says its collaboration with the University of Tokyo generates robot trajectories from a few demonstrations, then tests and revises them using simulation feedback. Only the selected trajectory goes to the physical robot.

2 Sources, 12d ago, first seen 12d ago

TLDR

Sakana AI introduced SAIL, a University of Tokyo collaboration slated for IROS 2026. The team says the method uses one vision-language model to generate robot movement plans and another to assess simulated runs, helping it refine plans without updating model weights. Across six manipulation tasks in simulation, raising the search budget from one candidate to 45 increased the average rate of finding a successful trajectory from 25% to 73%. The team also evaluated SAIL on a physical robot.

Combined views

—

2 Sources, first seen 12d ago

— likes— comments— saves— reposts

Combined views

—

2 Sources, first seen 12d ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Related

Sakana AI seeks an account executive to help build its sales team

Sakana AI says the role will focus on bringing products including Sakana Marlin, Namazu and Fugu to customers, especially large companies, while shaping its sales approach from scratch.

Sakana AI launches Fugu to orchestrate swappable AI agents

Sakana AI says Fugu coordinates models for complex tasks and claims its swappable pool of agents can match restricted frontier models like Fable and Mythos.

Sakana AI wins information and communications prize at Japan Startup Awards 2026

Sakana AI says it briefed Prime Minister Sanae Takaichi on AI applications in defense command-and-control systems and its approach of combining small, efficient models.

2 Sources

Sakana AI@SakanaAILabsIntroducing "Scaling In-Context Imitation Learning" (SAIL) to be presented at #IROS2026. This work is a collaboration between Sakana AI and the University of Tokyo. Blog: https://pub.sakana.ai/sail What does a robot need before it can tackle a new task? Teaching a robot something new usually starts with collecting demonstrations and training a policy. But foundation models have already learned from vast amounts of images, text, and robotics-related data. We wanted to see how much of that knowledge we could draw out for robot control without changing the model itself. Recent demonstrations suggest that GPT-6 Astra can operate physical robots alongside its general language and vision capabilities. Earlier work has also shown that LLMs/VLMs can generate entire sequences of robot movements from a few demonstrations. However, a foundation model does not necessarily produce a reliable robot trajectory in a single generation. Performance depends on the context provided, and a small error in a movement target can cause the entire task to fail. We propose SAIL, a method for more reliable VLM-based robot trajectory generation through test-time scaling. SAIL uses a policy VLM as a robot trajectory generator, conditioned on a few successful demonstrations. It tests the generated trajectory in a simulator and uses an evaluation VLM to review the resulting video and identify where progress stalled. The policy VLM then uses this feedback to revise the trajectory, with Monte Carlo tree search (MCTS) exploring alternatives while refining promising candidates. Only the selected trajectory is sent to the physical robot. Across six manipulation tasks in simulation, increasing the search budget from one candidate to 45 raised the average rate of finding a successful trajectory from 25% to 73%. We also evaluated SAIL on a physical robot. Our results suggest that robot trajectory generation can benefit from test-time scaling, with additional computation enabling the model to test and refine its proposed actions in simulation. We think there is more to learn about what existing models can do with this kind of feedback, and how far those improvements carry over to physical robots. Paper: https://arxiv.org/abs/2603.08269 🐟12d
Stefania Druga@Stefania_drugaRT @SakanaAILabs: Introducing "Scaling In-Context Imitation Learning" (SAIL) to be presented at #IROS2026. This work is a collaboration bet…12d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    Sakana AI

    2 Sources

    Sakana AI@SakanaAILabsIntroducing "Scaling In-Context Imitation Learning" (SAIL) to be presented at #IROS2026. This work is a collaboration between Sakana AI and the University of Tokyo. Blog: https://pub.sakana.ai/sail What does a robot need before it can tackle a new task? Teaching a robot something new usually starts with collecting demonstrations and training a policy. But foundation models have already learned from vast amounts of images, text, and robotics-related data. We wanted to see how much of that knowledge we could draw out for robot control without changing the model itself. Recent demonstrations suggest that GPT-6 Astra can operate physical robots alongside its general language and vision capabilities. Earlier work has also shown that LLMs/VLMs can generate entire sequences of robot movements from a few demonstrations. However, a foundation model does not necessarily produce a reliable robot trajectory in a single generation. Performance depends on the context provided, and a small error in a movement target can cause the entire task to fail. We propose SAIL, a method for more reliable VLM-based robot trajectory generation through test-time scaling. SAIL uses a policy VLM as a robot trajectory generator, conditioned on a few successful demonstrations. It tests the generated trajectory in a simulator and uses an evaluation VLM to review the resulting video and identify where progress stalled. The policy VLM then uses this feedback to revise the trajectory, with Monte Carlo tree search (MCTS) exploring alternatives while refining promising candidates. Only the selected trajectory is sent to the physical robot. Across six manipulation tasks in simulation, increasing the search budget from one candidate to 45 raised the average rate of finding a successful trajectory from 25% to 73%. We also evaluated SAIL on a physical robot. Our results suggest that robot trajectory generation can benefit from test-time scaling, with additional computation enabling the model to test and refine its proposed actions in simulation. We think there is more to learn about what existing models can do with this kind of feedback, and how far those improvements carry over to physical robots. Paper: https://arxiv.org/abs/2603.08269 🐟12d
    Stefania Druga@Stefania_drugaRT @SakanaAILabs: Introducing "Scaling In-Context Imitation Learning" (SAIL) to be presented at #IROS2026. This work is a collaboration bet…12d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet