Summary: Generalist AI’s GEN-1.5 robot foundation model can learn and execute new physical tasks from a single 3–12 second demonstration, marking a major step toward more adaptable AI-powered robotics.
A new robot foundation model called GEN-1.5 is pushing the boundaries of robot learning, embodied AI, and physical intelligence by enabling robots to learn new tasks from a single demonstration. Developed by Generalist AI, the system allows robots to observe a short 3–12 second example and immediately attempt the task without requiring retraining, gradient updates, or task-specific fine-tuning.

GEN-1.5 learns physical tasks from seconds-long demonstrations.
Unlike traditional robotic systems that often require large datasets and extensive optimization for each new task, GEN-1.5 uses what researchers describe as a “physical prompt.” A human operator or robot performs a task once, and the model infers the objective directly from the demonstration. The AI then generates robot actions in real time, processing video, language, sensor, and proprioceptive data through a multimodal architecture.
In tests involving 10 real-world manipulation tasks, including opening jars, retrieving money from a purse, stacking cups, sweeping debris, opening books, and unzipping pouches, GEN-1.5 achieved an average success rate of 59% from a single demonstration. When provided with just five minutes of task-specific data and a small number of optimization steps, performance increased to 83%.
The model maintains up to 30 seconds of contextual memory and generates motion trajectories at 100 Hz, enabling smooth robotic control. Researchers also demonstrated the ability to combine separate demonstrations into a new multi-step task. For example, after seeing how to unzip a pouch and separately retrieve money from it, the robot autonomously linked both actions into a continuous workflow.
GEN-1.5 further showed signs of robot generalization, adapting learned behaviors to different object positions, sizes, and even different robot hands. In one experiment, a demonstration performed by a human was successfully replicated by a robot using its own manipulators. The model also improvised novel solutions, using alternative tools when the original object was unavailable.
The results highlight a potential shift in robot programming and AI robotics, where future robots may learn new skills simply by watching brief demonstrations rather than undergoing lengthy retraining processes. For industrial automation, household robotics, and humanoid robot development, this approach could significantly reduce deployment time and expand robotic versatility.
