Robotics and Automation

Generalist AI unveils GEN-1.5 to teach robots new tasks in seconds

Generalist AI introduced the GEN-1.5 model, capable of learning new physical tasks from a single demonstration lasting between 3 and 12 seconds, without additional training or fine-tuning. The model achieved an average success rate of 59% on a set of short manipulation tasks, rising to 83% after adaptation with additional data.

2026-08-20
4 min read
13 views
certi.news
Generalist AI unveils GEN-1.5 to teach robots new tasks in seconds

Generalist AI unveiled its foundational robotics model, GEN-1.5, which enables a robot to learn new physical tasks from a single short demonstration instead of relying on specialized programming or task-specific training and fine-tuning processes. According to information published by the company, the model can analyze a demonstration lasting between 3 and 12 seconds and execute the task without incremental updates or an additional fine-tuning process.

Generalist AI calls this type of example a “physical prompt,” a context that combines the sensor data and motion sequences used in the demonstration. GEN-1.5 processes video frames alongside sensor data, language, and information about the robot’s position and movement. The model uses a 30-second context window and produces up to 100 motion commands per second.

Initial results with clear variation in performance

The company tested the model on 10 short manipulation tasks, and with a single demonstration it achieved an average success rate of 59%. When adapted using five minutes of data, including approximately 50 demonstrations per task and over 10 gradient steps, the rate rose to 83%. In another test based on one minute of data and a single gradient step, the model recorded a success rate of 66.5%.

These figures indicate that GEN-1.5 can benefit from only a few examples, but they also show that performance still depends partly on the amount of data and the degree of adaptation. Generalist AI acknowledges that the tasks it tested were relatively simple and short, and that current success rates remain limited.

From imitating motion to building a new sequence

The model is not limited to repeating the motion it observes. When separately shown a scene of opening a pencil-case zipper and a scene of removing coins from it, it was able to combine the physical prompts into a single continuous sequence. It also generated transitional movements that were not present in the demonstrations, such as repositioning and handling errors.

The company’s examples demonstrated the model’s ability to change its strategy when faced with unexpected conditions. In a task involving pushing an object into a container using a brush, GEN-1.5 used a banana as a temporary tool to achieve the same objective. When it encountered a shovel, it did not continue pushing with it; instead, it lifted the object and placed it in the container. It also dealt with a sheet covering the container and removed a Lego piece stuck to its finger using its other hand.

Learning from simulation and human movements

The model can also transfer a demonstration recorded in a simulated environment to a real robot without additional examples, even though the simulation data was not part of the pretraining stage. In some tasks, it was able to observe movements performed by people with their hands through the robot’s cameras and then reproduce the task using the robot’s hands.

What changes in practice?

The significance of this approach lies in its potential to reduce the need to program each robotic task separately. Instead of setting up a specialized system over the course of months, it may eventually be enough to demonstrate the desired action to the robot for a few seconds. However, this conclusion does not mean that the problem has been solved; the company says that pretraining for GEN-1.5 has been ongoing for more than eight months, and that the model is still being tested on limited tasks with incomplete success rates.

According to Generalist AI, after being trained on a broad range of physical interaction data, the model began adapting to new tasks using less data and computing power. The company says that in-context learning and improvisation capabilities emerged automatically during pretraining, without a special training method or an additional objective dedicated to these capabilities.

News source
c
Author

certi.news

In the same category

You may also like

View all news