Encord is testing the use of brain waves to improve the data used to train warehouse and humanoid robots, in a move targeting one of the most difficult problems in physical AI: the lack of sufficient, realistic training data. At the company’s facility in San Leandro, California, trainer Andrew Seja was removing pieces from a Jenga tower while wearing a helmet equipped with a camera and sensors measuring his brain waves.
The experiment is being conducted in collaboration with Zander Labs, a German neuroscience startup. The helmet is not intended to read thoughts, but to measure brain activity in order to infer mental states such as error, intent, and surprise. The partnership is still an early experiment: Encord wants to create a dataset labeled with brain-wave data and then feed it into customers’ robotic models to determine whether it improves performance before deciding whether to expand its use.
The Data Problem in Physical AI
Encord was founded to help companies developing computer-vision applications label data and evaluate models. But as its customers moved toward training general-purpose models for grasping and manipulating objects, the company realized that managing data alone was not enough and that it also needed to help produce it.
Vineeth Vel Morgan, Encord’s head of robotics learning, said that the required data simply does not exist. The ambition to bring to robots what generative AI has achieved for chatbots runs up against the need for real-world data from the physical world, which is difficult to collect and scale. Autonomous-driving companies collect real-world data themselves, but that approach is expensive and difficult to scale, while training on video alone may not provide the required precision.
Vel Morgan estimates that overcoming this obstacle could require a dataset roughly five times the size of a collection of YouTube videos. These estimates help explain the shift of robot-data production into an independent commercial activity rather than merely a research task inside laboratories.
From Human-Centric Video to Brain Waves
Companies developing robotic “brains” currently rely on two main sources of data. The first is human-centric video, recorded by workers wearing cameras, with additional viewing angles and various measurements. The second is data produced by robots controlled remotely by humans. Encord uses both approaches, collecting videos from several factories around the world and using its San Leandro facility to experiment with new methods such as brain waves or to create datasets for specific skills to fine-tune models.
During a visit to the facility, trainers were using leader-follower platforms consisting of two robotic arms: a human controls the first while the second mimics its movements. The platforms are used to record tasks such as pouring coffee into cups and stacking poker chips. The tasks also included connecting and disconnecting Ethernet cables at the back of a server, a task that could benefit data-center operators if robots can perform it with the required precision.
These tasks reveal the limitations of current robots: robotic grippers are less dexterous than human fingers and lack the degrees of movement we use naturally. For this reason, the facility stores a variety of materials and tools, from artificial flowers, books, and plastic vegetables to cat-litter trays and bags of cables, to train robots on household tasks and object-manipulation skills.
Muscle Sensors and Dense Labeling
Encord is also working on another method that uses sensors attached to the forearm to detect electrical signals in the muscles. Video showing the hands while they move objects may not capture the entire hand, so Vel Morgan hopes to use forearm signals to build a three-dimensional representation of the hand’s position at every moment, giving models a more complete understanding of movement.
Encord’s datasets are labeled with precise physical descriptions of what appears in the video, such as “the right hand tightens the screw,” to help models based on large language models understand what is happening. Vel Morgan estimates that this type of dense labeling is worth 100 times as much as low-quality “human-centric data” when training for specific tasks, while producing it costs only 20 times as much, which he considers a theoretically acceptable trade-off.
The Economics of Data Manufacturing
But the high cost remains a fundamental challenge. Collecting text from the internet—the approach that helped AI laboratories train language models using sources such as Stack Overflow and websites—costs almost nothing compared with producing physical-training data. Robot data, by contrast, must be manufactured by carrying out real tasks, recording them, and labeling them, fundamentally changing the economics of building models.
Encord believes that its position serving several robotics companies allows it to monitor which methods are gaining momentum in the sector and what succeeds or fails in improving physical-AI models. But the brain-wave experiment has not yet proved that it improves model performance, and the results of tests with customers will determine whether this data moves from an advanced experiment to a scalable part of the robot-training process.