In a warehouse in San Leandro, California, a data tools company called Encord is conducting a cutting-edge experiment: having "pilots" wear headsets equipped with cameras and EEG sensors while playing Jenga blocks to collect brain activity data. This EEG headset from German neuroscience startup Zander Labs can measure brainwave signals in real-time as operators perform precise tasks, attempting to infer mental states such as errors, intentions, and surprise, creating more valuable training datasets for robot models.
Encord believes the real bottleneck for humanoid and warehouse robots isn't model architecture but an extreme lack of real-world physical training data. As Vineeth Velmurugan, head of robot learning at Encord, said — "This data simply doesn't exist," and the required data volume is about five times the size of YouTube's entire video library.
The Economic Dilemma of Physical AI: From Manual Operations to Data Factories
Inside the warehouse, pilots are using "master-slave" robotic arm devices — one directly controlled by humans, and another mimicking its movements — to generate training data for tasks like pouring coffee and stacking poker chips. Shelves are filled with fake flowers, books, plastic vegetables, and cat litter boxes, all used as props for training household robotic hands. At another workstation, pilot Sofia Infante is controlling a robotic arm to plug and unplug server network cables — an automation task highly desired by data center operators, currently still out of reach due to insufficient flexibility of robotic grippers.
In addition to EEG signals, Encord is also developing new data modalities: attaching sensor arrays to the forearm to detect muscle electrical signals, thereby building 3D hand trajectories to compensate for missing information when fingers are blocked in videos. Each dataset comes with dense physical action annotations (such as "right hand tightening a bolt"), and Velmurugan estimates that this high-quality annotation is 100 times more valuable for specific task training than ordinary data, while production costs are only 20 times higher — theoretically a profitable deal.
