
The cutting-edge of physical AI resembles a Jenga game inside a warehouse located in San Leandro, California.
That warehouse is home to Encord, a firm specializing in data tools designed to train AI models. Andrew Ceja functions as a pilot—the term the company uses for its robotic trainers—and he is meticulously extracting wooden blocks from a precarious tower while equipped with a headset that features a camera to monitor his view. While this method for gathering robot training data is rather typical, this particular headset is outfitted with sensors that record his brain waves as he skillfully dismantles the block structure.
Encord is among a limited yet expanding group of startups betting that the next significant limitation for humanoid and warehouse robots will not be model architecture, but rather the significant lack of authentic physical training data. Instead of merely aiding robotics firms in managing existing data, Encord aims to cultivate a business model around generating the data that is currently absent.
The brain wave headset that Ceja is utilizing was developed by Zander Labs, a German neuroscience startup betting that gauging brain activity—to infer mental states such as error, intent, and surprise—can produce a more advantageous data set for training models. Encord’s collaboration with Zander is presently a pilot initiative; Encord claims that the objective is to create an initial brain wave-tagged data set, test it through customer robotics models, and assess whether it genuinely enhances performance before deciding to expand the project.
Lucas Gehrke, a neuroscientist from Zander overseeing the project, indicates that the level of brain activity registered at any given moment during a task can provide insights for model developers trying to determine when their highest-effort models should be deployed.
According to Vineeth Velmurugan, Encord’s head of robot learning, this is the “bleeding edge” of addressing the robotics data bottleneck. A former member of OpenAI’s robot lab and Berkshire Grey, a warehouse automation company, Velmurugan joined Encord to establish the company’s internal data-creation team.
Encord was established to assist businesses developing machine-vision applications in annotating data and evaluating models. As their clients—Velmurugan indicates that Encord collaborates with numerous leading robotics companies, although he cannot disclose their names—began implementing end-to-end learning in robotic manipulation tasks, executives realized they needed to create training data independently, not just manage it. “The data simply does not exist,” Velmurugan remarked.
The premise that generative AI can replicate its success with robots as it has with chatbots continues to encounter this same hurdle. Large language models (LLMs) were constructed from the text of the entire internet and beyond. Securing the equivalent raw materials for teaching neural networks about physical manipulation is difficult: self-driving car companies gather this data themselves, but scaling it is a challenge. Training from video can be effective, but it falls short in comparison to authentic real-world data. Velmurugan believes it will necessitate a data set roughly five times larger than YouTube’s video corpus to achieve a breakthrough—a scale that elucidates why data generation has evolved into a commercial endeavor rather than merely a research concern.
Satisfy your egocentric data requirements
Companies developing robotic intelligence are now relying on two primary sources: “Egocentric” video gathered by workers wearing cameras, often supplemented with additional viewing angles and metrics, and data collection from robots operated remotely. Encord engages in both, sourcing egocentric data from various factories worldwide and utilizing its San Leandro facility to test novel modalities, such as brain wave data, or assemble data sets around particular skills for fine-tuning.
During a TechCrunch visit, pilots were operating leader-follower systems—robotic arms working in tandem, with one arm directly controlled by a human and the other mimicking its movements—to gather data on tasks like pouring coffee from a pot into mugs (which tends to be quite messy) and stacking poker chips. “Every humanoid firm has requested these components,” Velmurugan notes.
Storage shelves were filled with boxes of artificial flowers in vases, books, plastic vegetables, cat litter trays and scoops, bags, and bundles of wires, the essential materials used for training manipulators in household tasks.
At one station, another pilot, Sofia Infante, skillfully manipulates robotic arms to connect and disconnect ethernet cables from the back of a server—the type of task data center operators hope could be automated, provided robots could achieve the necessary precision. Trying my hand at the controls, I understood why that remains elusive: robotic grippers are significantly less agile than human fingers and lack the versatility we take for granted in our arms.
Another innovative data modality that Encord is pursuing involves a set of sensors attached to the forearm to detect electrical signals in muscles. Videos of humans manipulating objects typically don’t capture the entirety of the hand, but Velmurugan aims to construct a 3D representation of the hand’s location at any point using the arm sensors, fostering a more comprehensive grasp for models.
Encord’s data sets are accompanied by physical descriptions of the content of each video—“right hand tightens bolt”—to assist LLM-based models in comprehending the actions unfolding. Velmurugan estimates that this type of thorough annotation holds 100 times the value of “poor-quality ego data” for training specific tasks, and it incurs only 20 times the cost to produce, which, on paper, seems like a reasonable exchange.
However, “20 times more” is still a substantial sum, and that’s the challenge: scraping content from the internet, as LLM creators did by sourcing information from Stack Overflow and other websites, cost frontier labs nearly nothing. Producing physical training data does not, and that represents the limitation of comparing physical AI to LLMs. This kind of data must be generated, not merely gathered, which alters the economics of constructing these models.
Velmurugan claims that advancements are underway—with Encord’s insight into programs across the industry, he can observe startups and frontier labs alike determining what is effective and what is not in enhancing physical AI models. This perspective—being positioned among many robotics firms simultaneously—is also part of Encord’s appeal. It can identify which data techniques are gaining traction across the industry prior to any individual customer.
This will keep the roughly dozen pilots at Encord’s facility engaged. Both Infante and Ceja are members of an emerging workforce constructing the foundations for neural networks; they previously worked at Scale, another AI data annotation firm, before joining Encord.
Ceja had experience at a waste management firm where his technological interests led him to oversee the maintenance of a robotic trash sorter. Now, as the Jenga tower teeters, he expresses his enjoyment of the challenges presented by training tasks for robots — “It’s something new every day!”
When you make purchases through links in our articles, we may receive a small commission. This does not influence our editorial integrity.

