The Frontier of Physics AI is a Jenga game in a warehouse in San Leandro, California.
That warehouse is occupied by Encord, a company that builds data tools used to train AI models. Andrew Ceja is a pilot (as the company calls its robot trainers), carefully pulling blocks of wood from a rickety tower while wearing a headset with a camera that tracks what he sees. Collecting training data for a robot would be common enough on its own, but this headset includes sensors that measure brain waves as it carefully dismantles a block tower.
Encord is one of a small but growing number of startups betting that the next real constraint on humanoid and warehouse robotics is not model architecture, but a complete lack of real-world physical training data. Encord doesn’t just help robotics companies manage the data they have, it’s built a business around manufacturing the data they don’t own.
The brainwave headset Ceja is wearing was developed by German neuroscience startup Zander Labs. Zander Labs is betting that it can measure brain activity to infer mental states such as error, intent, and surprise, creating more useful datasets for training models. Encord’s work with Zander is currently in the pilot phase. Encord says the goal is to build an initial brainwave-tagged dataset, run it on customers’ robotic models, evaluate whether it actually improves performance, and then decide whether to scale up.
Zander neuroscientist Lukas Gehrke, who is overseeing the study, says the amount of brain activity used at any given time during a particular task provides clues to model builders trying to decide when to introduce models that require the most effort.
Vineeth Velmurugan, head of robot learning at Encord, said this is the “cutting edge” of efforts to solve robot data bottlenecks. Velmurugan, a veteran of OpenAI’s robotics lab and warehouse automation company Berkshire Grey, joined Encord to build out its in-house data creation team.
Encord was founded to enable companies building machine vision applications to annotate data and evaluate models. As customers (Velmurugan works with a number of large robotics companies but declined to be named) started applying end-to-end learning to robot manipulation tasks, executives realized they needed to generate the training data themselves, not just manage it. “The data just doesn’t exist,” Vermurugan said.
Bets that generative AI can do for robots what it does for chatbots continue to hit the same wall. The LLM is built on texts and more from all over the internet. It is difficult to find the same raw materials to teach neural networks about physical operations. Self-driving car companies collect their own raw materials, but scaling up is difficult. Training from videos works, but lacks the fidelity of real-world data. Velmurugan says that to break through this would require a dataset about five times the size of YouTube’s video corpus. This scale helps explain why data generation itself has become a business rather than just a research problem.
Meeting self-centered data needs
Companies developing robotic brains are currently turning to two main sources. One is “egocentric” video collected by workers wearing cameras, often with additional camera angles and other indicators. The other is data collected from remotely operated robots. Encord is doing both, taking egocentric data from several factories around the world and using its San Leandro facility to experiment with new techniques like brain waves and collect data sets on specific skills to fine-tune.
When TechCrunch visited, pilots were using a leader-follower rig (a pair of robotic arms, one controlled directly by a human operator and the other mimicking its movements) to generate data on tasks such as pouring coffee from a pot into a mug (which can be very slippery) and stacking poker chips. “All kinds of humanoid companies have asked us for these parts,” says Velmurugan.
Storage racks held cartons of artificial flowers in vases, books, plastic vegetables, kitty litter and shovels, bags and bundles of wire, and household manipulator training inventory.
At one of these stations, another pilot, Sofia Infante, pilots a robotic arm to plug and unplug Ethernet cables into the backs of servers. Data center operators would like to automate these tasks if robots can operate with the necessary precision. When I rotated behind the control, I realized why it was still out of reach. Pliers are much less dexterous than human fingers, and they lack the freedom of the arms that we take for granted.
Another new data modality Encord is developing uses a series of sensors attached to the forearm to detect electrical signals in the muscles. Videos of human hands manipulating objects typically don’t capture the entire hand, but Velmurugan hopes to build a 3D depiction of where the hand is at any given time based on sensors in the arm to better understand the model.
Encord’s dataset is annotated with a physical description of what each video contains (“tightening a bolt with the right hand”) to help the LLM-based model understand what’s going on. Velmurugan estimates that this kind of dense annotation is 100 times more valuable than “junk ego data” for training a specific task, and costs only 20 times as much to create on paper, which is a good deal.
But “20x” is still real money, and that’s the catch. Scraping text from the Internet, which is how LLM makers build models by pulling from Stack Overflow and the rest of the web, costs Frontier Labs very little. This is not the case with physical training data generation. This is the limit of comparing physical AI and LLM. This type of data needs to be created as well as collected, which changes the economics of model building.
Vermurugan says Encord’s visibility into programs across industries now provides a similar understanding of what works and what doesn’t to improve physical AI models. Part of Encord’s pitch is the advantage of being located among many robotics companies at the same time. Identify which data technologies are gaining traction across industries faster than any single customer.
That will keep the dozen or so pilots at Encode’s facility busy. Infante and Ceja are both part of a fast-growing talent pool developing building blocks for neural networks. Before joining Encord, they worked at another AI data annotation company, Scale.
Ceja works at a waste management company where, due to her interest in technology, she was responsible for keeping a robotic trash sorter in working order. Now that the Jenga tower has come down, he says he enjoys the challenge of solving robot training problems. “Every day is new!”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
