The cutting edge of physical artificial intelligence is being forged in a San Leandro, California warehouse, where the delicate art of playing Jenga with a robot trainer is emblematic of the immense data challenges facing the robotics industry. This unassuming setting is home to Encord, a company at the forefront of developing sophisticated data tooling essential for training advanced AI models. Here, human operators, referred to as "pilots," are not just performing tasks but actively generating the intricate, real-world data that will power the next generation of humanoid and warehouse robots.
Andrew Ceja, one such pilot at Encord, meticulously extracts wooden blocks from a teetering Jenga tower. His actions are monitored not only by a headset-mounted camera, a common practice for capturing robot training data, but also by an array of sensors designed to capture his brainwave activity. This novel integration of neuroscience into the data collection process signifies a critical shift: the recognition that the primary constraint on robotic advancement may not be algorithmic innovation, but rather the sheer scarcity of high-fidelity, real-world physical training data. Encord is not merely managing existing data; it is pioneering methods to manufacture the data that simply does not yet exist.
This ambitious endeavor is a collaborative effort with Zander Labs, a German neuroscience startup. Zander Labs is developing technology that measures brain activity to infer mental states such as error, intent, and surprise. The current trial run with Encord aims to construct an initial dataset tagged with brainwave data. This data will then be fed into customer robotics models to rigorously assess whether it demonstrably enhances performance. The outcome of this evaluation will determine whether this pioneering approach to data generation is scaled up.
Lucas Gehrke, a neuroscientist at Zander Labs overseeing the project, explained the significance of this data. "The amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models," Gehrke stated. This granular insight into human cognitive load during task execution could be pivotal for robots learning to anticipate and respond to complex or demanding situations.
Vineeth Velmurugan, Encord’s Head of Robot Learning, describes this work as being on the "bleeding edge" of efforts to surmount the robotics data bottleneck. With a background at OpenAI’s robotics lab and Berkshire Grey, a prominent warehouse automation firm, Velmurugan joined Encord with the specific mandate to build the company’s internal data-creation capabilities.
Encord’s origins lie in providing data annotation and model evaluation services for companies developing machine-vision applications. However, as their clientele – which includes leading robotics firms that Velmurugan is authorized to work with but not name – began to embrace end-to-end learning for robotic manipulation tasks, a stark reality emerged: the necessary training data simply did not exist in sufficient quantities or quality. "The data simply does not exist," Velmurugan reiterated, highlighting the fundamental challenge.
The aspiration to replicate the generative AI revolution seen in chatbots for physical robots is encountering this significant data hurdle. While companies developing autonomous vehicles meticulously collect their own physical-world data, scaling this process is proving immensely difficult. Training robots using video data, while feasible, often lacks the nuanced fidelity of real-world interactions. Velmurugan estimates that a breakthrough in robot learning would necessitate a dataset approximately five times the size of YouTube’s entire video corpus. This staggering scale underscores why data generation itself has evolved from a research problem into a burgeoning business sector.
The Genesis of Egocentric Data and Beyond
In response to this data deficit, companies developing robot "brains" are primarily tapping into two key data streams. The first is "egocentric" video, captured by cameras worn by human operators performing tasks. This is often augmented with additional camera angles and other sensor data. The second source involves data generated by remotely operated robots. Encord actively engages in both, sourcing egocentric data from various global factories and utilizing its San Leandro facility for experimental data modalities, such as brainwave data, and for curating specialized datasets for fine-tuning specific robotic skills.
During a recent visit by TechCrunch, Encord pilots were employing "leader-follower rigs." These consist of paired robotic arms, where one arm is directly controlled by a human operator, and the other meticulously mimics its movements. This setup is used to generate data for tasks ranging from the notoriously sloshy act of pouring coffee into mugs to the precise stacking of poker chips. "Every humanoid company has asked us for these pieces," Velmurugan noted, indicating the high demand for such task-specific data.
The San Leandro facility itself is a testament to the diverse needs of robotic training. Storage racks were observed to contain an eclectic assortment of items: cartons of artificial flowers in vases, books, plastic fruits and vegetables, kitty litter trays and scoops, and bundles of wires. This seemingly random collection represents the essential inventory for training robotic manipulators designed for a wide array of household and industrial tasks.
In one workstation, pilot Sofia Infante demonstrated the intricate process of plugging and unplugging ethernet cables from the back of a server. This is precisely the kind of task that data center operators are eager to automate, but it requires a level of robotic dexterity that is still elusive. A brief hands-on experience behind the controls revealed the limitations: robotic pincers, while improving, still lack the fine motor control and degrees of freedom inherent in human fingers and arms, making such tasks exceptionally challenging for current robotic systems.
Novel Data Modalities and Annotation’s Value
Encord is also developing another novel data modality: a set of sensors worn on the forearm to detect electrical signals in muscles. While video footage of human hands manipulating objects often fails to capture the entire hand, Velmurugan envisions using forearm sensor data to construct a more comprehensive three-dimensional depiction of hand positioning over time. This could lead to a more robust understanding of manipulation for AI models.
The data generated by Encord is meticulously annotated. These annotations include detailed physical descriptions of the actions captured in the video, such as "right hand tightens bolt." This granular annotation is intended to significantly aid Large Language Model (LLM)-based models in comprehending the context and actions within the data. Velmurugan estimates that this dense annotation is approximately 100 times more valuable than what he terms "junky ego data" for training specific tasks, despite costing only about 20 times more to produce. This cost-benefit analysis suggests a favorable return on investment for high-quality, annotated data.
However, the phrase "20 times more" represents a significant financial commitment. This is where the economics of physical AI diverge sharply from the development of LLMs. The process of scraping text from the internet, which enabled LLM developers to build their models by drawing from vast repositories like Stack Overflow and the broader web, was virtually cost-free. In contrast, generating physical training data requires significant investment in infrastructure, human operators, and sophisticated sensor technology. This fundamental difference in data acquisition cost represents a critical limitation in directly comparing the development trajectory of physical AI to that of LLMs.
The Business of Data Generation and Industry-Wide Insights
Despite the financial hurdles, progress is being made. Velmurugan, with his unique vantage point across numerous robotics companies, observes a consistent evolution in the industry. Startups and advanced research labs are actively experimenting with and refining data generation techniques, sharing insights into what works and what doesn’t. This collective learning is crucial for improving physical AI models. Encord’s position, situated between many robotics companies, allows it to identify data techniques that are gaining traction across the industry before any single customer can. This holistic view is a key component of Encord’s value proposition.
The demand for this specialized data generation is keeping the dozen or so pilots at Encord’s facility consistently engaged. Both Sofia Infante and Andrew Ceja represent a growing workforce dedicated to building the foundational components for neural networks. Previously, they were employed at Scale, another prominent AI data annotation firm, before transitioning to Encord.
Ceja’s journey into this field is particularly illustrative. He previously worked at a waste management company, where his aptitude for technology led him to oversee the maintenance of a robotic trash sorter. Now, as he navigates the complexities of training robots by deconstructing a Jenga tower, he expresses satisfaction with the intellectual challenge. "It’s something new every day!" he exclaimed, highlighting the dynamic and problem-solving nature of his current role.
The current bottleneck in physical AI development is not a lack of theoretical frameworks or ambitious goals, but a fundamental need for vast quantities of diverse, high-fidelity real-world data. Companies like Encord, by embracing innovative data collection methodologies, including the integration of human cognitive data, are working to bridge this gap. Their efforts, though expensive and resource-intensive, are crucial for unlocking the full potential of robotics and bringing sophisticated AI capabilities from the digital realm into the physical world. The Jenga game in San Leandro, therefore, is more than just a demonstration; it is a microcosm of the intricate, challenging, and ultimately vital work being done to build the future of physical intelligence.
