
Work Here?
Encord provides an AI-native data infrastructure platform for Physical AI, helping enterprises build robotics, autonomous systems, drones, and related applications by curating, managing, annotating, and evaluating multimodal data needed to train real-world AI models. Its three products—Annotate for collaborative multimodal annotation, Active for data curation and quality evaluation, and Index for petabyte-scale dataset search and management—cover the full data lifecycle for Vision-Language-Action models, foundation models, and embodied AI. It differentiates itself by serving as a universal data layer for production-grade Physical AI, addressing the infrastructure gap left by cloud-era tools built for text and tabular data. The goal is to enable teams to efficiently build, evaluate, and deploy real-world AI systems with scalable data tooling across robotics, healthcare, life sciences, defense, and related industries.
Industries
Data & Analytics
Robotics & Automation
Enterprise Software
AI & Machine Learning
Company Size
201-500
Company Stage
Series C
Total Funding
$107.1M
Headquarters
London, United Kingdom
Founded
2020
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$107.1M
Meets
Industry Average
Funded Over
5 Rounds
Industry standards
Social events
Flexible hours
Hybrid work
Fresh equipment & gear
Health insurance
Generous PTO
Visa sponsorship
Team lunches
Encord bets on brain wave robot training to solve the physical AI data crisis. Brain wave robot training has arrived in a warehouse in San Leandro, California, and it looks like a Jenga game. Encord, a data-tooling startup that recently closed a $60 million Series C led by Wellington Management, is experimenting with headsets that measure brain activity while human operators disassemble block towers, pour coffee, and stack poker chips, all in the service of training robot models. The idea: that mental states, not just physical movements, might be the missing signal for physical AI. The headset in question was built by Zander Labs, a German-Dutch deep-tech company founded in 2016. Its hardware product, the Zypher Suite, is a mobile EEG system designed to monitor brain activity and provide real-time neuroadaptive data with local processing. Lucas Gehrke, a Zander neuroscientist supervising the Encord trial, says the volume of brain activity at any moment in a task reveals when a model needs to deploy its highest-effort processing, essentially flagging where the hard bits are. Encord's work with Zander is still a trial run. The plan is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually moves performance metrics before deciding whether to scale. That is a sensible way to frame a bet on genuinely unproven ground. Brain wave robot training gets serious, but the data problem is bigger. Vineeth Velmurugan, Encord's head of robot learning and a veteran of OpenAI's robot lab and warehouse-automation firm Berkshire Grey, describes brain wave work as the 'bleeding edge' of addressing the robotics data bottleneck. The bottleneck itself is less exotic: the physical world simply has not generated enough labelled manipulation data to train capable robot models at scale. 'The data simply does not exist,' Velmurugan says. He estimates it will take a data set roughly five times the size of YouTube's entire video corpus to break through. That figure helps explain why data generation has become a commercial business rather than a side project inside a research lab. Encord collects two main types of training data. The first is 'egocentric' video, captured by workers wearing head-mounted cameras, often augmented with extra camera angles and sensor readings. The second is data from leader-follower rigs, paired robotic arms where a human controls one and the other mimics it. When TechCrunch visited the facility, pilots were using these rigs to practise pouring coffee and stacking poker chips. Storage racks held fake flowers, plastic vegetables, kitty litter trays, bags of cables: the props list for teaching robots to handle household objects. Encord is also developing a forearm sensor that reads electrical signals in muscles, aiming to reconstruct a 3D model of hand position for moments when camera footage loses sight of the fingers. Each data set gets dense text annotation, 'right hand tightens bolt' style, to help language-model-based robot brains parse what is happening. Velmurugan estimates that kind of annotation is worth 100 times as much as raw egocentric footage for training specific tasks, and costs only 20 times more to produce. On paper, a good trade. In practice, 20 times more than near-zero is still a real number. The economics that make physical AI different from the LLM playbook. This is where the comparison between physical AI and large language models runs out of road. LLM builders scraped text from Stack Overflow and the rest of the open web for close to nothing. Physical training data has to be manufactured by people in warehouses with sensors on their forearms and, now, electrodes on their heads. The cost structure is categorically different. Encord's own growth suggests the market is real regardless of that friction. According to SiliconANGLE, the data volume on Encord's platform grew from just over 1 petabyte to more than 5 petabytes in the roughly 18 months before the Series C, and revenue increased more than 10 times over that same period. The Series C, which also drew in existing backers Y Combinator, CRV, N47, and Crane Venture Partners alongside new investors Bright Pixel and Isomer Capital, brings total funding to $110 million. Zander Labs, for its part, has broader ambitions than robot training experiments. The company signed a contract worth €30 million (approximately $32.9 million) with the German Agency for Innovation in Cybersecurity to develop neurotechnological prototypes under a project called NAFAS, which uses a passive brain-computer interface that does not require users to actively imagine actions. That contract runs until November 2027, according to MassDevice. Competition is forming. In June 2026, a rival physical-AI data startup called XDOF emerged from stealth with $70 million in funding from Thrive Capital, Spark Capital, a16z, Lux, and WndrCo. Founded in October 2024 by UC Berkeley researchers, XDOF already serves about 20 customers and released a 130,000-trajectory manipulation dataset with Berkeley's AI research lab, according to Business Model Analyst. The race to become the default supplier of physical training data is, apparently, fully on. Back in San Leandro, Andrew Ceja is rebuilding the Jenga tower and pulling it apart again. He previously worked at Scale, another AI data-annotation firm, before that at a waste-management company where he kept a robotic trash sorter running. Now he annotates his own brain waves for the benefit of robot models that do not yet exist at scale. 'It's something new every day!' he says. The brain wave data may or may not move the needle on model performance. The question Encord will answer first is whether the signal is real. The economics question comes after that. Marcus Hale has been filing general news for the better part of fifteen years. He started at a regional evening paper, moved to a mid-sized digital outlet covering UK news, and spent three years as a general assignment reporter before going freelance. He has covered inquests, council elections, infrastructure announcements, and the kind of stories that sit on page five but matter on page one. He writes about public services, housing, local government, and the institutional stories that take six months to develop and thirty seconds to read. He prefers facts to angles and considers that unfashionable. Marcus lives in Bristol. He still reads the local paper and thinks that makes him an endangered species. July 31, 2026 July 30, 2026
Brain waves: A new weapon for training AI robots. July 27, 2026 8 views In a warehouse in San Leandro, California, an employee sits wearing a helmet equipped with a camera and neural sensors, carefully pulling wooden blocks from a wobbling tower in full concentration. The scene looks ordinary at first glance, but it is in fact a next-generation experiment: measuring human brain waves with the goal of improving robots' ability to learn and understand the physical world around them. The data crisis: the real obstacle facing tomorrow's robots. The physical AI industry - the field concerned with building robotic systems capable of interacting with their surrounding environment - faces a fundamental challenge that has nothing to do with model design or processing power. The real bottleneck lies in the scarcity of real-world training data. Unlike large language models, which built their knowledge base from billions of web pages at minimal cost, robots require precise physical data that is difficult to collect or replicate digitally. UK-based Encord, a company specializing in data tools for training AI models, is tackling this crisis with a different approach: rather than simply managing existing data, it now manufactures data from scratch. Vineet Velmoorgan, Head of Robotic Machine Learning at Encord and a veteran of OpenAI's robotics lab, puts it plainly: "The data simply doesn't exist." Brain waves: data from a new layer. Encord is collaborating with German neuroscience firm Zander Labs to develop a helmet that measures the brain's electrical activity while human operators perform manual tasks. The idea goes beyond recording movements - it extends to capturing mental states such as error, surprise, and intent, converting them into an additional data layer that supplements traditional video footage. Lucas Gehrke, the neuroscientist overseeing the project, explains that brain activity levels at specific moments give models valuable cues about when to apply the highest levels of computational attention - which translates in practice to more efficient models that waste fewer resources. The experiment is still in its early stages; Encord plans to build a dataset annotated with brain wave data, then test its impact on the performance of actual robotic models before deciding whether to scale up. Multiple approaches to collecting real-world data. Encord's pipeline relies on two primary sources for generating training data: * Egocentric Video: Footage captured from the human operator's perspective via head-mounted cameras, supplemented by additional angles and complementary measurements, collected from factories at multiple locations around the world. * Leader-Follower Rigs: Paired robotic arms in which a human directly controls one arm while the other precisely mirrors its movements, used to capture data on fine-grained tasks such as pouring coffee and connecting cables. The team is also experimenting with a third emerging technique: forearm-mounted sensors that read electrical signals in the muscles, with the aim of building a full three-dimensional model of hand movement - something conventional video footage cannot achieve with sufficient accuracy. Cost: the decisive difference between robots and chatbots. Velmoorgan estimates that richly annotated data - such as "right hand tightening the screw" - is up to one hundred times more valuable for training than raw data, while costing only about twenty times more to produce. On paper, that's a winning equation. Yet "twenty times more" still means real money. This is the fundamental difference between physical AI and its language counterpart: language models built their knowledge wealth from mostly free internet data, while robot data must be generated hand by hand, trial by trial - radically changing the economics underlying the development of these models. Velmoorgan believes that reaching a true breakthrough requires a volume of data roughly equivalent to five times all the video content on YouTube. Strategic position: working with everyone at once. Encord's work with a broad spectrum of robotics companies - whose names Velmoorgan declines to reveal - affords the company a rare strategic vantage point on progress across the sector. It can observe which approaches prove most effective and which fall short, then leverage that accumulated knowledge for the benefit of all its clients. In a market where data is the rarest gem of all, this intermediary position amounts to a formidable competitive advantage. What is happening in the San Leandro warehouse is not merely data collection; it is an emerging model for an entire industry: the real-world data industry, which may prove to be the hidden engine driving the physical AI revolution in the years ahead. بقلم فريق دروب أيديا DROPIDEA. Dropidea hope this article has added real value to you. At DROPIDEA, Dropidea always strive to deliver high-quality content that helps you grow and evolve in the digital space. Follow Dropidea for more useful articles and guides. #الذكاء الاصطناعي الجسدي #الروبوتات #بيانات التدريب #موجات الدماغ
Brain wave data could solve the physical AI training gap. technology TrendPulse AI Analysis This article covers technology trends, sourced from TechCrunch. Its AI system has analyzed the key points and extracted the most relevant insights for decision-makers. Below is the structured breakdown of the original content. Quick summary. * Encord is testing the use of brain-wave sensors to tag training data, aiming to improve how humanoid robots learn complex physical tasks. * By capturing human intent and cognitive states, researchers hope to overcome the scarcity of high-quality, real-world data currently hindering physical AI development. * The industry is shifting from merely managing existing data to actively manufacturing synthetic and sensor-rich datasets to scale robotic capabilities. Key details. At a specialized facility in San Leandro, Encord is pioneering a new method for training robots by integrating neuro-data into the learning loop. Pilots wearing headsets equipped with sensors from the German startup Zander Labs perform delicate tasks, such as playing Jenga, while their brain activity is recorded. This data is intended to help AI models distinguish between routine movements and moments of high cognitive load, such as error detection or surprise. This initiative addresses the critical bottleneck in robotics: the lack of high-fidelity, real-world training data. While companies have traditionally relied on video footage or remote-operated robots, these methods often lack the nuance required for sophisticated manipulation. Vineeth Velmurugan, head of robot learning at Encord, notes that the industry requires a massive scale of data - potentially five times the size of YouTube's entire library - to achieve true breakthroughs in end-to-end robotic learning. The data simply does not exist. - Vineeth Velmurugan, Head of Robot Learning at Encord Why this matters. The integration of neuroscience into machine learning signals a maturation of the robotics sector. As the industry moves beyond basic automation, the focus is shifting toward "intent-aware" models that can navigate the unpredictability of the physical world. For investors and executives, this marks a transition where data quality - specifically data that captures human cognitive nuance - becomes a more valuable asset than raw model architecture. Companies that successfully master the creation of these specialized datasets will likely gain a significant competitive advantage in the humanoid and warehouse robotics markets. If Encord's trial proves that brain-wave tagging leads to higher performance, TrendPulse should expect a surge in demand for multimodal training data that goes far beyond traditional video capture. The bottom line. By bridging neuroscience and robotics, the industry is moving toward a new standard of data generation that could finally unlock the potential of physical AI. This article has been processed and analyzed by TrendPulse AI for informational purposes. Content may have been summarized or restructured for clarity.
Are brain waves the next unlock for physical AI? 5:19 PM PDT · July 26, 2026 The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California. That warehouse is occupied by Encord, a company that builds data tooling used to train AI models. Andrew Ceja is a pilot - the company's term for its robotic trainers - and he's carefully pulling wooden blocks from a tottering tower while wearing a headset with a camera that tracks what he sees. That alone is fairly common for collecting robot training data, but this headset includes sensors that measure his brain waves as he carefully disassembles the block tower. Encord is one of a small but growing number of startups betting that the next real constraint on humanoid and warehouse robotics won't be model architecture but instead the sheer scarcity of real-world physical training data. Rather than just helping robotics companies manage the data they have, Encord is building a business around manufacturing the data they don't. The brain wave headset Ceja is wearing was built by Zander Labs, a German neuroscience startup that's betting measuring brain activity - to deduce mental states like error, intent and surprise - can create a more useful data set to train models. Encord's work with Zander is currently a trial run; Encord says the goal is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up. Lucas Gehrke, a Zander neuroscientist supervising the work, says that the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models. This is the "bleeding edge" of the effort to solve the robotics data bottleneck, according to Vineeth Velmurugan, Encord's head of robot learning. A veteran of OpenAI's robot lab and Berkshire Grey, the warehouse automation firm, Velmurugan joined Encord to build the company's internal data-creation team. Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers - Velmurugan says they work with many leading robotics firms but that he's not authorized to name them - began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it. "The data simply does not exist," Velmurugan said. OpenAI's own model went rogue before Kimi had Wall Street sweating | Equity Podcast 0 seconds of 37 minutes, 12 seconds Volume 0% The bet that generative AI can do for robots what it's done for chatbots keeps running into this same wall. LLMs were built on the text of the entire internet, and more. Finding the same raw materials to teach neural networks about physical manipulation is challenging: self-driving car companies collect it themselves, but that's hard to scale. Training from video can work, but it lacks the fidelity of real world data. Velmurugan says it will take a data set something like five times the size of YouTube's video corpus to break through - a scale that helps explain why data-generation itself has become a business and not just a research problem. Feed your egocentric data needs. Companies building robot brains are now turning to two main sources: "Egocentric" video collected by workers wearing cameras, often augmented with additional camera angles and other metrics, and collecting data from robots operated remotely. Encord does both, drawing egocentric data from several factories around the globe, and using its San Leandro facility to experiment with new modalities, like brain waves, or collect data sets around specific skills for fine-tuning. When TechCrunch visited, pilots were using leader-follower rigs - paired robotic arms, one controlled directly by a human operator and one that mimics its movements - to create data about tasks like pouring coffee from a pot into mugs (very sloshy) and stacking poker chips. "Every humanoid company has asked us for these pieces," Velmurugan says. Storage racks held cartons of fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires, the stock in trade for training manipulators for household tasks. At one of these stations, another pilot, Sofia Infante, maneuvers robotic arms to plug and unplug ethernet cables from the back of a server - the kind of work data center operators would love to be automated, if only robots could manipulate them with the required precision. Taking a spin behind the controls, I was able to see why that's still out of reach: Pincers are far less dextrous than human fingers and lack the degrees of freedom we take for granted in our arms. Another new data modality that Encord is developing uses a set of sensors strapped to the forearm to detect electrical signals in muscles. Video taken of human hands manipulating objects typically doesn't capture the entire hand, but Velmurugan hopes to build a 3D depiction of where the hand is at any time based on the arm sensors, creating a more robust understanding for models. Encord's data sets are annotated with physical descriptions of what each video contains - "right hand tightens bolt" - to aid LLM-based models in understanding what is happening. Velmurugan estimates this kind of dense annotation is worth 100 times as much as "junky ego data" for training specific tasks, and it only costs 20 times more to produce, which is a good trade, on paper. But "20 times more" is still real money, and that's the catch: scraping text off the internet, the way LLM makers built their models by pulling from Stack Overflow and the rest of the web, cost frontier labs next to nothing. Generating physical training data does not, and that's the limit of the physical-AI-as-LLM comparison. This kind of data has to be manufactured, not just collected, and that changes the economics of building these models. Velmurugan says that progress is being made - with Encord's visibility into programs across the industry, he's able to see start-ups and frontier labs alike figure out what works and what doesn't to improve physical AI models. That vantage point - sitting between many robotics companies at once - is also part of Encord's pitch. It can spot which data techniques are gaining traction industry-wide before any single customer can. That will keep the dozen or so pilots at Encord's facility busy. Both Infante and Ceja are part of a burgeoning workforce developing the building blocks for neural networks; they previously worked at Scale, another AI data annotation firm, before joining Encord. Ceja had worked at a waste management company where his interest in technology found him in charge of keeping a robotic trash sorter in good working order. Now, as the Jenga tower topples, he says he enjoys the challenge of solving training tasks for robots - "It's something new every day!" When you purchase through links in our articles, we may earn a small commission. This doesn't affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@ techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $330 today!
Encord integrates NVIDIA Cosmos Reason and Embed. Co-Founder & CEO at Encord Summarize with AI Encord is excited to announce the integration of two NVIDIA Cosmos models - Cosmos Reason 2 and Embed - directly into the Encord platform. Both run natively on Encord's own infrastructure, so customers can use them in production from day one. The Cosmos Reason agent automates prelabels on physical AI video. Cosmos Embed makes that video searchable by behaviour, not just by scene. Together, they turn raw video into structured, searchable training data. Cosmos Reason 2 in the Agents Catalog Cosmos Reason: automated captioning and pre-labelling. Cosmos Reason is a vision-language world model. Given a video clip as input, it returns natural-language descriptions of the actions, objects, and scene context in the footage. Inside Encord, those descriptions arrive as prelabels attached to the right video segments. Annotators review and refine instead of starting from a blank canvas. Approved labels stream straight to the customer's training pipeline. What this means in practice: * A robotics team labelling dexterous manipulation no longer hand-types every grasp or release. * An AV team captioning camera footage gets a usable starting point on every clip instead of writing each one from scratch. * Industrial inspection teams get descriptions of anomalies, object states, and conditions without anyone watching the full reel first. This means less time describing video, and more time improving the model. Cosmos Embed: action-aware embeddings and behaviour search. Cosmos Embed is a video embedding model. Each embedding is calculated over an eight-frame window, capturing the action that happens across those frames. That makes a video dataset searchable by behaviour, using natural-language queries. In practice, you can query for what's actually happening in a clip - not "car on a snowy road" but "car turning in a snow storm," not "highway scene" but "car overtaking another vehicle." That makes edge cases easier to find - the rare scenarios that decide whether a model ships safely. How they work together. * Robotic manipulation. Reason generates action and state descriptions for dexterous manipulation footage. Embed lets teams find the specific behaviours that need more training examples. * Autonomous vehicles. Reason captions camera data at scale. Embed lets teams pull the exact driving scenarios that matter for their perception models - lane changes in rain, unprotected lefts, occluded pedestrians. * Industrial inspection. Reason describes anomalies and object states in operational footage. Embed surfaces every clip where a defect occurs in a specific context. * Vision-Language-Action models. Reason produces the grounded captions VLA training depends on. Embed makes the underlying dataset queryable by the behaviours those models need to learn. Get the data right. 300+ of the best AI teams in the world use Encord.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Robotics & Automation
Enterprise Software
AI & Machine Learning
Company Size
201-500
Company Stage
Series C
Total Funding
$107.1M
Headquarters
London, United Kingdom
Founded
2020
Find jobs on Simplify and start your career today