A robotics data startup that only recently came out of stealth mode is already in advanced discussions to close a major new funding round at a valuation of approximately $1.2 billion. XDOF, which specializes in gathering real-world teleoperation data to train general-purpose robots, is in late-stage negotiations for a Series B investment led by 8VC, according to multiple sources familiar with the matter. The talks are ongoing and the final terms have not yet been confirmed.
The speed of this fundraising push is notable. XDOF only exited stealth mode less than three months ago, and the company had not originally planned to return to investors so quickly after its previous round. What changed, sources say, is the company's rapid commercial traction - with annualized revenue now approaching $50 million - which drew venture capitalists to approach XDOF proactively about a follow-on investment rather than the other way around.
Neither XDOF nor 8VC responded to requests for comment. The total amount being raised in the Series B has not been disclosed, and it remains unclear whether the reported $1.2 billion valuation is pre- or post-money.
From Academic Research to a Billion-Dollar Startup
XDOF was founded in 2024 by two researchers from the University of California, Berkeley - Philipp Wu, who serves as CEO, and Fred Shentu, who serves as CTO. The company raised a $70 million Series A in June, drawing backing from a high-profile group of investors including Thrive Capital, Andreessen Horowitz, Lux Capital, and Spark Capital.
The intellectual roots of the company trace back to Wu's doctoral research, during which he was investigating how robots could learn effectively from large datasets. A core obstacle he encountered was the absence of sufficient real-world data at scale. To address this, Wu and Shentu collaborated on a project called GELLO - a low-cost teleoperation system that enables a human operator to remotely control a robotic arm and, in doing so, generate the kind of training data that robots need to learn physical tasks. The project produced an influential academic paper in the robotics field and ultimately became the technical foundation on which XDOF was built.
Filling a Critical Gap in the Robotics Data Pipeline
XDOF's core business model is built around solving one of the most persistent challenges in physical robotics: the absence of large, high-quality datasets that robots can learn from. Unlike large language models, which were initially trained on enormous volumes of text scraped from the internet, physical robots have no equivalent reservoir of real-world interaction data to draw from. This makes data collection a fundamental bottleneck in the development of general-purpose robotic systems.
The company positions itself as an outsourced data-supply chain for the robotics industry, building the infrastructure - including data pipelines, collection tools, and annotation systems - that frontier AI labs and robotics companies either lack the capacity or the resources to construct on their own. Investors have begun describing XDOF in terms that reference established data-labeling powerhouses, comparing it to Scale AI or Mercor, companies that played a significant role in enabling the broader AI boom by supplying structured training data at scale.
To gather its data, XDOF uses a dual approach. On one side, it deploys remote robot teleoperation, where human operators steer robotic arms from a distance to perform structured physical tasks. On the other side, it uses egocentric data collection, in which human participants wear body-mounted sensors that record everyday activities such as folding laundry or flattening cardboard boxes. The combination of these two methods is intended to capture a broad and diverse range of physical interactions that can be used to train robots across different environments and tasks.
A Major Dataset Release in Partnership with UC Berkeley
In a move that signals both its academic ties and its ambitions in the research community, XDOF is partnering with UC Berkeley's AI Research Lab to release what the company describes as the largest collection of high-quality robot training data ever assembled. The dataset, referred to as ABC, is intended to serve as a foundational resource for researchers and developers working on robotic learning systems.
To scale its data collection operations globally, XDOF plans to recruit and train dedicated teams of collectors in multiple countries. These teams will include both teleoperators - individuals who remotely control robots to generate structured training sequences - and egocentric operators who wear sensor equipment to capture natural human movement data in real-world settings.
Customer Base and a Competitive Landscape
XDOF has confirmed that it is already working with approximately 20 customers, a group that reportedly includes several frontier AI laboratories. This early commercial momentum appears to be a driving factor in the investor interest that has accelerated the Series B discussions.
The startup is not operating in isolation. Several other companies are pursuing similar opportunities in the real-world robotics data space. Mecka AI is among the startups focused specifically on collecting physical training data for robots. Meanwhile, larger human-data platforms that built their reputations supplying training data for language models - including Scale AI and Micro1 - are also expanding their offerings into the physical robotics domain, signaling growing competition as demand for robot training data continues to rise.
If the Series B closes at the reported valuation, XDOF will have achieved unicorn status in under a year of public operation - a trajectory that reflects both the urgency investors are placing on physical AI infrastructure and the relatively thin supply of companies capable of delivering data at the quality and scale that the robotics industry requires.



