reading
11 min read

The Physical World Has No Data Stack

Robots are about to become one of the largest data sources on Earth. The infrastructure to move, store, and query that data barely exists.

Tobi Coker

Venture funding for robotics and physical AI has doubled two years running

Robots have a data problem.

Autonomous vehicles, drones, humanoids, and industrial systems are moving from lab demos into deployed fleets. They generate data unlike anything the current stack was built to handle: LiDAR sweeps, video from multiple cameras, audio, GNSS, and proprioceptive state, all streaming at once and synchronized to the millisecond. A single vehicle can produce terabytes of it every day.

That data will feed every foundation model of the physical world. Yet most of it still lands in home-built pipelines, gets sampled down to a fraction of itself, or gets deleted because moving and storing it costs too much.

We have seen versions of this problem before. The relational era produced Oracle. Web-scale data produced Hadoop, followed by cloud warehouses and Snowflake. Real-time analytics, the most recent era, is producing companies like MotherDuck, TimescaleDB, and ClickHouse. These infrastructure companies were not side effects of their respective waves. They helped make the waves possible.

Physical AI is still missing that layer. Someone will build the Snowflake of the physical world. The harder questions are what that product needs to look like, who is best positioned to build it, and whether the winner can remain independent.

Why this data breaks the current stack

The current data stack just isn’t built for this world.

Start with the shape of the data. Snowflake and ClickHouse assume rows and columns: structured records, semi-structured text, and time series. A robot produces video, LiDAR point clouds, audio, location data, and joint-level telemetry. Each stream has its own format, and each is far less useful without the others.

The value comes from synchronization. An engineer needs to know what the camera saw at the exact moment a gripper slipped or a vehicle made the wrong turn. Columnar engines have no native concept of that moment across a dozen heterogeneous sensors. Teams rebuild the time-alignment layer themselves, often for every robot and every new sensor configuration.

Then there is the cost.

The cloud warehouse era relied on a convenient assumption: data was relatively small compared with available bandwidth, so you could ship everything to the cloud and rent compute next to it. Physical-world data reverses that equation. It is generated on vehicles, factory floors, farms, and construction sites, often over cellular connections. At fleet scale, uploading everything is a non-starter. Moving the data can cost more than storing it.

The winning system will need to decide at the edge what is worth keeping, move only what matters, and query the rest without hauling petabytes across a modem. Elastic cloud compute alone does not solve that problem.

The Cloud era assumed data was small and bandwidth was cheap.

Finally, there is portability. Structured data became an asset because it could move. SQL, CSV, Parquet, and shared schemas allowed data to pass between tools and organizations. Physical-world data has no broadly accepted conventions for clock synchronization, coordinate frames, sensor calibration, or units.

Without those conventions, one team’s logs can be difficult for another team to read, even inside the same company. Sharing data across companies is harder still. That limits the value of every dataset in the industry. Data that cannot be shared, pooled, or sold cannot compound.

Taken together, the product spec is much larger than a database. It includes an interchange format, edge-aware ingestion, petabyte-scale storage and query for multimodal streams, and the visualization and debugging tools engineers use every day. No incumbent covers more than a piece of it.

The value is in the moment, and the current stack has no way to find it

What we learned from the last data cycle

The structured-data stack was built by companies that did three things well.

They chose a type of data the incumbents handled poorly. They built around an open standard, or created a standard of their own. And they arrived just as a demand-side shift made the old way of working intolerable.

For Snowflake, that shift was cloud migration. In the data streaming era, platforms solved for the demand for real-time analytics that batch warehouses could not meet. In both cases, the customer pain moved from manageable to urgent within a few years. The companies that had spent the quieter years building were ready when that happened.

Physical AI is beginning to create the same conditions.

Every new shape of data has produced its own infrastructure company

Who is building it

Foxglove is the clearest category leader today. Its founders came from Cruise, where they had built internal tooling for robotics data and watched teams across the industry rebuild versions of the same stack.

The company’s open-source contribution, MCAP, is a container format for timestamped multimodal data. It has already won one of the standards battles that matters most: MCAP is the default logging format in ROS 2 (opens in new tab), the operating layer for a large share of the world’s robots. NVIDIA also made it the default recorder format in Isaac ROS 3.0 (opens in new tab).

Foxglove has built a commercial platform on top of that standard, with petabyte-scale storage, search, and query for sensor data, along with the visualization tools engineers use to replay and debug what a robot saw. Customers include NVIDIA, Amazon, Anduril, Wayve, and Dexterity (opens in new tab). In November 2025, Bessemer led a $40 million Series B (opens in new tab), bringing its total funding to more than $58 million.

Standard, platform, and workflow. Foxglove is the only company that currently has all three, which makes it the closest thing this category has to a Snowflake-shaped company.

Rerun is the open-source challenger. The Stockholm-based company started with a visualization library for spatial and embodied AI that spread through the research community. Meta, Google, and Hugging Face’s LeRobot ecosystem have all used it. Rerun is now building outward into a broader data platform after raising a $17 million seed led by Point Nine (opens in new tab) in early 2025.

Its bet will be familiar to anyone who followed the last generation of infrastructure companies: win developers first, then follow them into production.

Roboto and several smaller startups are going after narrower parts of the stack. Roboto raised a $4.8 million seed from Unusual Ventures (opens in new tab) to put natural-language search on top of robot logs. It is a useful product, but it looks more like a feature of the eventual platform than a platform itself. We expect many of these point solutions to become acquisition targets as the category consolidates.

The demand side may matter even more than the current vendor landscape. Humanoid companies, autonomous-vehicle fleets, defense primes, and industrial-autonomy startups are drowning in data they did not set out to manage. Every Figure, Wayve, and construction-autonomy company that scales a fleet could become an eight-figure infrastructure customer.

That is the demand curve these companies are racing to meet.

Why now

Venture funding for robotics and physical AI has doubled two years running


Felicis first looked seriously at this category in 2022. Our view at the time was straightforward: great engineering, small market, little urgency. There were only a few thousand potential buyers, most of them research-stage robotics teams. Bad data tooling was frustrating, but it was rarely a real budget line.

Today, the engineering is still great, but the market has mushroomed and the race is on.

Robotics and physical AI startups raised a record $27.6 billion in 2025 (opens in new tab), roughly twice the prior year’s total. The first half of 2026 has already surpassed it (opens in new tab). Humanoid programs at Figure, Apptronik, and a dozen other companies have moved from videos into deployments. NVIDIA has made physical AI a strategic priority.

Funding does not become deployment overnight. But over time, those dollars become fleets, and every fleet produces data that has nowhere good to live.

Cloud migration followed a similar path in the 2010s. A demand-side wave turned an infrastructure annoyance into an emergency. By the time customers felt the problem acutely, the companies that would define the category had already spent years building.

The standards for physical-world data are still unsettled and most of the large contracts are still unsigned. That makes this the period when the infrastructure layer gets decided.

The bear case

The counterarguments are real. They also explain why sophisticated investors, including us, spent years watching this category from the sidelines.

The market is still small. This is true today. The number of companies operating robot fleets at data-infrastructure scale is measured in the hundreds, not tens of thousands. The investment case depends on what happens next. If robots remain stuck in pilot programs for another decade, the market will not support a large horizontal platform. We do not think that is where this is headed. Data infrastructure is one of the most important bottlenecks for deployments, and so scaled deployments will pull forward the demand for better data infrastructure.

Customers may not feel enough pain to pay. This was the strongest objection in 2023. It gets weaker with every fleet deployment. A research team can tolerate sampled logs and home-built tooling. A company operating a thousand robots in customer environments and retraining models weekly cannot. The urgency rises with fleet size, and fleet sizes are beginning to inflect.

The platforms could eat the category. This is the most serious risk. NVIDIA has adopted MCAP rather than competing with it, but adoption can eventually become absorption. AWS and the other hyperscalers have every reason to add robotics ingestion to their existing storage products. An open standard is also, by design, something anyone can build on. MCAP is becoming largely standardized. However, if model providers adopt a data format, it is more likely for deployments that use the model to adopt that format as well. Platform level adoption is largely a positive for data infrastructure bets.

Snowflake answered a similar threat by becoming the neutral layer across clouds. Neutrality became part of its moat. The robotics equivalent would be a platform that every fleet trusts because it is independent of the chip vendor, the cloud provider, and any application company that might also be a competitor.

An independent company can occupy that position. Whether it can defend the position against hyperscaler pricing is the question we keep coming back to.

The market could fragment by robot type. A drone fleet, a humanoid, and a warehouse arm produce different data under different operating constraints. Each vertical could develop its own stack, bundled inside the application company. The best-funded companies may also decide to pay for it themselves. The first wave of developer tools for autonomous vehicles ran into exactly that: Waymo, Cruise, and their peers built most of the stack in house.

We think the commonalities are more important. The robots differ, but the underlying problems around synchronization, storage, edge ingestion, and debugging are shared. The same argument once suggested that every SaaS vertical would need a different data warehouse. That did not happen. And the in-house builds are where this category came from. Foxglove’s founders built the Cruise version before leaving to sell it to everyone else.

What the winner looks like

Remove the company names and the historical pattern gives us a fairly clear product spec.

First: The winner owns or stewards an open standard before monetizing around it. In a fragmented hardware market, the interchange format is the entry point. You cannot build a platform around data you cannot read. ROS is the cautionary version. The Open Source Robotics Foundation stewarded the standard, and ROS became the operating layer for most of the world’s robots, but nobody built the commercial platform on top of it, and the fleets that would have paid for one did not exist yet.

Second: The winner solves the economics of the edge. The defining cost in this market is often the connection, not the disk. A platform that assumes every byte will reach the cloud will break as fleets grow.

Third: They convert developers before expanding into large enterprise contracts. Robotics teams choose tools the same way software engineers have for years: they adopt whatever makes tonight’s debugging session shorter. The commercial expansion comes later.

Again, Foxglove currently fits that description better than anyone, which is why they have become the reference point for the category. But that doesn’t mean they will be a monopoly. Snowflake was founded in 2012 after BigQuery in 2010. ClickHouse spun out in 2021 after TimeScaleDB launched in 2018, and the structured-data market still created room for Databricks and a long list of adjacent winners. A market this large is unlikely to resolve around one company.

We expect physical AI data infrastructure to support at least one large horizontal platform, several vertical stacks inside major application companies, and a wave of consolidation across the point solutions in between.

The question we are watching is whether the horizontal winner remains independent or gets absorbed by the compute layer. The last independent company to solve a comparable data problem became worth more than $100 billion (opens in new tab).

That is why we think physical-world data infrastructure is one of the most interesting markets to be investing in right now.

--

A massive thank you to Eric Deng, Jacob Zietek, and Vishnu Mano for your thought partnership on this piece. If you’re in this space, please get in touch: tobi@felicis.com.

Authors

  • Tobi Coker

    Partner

Tags

    RoboticsInfraAI

Share

Newsletter

Get the latest news & insights

from the Felicis community.