Physical AI is one of the hottest sectors in venture investing, with companies raising billions to apply the tools that produced large language models to robotics.

That excitement delivered a big IPO for Unitree, China's leading robot maker, valued at $66 billion after arriving on China's equivalent of the Nasdaq. Shortly after, the bottom fell out and the company lost nearly half its value.

The reason analysts point to is single and plain: while the robots' physical capabilities are improving, they still lack the know-how to do value-creating work.

The industry's own diagnosis

At the Actuate conference, a gathering of developers building AI brains for robots, the excitement was visible. The event has tripled in size since it kicked off in 2023, reaching 1,500 attendees. But the risk was just as visible: a sign on an infrastructure company's booth promised to solve "the robotics data crisis."

That crisis is the lack of high-quality training data. Attempts to build generalized robots that can do any task are still far off, and using end-to-end learning for specific tasks has not delivered products with reliable, commercial performance.

For developers, the answer is to better mimic the advances of the frontier AI labs: find or create more diverse datasets, try different training regimes and build better reinforcement learning scenarios.

What "GPT-2 era" means

Harry Mellsop, a founder of Antioch, a startup building simulation tools for model builders, says physical AI is in its "GPT-2 era" — a reference to the OpenAI model that pre-dated ChatGPT.

The analogy carries two things at once. First, a capability level: something works, but it is not a product. Second, and more importantly, that the way out is known — more data and more compute. Mellsop notes the need for GPUs optimized for ray tracing in particular, which are used to create high-fidelity simulations.

Why autonomous vehicles are ahead

The furthest ahead branch is autonomous vehicles, for two reasons, both about data:

  • The data can be collected — relevant data comes from cars driven by people. There is no comparable source for how a robot arm should grip a glass.
  • The task is easier — the main job is to avoid contact, not to manipulate the physical environment. Not touching something takes far less skill than holding it correctly.

Much of the tooling for model-building also comes from autonomous vehicle companies. Foxglove, the conference organizer, was founded by former employees at Cruise, General Motors' erstwhile self-driving effort.

Now those car companies are increasingly betting that their investments in ML tooling will let them compete with dedicated humanoid makers. Tesla is trying this with its Optimus robot, and both AV-focused Wayve and rideshare giant Uber have launched robotics labs focused on humanoid form factors.

Together those two reasons produce an important consequence: the knowledge accumulated in autonomous vehicles becomes capital that transfers directly into robotics. Self-driving companies have spent years building infrastructure for collecting data, labeling it, running simulations and evaluating models. That infrastructure is largely indifferent to whether the machine has wheels or legs.

The strategic split: brain first or co-design

Here the sector divides in two.

Wayve CEO Alex Kendall argues for starting in vehicles: "I think you need to start in vehicles. Manipulation robotics is like self-driving five years ago." In his view the data infrastructure, simulation and ML ops infrastructure will probably be shared, but the specific world model for the simulator will require different post-training. Different embodiments will need some differences, but there will be far more commonality than not.

Kendall also thinks it is too early to commit to any one hardware platform: advances in sensors and other components are coming quickly, and a truly general model should be hardware-agnostic.

Théophile Gervet, CEO of Genesis AI, a vertically integrated humanoid robotics company that raised a $105 million seed round this year, disagrees: "We're too early in this wave for a brain strategy to work; our take is there's lots of opportunities to co-design hardware and AI."

Narrow or general: the real trap

Another hot topic is how to focus the business. The picture is now sharply split: robotics companies targeting specific tasks are getting their robots into the field — one is building solar farms, one is deploying robots in industrial settings, one is operating excavators autonomously. Meanwhile general-purpose humanoids are not getting out of the labs.

Gervet frames the dilemma sharply: "No customer cares about the general-purpose robot that works at 80% success rate. We see a lot of other players go general, but there is no value provided because there's no vertical focus. But then, if you're building for a narrow vertical on top of GPT-2, you're going to get crushed by the company building on GPT-4."

That sentence captures the sector's real squeeze. Going narrow brings revenue and real-world deployment data. But task-specific data may not have enough diversity to push general-purpose models forward. A model trained on excavator data operates an excavator well; it does not learn to empty a dishwasher.

The invisible layer: managing the data

Beneath this debate sits a layer that gets little discussion but decides a great deal: managing the data that gets collected.

Robot data poses a different problem from text. Visual and lidar data is dense; a single day of one robot can occupy far more storage than millions of words in a language model's training set. Finding the useful moment inside that pile — the segment where the model failed, behaved oddly or met a new situation — is a job in itself.

A product announced at the conference targets exactly that problem: built on top of Nvidia's Cosmos open-weight world model, it lets engineers search that data with sophisticated natural language queries. The goal is faster triage and debugging, and quicker construction of evaluations and simulations, so model builders can iterate faster.

The detail matters because it shows where the field's bottleneck sits. The problem is not only that there is not enough data; extracting what is instructive from the data you already have is a separate constraint.

Starting narrow and widening

The pull of vertical focus is not only revenue; it also brings real-world deployment data. Kevin Peterson, CTO of the company operating excavators autonomously, describes the approach: starting with excavation was a way to understand the challenges of "manipulation in the wild." But the plan does not stop there — the company intends to develop an intelligence layer that stretches across a series of construction machines.

That may be the way out of the dilemma Gervet describes: start in a narrow vertical, accumulate the data, then carry that accumulation into neighboring tasks. Whether it works is still unknown — there is no result yet showing that a model moving from excavator to crane holds a real advantage over one trained from scratch.

When will it count as having happened

Is there a "ChatGPT moment" the sector is waiting for? Three different answers emerge, and all three are instructive.

Kendall points out that the largest robot deployment in the world is still consumer vacuum bots. For him a ChatGPT moment would be something that excites consumers, not investors: "One example of that would be when you get eyes-off autonomy for less than $1,000 worth of hardware in a car."

For Gervet, the moment is manipulation working out of the box: "You can talk to a robot in natural language and have it do any basic task for manipulation, like say pushing, pulling, closing a laptop, cleaning up a table, whatever you want to do, and it works to some level of reliability, let's say 80% plus out of the box — that's roughly your ChatGPT experience."

Foxglove CEO Adrian Macneil rejects the question outright: "There will not be a ChatGPT moment for robotics. The thing that made ChatGPT a moment in time was the distribution — they went from zero to like a million active users in like a week. Distribution in the real world is way harder than that. I would be very excited for the Apple II moment in robotics or the IBM PC moment."

Why it matters

Macneil's objection offers the best frame for reading this field. In software, a breakthrough spreads at download speed. In the physical world it needs a production line, a supply chain, a service network and a price. The Apple II did not arrive in a moment; it spread over years.

The distinction matters for investors too. In a software company the link between user count and revenue is nearly direct; in robotics, production capacity, unit cost and service obligations sit in between. Selling a robot will never happen at the scale of downloading an app.

Unitree's loss of value is therefore not surprising. The market had priced robots at software speed; what the company offers is a business that moves at hardware speed. The gap showed up as the distance between expectation and delivery.