Sapphire Ventures
Partnering with expansion-stage, enterprise software companies that we believe can become category leaders.
Menu close
Sapphire

Physical AI: The Countdown to Robotics’ ChatGPT Moment

Table of contents

Physical AI: The Countdown to Robotics’ ChatGPT Moment

Advances in AI and Hardware are Beginning to Collide with Real-World Demand. A Market Map of the Top Companies Making Robots Useful.

Robots have been “two years away” for about 20 years. Physical AI, AI systems that let machines perceive, reason and act in the physical world, is the shift now closing that gap.

Perhaps it’s wishful thinking, but the enthusiasm is easy to understand. Generations have been captivated by the sheer potential. And the prize is hard to overstate. Global labor automation is a $40+ trillion annual market (Morgan Stanley “The Rise of the Humanoid Economy); if machines could perform even a small fraction of physical work, physical AI would be one of the largest markets (and value drivers) ever created.

So it’s no surprise that each generation has expected robots to break out of factories and warehouses and proliferate across the rest of the economy. However, that kind of ubiquity hasn’t landed. At least, not yet…

“We’re talking about the biggest market that any of us are going to see in our lifetime. It’s not a choice for these big tech companies of whether or not they’re going to play in it.”
Jeff Cardenas

We’re cautiously optimistic that this time can be different. ChatGPT changed what people believed machines could do. We all remember pre-ChatGPT, when “AI” was just a ranking & recommendation engine, a fraud detection system or a simple voice assistant.

Now, AI can reason, plan, respond and act all on its own, and over a billion people across the world have seen it happen. The question is no longer whether machines can be intelligent, but what else they can do. Can the same breakthroughs we experience every day with language models be pointed at the physical world? How long until that intelligence is given a body?  

Robots are Growing Up, Fast

We believe the answer is: sooner than people think. A few shifts have converged over the past few years that move physical AI from a distant promise to a tangible inflection point, and understanding them is key to understanding where the market goes next.

Intelligence is the first area where physical AI has seen real progress. For decades, robots ran on hard-coded rules and narrow policies. While effective for certain use cases, these were brittle systems that broke the moment the world looked slightly different from their training. Show them an object from an unfamiliar angle, give them an instruction outside their scripted range or shift the production line three inches from its original position, and they simply failed.

This is changing quickly. Models can now reason across novel objects and interpret natural language commands. Breakthrough architectures like vision-language-action (VLA) models take in what a robot sees, interpret what it’s told to do and output real-world actions. These models replace the traditional perception-planning-control pipeline with one model that’s full stack. And because they’re built on top of models pretrained on internet-scale data, they inherit a base of world knowledge that no hand-engineered system could match. In this way, a robot that has never seen a cup before can still pick it up because the robot understands what a cup is and what it can do. Researchers are still debating specific approaches to training and form factor, but the results are starting to speak for themselves. Generalist released GEN-1.5 recently and the model can learn a new task from a single demonstration of just 3 to 12 seconds, with no training or fine-tuning at all. One-shot success rates are 59%, but rise to 83% with a few minutes of extra data. This is the first model to show that one-shot learning of physical skills can emerge at scale. A robot brain with intelligence is starting to emerge.

Hardware has advanced in parallel. Actuators, sensors, lidar and onboard compute that once cost hundreds of thousands of dollars have dropped by an order of magnitude, making it viable to build capable robots at price points that actually pencil out. Robotics-specific components have matured alongside the cost curve, approaching the performance needed for real-world manipulation, though onboard power budgets still cap how large a model a robot can actually run in real time. Robots also inherit innovation from adjacent industries as the supply chains built for EVs and drones, spanning motors, batteries and cameras, adapt directly to physical AI.

Additionally, simulation environments are now good enough that policies trained in simulation can transfer to the real world with far less of a gap and far less fine-tuning than before. Large-scale teleoperation setups now let humans generate high-quality demonstration data at a pace that wasn’t possible a decade ago.

Demand is the third shift. Reshoring, aging workforces and persistent labor shortages in manufacturing, warehousing and construction have turned physical AI from a nice-to-have into an operating necessity. Deloitte predicts 1.9M US manufacturing jobs could go unfilled through 2033 if talent challenges persist. The US government projects shortages of ~350K+ nurses by 2038 and 92% of construction firms currently report difficulty filling open positions.

Expand AI physical market map_projected global robot revenue_
Source: Morgan Stanley (Data as of July 15, 2026)

This demand is translating into behavior. In two years, executives watched language models go from curiosity to a daily tool. They no longer dismiss automation as science fiction. The proof points are everywhere. Amazon runs more than a million robots across its fulfillment network, construction faces a shortfall of hundreds of thousands of workers every year and defense budgets are pouring into autonomous systems.

Across both commercial and government end markets, buyers are graduating from pilots to production deployments, and that shift from experimentation to line-item budgets is what makes this cycle different from previous ones.

Capital, of course, follows innovation. Investors have seen the same shifts converge, and venture dollars are flowing in at a pace the sector has never seen. Funding into physical AI grew from $12.7B to $21.8B from 2024 to 2025 (+73%), with recent mega rounds flowing into Skild ($1.4B Series C invested at a $14B valuation), Figure ($1.0B Series C invested at $39B valuation), Physical Intelligence ($1.0B Series C at $11.7B valuation) and Generalist ($400M Series B invested at $2B valuation).

“Humanoid robots will bring physical AI to the world’s largest industries, opening a multitrillion-dollar economic opportunity.”
Jensen Huang

Venture funding is critical for physical AI. Hardware development cycles are notoriously long and capital-intensive, and underfunding challenged many of the last generation’s most promising companies. For the first time, the war chests exist to survive the journey from prototype to production. The potential is hard to overstate, given the physical AI industry is projected to grow from ~$138B in 2026 to $500B+ by 2030, with longer-term forecasts for humanoids alone stretching toward the trillions by 2050 (Morgan Stanley).

Expand physical AI market map_VC funding
Source: PitchBook (Data as of June 30, 2026)

The Road Ahead: What Still Has to Get Sorted

Nonetheless, physical AI has humbled optimists many times before. And for all the momentum, there will be road bumps with real obstacles and open debates still remaining. How the industry navigates the dynamics below will separate this cycle from the prior false starts.

Training & Data

GPT-class models have worked because they have a large, internet-scale training corpus. Trillions of tokens of human-generated content and data, mostly free to harvest. In physical AI, the working assumption has been that no such corpus exists. The raw video is arguably out there, though, and the bigger constraint may be that so little of it is action-labeled. If that’s the real bottleneck, then annotation becomes an advantage for whoever solves it, and most of the approaches in the field today are attempts at exactly that. Egocentric human video is one of the more promising, since the action is captured alongside the footage, and early evidence suggests it can serve as a scalable training signal and open a new power law for the field. None of the methods have proven sufficient on their own yet, which is why the industry continues to rely on a mix of them.

Form Factor

The industry is still debating the fundamentals, down to something as basic as the hand. On one hand (no pun intended), builders are pushing for five-fingered dexterity. In their view, the physical world is built around the human hand, so a hand with the same morphology has the highest ceiling of any design and would be the most generalizable. A five-finger hand also narrows the embodiment gap, since training data (video and teleops) is the most transferable and can be most easily adapted.

On the other hand is the case for simpler form factors like grippers, parallel jaws and suction cups. These form factors are cheaper and less complex, requiring materially fewer actuators, tendons and sensors, but offer fewer degrees of freedom (DOF) vs. a five-finger hand.

The same debate even plays out a level higher at the body. Humanoids get the headlines and lots of funding, but wheeled bases, mobile manipulators and fixed-arm cells reflect much of the deployments today. Bipedal locomotion is often overkill as it’s more expensive, energy-hungry and simply not needed for the majority of tasks.

Ultimately, the answer likely depends on the end market and ultimate use case: Warehouses may standardize on wheeled manipulators, homes and construction sites may demand legs and factories may keep running fixed arms. The bet a founder or investor is really making is less “what will a robot look like” than “which form factor wins which market,” and the honest base case is likely that no single design wins it all.

Business Model

The economic model behind robots is still getting sorted. The sharpest divide is between Robotics-as-a-Service (RaaS) and asset sales. RaaS turns a large capital expense into a predictable monthly line item, which makes it easier for customers to say yes and for vendors to land initial deployments. But it forces the vendor to carry the balance sheet, absorb depreciation and take on the ongoing cost of maintenance and uptime. That works if the robot is reliable and the fleet scales, but it can be challenging for young companies if utilization drops or units break often. RobCo is a well-known example of the RaaS approach. Outright sales flip the equation. The manufacturer gets upfront revenue, but the customer inherits the service and support burden, and often needs to build internal expertise to keep the robot running. Standard Bots is a well-known example of the asset sale approach.

We think that the ‘right’ answer will depend on who is buying. Large enterprises with sophisticated ops teams may prefer to own the asset outright. Smaller operators and newer end markets tend to prefer RaaS, because they want the outcome without the operational headache. A growing middle ground sells the hardware but bills software, updates and remote support as a recurring subscription. That structure gives vendors both upfront cash and the recurring revenue public markets reward, but it only works if the software layer is genuinely valuable on its own. The framing itself may also shift with scale, and once F2000 companies are replacing labor with robots, the decision may look less like CapEx versus OpEx and more like the cloud versus on-prem question enterprises worked through a decade ago.

Safety and Security

Safety is a gating issue for physical AI companies on three fronts. The first is regulatory, where regional compliance requirements set the floor for market access, and liability in shared human-robot workspaces is still a legal gray area that slows deployment. Second, there is cultural perception; people need to feel comfortable working next to a few-hundred-pound autonomous machine, and one bad incident on camera captures the headlines and can set the category back years. The third is security. A robot is essentially a networked computer with actuators, so its attack surface only widens as fleets scale. Physical AI companies are navigating these safety and security challenges in real time as the industry rapidly evolves.

China: A Competitive Force to Watch

And finally, China is a tricky dynamic overhanging the entire category. On one side, a lot of the momentum in physical AI is coming from China. Beijing declared humanoids a “new frontier in technological competition” and set national mass-production targets, backed by deep manufacturing capacity and fast iteration cycles that let companies ship novel designs at prices Western manufacturers struggle to match. The culture is also aggressively experimental.In April 2026, Honor’s ‘Shandian’ (Lightning) robot won Beijing’s E-Town Humanoid Half-Marathon in 56:51, faster than the 57:20 human men’s world record, with Honor sweeping the top three spots (DOGO News). Four months later at the 2nd World Humanoid Robot Games, X-Humanoid’s Tiangong Ultra ran a 100m in 8.86 seconds against Usain Bolt’s 9.58 record and cleared 2.88 meters in the standing high jump (Global Times). That event drew 666 teams and 2,056 robots from 16 countries. Public spectacles like this compress iteration cycles and drive innovation in ways lab demos don’t. Public spectacles like this compress iteration cycles and drive innovation in ways lab demos don’t.

Supply chain and dependency are another key challenge to monitor. China dominates both robot production and the upstream supply chain for critical components like rare earths, batteries and actuators. Every US physical AI founder we’ve talked to has flagged this as a dynamic to watch, and only a handful are building alternative supply chains from scratch.

Chinese Robots Racing Humans in the Half Marathon
Chinese Robots Racing Humans in the Half Marathon

Our market map groups the ecosystem into six layers:

  1. Hardware: the underlying components that make robots possible
  2. Embodiment and Form Factor: the physical robot itself
  3. Applications / Verticals: the industries putting them to work
  4. Development: the training layer that teaches them
  5. Infrastructure and Enablement: the runtime layer that keeps them running
  6. Intelligence and Models: the models giving them intelligence

We walk through each, covering what’s driving activity today and sharing our perspective on every subsegment.

Hardware: The Building Blocks

Every robot on the market map rests on a foundation of physical components. Sensors, actuators, motors, gearboxes, cameras, lidar, edge compute and batteries are the ingredients that determine what a robot can physically do, how much it costs and how reliably it can run in the field. For decades, the hardware layer was the primary bottleneck in physical AI. Building a capable robot meant assembling expensive, specialty parts sourced from a fragmented supplier base, and the resulting economics only worked for a narrow set of use cases.

That’s changed materially over the past few years. Actuator, sensor and compute costs have all come down by an order of magnitude, driven by scale in adjacent industries like electric vehicles, smartphones and consumer electronics. Off-the-shelf components that used to cost hundreds of thousands of dollars are now available at price points that make broader classes of robots economically viable. The hardware layer has quietly made the rest of the map possible. Without the cost and capability improvements at the component level, none of the intelligence, development or application activity above it would be commercially feasible.

Embodiment and Form Factor: What Robots Look Like

If hardware is what a robot is made of, embodiment is what those components add up to. The physical shell determines what a robot can pick up, where it can go, how it interacts with the world and what it costs. It’s the layer where design choices become fate, and the ecosystem is still evolving.

Most of the attention today goes to humanoids, but purpose-built non-humanoids quietly generate a lot of the revenue in physical AI today. A robot that picks items in a warehouse doesn’t need legs, and a robot that inspects pipelines doesn’t need arms. These designs are simpler, cheaper and more reliable because they solve a narrower problem set, with real deployments and measurable ROI already on the books. Buyers in warehousing, logistics, agriculture and industrial inspection are spending real money on these systems today because they work, they’re financeable and the payback periods are short enough to justify. The tradeoff is scope. A robot built for one task rarely does another without re-engineering, but for buyers who need automation today rather than a platform tomorrow, that’s a reasonable trade.

Humanoids are the bigger long-term bet. The world is built for humans, so a robot with roughly human proportions can operate on factory floors, warehouses and homes without redesigning any of them, and that same platform, if it can learn to do many jobs, represents a huge market opportunity. That’s the bet companies like Figure, Tesla and Unitree are each making with different flavors, from commercial-first industrial deployments to vertical integration plays to low-cost approaches. Balance, dexterity and safety around humans remain open challenges, unit costs are high and reliability in unstructured environments is still years away. What makes the category worth the difficulty is the vision of generalization.

Generalization matters differently depending on the environment. Homes look more like the chatbot world, where broad training data and a flexible form factor are what unlock the use case, while factories reward nines of reliability, repeatability and throughput, pointing toward heterogeneous systems of specialized machines with a layer of broader intelligence on top. Consumer robots sit on a different curve altogether. Adoption has been slower because the price-to-value equation is harder in the home and consumers have less tolerance for robots that break, but better models and cheaper components are starting to open up new use cases in home cleaning, companion robots and personal assistants. Sunday is one of the companies pushing into this space, developing a wheeled home robot trained on real household demonstrations from wearable gloves.

 

“In the long term, we imagine everyone having a personal robot doing anything they need.”
Sam Altman

While humanoids, purpose-built systems and consumer robots are still finding their footing, two form factors have already reached commercial maturity. Aerial systems and autonomous vehicles are each big enough to warrant their own market map. Drones are already deployed across defense, inspection, delivery and mapping, while robotaxis and autonomous trucks are operating commercially in cities and freight corridors. Both categories have influenced the broader stack since many of the underlying technologies (perception, planning, sim-to-real, fleet management) are directly transferable to other form factors. Key players span defense (Auterion, Firestorm), enterprise inspection (Skydio), commercial delivery and autonomous driving (Waymo).

Applications and Verticals: Where Robots Go to Work

Every robot has to earn its keep somewhere, and this is the layer of the map where that happens. It’s the biggest layer by volume, and the one where the industry’s claims get tested against real budgets.

The initial proving ground has always been indoors. Factories and warehouses offer structured environments, repeatable tasks and clear unit economics, which is why they’ve hosted robots for decades. What’s new is the kind of work getting automated. Legacy systems handled the narrow, rigid jobs while anything requiring dexterity or judgment stayed human, and that boundary is finally moving. AMCA is using AI-driven design tools to compress aerospace qualification timelines from years to months, RobCo builds the whole stack (software, hardware and AI), selling into factories from mid-sized manufacturers to enterprises, with 2.5M+ operating hours already under its belt. Theker is also building general-purpose industrial robots that reconfigure across tasks without reprogramming. One floor over in the warehouse, Dexterity is teaching robots to load and unload trucks where every box is different, and Standard Bots is bringing AI-enabled arms to buyers who need one system that flexes across jobs.

Step outside the building and the problem gets harder, though the two biggest outdoor industries have taken very different paths. Construction is unstructured and different every day, which helps explain why one of the world’s largest industries is also one of its least automated. The natural entry points are the workflows that look most like factory work, repetitive and physically demanding with well-understood equipment. Bedrock Robotics bolts autonomy onto existing excavators and dozers so contractors can keep the fleets they already own, while Built Robotics went purpose-built with autonomous pile drivers for utility-scale solar. Agriculture, an even messier outdoor environment, actually got there first. GPS-guided tractors and auto-steer combines have been standard equipment for over a decade, and modern row-crop operations largely run themselves. The frontier has moved to the harder problems traditional automation couldn’t touch, like Burro‘s people-scale robots that carry, tow and follow workers through specialty crop fields, or Carbon Robotics‘ AI-driven laser weeders.

Follow the food off the farm and the environment gets tricky in different ways. Restaurants and food producers run on thin margins, change menus constantly and handle ingredients that behave differently every single day. AI can finally cope with that variability, and with labor getting harder to hire every year, operators are running out of reasons to wait. Chef Robotics is portioning and assembling meals in high-mix food production, adapting to new menus without reprogramming.

The most delicate handling problem of all is the human body, and it’s also one of physical AI’s longest-running commercial successes. Intuitive Surgical has spent decades building its da Vinci franchise around moats of regulatory approval, reimbursement and surgeon training. AI is also now moving into camera control, workflow optimization and real-time guidance, expanding surgical robots beyond just a handful of procedures.

Defense moves at a different speed entirely. It has become one of the most active corners of the market, driven by a shift from small numbers of exquisite platforms to large numbers of cheaper autonomous systems that can be produced, coordinated and replaced at scale. Anduril is building the end-to-end version of that vision, spanning air, sea, space and command-and-control through its Lattice software platform, while Saronic is going deep on maritime, building autonomous vessels along with the shipbuilding capacity to produce them in volume.

And the newest vertical is a product of the AI boom itself. Data center capacity is being built out at a breakneck pace, and every rack has to be physically assembled, cabled, tested and maintained. Doing that by hand doesn’t scale, so hyperscalers are pulling robots into the job. Gradient Robotics is building general-purpose autonomous robots for demanding industrial environments, and Watney Robotics is partnering with hyperscalers on physical tasks inside live facilities. Robots are now building the infrastructure that trains the next generation of robots.

Development: How Robots Learn

A robot’s body determines what it can do, but its intelligence determines what it can figure out. Development is the layer of the market that produces that intelligence, spanning the data that teaches robots how the world works, the simulated environments where they can safely practice and the world models that let them predict and plan. For most of physical AI history, this layer barely existed as a category, with behaviors hand-coded, testing done on real hardware and every new task starting over. That approach doesn’t work in a world where robots need to generalize across environments, adapt to novel situations and be trained at a pace that keeps up with hardware progress.

The reason this layer has become so critical is that language models had an advantage physical AI doesn’t. The internet was already sitting there, waiting to be trained on. Physical AI has plenty of raw video, but very little of it carries the action labels a robot policy needs, so the development layer has to build the data pipeline before it can build the models. That work is expensive, but it pays off across the entire fleet, because every improvement at this layer multiplies across every robot in the field. A model that gets 10% better makes the whole fleet 10% better overnight, and that kind of leverage doesn’t exist at the hardware layer where every improvement has to be manufactured, shipped and physically installed. Whoever solves the data problem sets the ceiling for what every robot in the field can eventually do.

The most direct way to solve it is to go collect the data. Traditional labeling players like Scale came up serving language and vision models and have since expanded into robot-specific tooling, but the pure-play physical AI data vendors are a different segment altogether. They aren’t labeling static images. They’re building end-to-end pipelines that capture continuous multimodal data streams synchronized down to the millisecond, often through custom hardware. XDOF is building a three-tier data pyramid spanning teleoperation on the exact robot being deployed, more general teleoperation via low-cost interfaces and egocentric human data captured through wearable sensors, while Mecka is taking a different bet, collecting human motion data through body sensors and iPhones on the thesis that human-sourced data will generalize better across robot morphologies.

Collecting data in the real world is slow and expensive though, which is why simulation has become the other half of the equation. Running billions of hours of scenarios in a virtual environment is orders of magnitude cheaper than running the same tests on physical robots, and modern sim environments have gotten good enough that policies trained in simulation now transfer to real-world hardware with meaningfully less fine-tuning than in past cycles. That sim-to-real gap closing is one of the more important shifts in physical AI over the past few years. Applied Intuition is the incumbent leader with a simulation stack originally built for autonomous vehicles that now extends into trucking, defense, mining, construction and general robots. Newer entrants like Zeromatter and ReSim are building cloud-based sensor simulation and virtual testing platforms that let embodied AI teams run thousands of simulations in parallel and catch regressions before code hits real hardware.

A very ambitious bet in the layer is that neither collection nor simulation is enough on its own, and what robots really need is a world model that can predict what happens next. A world model takes in a robot’s sensor input and generates an internal representation of the environment, letting the system reason about how objects will move, how physics will play out and what the consequences of a given action will be. That predictive capability is what lets an agent plan multiple steps ahead rather than react frame by frame, and it also closes the sim-to-real gap from the other direction since a good world model can generate synthetic training environments that behave the way reality does. The category has been one of the most active segments of AI over the past year, with world model labs raising some of the largest rounds in the ecosystem. Advanced Machine Intelligence is Yann LeCun’s lab developing action-conditioned world models that learn abstract representations of sensor data. Decart is building real-time world models that generate photorealistic environments frame by frame. General Intuition is training world models on billions of hours of gameplay data to build agentic systems that can transfer intuition about movement, physics and spatial reasoning across environments.

Infrastructure and Enablement: How Robots Run

Development is what teaches robots how to act. Infrastructure and enablement is what keeps them acting as expected once they’re in the field. It’s the layer that connects a fleet of robots to the systems running them, the data pipelines watching them and the networks tying them together, and the maturity of the tooling here is a major determinant of whether a physical AI company can actually operate at scale.

For most of physical AI history, this layer was built in-house, with every deployment coming with its own bespoke fleet management dashboards, custom telemetry pipelines and hand-rolled communication protocols. That worked with a few robots when use cases were narrow, but it doesn’t scale to a world where fleets are growing into the thousands and running mission-critical workloads. A commercial ecosystem is now forming around the operational stack, creating the equivalent of what DevOps, observability and cloud infrastructure did for software.

Most of that ecosystem is being built around the data robots produce. A single autonomous vehicle can generate terabytes of sensor data per day, and a humanoid running a full multimodal stack (video, point clouds, force feedback, teleoperation trajectories) is in the same territory. Making sense of that data is the job of observability tools like Nominal, which focuses on test data analysis and validation workflows for hardware engineering teams, and Revel, which is replacing decades-old industrial control tools with a Python-inspired programming language and real-time telemetry system. Moving the data reliably is a separate problem, especially when robots operate in environments with unreliable connectivity like warehouses, construction sites and offshore vessels. LiveKit is building the real-time communication layer with a WebRTC-based platform that powers video streaming and teleoperation.

Running the models themselves is its own layer. Robots need fast, reliable inference to act in real time, and even small delays between perception and action can be the difference between a smooth pick and a dropped part. Reactor is building real-time inference infrastructure purpose-built for world models, and is already running them in production for robotics customers. These models have different compute and memory profiles than the language models most inference providers were designed for. That specialization matters as the field shifts toward world-action models, where inference has to keep pace with continuous physical control rather than the batched request patterns most serving stacks assume. 

Once a deployment moves beyond a handful of machines, the challenge shifts from any single robot to managing the fleet as a whole. Teams need a way to onboard new robots, push software updates, monitor status across sites, coordinate multi-robot workflows and step in when something goes wrong. Formant is the leader here, providing cloud infrastructure for fleet management, teleoperation, incident management and AI-driven operations analytics across heterogeneous deployments. Sitting alongside all of this is perception and navigation, the software layer that turns raw sensor data into a usable understanding of the environment. Historically dominated by specialized vision systems tightly coupled to specific hardware, the space is shifting toward AI-driven models that generalize across environments, robots and tasks.

Intelligence and Models: The Robot Brain

Every layer of the market map eventually points back to intelligence and models. The intelligence layer is where robots get the ability to reason, plan and act, and it’s the layer that has changed most dramatically over the past two years. What was a research curiosity in 2023 has moved into production infrastructure in 2026. Several labs are now training models on trillions of tokens of internet-scale video, robot trajectories, simulation output and human demonstration data.

Capability is improving fast. When Generalist introduced GEN-1 in April 2026, average task success rates jumped from 64% to 99% on the tasks the field had been benchmarking, completed those tasks roughly 3x faster than the previous state of the art and required only one hour of robot data to hit those results. Physical Intelligence’s π0.7 release the same month showed meaningful cross-embodiment generalization, letting a single policy fold laundry, pack boxes and operate coffee machines across different robot bodies. Google DeepMind’s Gemini Robotics is now running on hardware from Agility, Apptronik and Boston Dynamics. NVIDIA’s open GR00T family has become a common starting point for developers, providing pre-trained foundation weights that any team can fine-tune on their specific hardware and use case.

The dominant architecture behind these results today is the vision-language-action model (VLA). A VLA takes in what a robot sees, interprets what it’s told to do and outputs a real-world action, all in a single end-to-end trainable stack. VLAs matter because they let a single model transfer across tasks, embodiments and environments in ways that the hand-coded policies of previous eras never could. Not every lab is building a pure VLA, and the field hasn’t converged on which specific approach wins. But most of the visible progress in the past year has come from teams pushing on some version of this architecture. The frontier is now moving toward fusing a world-modeling objective into the action stack, producing a hybrid world-action model. The idea is that the next wave of systems will learn how the world works, not just how to move in it.

Another question shaping the space is whether companies should stay in research mode or push robots into the field. The research-mode camp, exemplified by Physical Intelligence, argues that building a truly capable foundation model requires long, uninterrupted scaling work, and that commercialization pressure forces teams to compromise on the model to chase near-term revenue. The deployment-first camp, exemplified by Skild AI, argues that real-world deployment is the only reliable source of the training data foundation models need, and that scale can’t be achieved from teleoperation and simulation alone. Skild has already deployed its models on the robot assembly lines that build NVIDIA’s Blackwell GPU server systems at Foxconn. Physical Intelligence has taken the opposite approach, telling investors it has no fixed commercialization timeline. Both approaches have merits. Deployment produces data and revenue but exposes teams to real-world edge cases early. Research mode preserves optionality but risks losing the data flywheel.

Funding at this layer has kept pace with the ambition. Earlier this year, Skild AI raised $1.4B at a $14B valuation to build what it calls an omni-bodied robot brain. Physical Intelligence is reportedly raising at an $11B valuation, on top of the $600M it raised in November 2025. Generalist recently raised $400M at a $2B valuation, bringing its total funding above $500M. Startups like RLWRLD (a former Google DeepMind team), Field AI, Flexion and DYNA have raised at valuations that most enterprise software companies would take years to reach, on the thesis that whoever cracks the robotics foundation model becomes the next OpenAI or Anthropic of the physical world.

Several companies are shaping the space.

  • Physical Intelligence is developing the π family of general-purpose VLA models designed to control any robot on any task, taking a research-first approach with no fixed commercialization timeline.
  • Skild AI is building the Skild Brain, an omni-bodied foundation model designed to work across quadrupeds, humanoids, tabletop arms and mobile manipulators.
  • Generalist is training embodied foundation models on large-scale real-world human data, and its GEN-1 model is one of the most cited capability step-changes in the field.
  • Elorian is building foundation models for visual thinking: grounded, spatially aware models that natively understand depth, structure and relative position, rather than inferring them from text.
  • Genesis AI is advancing dexterous hand manipulation capability by learning from human data.
  • Field AI is building risk-aware foundation models designed for robots operating in unstructured outdoor and industrial environments.
  • DYNA recently released DYNA-2, a world-action model pretrained on one million hours of egocentric human video, and is an early proof point that human video can serve as a new scaling axis for robot foundation models.
  • NVIDIA sits at the center of the ecosystem, providing open foundation models (GR00T, Cosmos), simulation infrastructure (Isaac) and onboard compute (Jetson) that most physical AI companies build on top of.
  • RLWRLD, Flexion and Covariant are among the other labs pushing on different architectural approaches to physical AI.

These are just a handful of the companies working on the problem, and the space is expanding as more frontier labs commit to building foundation models for physical AI.

Physical AI Is Nearing Its Inflection Point

So, will physical AI have its ChatGPT moment soon? We think so.

For the first time, the ingredients coexist. Foundation models are meaningfully more capable than they were even a year ago, unit economics are improving across the stack and the demand-side pressure is increasing. The past two years have moved physical AI from a research bet to a category where multiple companies are shipping fleets, generating revenue and building the infrastructure that will define the next decade of physical AI.

“I think Optimus will be our biggest product, not just Tesla’s biggest product ever, but probably the biggest product ever.”
Elon Musk

Unlike ChatGPT though, the moment won’t announce itself with a single product launch. It’ll look more like a cluster of signals landing at once. A general-purpose foundation model that a developer can point at their robot and have it do useful work out of the box. Humanoid unit costs falling into a range where broad commercial deployments actually pencil out. A vertical crossing the tipping point from experimental pilots into core operational infrastructure. And a data flywheel that keeps compounding, whether through teleoperation, simulation, human-sourced data, or some combination of all three.

We’re still in the early innings, and much of what will define the winners hasn’t been decided yet. A few of the themes and open questions we’re tracking most closely:

  • The data flywheel. Whoever solves data collection at scale will have a structural advantage that hardware and compute alone can’t replicate. The debate between teleoperation, simulation and human-sourced data isn’t settled yet.
  • Generalist versus specialist models. The industry is watching whether a single foundation model can control any robot on any task, or whether specialist models fine-tuned to specific embodiments and workflows will continue to outperform.
  • How capability gets measured. There’s no standard benchmark in physical AI, and results are self-reported on task sets each lab picks itself. Whether the field converges on a shared eval, and who defines it, will determine how easily buyers can tell real progress.
  • Humanoid unit economics. Whether the platform can hit the cost and reliability thresholds that make broad commercial adoption possible, and how quickly.
  • The transition from pilots to production. Which verticals cross the tipping point from experimental deployments to core operational infrastructure, and how quickly the ecosystem around orchestration, observability and fleet management scales to support that transition.

There’s much more to be excited about in physical AI than any single blog can capture, and the ecosystem is expanding faster than the market map can keep up with. If you’re a founder building in physical AI, a customer deploying robots at scale, or a fellow investor thinking through the space, we’d love to hear from you. Please reach out to us at [email protected] and [email protected].

Also special thanks to Jai Das, Anders Ranum, Lindon Gao, Carla Gómez, Roman Hölzl, Gesa Biermann, Alberto Taiuti, Zhou Xian, Michelle Ling, Andrew Dai and Arash Tajik for their thoughtful insights as we wrote the blog!

Key Takeaways

  • Three shifts are converging to make physical AI commercially viable for the first time: smarter models (vision-language-action architectures), cheaper hardware (component costs down an order of magnitude), and real demand (labor shortages across manufacturing, construction and healthcare).
  • Investors are betting this convergence is real: funding into physical AI grew from $12.7B to $21.8B in a single year (+73%), and the industry is projected to grow from roughly $138B in 2026 to $500B+ by 2030.
  • The resulting market spans six layers: hardware, embodiment and form factor, applications and verticals, development, infrastructure and enablement, and intelligence and models, each covered in our Physical AI market map.
  • Five open questions will determine which companies win as this market matures: the data flywheel, generalist vs. specialist models, research mode vs. deployment mode, humanoid unit economics, and the transition from pilots to production.
Legal disclaimer

This article is for informational purposes only. Nothing presented within this article is intended to constitute investment advice, and under no circumstances should any information provided herein be used or considered as an offer to sell or a solicitation of an offer to buy an interest in any investment fund managed by Sapphire. Information provided reflects Sapphires’ views as of a time, whereby such views are subject to change at any point and Sapphire shall not be obligated to provide notice of any change. Companies mentioned in this article are a representative sample of portfolio companies in which Sapphire has invested in which the author believes such companies fit the objective criteria stated in commentary, which do not reflect all investments made by Sapphire. A complete alphabetical list of investments made by Sapphire’s Growth strategy is available here. No assumptions should be made that investments listed above were or will be profitable. Due to various risks and uncertainties, actual events, results or the actual experience may differ materially from those reflected or contemplated in these statements. Nothing contained in this article may be relied upon as a guarantee or assurance as to the future success of any particular company. Past performance is not indicative of future results.

 

Sources:

Morgan Stanley “The Rise of the Humanoid Economy” (Thoughts on the Market podcast with Adam Jonas and Sheng Zhong) — this is the primary analytical source with the actual derivation: https://www.morganstanley.com/insights/podcasts/thoughts-on-the-market/humanoid-robot-market-rising-adam-jonas-sheng-zhong

Evercore’s “A Primer on Automation Tech & Physical AI” (July 2025)

Morgan Stanley’s “The Humanoid 100: Mapping the Humanoid Robot Value Chain“ (February 2025)

https://themanufacturinginstitute.org/manufacturers-need-as-many-as-3-8-million-new-employees-by-2033/

https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-2025-letter-to-shareholders

https://www.dogonews.com/2026/5/8/robot-outruns-humans-in-beijing-half-marathon

https://www.abc.org/News-Media/News-Releases/abc-construction-industry-must-attract-349000-workers-in-2026-despite-macroeconomic-headwinds

Physical Intelligence is reportedly in talks to raise $1B, again (TechCrunch, March 2026)