Physical Intelligence Transcript: Building Universal Brains for Robotics Through General Foundation Models and Reinforcement Learning
September 2, 2026 | In-Depth Discussion with Sergey Levine on Embodied AI
Introduction and Defining Physical Intelligence
Host: My guest today is Sergey Levine, one of the co-founders and researchers at Physical Intelligence. As a disclaimer, I am an investor in Physical Intelligence because I believe it is one of the most important companies tackling the problem of robotics. As you hear us discuss today, robotics has what I would call a scarecrow problem. All of these amazing physical devices are becoming ever more possible in all sorts of cool permutations, but what they all really need is an intelligence, a brain, and that is what they are developing at Physical Intelligence. They are trying to develop foundation models that can make any physical robot do any task in any environment. That challenge is daunting and has required many of the world's best researchers, Sergey being one of those leaders coming together to try to solve this problem. The nature of our conversation today covers all of the problems facing robotics and all of the promise of solving these problems across the world. I hope you enjoy this great conversation with Sergey Levine.
Host: This is going to be a real treat and a blast to learn about possibly the most exciting, impactful area of technology being developed. Just to set the stage before we go back in time, maybe you could define Physical Intelligence as you see it.
Sergey Levine: Fundamentally, the goal of Physical Intelligence is to develop robotic foundation models that can control basically any embodied system to do any task. Broadly speaking, you could imagine that in the same way that a language model is rapidly evolving towards a system that can do any task that can be expressed in language, what we would like is to build a new class of models that can do any task that can be done by a physical actuated device. Part of the thesis of this company is that we believe doing it at the full level of generality might actually in the long run be easier than trying to special-case very specific narrow application domains. Again, in much the same way that for language models, it turned out to be easier in some ways to solve natural language tasks in their full generality than to narrowly target machine translation or sentiment analysis.
The Trade-offs of Generalization vs. Domain-Specific Systems
Host: It may not be obvious why you would make that bet versus a robot that just does your dishes or something. What are the key trade-offs to understand, and why make the decision that you made?
Sergey Levine: Maybe I can give you a two-part answer for this. First is the analogy to language models, and the second is what that means in the robotics world. The first one is informed by evidence. In the world of natural language, we saw that there were a lot of efforts to develop domain-specific solutions that tackled specific problems. For example, somebody would spend a lot of time thinking about how English differs from French and then build a machine translation system. The reason that language models actually took over for all of those different application domains is because they can leverage much broader sources of data.
Sergey Levine: It is not even as simple as saying we had this data for this application and this data for that application and merged everything. It is actually more than that. When you can leverage weakly labeled data, like data that you mine from the web in the case of language models, you actually learn more about the world. You establish a foundation of world understanding, and on top of that foundation, it turns out to be much more effective to build out different applications.
Sergey Levine: To bring this into robotics, the calculus does not look quite the same because in robotics we do not have an internet-sized dataset that we can just draw on. But this notion of understanding the world is, if anything, actually more important in robotics. If you have many different tasks and many different physical systems, then you can go from training individual dishwashing specialists or laundry-folding specialists to instead training a model that actually understands physical interaction. People can master new skills very rapidly because we understand physical interaction. We can intuitively grasp what is going to happen in a new unfamiliar situation, which lets us bootstrap things really quickly. If we can draw on data from many sources, many applications, and many robots, then we can have a model that has physical understanding, making it much easier to put new applications on top of that platform.
The Challenge of General Models and the Trap of Narrow Demos
Host: What is the hardest part about building in this way when you see other approaches that are perhaps more legible to the average person, like a robot moving around doing one specific thing?
Sergey Levine: This has been an issue throughout my career because when you work on robotic learning, effective generalization is not actually the optimal way to produce a really exciting demo. The way to have a really exciting demo is to pick a really cool task, control everything else in the environment, set it up so that it is perfectly clean and pristine, and just make it work in that one setting. That is how you make a classic robot demo.
Sergey Levine: Generalization cannot just be shown in one spot. The point of generalization is that it does something relatively mundane that any human could do, but it does it in any situation. We released some demos where we showed our robot cleaning kitchens. If you watch an individual video out of context, it just looks like it is picking up plates, which anybody can do. But the point was that we put the robot into that home just for that demo, and it had never seen training data from that setting. You have to understand what is going on behind the scenes to appreciate why that is pushing the frontier.
The Long-Term Stakes and Platform Shift
Host: What is your model for the stakes of what you are doing? If you are successful and cross that chasm of general physical intelligence, what does that enable?
Sergey Levine: One of the things that would be really exciting is the ability to unlock people's imagination in how they build robots and other embodied systems. Personal computers were a really big deal because they made it possible for lots of people to hack together all sorts of cool applications. There was this Cambrian explosion of amazing software that started in the 1990s and was further accelerated by the internet. I think something like that might happen in the world of robotics, but it cannot happen today because if you want to put together a cool new robotics application, you have to build a monstrous technical stack and basically solve the intelligence problem yourself.
Sergey Levine: If there is a foundation model that someone can build on top of, which you can prompt to provide basic functionality and then fine-tune or adjust to your application, it makes it far more tractable for lots of people, companies, and individuals to try out all sorts of different things. Sometimes we think robots are going to be just one thing, like metal people. But no technology has developed like that. It is going to be more like a toolkit where you can put together all sorts of applications. You can get creative, like building a robot with five arms that hangs from the ceiling. You need the right software foundation platform on top of which to do that, and the foundation model can be that platform.
Pros and Cons of Humanoid Robots vs. General Embodiments
Host: What are the pros and cons of the humanoid approach to robotics in your mind?
Sergey Levine: One pro is that it is really cool. You show it to somebody, and they immediately get it. There is a lot of value in capturing the imagination and getting people to think about what the future might look like in an understandable way. But in my mind, humanoids are one of many possible kinds of robots that we are likely to have. Fundamentally, the challenge of intelligence looks very similar for all these different robots. I do not think we should be tackling intelligence solely in the context of one specific body. We should handle it in a general way because we need lots of data across diverse systems.
Sergey Levine: The cool thing about building robots is that ultimately they do not have to be constrained to look like humans at all. You can build the right tool for the job. You could imagine building a house with a swarm of 10,000 quadcopters. In the future, we will have a robotic foundation model that can be adapted to all sorts of applications, running the gamut from bulldozers to humanoids to robotic arms. The fundamentals of how you interact with objects, how things move in the world, and how causality works are conserved across all of these different systems.
Host: Do you have a favorite example of what might be possible with true general intelligence that might not be possible with a humanoid-only intelligence?
Sergey Levine: There are a few things worth thinking about. One is that we can make machines that are very big and machines that are very small. In the long run, there are exciting applications in medicine and surgery where we might not be limited to robots that look like humans or even robots that can be controlled directly by humans. Currently, robotic surgery is done through teleoperation, requiring real-time human control with high dexterity. In the long run, autonomous systems could address tasks beyond direct human motor capabilities.
Historical Milestones in Robotics Research
Host: If you think about the most important milestones on the timeline of robotics research that have brought us to this point, how would you walk us through that history?
Sergey Levine: At some level, doing end-to-end control for robotic systems is an old idea. The first autonomous driving systems that used end-to-end learning existed in the 1980s. ALVINN, developed around 1986 or 1987, was a driving system demonstrated on highways controlled by a tiny neural network receiving camera input. While the concepts are venerable, historically the difficulty in robotic learning has been satisfying several requirements simultaneously: it must handle the application cost-effectively without requiring massive data collection for every single task, it must handle long-tail scenarios with common sense, and it must remain robust, fast, and reliable.
Sergey Levine: Machine learning works best when there is a lot of data. If you naively approach a problem like washing dishes by collecting an enormous amount of dishwashing data, it is not cost-effective because you have to repeat that entire process for the next task. Training general-purpose models reduces the data needed per task. What has changed most recently is the ability to handle unusual scenarios. In edge cases where you lack prior experience, you must rely on knowledge acquired from other sources and ground it in the new situation. Multimodal language models are adept at pulling in broad web knowledge. While they are not inherently grounded in physical situations, they provide a path to import common sense into robotic control.
Host: Are there milestone equivalents on the timeline comparable to AlexNet or the transformer in robotics?
Sergey Levine: It is early to answer that definitively. Looking back, the first end-to-end learning systems in the 1980s were a milestone. The first deep reinforcement learning systems in the early 2010s were another major milestone because deep RL gives us a way to surpass human-level performance. More recently, the advent of multimodal LLMs adapted for robotic control to supply common sense is a critical advance, and we will likely see several more major breakthroughs over the next few years.
Sergey Levine's Research Journey: From Blank Slates to Large Models
Host: Can you share your personal journey of approaching this problem, from when you first became interested to how you decided what to focus on?
Sergey Levine: I started working in robotics in 2014 after finishing my graduate degree and beginning a postdoc with Professor Pieter Abbeel at UC Berkeley. Before that, I worked on computer graphics and character animation. The core question I always wanted to figure out was how to build AI systems that continually improve the more they interact with the world.
Sergey Levine: Initially, I approached it from a blank slate perspective, where an agent starts with nothing, practices a skill, and gets better. That works in narrow, controlled settings, but it is hard to turn into a general system for open-world settings because any slight change requires retraining from scratch. Later, when working at Google, I explored parallelizing this across many robots—collective learning, where you place 20 robots in a room to learn simultaneously. That improved generalization, but still produced narrow specialists that struggled with edge cases.
Sergey Levine: The next step is combining the ability to practice skills with large amounts of prior knowledge. That is a hard problem across all of AI. The two major achievements in AI over recent decades have been generative AI, epitomized by LLMs reproducing human capabilities, and deep reinforcement learning, epitomized by AlphaGo finding superhuman solutions like Move 37. The challenge at Physical Intelligence is combining those two paradigms: bringing in vast generative world knowledge while using reinforcement learning to exceed human-level performance.
Vision-Language-Action Models and Chain of Thought
Host: What specifically are you doing to make that combination happen?
Sergey Levine: Over the past few years, we started by developing Vision-Language-Action (VLA) models. A VLA model can be understood as an LLM adapted for robotic control. It is pre-trained on text, adapted with extensive web image data to understand visual scenes, and then fine-tuned on diverse robotic data. That serves as the foundation for bringing web-scale knowledge into robot control.
Sergey Levine: On top of that, we explore two primary capabilities: handling unusual situations with common sense, and improving through reinforcement learning. To achieve common sense, we utilize chain-of-thought reasoning. When the robot encounters a scene, instead of immediately executing motor commands, it reasons through the task. If asked to clean a kitchen, it analyzes the visual input, generates an intermediate semantic thought like picking up a specific plate, and then acts. These intermediate inferences benefit from web-scale pre-training.
Sergey Levine: Reinforcement learning is applied when practicing tasks repeatedly to improve robustness, speed, and precision. In our espresso-making demonstration, the system practiced repeatedly to optimize throughput and reliability. We are continuing to build upon this dual foundation.
Sensor Requirements and Data Flywheels
Host: When looking at these robotic platforms, data is collected through sensors placed on the hardware. How do you think about the necessary sensor suite?
Sergey Levine: You can often get away with fewer sensors than people think. Our primary experimental platform uses three cameras: one on each wrist and one on the base. It does not have dedicated tactile or force sensing. A capable learning algorithm can compensate for hardware deficiencies. For example, wrist cameras act effectively as touch sensors because the visual feedback captures local physical deformations upon contact.
Host: In language models, massive internet scale drove the breakthrough. How do you create the necessary data reservoir for embodied AI?
Sergey Levine: Nobody knows exactly how much robot data is needed to achieve universal generalization, but we may not need to know in advance. What matters is getting the systems useful enough that they can be deployed to gather data autonomously. Tesla does not worry about collecting insufficient driving data because their fleet is already useful and actively collecting. The objective is to deploy systems capable of performing diverse tasks and establishing a data flywheel.
Host: Since starting Physical Intelligence, what has surprised you most about the trajectory of the research?
Sergey Levine: We made far more rapid progress on dexterity than I initially anticipated. Based on prior work, I expected generalization across diverse objects and scenes to scale steadily with more data. However, the models also learned highly dexterous behaviors without requiring complex, specialized architectures.
Sergey Levine: Furthermore, the models generalized across different robot embodiments—including multi-fingered hands and arms with varying degrees of freedom—without needing architectural modifications or explicit prompting about the hardware configuration. Fine-tuning on embodiment-specific data was sufficient.
Moravec's Paradox and Semantic Coaching
Host: Where are systems today more advanced or less advanced than people might expect?
Sergey Levine: This relates directly to Moravec's paradox. Humans intuitively assume that tasks difficult for us, like advanced calculus, are hard for machines, while tasks easy for us, like grasping a cup, should be simple. In reality, human brains have evolved massive specialized machinery for physical interaction and visual perception. Programming those physical skills manually is extraordinarily hard.
Sergey Levine: Machine learning shifts this dynamic. Where data collection is straightforward, tasks become tractable even if they are physically intricate. Conversely, tasks that require abstract common sense, multi-level reasoning, and linking physical execution to broad world knowledge remain difficult.
Host: How do you define common sense in the context of robotics?
Sergey Levine: Common sense in robotic learning means applying semantic inferences learned from other domains to the physical task at hand. It is the opposite of pure muscle memory. Muscle memory operates automatically through repetition. Common sense applies when you encounter a situation, recall a relevant fact or principle learned elsewhere, ground it in the immediate environment, and select the correct action.
Host: Language models have advanced from single-turn chat to long-horizon agentic reasoning. What is the equivalent long-horizon capability in robotics?
Sergey Levine: We are working extensively on this. Using chain-of-thought reasoning, our models can execute extended tasks, such as unloading dishes from a dishwasher, placing them in cabinets, and wiping down counters. An important discovery we made roughly 6 months ago is that model performance on long-horizon tasks can be improved simply by supervising them with high-level semantic instructions.
Sergey Levine: When a robot failed during a kitchen cleaning task, instead of collecting additional low-level teleoperation trajectories, we provided semantic labels for the intermediate steps. This improved generalization. The operational bottleneck had shifted from low-level physical actuation to mid-level scene interpretation and step selection. This allows human operators to improve the robot's performance through linguistic coaching rather than manual physical demonstrations.
Technical and Practical Risks to Broad Deployment
Host: If we look ahead to 2050 and general household robotics are still not widespread, what would be the primary explanation?
Sergey Levine: The bottleneck is likely the intersection of technology and human comfort. Autonomous vehicles faced similar adoption dynamics. Deploying systems in homes involves edge cases and questions of fault tolerance. Are users comfortable with occasional broken dishes or deploying robots around small children? Navigating those societal and safety factors will dictate deployment timelines across different environments.
Sergey Levine: From a purely technical perspective, the primary risk is managing open-world breadth. Structured commercial environments, such as cleaning hotel rooms or assisting in commercial kitchens, have bounded variations. Unstructured domestic environments feature unpredictable events that require robust inference and safe failure modes.
Simulation vs. Real-World Data and the Robot Olympics
Host: How do you view the debate between using simulation versus real-world data in robotics?
Sergey Levine: There is a clear dichotomy in robotics research today. In humanoid locomotion, the dominant pipeline relies heavily on physics simulation with minimal or zero real-world data. In robotic manipulation, the dominant approach uses large real-world datasets combined with foundation models, with little simulation. It remains an open research question whether one approach will dominate or if a synthesis will emerge.
Host: How do you balance developing impressive demonstrations versus building useful products?
Sergey Levine: Our approach is to make systems as exciting as possible subject to the constraint that they must be useful. We prioritize fundamental research that advances general robotic foundation models, but we stress-test models against difficult manipulation benchmarks.
Sergey Levine: For instance, former Everyday Robots researcher Benji Holson proposed a practical Robot Olympics consisting of everyday manipulation tasks that humans find trivial but robots struggle with—such as opening doors, scrubbing greasy pans, or using a bag to pick up waste. We used this list to test our general onboarding pipeline. Without developing specialized algorithms for individual tasks, our foundation model successfully solved nearly all of them. The only exceptions were turning a dress shirt inside out, which was physically constrained by gripper size, and peeling an orange purely with fingers, which required a small tool due to grip strength limits. This demonstrated the power of a general model to onboard diverse tasks rapidly.
Embodiment Agnosticism and Research Philosophy
Host: You mentioned the concept of tool use in biological systems. How does that relate to robot intelligence?
Sergey Levine: Studies in neuroscience demonstrate that when monkeys use tools, neural representations in the brain adapt so that the tool tip is treated as an extension of the physical hand. This reinforces the view that intelligence should be embodiment-agnostic. A universal foundation model should adapt to control arbitrary kinematic chains and tools. Manipulation, locomotion, and industrial control are unified under a single core intelligence problem.
Host: What are the main philosophical debates within the robotics research community today?
Sergey Levine: Historically, the debate was whether machine learning had any place in robotics compared to classical engineering and physics models. While learning is now widely accepted, there is still debate over end-to-end learning versus modular systems incorporating hand-coded physics. The Bitter Lesson posits that leveraging general methods with data and compute outperforms hand-engineered domain knowledge over time. While incorporating explicit physical models offers short-term structure, end-to-end learning provides the scalability and self-improvement necessary for true open-world generality.
Host: What types of tasks do you expect to be among the last to be automated by robots?
Sergey Levine: Tasks involving direct, nuanced physical interaction with humans—such as changing a child's diaper, eldercare, or helping an injured person out of bed—will be exceptionally challenging. These tasks require precise force modulation, high safety margins, and deep social-physical common sense.
Research Execution, Hardware Costs, and Future Outlook
Host: What differentiates exceptional AI researchers when navigating these difficult problems?
Sergey Levine: Research requires making critical decisions about when to persist with an approach versus when to pivot. Pivoting too early means missing breakthroughs right beneath the surface; persisting too long on a flawed premise wastes years. Successful researchers develop strong intuition for balancing depth of focus with exploratory experimentation.
Host: Hardware costs have also dropped dramatically. How does that impact the ecosystem?
Sergey Levine: A decade ago, a research platform like the PR2 cost around $400,000. When I established my lab at UC Berkeley, standard arms were around $30,000. Today, low-cost robotic arms can be sourced for roughly $3,000, and costs continue to decline. Because learning-based software can compensate for hardware compliance and sensing limitations, low-cost hardware becomes viable for complex tasks, radically lowering the barrier to entry.
Host: What is the primary focus for Physical Intelligence in the near term?
Sergey Levine: Our primary focus is advancing mid-level semantic and spatial reasoning. Low-level physical control is largely manageable, but linking those motor primitives with high-level world knowledge requires internal representations tailored for physical interaction rather than purely linguistic tokens.
Host: Looking back across your career, what is the kindest thing anyone has ever done for you?
Sergey Levine: Several key opportunities shaped my trajectory. When I was a sophomore in college, a hiring manager at Nvidia took a chance on me with an internship. Later, Pieter Abbeel took a bet on my potential by accepting me as a postdoc at UC Berkeley despite my having no prior robotics background. Finally, during my time at Google, leaders like Jeff Dean and Vincent Vanhoucke supported our proposal to deploy dozens of robots simultaneously for collective learning—the arm farm project—empowering early-career researchers to pursue ambitious ideas. Having leaders who provide agency and take bets on people makes a profound difference.
Host: Sergey, thank you so much for sharing your insights today.
Sergey Levine: Thank you.