DruckFin

Physical Intelligence Transcript: Sergey Levine on Humanoid Robotics Scaling, Generalization, and Ecosystem Challenges

September 2, 2026 | Sergey Levine in Conversation on the State of Physical AI and Humanoid Robotics

Introduction

Sergey Levine: To me, that is kind of mind-blowing because the base model was not trained on any human data at all.

Host: This is Sergey Levine, one of the world's leading robotics researchers, and I asked him all about the current state of humanoid robotics. What are the most astonishing emergent capabilities you have seen so far?

Sergey Levine: I do not think anybody watching that eval thought that the robot was going to do this.

Host: I have some questions here on China. There is no avoiding it sometimes. If humanoid robotics did not succeed in single-digit years, what do you think would be the most likely reason why humanoid robotics failed? Here is the full episode.

Where We Are Today in Robotics Scaling

Host: Large language models have caused unprecedented impact and investment in that area. Physical AI and humanoid robots could potentially be even bigger. I wanted to ask you today about where we are today with humanoid robotics and how you foresee this technology actually being deployed into the world.

Sergey Levine: I guess with machine learning, what we have learned over the last few years, or over the last decade rather, is that it works when you do it at scale. This is very obvious now, but it was not always obvious. There is a caveat, which is that you have to scale the right thing. Initially, when people started working on models for language, for example, the dominant design was LSTMs. Some people remember what those are. They were kind of okay. They were a lot better than what came before that, but they did not really scale as well. The big thing with transformers was not that transformers were somehow particularly mathematically elegant or anything like that; it is just that they scaled better. They were easier to train on very large amounts of data with lots of parameters.

Technology proceeds in phases. First, you figure out what you can scale, basically what is the scalable technology, and then you pour on an industrial-scale effort, adding lots of data and adding to model size. That is when the magic happens. When we are doing more fundamental technology development, the key is to understand what those scalable levers are. You figure out the design, figure out roughly the mixtures, and that by itself does something pretty cool, but that is not the thing that actually changes the world. It is when you start pulling that lever that things actually change.

With LLMs, when the first GPT models came out with GPT-2, it did some stuff, but it was sort of like a parlor trick. You could get it to synthesize a story about unicorns in Peru or something, and it was coherent English, but it was not a thing that would solve lots of real-world problems. The folks that worked on this kind of recognized that there is something magical that happens because as you add more data and make the model bigger, this stuff gets more coherent and more effective. They could see that if we do a lot more of that, then it will become a lot more powerful.

To come back to your question, what I would say about robotics is that it is not in the GPT-4 to GPT-5 stage where it is an industrial-scale effort to make the model bigger and get more capability out of it. It is in that stage where we are establishing the fundamental technologies. Because of that, what one should expect to see right now is not necessarily that each month the model gets bigger and more powerful by some predictable scaling curve. It is that the scaling properties themselves are evolving as we develop the right technologies.

To bring this back to something closer to reality, I am very happy with the demos that we are doing here at Physical Intelligence, and I think that a lot of the results that other people are coming out with are really cool. But to put them in context, we should not expect these to be the things that are actually illustrating the power of scale. We should expect them to be developing the fundamental technologies that will be scaled up after that. Where I think we are at now is that we are actually getting all those puzzle pieces in place, and I think it is actually very close. A lot of the puzzle pieces are falling in place. What makes it so hard to prognosticate about where the technology is going to go is that it is not yet at that predictable scaling stage. It is at the stage where we are figuring out the puzzle pieces, which I think is really exciting, but it means that it is also very hard to foresee what the coefficients on that will be.

Most Surprising Capabilities and Semantic Mistakes

Host: What are the most astonishing emergent capabilities you have seen so far?

Sergey Levine: That is the thing that is the most fun, and certainly we have seen a lot more of that happening as we progress. In the very beginning, it was little things, but they were kind of magical because in robotics, basically prior to 2024, this stuff never happened. The little things that happen, maybe about 2 years back, we would see things like training our policy for folding laundry, and it takes individual shirts out of the hamper and tries to fold them. One very vivid memory I have in late 2024 is watching one of the evals where it takes out two shirts at the same time. I am watching this thinking it is done for and there is no way it can possibly do this. Then it puts the two shirts on the table, disentangles them, puts one of them back, and starts folding the other one. In retrospect, you can do some detective work and figure out where it got that from some piece of training data, but that was one of those moments where I do not think anybody watching that eval thought that the robot was going to do this. It is exhibiting the common sense you expect people to have.

Actually, the thing that I find more interesting recently is some of the mistakes. One of the things that was pretty remarkable about LLMs is that once they got good enough, even the mistakes kind of made sense in the sense that they were not crazy mistakes where it just outputs random characters all the time, but mistakes that are semantically sensible. We had an evaluation last year for π0 where the robot was cleaning up a kitchen, and it is told to put away all the utensils, like some spoons and spatulas. It tries to open the drawer where it thinks the silverware goes and it cannot get the drawer open. So it slides over, opens the oven right next to it, and starts putting the stuff in the oven. You can imagine that if you ask a child to clean stuff up and put it away, they might decide to do that because it is a container, you can put stuff there, and nobody sees it.

Another experiment we had was washing all the plates. It would pick up the plates, wash them with a sponge, and put them on the drying rack. This was an experiment on memory because it has to keep track of everything that it is doing. It has a scratchpad kind of memory where it is writing down that it had 3 plates, cleaned the gray one, and cleaned the green one. Then it drops one of them on the floor and drives the base over so you cannot see it, and concludes that it has cleaned the gray plate and it is done. Obviously, these are not the things that we want to see, but it is interesting that some of the mistakes are almost like what you would associate with a child trying to do the task. Now it just needs to grow up.

Lessons from Autonomous Driving

Host: You mentioned the advancements happening all over the industry. I know there are a lot of people building humanoid robots. There is Figure, there is Tesla, and there are many other competitors. You have the expertise of what is hard and what is not. When you look at the competitors, has there been any advancement or achievement where you think that is really admirable and impressive?

Sergey Levine: I actually think that one of the most inspiring things to me in the industry is to see the kind of takeoff that autonomous driving systems have had. One of the criticisms that is sometimes leveled against robotics researchers is that it is like nuclear fusion, the technology of the future that is always in the future. That is what people said about autonomous driving too. Now, we are in San Francisco, you can go outside, take a Waymo, and it will actually take you to your destination with no driver sitting there.

Without getting too much into the technical details, what is really inspiring about that is just this case in point that you can actually have one of these technologies of the future land. I think it is not an accident that it is landing now in the mid-2020s, because a lot of the puzzle pieces for large-scale machine learning are getting to the level where we can put them together with actual physical systems. There are a lot of differences between driving and robotic manipulation, of course, but the illustration that we can actually land learning-based technologies in the real physical world is really inspiring.

Impact of Frontier AI Labs Entering Robotics

Host: If OpenAI or Anthropic started investing more heavily into robotics, how do you think that would impact the industry? Do you think competitors would be worried about that?

Sergey Levine: Robotics is an area where, to be fair, the ecosystem has not been as healthy as it has been in other areas of machine learning. What I mean by that is that computer vision and NLP are things that naturally lend themselves to a machine learning-based ecosystem because there is freely available data. People have a general acceptance that they are going to be using learning. There are not as severe concerns about physical safety. People rightfully are concerned about AI safety, but it is not the same as a physical device causing physical harm. Because of that, it is a bit easier to spin up a very serious large-scale machine learning effort in those areas.

Robotics traditionally is not a discipline that really embraces the sharing of data, for example. The more activity there is around learning in robotics, the more I think it will shift people's thinking towards this future where we accept that robots will be controlled by learned models, not by hand-designed controllers. There will be data, and that data will need to be shared, because there is no way that somebody can build out a true foundation model in a single vertical. It will basically shift the entire thinking around robotics to look more like how we think about vision and NLP, as opposed to traditional factory automation. In that sense, much as I am proud of the work that we are doing at Physical Intelligence, I think it will take more than one company to shift everyone's thinking in that direction.

Chinese Robotics and Ecosystem Dynamics

Host: As a bystander, I see on Twitter these really impressive demonstrations of Chinese robotics. It almost feels like they are ahead in some sense, but I do not have that deep domain expertise. I was curious about your thoughts on China's robotics and whether they are further along.

Sergey Levine: I think that one thing that is very useful and constructive for those of us working on these things in the United States and in Europe to do is to ask what the lesson to learn is. To me, one lesson is that it is important to have a healthy ecosystem. An ecosystem means that there should obviously be good researchers and good engineers working on these things, and there should be healthy open source. But it also means that the different industries that contribute to robotics need to individually be very healthy. Those industries are not just computer science, ML, and model building; it is also supply chains, manufacturing, and hardware R&D. These are all very important things.

Aspects of those things are things that the United States does quite well, while other aspects are things where the United States has let things go a little bit. We should look at what is going on in the world, look at some of the excellent results that Chinese labs and labs in other countries are producing, and take away the lesson that we should strive to build a healthier ecosystem. That means investing in all the different facets that contribute to this. I am not much of a business person or an investment person, so I cannot claim to know how to do this, but it is important to fully embrace that this is a holistic effort and not something where we can do just one piece of it and outsource everything else.

Host: Of all the pieces of the ecosystem, are there parts of it that, if improved in the US, would have the biggest impact on the advancement of robotics?

Sergey Levine: Certainly, the availability of reliable, low-cost hardware is a big deal. Right now, for hardware used in robotics research, a lot of that does come from China. It is good, relatively inexpensive, of high quality, and meets the standards that people generally need. It would be awfully nice to be able to source all that domestically as well. I do not think there is anything impossible about that; it is just a matter of embracing the fact that the entire ecosystem needs to be supported rather than just one piece of it.

The Data Flywheel and Scaling Milestones

Host: You can imagine that the lab that is first to get to scale will be the first to break out. There may be this exponential growth effect where deployment helps you grow faster, and growing faster helps you deploy, creating this flywheel. Do you believe that will happen in this industry, where one lab, whether in the US or China, hits some breakout point and jumps ahead of everyone else?

Sergey Levine: There is a lot of truth to that. There is an important detail to keep in mind: you have to scale the right thing. Having an effective positive feedback loop where more deployed robots translates to more model capability is the key. The trick is that there are lots of ways that it can be done wrong.

One obvious example is if I am a car company and I have a robotic arm that is welding cars on the assembly line. It welds cars every day and gets 1,000,000 welds every month. If I just use that as my data flywheel, I am unlikely to get something more capable than a robot that welds cars. Since the robot is already there and already welding cars, the marginal improvement for that is not all that valuable. That is an example of how you have to scale the right thing, because data is not quite fungible like electricity or oil. You cannot just buy more of it; it has to be heterogeneous. You have to scale the right thing with the right technology and the right source of diverse learning. Data is more like an education program for your robot than it is a fungible commodity.

Host: If you think about that data flywheel and putting out the proper platform that generates data that matters to improve that platform, when do you first see this kind of happening in the world?

Sergey Levine: On the technology side, things are advancing very rapidly towards that. A particular balancing act to strike that would significantly determine that timeline is how much structure somebody is willing to admit. On one extreme, you could imagine jumping straight to fully unstructured deployment domains like home robots. That could be really exciting because you have a lot of diversity right off the bat, but the bar is a lot higher to be effective and safe enough, because safety is a much bigger issue around people in their homes.

On the other extreme, you could imagine much more structured tasks. Maybe not quite the welding robot, but the robot down the hall that does something slightly more unstructured. That could be an easier domain in the sense that there are fewer safety concerns because it might be around trained humans and the task might be more predictable, but the marginal value of each bit of data you get in that domain is lower because there is less variety. You could take off earlier but with a smaller slope, or later but with a larger slope, and you have to calibrate that. My sense is that whichever end of that extreme we are talking about, it is in single-digit years rather than double-digit years at this point. The more structured ones might be happening now or next year. The less structured ones might be a few more years out, but probably not a decade.

Roadmaps, Generalization, and Real-World Testing

Host: A lot of people talk about timelines and it is nebulous. If you were to put down milestones towards that North Star of a robot in the home, what would that roadmap look like?

Sergey Levine: There is one thing that is very important to say about robotics that is very easy to miss from the hype cycle and the demos that people put out, which is that the hard thing in robotics was always generalization. When somebody shows a demonstration of their system, the demonstration alone usually does not make it clear what level of generalization is being shown. Highly acrobatic robot demos, for example, are really exciting to look at. But typically, if it is something rehearsed on a stage, that is literally just a show. It is not the same as doing a task reliably every time in any home.

The generalization piece often does not look that impressive when viewed in isolation because generalization is a property of many trials, not of one trial. You might see the robot doing something fairly mundane and unimpressive, but what is exciting about it is that it is doing it with an object that it has never seen before in an environment that has never been tested in before. That is actually harder than doing an acrobatic backflip that it has practiced 1,000,000 times.

The roadmap is all about achieving better generalization and having a mechanism to get more generalization as you generalize. One step on that roadmap is to have a very concrete demonstration of a robotic system that gets better with autonomous experience collected in a setting that it was not originally trained for. You train the model on the back end, put it in a new setting, whether a home or a factory doing something real, and it does okay, but over time it gets better and better until it reaches practically relevant levels of robustness without capping out at 50%. That would be a major milestone. If that is truly an automated process that improves as it collects useful experience, you can put it in lots of different domains to collect experience, do something people want, and improve the model.

Another major step is to demonstrate a concrete and practically useful way to transfer common-sense knowledge to achieve robustness. That is the scenario where you do not get to practice. If you are driving on the road and see a fire truck and traffic cones, even if you have never been in that situation before, your common sense tells you to slow down rather than barrel through the cones. If you can apply that common sense to effectively recover from unexpected situations, like when the robot put the spatula in the oven and should realize it needs to take it out, that tells us we can use common sense to fix mistakes.

Host: I could imagine a more narrowly scoped robot, like a humanoid robot doing one step on an assembly line. In my understanding, you are less interested in that because that is not really a step towards general intelligence.

Sergey Levine: It is not that I am less interested in it. It is that the real world has leaky abstractions that make that kind of stuff a lot more complex than it seems. In the 1990s, when people started working full steam on autonomous driving, there was an idea that we could avoid a lot of the hard problems by instrumenting the environment, putting magnetic sensors along the highway and transmitters on cars so they know where everyone is, similar to aircraft. People thought we did not need fancy AI, just sensors, and it would work. That did not go anywhere because the real world has so many messy exceptions and special cases. Even if magnetic sensors work 99% of the time, that 1% when someone steps into the road or there is a piece of trash messes everything up.

The thing that actually worked was deciding not to avoid the hard problem, deploying cars directly in messy environments like San Francisco, and dealing with it head-on. Robotic manipulation is going to be the same way. Past the fully structured world of the factory, if you want to go even a little bit outside of that, even if 99% of the time it is straightforward, the 1% when something weird happens means the full scope of the problem needs to be addressed.

Host: This reminds me of when Figure had a demo where they live-streamed the robot sorting packages. Does that demonstrate generalization in your opinion?

Sergey Levine: Yes, I think it does. This is something I find very encouraging. Because it is hard to show generalization in a video, lots of people are thinking creatively about how to present something at a glance, using live demos and long time lapses. That is a great way to elevate the importance of generalization in people's consciousness.

When we were working on the π*0 RL project late last year, we wanted to do some longer-horizon experiments. We had our robot assembling boxes at the Dandelion Chocolate Factory for several days. We also had a coffee task where the robot used an espresso machine to make espresso, running for 13 hours making drinks. We tried to be environmentally conscious and not throw out the coffee, so after 13 hours everyone in the office was wired from drinking espresso. It screwed up a few times, like spilling coffee grounds and having to wipe them down with a cloth, but nothing exploded over 13 hours.

Host: When it spilled the coffee grounds and cleaned it up, did it do that by itself?

Sergey Levine: The way that experiment was done is that there is high-level prompting updated roughly every 5 minutes in between semantically coherent tasks. You tell it to make espresso or clean up the machine, so the command to clean up was prompted by a person. In principle, we could automate that. We are spending a lot of effort now improving our high-level policy that handles those commands. For that experiment, every 5 minutes someone updated what it was being asked to do, just like ordering a latte at a coffee shop, except you also had to tell it to clean up before making the next one.

Host: OpenAI, Anthropic, Cursor, and Vercel all use this product to make their lives better. The problem it solves is when you are building SaaS or an AI product and you want to sell to other companies, there are all these requirements you need to meet: SSO, SCIM, RBAC, audit logs. These take time to integrate but are not the main focus of your app. WorkOS is an API layer that lets you meet all of these requirements in just a few lines of code. Check them out at workos.com to learn more and get started. I appreciate them sponsoring this podcast.

Types of Data and Cross-Embodiment Transfer

Host: On the way to generalizing, data is a very important part. There is simulated data and physical interactive data. What is your take on the best data to get, the worst, and the pros and cons?

Sergey Levine: There is a lot of discussion in the robotics community about this, and people hold opposite opinions. My take is that a lot of different data sources are easier for the model to internalize if it can ground them in a thorough physical understanding of the world.

If you want to learn to fly an airplane, you will probably use a simulator during part of your training. The simulator makes sense to you because you bring a lot of world knowledge to ground what is going on. You know you are acquiring knowledge to use on a real airplane. Similarly, if you watch someone cooking a meal, even though you do not feel their exact muscle movements, you have prior knowledge to file it away at the abstract level of adding salt without having to calculate their joint trajectories. Once you understand how to do things physically with your own body, all other sources of knowledge can connect to it because your embodied foundation grounds everything.

If we have a robotic foundation model trained on lots of real embodied data that provides grounding, it might actually be much better able to absorb other sources of knowledge. Looking at the success of internet data for LLMs, it is tempting to start with YouTube videos and put robot data on top of that, but I think it is the other way around. My colleague Sudeep together with Samarth from Georgia Tech took our robot foundation model and added human video data on top of a model initially trained on robot data. When using a small model with minimal robot data, human experience and robot experience were fully separated in feature representations. But when trained on lots of robot data from many different robots, the features grouped entirely by task identity rather than embodiment. In the t-SNE embeddings, cranked up to 100% robot data, task identity aligned cleanly with minimal sensitivity to embodiment. The base model was not trained on human data at all, but once you add human data, it represents it the exact same way.

Host: Is it important that the base model has data collected using that specific set of motors and joints?

Sergey Levine: We have put a lot of effort into cross-embodiment models that handle many robot types, but generally you do need some data from the target robot to get good performance. The metric of generalization is not zero-shot transfer to a new robot, but getting away with less experience on the new robot by transferring skills from others.

The amount of special architecture needed is surprisingly minimal. Initially, I had a long list of research ideas to accommodate different morphologies, like factorizing representations for a 6-degree-of-freedom arm versus a 7-degree-of-freedom arm. We did not do any of that. The model outputs a vector of numbers, and if the robot has fewer degrees of freedom, it simply zero-pads the output. It trains across all robots and outputs actions based on what it sees through the camera.

To transfer skills between robots, intermediate thinking helps. Thinking in text is good for transferring high-level behavioral structure, like knowing to open a drawer before putting away silverware. For lower-level physical tasks, we ran an experiment where the thinking stage is expressed in images, generating an image of the next milestone in the task. With that, we got a UR5 robot to fold a T-shirt even though we had zero T-shirt folding data on that specific robot. Cooking up an image of what a folded shirt looks like is straightforward with a generative model, and backing out the required joint angles from that synthesized image is relatively simple. Introducing intermediate thinking in the right modality makes cross-embodiment generalization much easier.

Robotics Stack Simplicity vs. Autonomous Vehicles

Host: Earlier you mentioned Waymo was inspiring as proof of generalized robotics. In the case of humanoid robotics, what makes you confident in a single-digit year timeline versus the long tail of edge cases and policy challenges that autonomous vehicles faced?

Sergey Levine: One big difference between robotic foundation models and traditional engineered systems is that the software stack is really thin. Training a foundation model is hard due to data curation and labeling, but the actual software running on the robot is very simple. In terms of raw lines of code, it is much smaller than a traditional autonomous vehicle stack. Autonomous vehicles started earlier with very different legacy technologies and have severe safety-critical constraints. Dropping a fragile object with a robotic manipulator is far less catastrophic than a car hitting someone.

Safety challenges in robotics are real, but they are not as much of a hard stop to practical deployment. You can select tasks, environments, and hardware where those risks are manageable. The combination of a radically simpler software stack, less drastic safety hurdles, and the positive data flywheel makes single-digit timelines realistic.

Premortem on Humanoid Robotics Failure Modes

Host: If humanoid robotics did not succeed in single-digit years, what would be the most likely reason why it failed?

Sergey Levine: For these systems to be truly useful, they need to reach a level of reliability and generalization higher than what we expect from LLMs or generative video models. LLMs are human-interactive tools where a user can iterate on prompts until a task is solved. With a robot, the entire value proposition depends on autonomous execution. Having a human constantly intervene defeats the purpose.

The biggest risk is the difficulty of reaching that high bar of robustness. Reinforcement learning techniques that leverage autonomous experience can help close the gap from 95% to 100%, but crossing that final stretch remains an open technical challenge requiring further algorithmic innovation.

Foundation Models vs. Narrow Specialists

Host: In robotics modeling architectures, is everyone doing the same thing, or are there divergent hot takes?

Sergey Levine: There is more heterogeneity than it seems. The primary dividing line is between fully embracing the foundation model ethos versus focusing on narrow vertical areas. The foundation model ethos states that to solve a specialized task, you are better off training a general model across a wide breadth of data, which will ultimately outperform a narrow specialist on edge cases.

In robotics, that idea makes people uncomfortable. If you want to build a warehouse automation system, collecting data on putting away silverware in kitchens seems counterintuitive. However, collecting diverse data builds generalized skills that transfer to unpredictable edge cases in the warehouse. Accepting that breadth beats narrow vertical optimization is still difficult for traditional robotics engineers.

AI Safety in the Physical World

Host: If Physical Intelligence develops an extraordinarily capable general model, how do you think about AI safety, risks, and regulation in physical systems?

Sergey Levine: AI safety has been studied for a long time, but when technology moves quickly, the real issues depend heavily on how society adopts the tools. Our philosophy is empirical experimentation: deploy systems in the real world, observe what succeeds and fails, and adjust iteratively. If software-only AI creates significant safety concerns, physical AI systems capable of interacting with the physical world will naturally require even greater scrutiny.

Economic Impact and the Future of Labor

Host: If everything goes well over the next 10 years, does humanoid robotics mean the end of human labor?

Sergey Levine: It is a mistake to think of robots simply as mechanical people. When personal computers took off, they did not replace human brains; instead, we saw ubiquitous computing proliferate into desks, pockets, refrigerators, and cars. By analogy, physical AI will likely introduce ubiquitous physical actuation, providing automated assistance across everyday tasks rather than acting as a direct one-to-one replacement for humans. Modern coding agents serve as a good parallel: rather than eliminating software engineers, they provide leverage to amplify human productivity.

Essential Research Papers: Aloha and ACT

Host: If someone wants to understand the lineage of breakthroughs in robotics, what foundational papers should they read?

Sergey Levine: I would point out the original Aloha and Action Chunking with Transformers (ACT) paper led by Tony Zhao, which I co-authored. The paper demonstrated that using low-cost $7,000 hobbyist arms from Trossen Robotics in a bimanual teleoperation setup, combined with a straightforward transformer model, could solve highly dexterous tasks like replacing remote control batteries or putting a shoe on a mannequin foot.

Academic research often prioritizes mathematical complexity, but the insight here was that simple, accessible hardware paired with the right end-to-end learning setup could achieve remarkable results. The open-source ACT codebase has become a standard starter kit for modern robotic learning because it proves how far simple building blocks can go.

The Evolution from Classical Controls to Learned AI

Host: Prior to modern machine learning, Boston Dynamics commanded massive attention with impressive demonstrations, but we hear comparatively less about classical approaches today. What drove that industry shift?

Sergey Levine: Robotics requires full-stack integration across mechanical design, actuation, and high-level decision-making. Classic Boston Dynamics demonstrations showcased exceptional hardware and sophisticated hand-designed control engineering. At the time, that was essential because without functional hardware, decision-making software is irrelevant.

Today, hardware is good enough, and the primary bottleneck is closed-loop decision-making in unstructured environments. The dividing line between classical controls and modern AI comes down to whether the robot only needs to control its own body or whether it must react to the external world. Doing a backflip on flat ground is primarily a controls problem involving the robot body. Picking up a coffee cup requires understanding the external environment. If a human engineer can hand-design a control law for a behavior, that serves as existence proof that a compact, generalizable policy exists and can be learned directly through machine learning.

Advice to Younger Self on Prior Knowledge

Host: If you could go back to when you entered the industry, what advice would you give yourself?

Sergey Levine: I would tell myself to take prior knowledge much more seriously. Early on, many researchers, including myself, believed robots should learn entirely from scratch. During the Arm Farm project at Google, we had arms grasping millions of objects from a blank slate. They learned to grasp, but that data did not easily bootstrap more complex skills.

Animals and humans do not learn in a vacuum; they rely on observation and existing models of the world. Incorporating broad prior knowledge, whether through language, internet video, or cross-embodiment data, provides the necessary scaffold for efficient learning. Attempting to learn without prior knowledge forces the robot to solve a strictly harder problem than nature ever had to solve.

Host: Thank you so much for your time, Sergey.

Sergey Levine: Thank you for the questions.

Host: Thank you for watching this podcast. If you enjoyed it, please leave a comment and a like. I am also working on building an ergonomic split keyboard prototype. We launched on Kickstarter and hit our funding goal within 8 hours. Tooling is currently underway, and late pledges remain open. Thank you for watching, and I will see you in the next episode.

Disclaimer: This article is for informational purposes only and does not constitute investment advice or a recommendation to buy, sell, or hold any security. Our analysts provide detailed coverage of corporate events but can make mistakes, always conduct your own due diligence. The views and opinions expressed do not necessarily reflect those of DruckFin. We have not independently verified all information used herein, and it may contain errors or omissions. Before making any investment decision, consult a qualified financial advisor. DruckFin and its affiliates disclaim any liability for any losses arising from reliance on this content. For full terms, see our Terms of Use.