{
  "video_id": "4EsUaur0nsQ",
  "transcript": "The equation I think for starting a robotic business has changed and will continue to change at an accelerating pace. Because the upfront cost is not that high anymore. Everyone's sort of spending a lot of time in the digital world. And it feels like now is the time to start thinking about the world of atoms. You literally just gave people the playbook for how to build a vertical robotics company. This has really been our mission from the start is to create that Cumbrian explosion. Es still blows my mind. Ich wusste nicht, dass das überhaupt in meinem ganzen Leben existiert. Willkommen zurück zu einer anderen Episode von The Light Cone. Heute haben wir eine sehr spezielle Guest, Kwan Vuong. Er ist ein der Co-Founder von Physical Intelligence, das wir denken, der Robotics AI-Lab, die über die GPT-1-Möme für all die Robotik. Kuang, danke für uns. Ich bin ein sehr lange an, ich bin ein Admirer von YC. Und unsere Mission ist ein Model, das kann jeder roboten, das man sich physisch zu tun, und so als eine high-level performance, das wird für Menschen in all walks of life. Und so GPT-1 für Robotics, was ist es? Ist die ChatGPT Moment für Robotics real? Our perspective here is that we want to build a model that's really intelligent. We want to build a platform that allows us to externalize that intelligence to the rest of the world and allow them to use it to build very interesting applications in all sorts of vertical and robotics. And we think that it's going to be more like a peeling an audience analogy where you start from a really strong base model that has all sorts of common sense knowledge and already works to some extent on your robot, you have then a mixed autonomy system, very similar, for example, to an autonomous driving car today. And then you actually deploy that system to do a real job. That system might make a mistake. It's okay. And then over time, by actually exposing the system to the complexity and the edge case of the real world, that system gets incrementally, even just slightly better over time every day. And, you know, one day you wake up and you suddenly have a system or just fully autonomous and provides tremendous value. Might be helpful to give the audience a bit of a mini history lesson on why robotics is so hard. And there's been a lot of breakthroughs in the last two years. And I mean, just to simplify the robotics problem is 3 pillars. Symantec, which I think we got out of the logs and with language models that somehow we poured it into robotics. Then you have the planning and the last thing is control, which needs to be done in real time and interact with the environment that changes. Walk us through the seminal papers that a lot of the team of Py Robotics published that gave you the inkling that the GPT-1 moment is near. And that started in 2024. So the dream to build general purpose robots has been a long time dream, I think, in humanity. We're not the first to say that our mission is to build a model that can work on any robot. Und wir sind wirklich in dieser Moment in Zeit, in der wir fühlen, dass es möglich ist. Zu zurück ein paar Jahre, da war ich ein bisschen, die erste Seikan, die zu mir war, die erste Demonstration der Langemodel und wie man kann all die Kommun-Sense-Knowledge in Langemodel in Robotics. Und, therefore, das Significantly reduziert die need zu collectivitäts-Daten. So for example, if you have a task of, oh, I want to go to the YC office to record a podcast, you know, what a step I need to take. You can ask a language model, you know, just show me the step and show me the plan. And that worked incredibly well. And then the way kind of language model infiltrate, if you will, in robotic is it start at the planning level, at the semantic level. And then, but there's still the control problem, you know, at the end of the day, you still need a mechanism to convert the plan into low level action that can actually accurate Das bringt uns zu POM-E und das bringt uns zu RT2, die für Robotic Transformer 2. Was diese 2 wirklich showt, ist, dass wenn du start von einem Vision-Language Model, das ist wirklich powerful, und du benutzt Robotic Data, um das Modell zu sprechen, zu sprechen Robot-Language, wenn du, dann sieht man viel transfert aus dem der kind of knowledge, der Existen der Vision-Language-Model nach dem Low-Level Action. One of my favorite examples, when we did the RT2 project, was you can have a picture of a celebrity on the table. You have a picture of Taylor Swift, you have a picture of the Queen of England, and you can ask the robot, you know, pick up the Coke can and move it to Taylor Swift, even though the concept of Taylor Swift, it just doesn't exist in the robot data at all, and that works. You can do other examples such as spatial reasoning that doesn't exist in the robot data at all. For example, move the dinosaurs next to the red car. And these are all just completely unseen objects in robot data. And so that was RT2 and that was POM-E. Now, RT2 and POM-E are single embodiment exercise. Just for the audience, single embodiment, meaning it worked for a very specific robot. It worked for a very specific robot. In Robotics, you can ask the question, how do you scale? Especially, how do you scale data collections? And one of the insights that we had back then was, you know, maybe the data from one robot is not that different from another robot anyway. If you have enough robots in your training data, maybe what the model learned isn't to control one specific robot. What the model learned is something that's more abstract, which is how do I kind of learn a general notion of what it means to control any particular robotic platform. And therefore, I will be better at controlling any particular platform. And that brings us to what we call OpenCross Embodiment and Robotic Transformer X. That was a big paper because it was the first that showed potential scaling laws that apply to robotics because now you could start training all these models across multiple kinds of hardware, not just one, which has never been done in robotics ever before. Because from all the research labs, they would all train with a very specific set of sensor actuators and motors. And it was all very finicky with that particular hardware, right? Yeah. One of the really interesting results from OpenCross Embodiment, and let me provide the context here, is that you can take, let's say, 10 different robot platforms, collect data from them, train a policy, and really optimize the policy to work well on that platform. So let's say you have that, you have 10 different platforms, 10 different policies. Und jetzt, wenn du einfach den Daten ab und es in einem Modell mit einem hohen Kapazität genug zu wirklich absorbieren, und du kannst es compare. Du hast diese Generalist, das lernt die 10 verschiedene Robots, du kannst es mit den Specialist, die auf den Beispielen zu arbeiten, auf eine bestimmte Embodiment. Wie ist es compare? Und die interessante resultat von OpenX ist, dass es 50% besser ist. Wow. Und das war wirklich sehr surprising. because in Robotic, it's hard enough to get your model to work on one particular robot platform. And one of the reasons why I say that we're really fortunate to be in this moment in time in Robotic is because OpenAC was really only possible because of the support that we received from the Robotic community. It was a huge collaboration across the Robotic community. And the reason why that's really important is there is this joke in Robotic grad school dass wenn du zwei Jahre anstatt zwei Jahre zu deinem PhD, dann einfach nur auf eine neue Robot-Platform. Das ist das Logik. Wenn du eine zehn Robot-Platform hast, dann ist 20 Jahre. Warum ist das? Es dauert ein Jahr oder zwei, um die Platform zu einem neuen, um zu collectieren? Ja. Ist es fair zu sagen, dass die Datasette von Embodiment X ist ähnlich wie die scale und impact das ImageNet für Vision gemacht hat? Vision, because it was huge and it was the first large dataset across multiple hardware, huge collaboration. I still think that ImageNet was more impactful in the Vision community. And the reason for that is a few. The first is that ImageNet also allowed for reproducible evaluation. Right. You know, OpenX as an effort was more about making data available for kind of people to use. Und Evaluation ist ein wirklich schwieriges Problem in Robotic that OpenX nicht lösen. Und der zweite ist, I think OpenX ist ein drop in den Bucke, in der Robotic-Community. Wenn man sich die scale, die volume und die diversität der Daten, die Community ist collecting. I think OpenX, at this point, ist ein drop in den Bucke. Ich glaube, wir haben angefangen über GPT-1, aber auch GPT-1, das war das Moment, where you can prove, Alec Radford figured out that there was a neuron based on a very specific input and output, and then that allowed the scaling laws to sort of take hold. The biggest problem in robotics I've heard is basically actually exactly what we've been talking about. It's the data problem. Language you could bootstrap off of the sum total of what you could get off the internet, which is actually quite a lot. Can you give us a sense for scale? Is it petabytes? What do you think is necessary as an input to the true GPT-1 of robotics? Yeah, so the data scarcity problem in robotics, there's a few ways to look at it. The first way is that it's really two problems in disguise. There is a generation, data generation problem, and there's data capture problem. And the difference is that the data capture is that there might already be lots of robotic data that is being generated, Aber da emulate war überhaupt einen incentive für zu Indizetz in die Hand scammert, lecturer wird es in Qualgebiten 1ostenkon seizure und einen iT 그런데, weil hier free Roboter es ein echt guterWissen兩 individuals zu diskutieren undları stein. Es ist ein sehr vieles, sehr vieles zu veröffentlichen, zu collecten Daten. Und da ist die Frage, ob es es zu werden? Well, die Weise ich es sehe, ist, dass wir US-GDP, 24 Millionen US-Dollar. Wir sagen, wir haben eine Modelle, das kann eine Roboter zu tun, Napkin Maths, vielleicht 10% zu US-GDP. Das ist bereits ein Massive-Number. Und ich glaube, das Promis ist eine der�inetlichen Umgebung. Das ist ein Beispiel für die Unterstützung für einen hull solchen Robotischen Concept. Und das sieht man, so wie auf Sie fokuss應 auf Cross-Embapiations perché es sehr viel Và steepens isto. Und in der Cross-EmbARR remembrance gibt es auch den Daten Zweck, ist es sehr sicher – dass Ihr ihr hazgelöst und Ihr organisiert Euch keine Arbeit umgeben abzudeтесь zu vieleтиen von verschiedenen Schößen von dann ein anderekem發現. That actually allows you to scale easier. For example, if I were to contrast our approach compared to, let's say, a company that have a particular hardware platform that they optimize for and they scale, it's not an approach that has really allowed people to scale because it's just much harder to figure out how do you manufacture like a thousand units of something for now compared to making sure that you yourself are ready to absorb data von 1000 verschiedenen Robots in der Community. Das ist ein crazy Problem, nicht es? Die Hardware itself, in dem Zusammenhang mit Embodiment, wenn es ein Hardware run geht, oder eine der Servos ist etwas anders, dann sieht man es in den Daten Und dann wie du das f das controlst Wir waren ein ein Inventory Robots in der Firma Wir waren so shocked dass da keine roboten Platformen sind. Und wenn du Leute in der Royal Community fragen, dann gibt es eine Debatte über Multi-Robot versus Single-Robot. Und die Argument ist, dass Single-Robot ist simpler zu scale. Und eigentlich, das ist nicht wie es in Praxis ist. Wie es in Praxis ist, ist, dass wenn du eine Single-Robot dass du das Robot, das du optimiert hast, über die Zeit, das Platform wird, vielleicht auch, wenn du, vielleicht, wenn du, vielleicht, wenn du, oder du, oder, oder, du, in der Situation, dass es viel harder für dich zu re-use old data. in Machine Learning, wenn du, wenn du, wenn du, viele Sample von der Distribution und wenn du, einfach nur eine Robot Platform mit einem Veränderungsmöglichen für drei Monate, vielleicht, vielleicht, wenn du, ein paar Daten von der Distribution. Wenn du, die, dass, wenn du, robot platform in your fleet, your model is going to learn something more abstract, which is how do I control a robot, not any particular robot, then the model will be able to ingest data from a slightly different robot barrier. And actually we're starting to see emergent property in this kind of robot large foundation model. That's good news. Where you start to see interesting transfer between different data sources. So, for example, today it's possible to perform tasks zero-shot. Zero-shot meaning you don't collect any data. And these are the tasks that last year might have required hundreds and hundreds of hours. What are some examples? Yeah, do we have any videos we can see that show it? So, you know, I get some flack when I come back because this is not published results. Hopefully this will come out soon. So I want to reserve the excitement for that and I'm building up the excitement a little bit. So hopefully this will come out soon. All right. Das sind einfach nur die Herausforderungen, die letzten Jahrzehnte, die überhunderten Jahre von Daten zu erfunden haben. Du hörst hier an Lightcone first, dass es emergenten Proberen werden, die werden aus Pi kommen. Kann man das auch sagen, wie die Flavor der Tasks? Es ist wirklich einfach zu foolern. Und so wir wollten die Tasks über viele Tasks von verschiedenen Flavoren. Tasks, die Positionen sind, die Tasks sind, die Reisung mit Multiple Objekte in der Szene sind. Es all scheint zu haben, diese Property zu haben. Das ist wirklich nett. Es scheint, dass das eine mehr General Property zu erzeugen, eher als wir einfach nur luckieren und die Models starten auf einen bestimmten Task. Können Sie uns verstehen, wo wir jetzt sind, in terms of wie es funktioniert und wie es funktioniert? Wir sind nicht in der ChatGP moment, aber wo wir sind? Und ich glaube, du hast einen Videos, um die Leute zu zeigen, was die current State-of-the-Art eigentlich aussieht. I think where we are is I think if you have a test where it's okay for the robot to make a mistake and it's possible for you to set up a mixed autonomy system where you have a person that takes over when the robot makes a mistake and provides corrections, it is possible to get to a level of performance where it starts to make sense to think about scaling robot deployment. And the example that I specifically want to highlight here is this blog post that we did with Weave and Ultra. And, you know, it's great that these are both YC company. I want to provide a little bit of context here first. The context is that Pi is a primarily research organization. We want to focus on building the best model. But we also want to not be tunnel vision. Wir wollen, dass die Modell wir gebaut sind, ist tatsächlich gut, und tatsächlich performt die Menschen in der Gesellschaft interessiert. Und eine gute Art von uns zu tun ist, dass wir wirklich mit einem Unternehmen mit einem einen Unternehmen mit einem Unternehmen, das wir heute machen. Und die Art von diesen Relationshipen arbeiten ist, dass wir uns auf der gleichen Team wie wir auf der gleichen Team sind. Wir haben eine sehr freie Fläche, die Informationen zu machen. Und wir haben eine Systeme, die wir uns bestellen, die bester Leistung für die Das ist die Unternehmen, die wir über die WI-FUS haben. Was du sehen in diesem Video ist ein System, das wir zusammenbauen, um eine wirklich diverse Artikel von Läumen in eine Real Laundromat. In der Mission, du kannst du Menschen auf die Wandern gehen. Und warum diese Artikel ist schwierig, ist es, dass es einfach ist, dass es infinite Möglichkeiten der Observation-Space ist. Die Kleidung sind deformable und zwei Pläne hier sind die gleiche. Und diese sind auch nicht un-sehn. Diese sind nicht Pläne, die sind in den Training Daten. Ja, ich liebe diese Team. Sie sind die meisten von Apple aus der ersten Welt. Gary war die Partner für Weed. Vielleicht wollen Sie erklären, was Weed ist und was sie ist. Ja, ich meine, sie sind eigentlich, Sie sind ihre ersten Roboter in die Home. Wir haben es als, wie es zu tun, wie es zu tun, wie es zu tun, wie es zu tun. Und ich denke, sie waren sehr inspiriert von Physical Intelligence's ersten Demos mit Laundry Folding. Es ist eigentlich ein total trip zu hören über es. A Jahr ago, wir waren über sie zu tun, und dann zu sehen sie es, arbeiten sie Hand in Hand mit Ihnen ist wirklich awesome. Ich denke, das ist ein tolles Beispiel. Du hast die Model Smarts, du hast die Daten Collection, und dann die Hardware und die System Integration all working together ist es einfach zu knallen. Ja, und zu gehen zurück zu deinem Frage, warum Robotic ist hart, es ist wirklich ein wirkliches System Problem. Du musst alles gut und gut zusammenarbeiten, um das Ergebnis zu machen. Und es ist eine große Team für uns, um das Ergebnis zu arbeiten, um das Ergebnis zu machen. Und es eigentlich nicht wirklich das lange Es war es, wir haben einen Grund, und vielleicht zwei Wochen später, wir haben einen Model und einen System, das war gut genug, um das zu tun. Es ist immer so, wie es mein Herz macht, um, dass ich eigentlich das Foto von Laundry. Weil ich, bis ich, bis ich, bis ChatGPT hatte, ich wusste, dass das auch in meinem ganzen Leben existiert. Weil, Foto von Laundry, ich meine, es hat immer immer den Turing-Test für Robotics, Because there's no way to deterministically program a system the way that you did pre-AI to do this, because the space is so infinite. And we've shown that it's possible for us to do. Basically, if everyone can do this, robots will be able to do everything. It's only a matter of improving it from here. There was a funny story where when we first published Py Zero, people thought of us as the laundry company. Because the demo was just focused on laundry. und eigentlich picking home tasks, especially tasks that has to do with deformable objects, is a very intentional choice on our end. We're not just after the home. We really want to make it broadly applicable. But picking home tasks for us to start with has a few benefits. One is relatable. You can see the laundry folding demo and you can kind of grog how this is going to be useful and you can get a sense of why it's hard. And the second is that it's really easy to set up to test generalization. Wir können über Ultra, die Ihr Unternehmen, Jered. Ja, das ist Ultra. Das ist das Video. Das ist es, dass es aus dem Video ist. Und du sehen, das ist 4x speed und es ist 100 Minuten. Wenn ich auf den Ende scrollen, die Sonne hat sich auf. Oh, wow. Das war eine der großen Probleme in Robotics, wo es so sensitiv zu den Environment ist, in der Licht messen, die Vision System, die Semantics und die Parten. Ja, und die interessante Sache hier ist, dass es möglich ist, dass die Autonomie der Robotin ist, dass die Task ist. Das ist Autonomie auf der Scale. Das ist bereit zu werden. Kwan, weil diese Task ist weniger familiar als Laundry Fielding ist, wie die Robotin ist hier und was Ultra ist als ein Unternehmen? Ultra ist ein Unternehmen, die wirklich einfach zu adaptieren, die Robotin zu neuen Taschen. und jetzt sind sie auf logistisch-baser, das ist wirklich wichtig, weil es vieles ist in Logistik ist. Und die Task, die wir zusammenfassen haben hier, ist, wenn du ein Item aus Amazon auskommst, du manchmal hast du das Pouch, das Item wird von, und die Task hier ist, du hast ein Tray, diese Items hier und die Robotin ist supposed zu picken, einen an den Tag und placeen in dieser Pouch. Die Maschine wird dann close es, und dann pick up die Pouch und put es auf der leften hier, um es bereit für die Schuppung zu werden. Das ist hart, weil es viele verschiedene Objekte gibt, die in dieser Tray stehen. Und die Öffnung hier ist sehr narrow, so Sie sehen hier eine interessante Art, der Robot nötig die Item, um die Pouchen zu gehen. Und das ist wirklich hart. Das braucht eine sehr gute Entstehung der Scene und eine sehr precise Möchtung und nudge die Objekte in die Pouch. Die andere Sache ist das hart über diese Task ist die Lever der Autonomie, die es gebraucht ist. Das ist für eine ganze Zeit. Es gibt noch immer mehr Intervention, ich sage, in dieser Full-Day-Operation. Aber die Lever der Intervention ist eigentlich ziemlich minimal. Das ist nicht nur eine Demo-Station, das ist in einer e-commerce Warehouse, wo sie realen Produkte zu realen Kunden sind. This isn't just like a lab. This is packaging real customer, real order for customer to be shipped out in a real warehouse. So this is real operations. So I think this is really cool because I think when people think about robots, they tend to think of the consumer use cases like Weave because that's, you know, what we're familiar with in our daily life. What I find really interesting is that there's like a million applications like this Ultra thing that you wouldn't think of as obviously like, oh, who packs the like soft pouch of things that you get from like Amazon? There's some person who does that and this is a job that we could not build a robot to do. The interesting thing about the approach is that you're converting it from a very difficult engineering problem into an operation problem of how do I identify the use case and how do I collect the right data, which is in some sense more scalable because you can build a system that allows you to collect data for many different tasks. So it's now a problem of how do I scale data collection rather than for every new task, How do I design a really difficult engineering system to solve it? YC Start a School is back. We're hand-selecting the most promising builders in the world and flying them out to San Francisco for July 25th and 26th to discuss the cutting edge of tech. Apply now for a spot. Okay, back to the video. I think one thing that the audience may not know is that you have a very unique technical insight that in the past, robotics folks would have kind of gasped and be shocked because robots need to run in real time. A lot of times all of the compute runs on device, but you guys have done something very different. Can you tell us more about that so that this works in real time with large models and really well? So the context here is that we talked to many companies that would like to deploy robots and one of the first questions we get is what compute units should we get on the robot? You know, it's expensive, it's going to increase the bomb cost und ich bin, dass es sich in der Fall sehr schnell ausgehen wird, weil die Modell verändert, die Modell wird größer. Wie soll ich sicher, dass die Hardware, dass ich heute heute mit dem Wissen verwendet bin, für ein paar Jahre. Es gibt eine sehr schwierige Frage. People sind oft wirklich surprised, wenn ich sie mir sagen, dass fast all die Robot Evaluation, dass wir bei Pi heute haben, die wirklich kompliziert Demo, wir haben schon gemacht, machen, die Foto, die Möbel, die Robots navigieren, die Modell actually hosted in der Cloud. Und das ist nicht, like a cloud as in a server in the office, it's a real cloud. The model is hosted in a data center somewhere. And within this high frequency control loop that is controlling the robot, the robot is actually querying an API endpoint that hosts the model, sending it images and language command and getting back action that then execute it directly on the robot. And this is surprising because of precisely the reason that you mentioned, how do you actually make it work? Das ist wirklich wichtig f Pi zu verbinden hardware und model development and research very tightly together because it allows us to solve for this problem So for example one of the insights we have here is that you can actually bury the inference time within the robot control loop because if I'm a robot, I have enough action for me to execute for the next 100 milliseconds. There's no reason for me to wait until I finish executing that action to ask my model for a different action. I can do it as fast as inference, essentially. And so maybe when I only have 50 milliseconds of action worth left, I can ask for the next sets of action. And when the current 50 millisecond is over, I have something that's ready for me to continue with my next 100 milliseconds. So that's one of the insights, the other kind of algorithmic improvement. We refer to them as real-time chunking. Desire inference in such a way that you know there's going to be a delay in how long it takes to query the model on the cloud, basically. Like the problem here, if I get a little bit more technical, is an action chunk is a sequence of action that I can execute on the robot. So, you know, it's not just one action. Und wenn ich eine ActionChunk, die ich kann für 100 ms. Und 50 ms in, ich will eine andere ActionChunk. Und ich will das neue ActionChunk, nachdem ich das 50 ms. ist es über. Wie soll ich das zwei tun? Wie soll ich das, wenn ich das so, dass ich das nächste ActionChunk, die nächste ActionChunk wird, wird es weitergehen, das so, wie das so. Du kannst pre-compute. Ja, du kannst pre-compute. Und das ist eine der Algorithmus-Impfunkungen dass wir mit den Infernfern aus dem Modell in den Cloud können. Ich studiere Computer Engineering, so ich bin nicht wirklich ein Algorithms-Person. Aber wenn es zu Systeme gibt, wie Pipeline, das ist gut, das ist so interessant. Das ist so interessant. Es ist eine Brilliant-Schutz, weil es so vieles für die Robots ist. Man braucht all diese Clunky, die zwei Operating Systeme sind, für Robots, die Embedded-Artaus, und dann die regularer und all diese complex giant compute und power. Und das ist was die initiale Versions von Waymo, die wir eigentlichen haben, um die Server auf den Trunk. Und du kannst du das mit General Day Robotics, which ist toll, dass du das so gut wie es. Ja, du musst du. Ich kann es tun. Some of es, obviously, hat es zu sein, dass es ein Compute da ist, aber ein paar von den Compute kann passieren. Und dann ist es, da muss ich ein Video, das Ding, das wir in der Top-Left, wie wie das, wie das, Video feedback, how much of it is local processed. Is there any computer locally on this robot, or is it just a dumb video camera that streams data to the cloud? For this, I'm not 100% sure, but I am inclined to believe that it's just a dumb computer. For this specific video, I don't remember, but I'm just 100% confident that we can make this work with a dumb computer and a robot. One other interesting thing about our collaboration with Weave and Ultra ist, dass ich nicht der Robot in person gesehen habe. Oh, wow. Ich habe sehr wenig idea über wie der Robot wirklich funktioniert. Interesting. Das ist ein sehr Intentional-Chall. Ich möchte, dass das so schnell wie möglich. Ich also nicht wie sie collecten Daten. Ich nicht, dass sie diese Frage nicht. Ich verstehe, dass es möglich ist, für ein Organisation wie Pi zu parachutieren in der System und zu arbeiten wirklich mit ihnen auf die Dinge, die eigentlich zu machen, um die System zu arbeiten und nicht zu lernen, wie sie auf der System setzen. Denn in einem Weise ist das eine mehr Scalable Recipe. Ja, du komplett decouple die Hardware Control Loop von den Semantiken und Planning, das funktioniert. Das ist wirklich toll. Ja, ich bin wirklich froh. Es funktioniert. Und wenn wir die Firma angefangen, Wir haben das Real Deployment, das ist nur in der Conversation, fünf Jahre um die Leben, um, in der Welt der Firma zu verändern. Weil die Problematik ist, ist wirklich hart. Und wir haben zwei Jahre in, und das ist die Resulte, dass wir haben. Und Real Deployment und Skilling, die number of Robots, ist ein wirklich sehr vieles Consideration heute. Und so die Prozesse der Prozesse hat sich sehr viel, sehr viel, viel schneller als wir dachten, davor. Often, auf diesem Podcast, wir sprechen über, wie, was das all bedeutet für Startup Founders bedeutet. Ich finde das könnte eine 마음에 Let me provide a few more contexts. The first is that robotic is traditionally really hard because it's an extremely vertically integrated business. You need to have your own customer relationship, your own hardware, your autonomy stack, your own safety certification, your own everything. And the barrier to entry is just really high because of that. And one of the things that we're trying to change is that we're trying to provide a foundation of physical intelligence that the community can build on top of that allow them to onboard autonomy onto their robot and their task much quicker than before. So that's the first, you know, we want to provide that kind of seat of intelligence that allow people to move much faster so that they can focus on other problems. The second thing is that I think the recipe for starting a vertical robotic business today is one, have a really good understanding of the existing workflow because the robotic system needs to fit into the same workflow. And the second is to be very meticulous about identifying where the opportunity is. You know, if there's a workflow that needs X number of work today, you know, where is the robot when you insert it's going to make the biggest difference. And two is to really be scrappy when it comes to hardware and data collections. You don't need an incredibly expensive robot that is capable of very precise motion today to be able to do these tasks. Und die Grund ist, dass diese Modelle wirklich Reaktiv sind, sodass sie die Inaccuracy in der actualen Robotmove. Und das zu ermöglichen, dass du die Abilität zu collectieren und zu runnen, insbesondere in realen Deploymenten. Die nächste Schritt nach dem ist, dass ein Mixed Autonomie System, das ermöglicht du, dass es brecheven ist. Das ist brecheven, economically. Das ist brecheven, economically. Der erste Punkt ist Richtig, weil sie này dann sind, Durch den ein Ebene für startupsfl Dlatego hat. What is the upfront cost? The upfront cost is much cheaper hardware, ability to collect data, ability to collect evaluation, and ability to understand the use case to see where they should insert the robot. It's not about having incredibly expensive hardware. It's not about having your own proprietary autonomy, classical stack anymore to be able to do this task. So it allows companies to focus on the component that will actually allow them to differentiate themselves from the rest of the space. Now that you sort of unbundled it and you no longer need to build this fully vertically integrated company in order to build a robotics company, are we on the precipice of a Cambrian explosion of vertical robotics companies where there's going to be like a thousand companies like Ultra going after, you know, every like menial job in the economy and like getting a deep understanding of the customer, building a robot that can solve that problem, doing mixed human machine deployment until it can run fully autonomously, and building a company in every sector. Is that the future that you see people building on top of Pi? It's funny that you mentioned Cambrian Explosion, because when we wrote this blog post, there was that term that was very hotly debated. We are, I think, academics at Hurt, and we want to be very major when we communicate. But, you know, myself personally, I believe there's going to be a Cambrian explosion of robotic company across the entire world and across many, many different verticals. Just because it's just so much cheaper to build. And it doesn't require, you know, someone with 20 years of experience in robotics to start anymore. Es braucht jemanden, der ist wirklich scrappig, kann man schnell, kann man den System integriert, kann man verstehen, was sie wollen, um zu starten, was sie zu starten. Ich meine, was kommt für mich ist, wir arbeiten mit einigen Robotics-Companien und haben viele Founder-Einrichtungen. Es fühlt sich so, dass es eine Art, eine Art, an die Anlage von Personal Computern ist. You could argue that industrial robotics today is basically like mainframe for mini computer level. Like, you know, if you look back in the 70s, huge public companies like Digital Computer that, you know, just did like these sort of very, very expensive deployments. But like they were very, very specialized and it was all extreme enterprise. Like, you know, the idea of a personal computer was ridiculous, right? Es hat die Altair, Apple I und Apple II und IBM PCXT zu kreieren, Personal Computing. Und dann die Tradition advice für Robotics für viele Jahre ist, go after Dirty und Dangerous. Und dann, of course, das sind die Industrieal-Case. Du hast diese Giant Tesla-Robots in den Gigafactory und Dinge wie das. Es ist wie was du gesagt, Profitabell ist wirklich, wirklich groß. So, you know, does that mean that the people who do the vertical robot Cambrian explosion sort of moment, the people who are sort of first in that, like, it sounds like they would be the first to be profitable and not dirty and dangerous? I think this is already happening today. I think we have the fortune of having lots of visibility into the robotic community because people would like to talk to us, people would like to learn what it's like to build a foundation model for Robotic and people would like to know, how do I get the same level of autonomy? And there's so many companies and businesses that we talk to that would love to put a robot into their space that it's okay for the robot to make a mistake and they just need it so much. I really believe that the recipe that I mentioned earlier of identifying where the robot can fit in, focus on cheaper hardware, collect data, run evaluation, mixed autonomy, break-even scale robots, will work across many different verticals. And I'm seeing it play out today, and it's just incredibly exciting to see. And this is pretty cool that you literally just gave people the playbook for how to build a vertical robotics company. Like, this is a playbook that could possibly be followed successfully hundreds or thousands of times. Und die Grunde ich möchte, dass ich es, ich möchte, dass die Cumbrianen sehen. Und wir wollen es helfen, dass es. Für Pi, wenn wir über die Frage, warum Pi ist es zu verlieren, wird es wahrscheinlich sein, weil die Probleme ist, weil es zu schwer ist. Vielleicht wird es 50 Jahre mehr, um die Roboik-Problem zu lösen. Und nicht ein paar Jahre, fünf, zehn Jahre. Und so, wir wollen die Community, wir wollen die Prozesse, und das ist warum wir sehr open. Wir haben unsere Forschung, wir open source Pi 0 und Pi 05. Und Leute waren auch so shocked, wenn sie mich gefragt haben, ist es eine Unterschiede zwischen Pi 0 und Pi 05 wie du open source versus die Modell wie wir internally in Pi 0 und Pi 05 benutzen Und answer was actually no it the same Model The pre Model weights that you using that we Open Source is also the pre-trained Model weights that our researcher internally use for Pi Zero and Pi Zero. And so we really want to help accelerate progress in the community and to create that Cambrian Explosions. Yeah, that's very inspiring. I mean, I feel like everyone's sort of spending a lot of time in the digital world, Und es fühlt sich jetzt die Zeit zu starten über die Welt der Atoms zu denken. Und das ist die perfecte Mix. Wie du dich die Elektronen und in die Atoms zu tun, in der Atoms-Welt. Und ich denke über Dario Amadei's essay, All Watched Over by Machines of Loving Grace. Und wenn du wirklich über die perfecte Manifestation von der Atomkrieg, Es ist nicht wie, you know, perfect agents that look over you just like in the electronic world. Es ist, you know, eigentlich something a little bit more akin to what we're seeing here. Yeah, and this has really been our mission from the start is to create that Cumbrian explosion. And, you know, this is why we choose to focus on the model because we believe that is the bottleneck to just really make robot useful across many different tasks in the world. Und das ist auch wir auch auf cross-embodimenten. Success für uns ist nicht nur unser Modell auf den Robotern, die es wichtig ist. Der Surface-Area für Success ist eigentlich viel größer, das ist unser Modell, die wirklich wichtig ist, auf jemandem der Robotern. Vielleicht, dass wir nicht even wissen, was das Robotern ist, in einem Weg zu den End-Consumer. Können wir vielleicht ein bisschen über die Humans behind den Robotern hier reden? Wie ist die Firma? Wie sind die Co-Founder? Wie sind die Zusammenarbeit? Was haben Sie sich mit solchen Problemen gebracht? Manchmal die joke ich mache hier ist, dass die Menschen behind den Robots sind auch Robots. Also, nicht wirklich. Ja, so Pi ist ein sehr, ich würde sagen, ein traditionaler Unternehmen. Wir haben eine Lager-than-average-Foundings-Teams. Und wir haben uns sehr gut zusammengekommen, als wir die Robotic-Team an Google waren. Google. The RoboRX team at Google was, I think, a really, really great environment for seeing the sign of life and creating the relationships in the community that allow the RoboRX community and these advances to flourish. There is Locky, which we met when we were thinking about starting the company and has just been really instrumental in making sure that we're a good business. And There is Adnan, our hardware lead, that came over from Andro. And Adnan has a really difficult job because if you want to work on cross embodiment, you remember my joke about how if you want to add two years to your grad school, you bring on one more robot. The hardware problem and the operational problem for us is how do we build, improve, and scale a fleet of heterogeneous robots. It's just not one robot platform. und weil wir die Organisation von Beginn aus dem Beginn zu unterstützen, ich glaube, wir können das aber es ist ein wirklich hart Problem. Weil es nur zwei verschiedene Roboter in der Fleet gibt, wie man sich alles gut macht, alles gut funktioniert. Wir sind wirklich gut an Divide & Conquer, wenn du fragen. Aber wie viele Co-Founders sind in total? Wir haben Brian, wir haben Chelsea, Sergey, myself, Laki, und Adnan. P scusaires zu haben, dass wir sicherlich keine Co- conviction haben, wie es ein Problem dasselbe? Oder ist das Ärger weil das eine Unterinstall Acqu Entwicklung kennengend ist? Whatever du damals mit Fire frontstellst werden erwartet? Ja, eine kleine Frage von Leuten ist, warum arbeiten wir zusammen? Übrigens, wir unterstützen uns ein Unternehmen der Schule. Wir spenden einhalb von viel Zeit im Beruf at Work und in einem anderen Fall in manches Gesicht induegal. Und so sch ningún Ziel wir wichtige Leiter ע liest. And the second is that any one of us could have started a company and be successful. But the problem is just so incredibly hard. And the chances of success is just so much higher that we band together and we can divide and conquer the problems. And that's, I think, one of the main reasons why the progress has been much faster than we expected. What were the differences of you working before in either academia or a big industry, big company like Google and as opposed to now in a startup? This is the first time for a lot of you doing a startup, right? Yeah, this is the first time for a lot of us. One of the really surprising things that we learned when we started the company is that the infrastructure for supporting large-scale general-purpose robots were just not there. And, you know, this starts from the software itself. How do you collect data? What device do you use to collect data? How do you manage the data? How do you annotate the data? How do you get visibility into the data? How do you run evaluation? How do you build operational process? Like, there wasn't company that offered this kind of services, which is very different from software, and we were really surprised to find out. And so we end up writing a lot of the software at Pi ourselves, I think this is another area of incredible opportunity of kind of building services for robot company. Like, you know, if you can offer remote tele-offer, for example, if you can offer data collections, if you can offer annotation service, because, you know, these are functions that doesn't need to be repeated from one company to the next. So I think there's lots of opportunity to build kind of support for growing robotic business. So that's one thing, like one surprising thing that I learned. And the second is, I think one of the reasons why we have managed to achieve such progress is that there is a really tight loop of collaboration in the entire lifecycle of model development. Going from what task do you collect data for? You collect data for the task, how do you do it? What hardware do you use? Once after you collect the data, how do you get visibility? How do you ensure data quality? How do you then make sure that you can easily train on that data? After you train on that, how do you run evaluation? Evaluation is a really hard problem in robotics because it scales super linearly to model capability. Let's say you have a model that can perform a two-minute task. Running evaluation for that is very different from running evaluation for a task that's 20 minutes. It's not 10 times harder. It's more than 10 times harder. After you run evaluation, how do you can discover und meine Studien ακ precedent darétat von den So st squeeze! der im Herzen einen-겠지만33kt clovesprojekt and curls dedi기�acoatr solic 커 argues ist eine Software-Rot 36 September Actor esempio auf dem Bez MILLIONSe gör ich hier also ist es schwierig und UnsaciFuffle ... Baby ist es gewesen an hunting ¡ Falte Analyze Fehler modes, understanding, oh, is the robot performing this way because of the data that was collected or the way that it was annotated or the way that we train the model? And then suggest ideas and actually try them to figure out if those hypotheses are correct. So that's something that I would love to have and would dramatically unlock us. So sometimes I make the joke in the company that we should record all of the meetings and then train a model to basically just make prediction about what is the next set of experiment. Oh, you could. You totally could. Was ist OpenClaw und Obsidian und Markdown files und, you know, a Brain.md mit Ontologie custom zu deinem Use Case? Und was ist 100 OpenClaws in den Backgrounden, die du orchestratieren? Ich denke, dass zwei Sides sind. Die erste ist, dass wir ein bisschen ein Sides-Life haben, wo für einfach Fehler-Modet, während der Evaluation, wenn du das Wort beschreiben, die Art der Robotnäu in Text sehr, sehr und sehr, sehr, start – then you can ask the language model to make very reasonable deinen Moreover recommendation about what the next step is. But, the flip side is this only work for simple cases today. And the reason why that's the case is because I think it's pretty fundamental limitation of the model that we have today, which is that they're not at the core model that takes action in a world and sees the consequences of it's own action, especially action that really changed the physical world, and so I think this kind of very fundamental understanding about how the physical world works is missing from the really large foundation model. I think that's one of the ingredient that's missing, to be able to build this automated robot-research advocacy, how can Klaue hearer Name decir, Liebe者 Siemens, Amy Luway can just offer you things, which is interesting. At that point, it's on the research lab to provide CLI, MCP, endpoints to the things that might control robots or reconfigure rooms. I think Karpathy feels like he's starting to talk a bunch about this, where if you mix auto-research plus what he's been talking about with markdown files, it might just happen in the open. There's a sort of sense that you have to make something much, much more complicated to make it work. But what if that's just wrong? What if we just have Markdown files and agents and you could make it yourself with literally Claude Code and MCP today? What if it's not an algorithm problem? It's just literally an integration challenge. We have a version of this internally that I use a lot. Das war ein Punkt, wenn ich ein ganzes Mal an viel Geld an API Quere. Und mein Team war, Quaden, was du? Oh, ich bin das Typ bei Y Combinator. So, ich will dir einen Beispiel. Wir haben eine Klotzkilde, das ist eigentlich die Rolle der Pre-Training-On-Call. Wir haben diese Pre-Training-Runs, das sind wirklich großartig. Es ist ein sehr, ich denke, ein sehr schwieriges exercise zu keepen sie alive, für sie zu continue zu churn, weil es so viele Dinge mögliche zu tun. Und wir haben, ein Prototyp, ein Pre-Training On-Call, das kind of babysitzt die Run, und haben die Permission, zu take action, zu remedy Error, dass sie sie sehen. Und eine der Überraschung der Ausstellung von der Übersetzung ist, dass es über 50% improvement in compute usage, Das ist einfach für das Lashby-Training, das ist groß für uns. Und das ist einfach nur ein kleines, kurzer Prototypen. Ich glaube, es ist viel mehr zu tun. Kwon, das ist unglaublich. Vielen Dank für alles. Vielen Dank für das Physikal-Intelligence. Vielen Dank für die Showen uns diese incredible Demos. Honestly, das meiste, die mir die meisten Hoffnung ist, ist das Idee, dass es ein Entity, ein Research Lab, das ist auf der Welt geboten, about to create this Cambrian explosion of robotics startups. So someone watching right now will be inspired by this and start playing with your models and they might create a robot that touches billions of people's lives for the good. Thank you for having me. It's been a pleasure. To the listeners, the one takeaway that I want you to have is I think robotics has changed a lot and the cost of building in robotics has decreased ist, und auch weiter 나온 thực hiện eine sehr Kommitte, saturated der Danke.",
  "transcript_chars": 45982,
  "transcript_filled_at": "2026-05-25T02:22:47.362590+00:00",
  "transcript_filled_by": "groq"
}