{
  "video_id": "ekgvWeHidJs",
  "channel_slug": "machinelearningstreettalk",
  "channel_handle": "machinelearningstreettalk",
  "title": "Type a Sentence, Get a Playable 3D World in 3 Seconds - Shlomi Fuchter & Jack Parker-Holder",
  "duration_seconds": 3503.0,
  "url": "https://www.youtube.com/watch?v=ekgvWeHidJs",
  "upload_date": "",
  "transcript": "By the way, look at this dog. This is\namazing. This is insane.\nWhat was the prompt to create that?\nToday is a world exclusive of what is,\nin my opinion, the most mind-blowing\ntechnology I've ever seen and the most\npoggers I've ever been. You're not going\nto believe what Google DeepMind showed\nme in an exclusive demo in London last\nweek. This technology might be the next\ntrillion dollar business and might be\nthe killer use case for virtual reality.\nGoogle DeepMind has been slaying so hard\nrecently that even Gemini Deepth think\ncan't count the number of wins in the\ncontext window. Let me explain.\nToday we're going to talk about a new\nclass of AI models which are called\ngenerative interactive environments.\nThey're not quite like traditional game\nengines or simulators or even generative\nvideo models like VO, but they do have\ncharacteristics of all three. They're\nbasically a world model and video\ngenerator which is interactive. You can\nhook up a game controller or any kind of\ncontroller for that matter. Deep mind\nsay that a world model is a system that\ncan simulate the dynamics of an\nenvironment.\nThe consistency is emergent. There is\nnothing explicit. the more the model\ndoesn't create any explicit 3D\nrepresentation.\nHow do you square the circle between\nlike a stochcastic\num neural network and yet it has\nconsistency? Right? So I look over here,\nI look back, I look there again, the\nthing is back like isn't it a bit weird\nthat a subsey symbolic stochcastic model\ncan give us apparently consistent like\nsolid maps of the world? Do you remember\nthe quake engine in 1996? It required\nexplicit programming of the physics and\nrules and interactions. But this new\ngeneration of AI systems learn real\nworld dynamics directly from video data.\nYou can control an agent in the world in\nreal time. The move towards generative\nworld models was born from the\nlimitations of handcoded simulators.\nEven their most advanced platform XLAND\nwhich was designed for general agent\ntraining. It was the frontier for\nembodied agent training with curriculum\nlearning. But it felt far from the real\nworld. It was almost cartoonike. It\ncould model 25 billion tasks, but it was\nstill handcrafted. It was constrained to\nthe rules of that particular domain. And\nit was janky. Imagine if you could just\ngenerate any interactive world you\nwanted to train your agents on with a\nsimple prompt. Now, cast your minds back\nto last year when I interviewed Ashley\nEdwards at ICML. This was the first\nversion of Genie, which was trained on\n30,000 hours of 2D platformer game\nrecordings.\nWhen we're generating next frames, the\nobjects that are further away are moving\nmore slowly than objects that are\ncloser. And this is a sort of effect\nthat you would often see in games such\nthat you can kind of uh simulate depth.\nIt's something that we also have it, you\nknow, like when we observe things\nmoving, we see things moving slowly when\nthey're further away. So yeah, the model\nlearned that\njust being able to be that good at\nunderstanding the physical world was not\nsomething uh we were expecting it to be\nthat good at that quickly.\nThe core innovation of Genie 1 was a\nspatial temporal video tokenizer that\nconverts raw footage into processable\ntokens and a latent action model that\ndiscovered meaningful controls without\nlabel data and an auto reggressive\ndynamics model which predicted future\nstates. The latent action model, a form\nof unsupervised action learning, was the\ncore innovation. Genie discovered eight\ndiscrete actions which remained\nconsistent across different environments\npurely by analyzing frametoframe changes\nin game recordings. This means it knew\nwhat jump meant or what move left meant\nwithout being explicitly trained on\nthose actions. This was an OMG moment\nfor me. I mean, how was that even\npossible from training on offline game\nepisodes? Even more surprising was how\nit seemed to have emergent capabilities\nlike 2.5D parallax. Just 10 months\nlater, Genie2 arrived with 3D\ncapabilities and near realtime\nperformance. The visual fidelity was\nmuch higher. Now, it can simulate\nrealistic lighting like the Unreal\nEngine. you know, things like smoke,\nfire, water, gravity, pretty much\nanything you might see in a real game.\nIt even had a reliable memory. You know,\nyou could look away from something and\nbring it back into view and it would\nremember the thing. This is Gigachad\nJack Parker Holder. He's a research\nscientist at Google Deep Mind in the\nopen-endedness team talking about Genie2\nwith Demis, no less.\nThis is a photograph taken by someone in\nour team somewhere in California. And\nwhat we then do is ask Genie to convert\nthis into an interactive world. So we\nprompt the model with this image and\nGenie converts it into a game-like world\nthat you can then interact in. Every\nfurther pixel is generated by a\ngenerative AI model.\nSo the AI is making up this scene as it\ngoes along.\nExactly. Yes. Someone from our team is\nactually playing this. They're pressing\nthe W key to move forwards and then from\nthat point onwards, every subsequent\nframe is generated by the AI. Around the\nsame time last year, uh you'll probably\nremember this by the way, Deep Mind's\nIsrael team led by Schlomi Frutter\nshowed diffusion models simulating the\nDoom engine. The system was called game\nengine. It's almost a meme at this point\nhow Doom runs on, you know, calculators\nand toasters. But here is a neural\nnetwork confabulating a Doom game frame\nby frame in real time. Like look at how\nit just knows what the health is. You\ncan shoot characters. You can open doors\nand navigate around maps. You know,\noccasionally it was slightly glitchy,\nbut this is just unreal. You know, you\ncan just simulate Doom at 25 frames a\nsecond on a single TPU. The only\nlimitation, of course, was that it could\nonly do Doom and nothing else. So, last\nweek we waltzed our way into London and\nJack and Schlomy gave us a demo of Genie\n3. Honestly, I couldn't believe what I\nwas seeing. The resolution is now 720p\nwhich is firmly in the good enough\nterritory to suspend disbelief. It's\nreal time. It can simulate real world\nphotorealistic experiences which can\ncontinue for several minutes before\nrunning out of context. Schlommy had his\nhands all over VO3 by the way and they\nseem to have combined elements of the\nGenie architecture with VO producing\nsomething I can only describe as VO on\nsteroids. Unlike Genie 1 and 2, the\ninput is now a text prompt, not an\nimage, which they argued is a good thing\nfrom a flexibility perspective, but it\ndoes mean that you can no longer take a\nphoto of a real place and generate from\nthere. One of the main features of Genie\n3 is that it has a diversity of\nenvironments, a long horizon, and\npromptable world events. Now, on the\nworld events, let's take this ski slope\nexample. We might type in another skier\nappears wearing a Genie3 t-shirt or a\ndeer runs down the slope and there you\nare. Things just happen in the world.\nThey say that this might be very helpful\nfor modeling things like self-driving\ncars where you can simulate rare events.\nBut I was left thinking that this is\njust turtles all the way down. How can\nwe write a process to prompt the\npotentially infinite number of rare\nthings which could happen in a scene?\nThere was an example they showed of\nflying around a lake and it was amazing.\nBut I was like thinking where are the\nbirds mate? Like can can you can you\ntype the birds into the prompt? The team\nbelieves that we haven't yet had the\nmove 37 moment for embodied agents. You\nknow where an agent discovers a novel\nreal world strategy. They see Genie 3 as\nthe key to enabling that. But the real\nworld constantly surprises us because\nthe real world is creative. Creativity\nsimply means that the tree of things\nwhich can happen keeps growing new\nbranches and leaves just keep appearing.\nPerhaps in the future we might have an\nouter loop which makes the system more\nopen-ended. But right now in my opinion\nGenie 3, like all AI gives you exactly\nwhat you ask for in the prompts and\nisn't creative on its own. Currently,\nthe system only supports a single agent\nexperience, but imagine how cool it\nwould be if you could extend that to a\nmulti- aent system. Apparently, they are\nworking on that. I mean, personally, I'm\nmost excited about a new modality of\ninteractive entertainment. You know,\njust imagine YouTube version two. Deep\nMind sees the main use case of being\nable to train robotic simulations as\nbeing the real gamecher. This seems\nplausible to I mean like the miracle of\nhuman cognition or in brains is that we\nhave evolved to simulate the world\nwithout direct physical experience which\nis expensive. This is basically the same\nidea right why train in the real world\nif we can just simulate any possible\nscenario in a computer just like that\nblack mirror episode. Here's a couple of\nexamples they gave of using simulated\nenvironments to train an agent to do\nsome specific language tasks. Now, with\nGenie2, they said they were happy if it\nwas consistent even for 20 seconds. But\nnow, when you notice something\ninaccurate, it's very surprising. The\nkey thing is that it now extends beyond\nthe prediction horizon of the average\nhuman, and the glitches are getting\nharder and harder to spot. They said\nthat Genie 2 wasn't actually real time.\nYou had to wait a few seconds between\ntaking different actions. you know, it\nwas low resolution, had limited memory,\nyou know, I mean, it was superficially\nreally good, but it didn't look\nparticularly photo realistic. Genie 3\nchanges all of that. So, Genie supported\naround 10 seconds of generation. Genie 2\naround 20 seconds. Genie 3 is able to\nsimulate interactive environments for\nmultiple minutes. This time around, they\nwere a little bit more tight- lipped\naround the architecture. They wanted to\nfocus on capabilities in the interview,\nand that's fair enough. I mean, it's\nunderstandable given that this is\npotentially a trillion dollar business\nand Zuck will be sniffing around like a\ntruffle hound. My my biggest concern\nwith this is that as soon as Zuck gets\nwind of this, he is going to be getting\nout his checkbook. He's going to go\nstraight to Jack and Schlomi and he's\ngoing to be like, \"Come on, boys. $100\nmillion. Come to work for me.\" Um, Zuck,\nmate. Seriously, no. Don't do it. These\nguys, they they're doing God's work over\nhere. You need to just let them let them\ndo what they're doing. You can make it\nyourself if you want, Zach. Leave them\nalone. I should say I did joke at the\nend of the interview that if you are\nlearning Unreal Engine right now, you\nmight want to pivot to a different\ncareer. But the Google guys were quite\ngrounded. They argued that this is a\ndifferent type of technology. There are\npros and cons, you know, which is fair.\nI should stress that as amazing as this\ntechnology is, it's still a neural\nnetwork and it still has many important\nlimitations. Certainly though, just\nimagine how easily you could generate an\ninteractive motion graphics with this\ntechnology. You know, that's something\nthat Unreal Engine has been leaning hard\ntowards in version 5.6. So, do I need to\nfire my motion graphics designers?\nVictoria,\nwill users be able to use this? Not\nanytime soon. This is still a research\nprototype and given the obvious safety\nconcerns, they're going to open this up\nprogressively through their testing\nprogram. One question did come up in the\npress conference yesterday though like\ncould it generate an ancient battle and\nSchlomi said that it's not trained on\nthat kind of data wouldn't be able to do\nthat yet. So I mean certainly not a\nspecific historical battle anyway. So it\ndoes sound like there are still some\nlimitations. How can a system like this\never be fully reliable? Well they did\nsay that with better models the trend is\nthat they get more and more accurate.\nThe glitches become fewer and they\nexpect to see further improvements. You\nknow, there's this annoying phrase like\nthis is the worst the model will ever\nbe, but even as I said, they can\ngenerate some edge cases using a whole\nbunch of prompt augmentations, but it\nmight just be turtles all the way down.\nYou know, how do you come up with all of\nthe rare black swan events that might\nhappen? So, what data was it trained on?\nThey were quite cy about this as well.\nIt's probably safe to assume that it's\nbeen trained on all of YouTube and lots\nmore besides that.\nHow much compute does this thing need?\nWell, I asked them that and they were a\nlittle bit vague about it. Um, they said\nthat it ran on their TPU network. So,\nI'm inferring from that it needs a crap\nton of compute. However, I can say that\nit was demoed in front of me. It was\nvery responsive. You put a prompt in, it\nthinks for about 3 seconds, and then\nyou're just in and it just works. They\nalso mentioned some cool stuff about\nhow, you know, like Genie can be used to\ntrain agents as we said, but the agents\nthemselves could be used to better train\nGenie 3, creating this virtuous cycle of\niterative improvement. If you're in a\nworld walking around and say you go to\ncross the street, you sort of check the\ncues of the of the drivers, for example,\nmaybe there's not a crosswalk um and you\nneed to know when to stop. You can see\nthat they're slowing down, so that's\nwhen you would go and the other agents\nshould be simulated in that fashion.\nGen3 and other similar models would be\nimpossible without at least some human\nfeedback in the training loop or the\ndata curation or the evaluation.\nProlific is a human data platform and\nthey are sponsoring this video today.\nMy name is Enzo. Um I uh I work at\nProlific. I'm the VP of data and AI. I\nsupport everything from AI data research\nuh and the likes. Um, for those\nunfamiliar, Prolific is a human data\nplatform for working with everything\nfrom academic researchers, but also\nsmall and large players in the AI\nindustry. Visit prolific.com.\nYes, this is the demo where they've got\nG3 memory test on a blackboard. you see\nthere's like an apple and a cup and then\nyou kind of like go out you you look out\nthe um the window you see there's like a\nfew cars and the purpose of this test is\nto kind of say they've got such a long\ncontext window you know similar concept\nto a large language model that it still\nremembers all of the things that it\ngenerated you know even if it's minutes\nago we've got the the blackboard over\nhere we look up and there it is it\nremembered it gen3 memory test I've also\nnoticed that this model is is even\nbetter than V3 at things like text. It's\nbecause you you would think that they\nwould they would dumb down the model to\nmake it interactive and to make it this\nsophisticated. But even as a video\ngenerator model, it seems almost better\nthan V3 for doing a whole bunch of\nstuff.\nAll right. So, I'm Shomi Fer. I'm a\nresearch director at Google DeepMind.\nI'm the Veo Collid. Um and um I'm\nbasically working at Google for about 11\nyears. um recently on diffusion models\nin the various modalities, image uh\nvideo and we'll tell you more about what\nwe're working on right now.\nHey, I'm Jack Parker Holder. I'm a\nresearch scientist at Google Deep Mind\nin the Open Endlessness team. Um\noriginally working on open-ended\nlearning and open-endedness and more\nrecently working on world models. We are\nhere at Google Deep Mind in London and\nyou guys have just demoed to me\nsomething which I think I'm more\nimpressed with this I think than\nanything I've seen probably ever before.\nI think it's a paradigm changing moment.\nShomi, can can you can you tell us a\nlittle bit about this new version of\nGenie?\nSure. So, Genie uh is our most capable\nworld model. And by a word model, what\nwe mean is basically a model that is\nable to um predict how an environment\nwould evolve and also how different\nactions of an agent would affect this\nenvironment. So with um Genie3 we are\nable to basically push the capabilities\nof a world model um to a new frontier um\nthat means high resolution much longer\nhorizon and better consistency and all\nthat real time in real time basically\nallowing wherever if whether it's an\nagent or a person that interacts with\nthe system to walk around it navigate it\naffect it while the generation happens\nin real time.\nGenie 3 is just ridiculous, right? It's\njust on a completely different level.\nBut maybe we should just contextualize\nthat around Genie2. So what was Genie 2?\nYeah, it's a great question. Um, so\nGenie2 was sort of the culmination of\ntwo years of research in what was quite\na new area which is foundation world\nmodels we called it at the time. So\nessentially in the past world models had\nmodeled a single environment. So the\ncanonical world models paper in 2018\nfrom David Haren Huber modeled the car\nracing environment uh which is a major\nenvironment and it could just model that\none environment. It could predict the\nnext states given any actions in that\none world. We've seen with the dreamer\nseries um also from from Google mind\ndagar uh with Atari games and other kind\nof environments as well but no one had\never done something that could create\nnew worlds. So with Genie 1, the real\nnovelty there was that we had a model\nthat for the first time could be\nprompted to create completely new worlds\nthat didn't previously exist. But that\nbeing said, they were fairly\nrudimentary. Like they were um low\nresolution. You could only play with it\nfor a couple of seconds. Um so agents\ncouldn't really learn long horizon\nbehaviors that wanted to and the\ndiversity was still fairly constrained\nand it required some form of image\nprompting. With Genie2, we really pushed\nthat um um to the next level, right? So\nwe trained it on a much light larger\ndistribution of 3D environments. We\nmoved to 360p from I think 90p before.\nSo it was it was more like close to what\nwe see now, but it was still sort of\nscratching the surface because we didn't\nreally know that this approach could\nscale the way we've seen other methods\nhave. So we wanted to really like test\nthis from a research standpoint. Um but\nthen I think for this year we wanted to\nreally take that to the to the next\nlevel. Uh and that's what we think we've\ndone.\nYes. And it's now 720p. It's\ninteractive. So Genie 2 wasn't wasn't\ninteractive. It was it wasn't fast\nenough. And you know, Steve Jobs said\nthere's something magic about the\ntouchcreen, right? There's something\nmagic about it. And of course, the magic\nhappens when it's interactive. And some\nof the demos you showed me were just\njust insane. Right. So um photo\nrealistic I mean it's kind of like a\nfusion of VO I suppose that you can now\nunderstand the real world and you can\nbuild essentially a foundation model for\nthe real world which is interactive\nthat's mind-blowing and and just tell me\nabout some of the examples you showed\nyeah so I think what you said about veil\nor more generally about video models is\nright like we there is a way we can\nthink about them as somewhat a world\nmodel but it's not really it doesn't\nallow us to actually um navigate or\ninteract with it um completely\ninteractively. And I think that's that's\none of the limitations of of video\nmodels that with Genie Free we're trying\nto address. Um and basically in the\nexamples that you've seen um we're able\nto because Genifreeze generates the\nexperience and what the what we see\nframe by frame. Um, it lets the user or\nthe agent that is using it basically\ncontrol where it wants to go in every\nlike in a very low latency that allows\num uh basically exploring the\nenvironment and and creating new\ntrajectories that are not predefined\nlike video models. Um so in the examples\nthat you've seen for example um you can\nsee the character or the agent in this\nvideo moving around maybe going back to\nthe place there already been in uh\nbefore and everything remains consistent\nand I think that's a very remarkable\nproperty or capability of the model the\nability to preserve and the consistency\nof the environment along very long\ntrajectories.\nYes. and and even Genie2 had some kind\nof object permanence and consistency,\nbut nowhere near as much as we have now,\nbut we'll come back to that in a second.\nWe can't say too much about the\narchitecture for Genie 3. But in in\nGenie2, there was a um an ST\ntransformer, so a special temporal\ntransformer, which was conceptually\nquite similar to like a VIT, and there\nwas um a latent action model, which\nmeans even from like, you know,\nnon-interactive data, you could infer\nsome low cardality action space, and\nthen those went into a dynamics model. I\nI think we can what we can say about the\narchitecture that you know might be\ninteresting is that definitely because\nof the interactive nature of the problem\nor the setup then the model is not\nregressive. So what it means that it\nmeans that um the model generates frame\nby frame and has to refer back to\neverything that happened before right.\nSo if for example uh we're walking\naround some auditorium or what some some\nother environment um basically if we\nrever if we revisit an a place that we\nalready been to uh the model has to look\nback and and understand that this\ninformation has to be consistent with\nwhat's happening uh um in in the next\nframe. Um so I think the interesting the\ninteresting point here is that\neverything here like the consistency is\nemergent. There is nothing explicit. the\nmore the model doesn't create any\nexplicit 3D representation and I un\nunlike you know methods other methods\nlike nerves and goian splatting um so I\nthink that's that emergence kind like\ncapabilities very interesting and\nsurprising for us yes\nyes and and even G2 had emerging\ncapabilities like parallax and you know\nit could model certain forms of lighting\nand so on but this just blows my mind\nyou're you're involved in that Doom\nsimulation last Yeah. And even that just\nblows my mind. Right. So, we all played\nDoom in 1993. It was one of John\nCarmarmac's finest. And now you're\nsaying that I mean certainly the work\nthat that you folks did last year.\nYou've got a neural network model which\nis subsemb. So there's no explicit model\nof the world. You don't know where the\ndoors are. You don't know where the\nlakes are, where the maps are, and so\non. You just kind of take a, you know, a\nsample, a traversal through this space,\nand and it just produces the game in\npixel space. I mean that's\nyeah you know this is really you know\nI've been playing games obviously\nincluding you know Doom and others and\num I also worked some on on game engine\ndevelopment at some point uh very early\nin my teens and I I think what what I\nreally like about this project is that\nwe now are now able to run models that\nactually generate consistent 3D\nenvironments as in you know game engine\nand the Doom simulation and they run on\nGPUs or TPUs uh while in the past we\nwere running you know this like uh game\nengines on the same hardware. So I think\nthere is it's really something very\ninteresting and it kind of closed this\ncircle for me. Um and in particular in\nthe case of game engine we try to push\num on the real time interactive kind\nlike aspect. So uh we basically said\nokay would uh a diffusion model be able\nto simulate a game environment end to\nend with nothing explicit no code\nnothing except for actually generating\nthe pixels and getting the inputs from\nthe user and you know we weren't sure if\nit's going to work so I think there is\nlike with this kind of research we try\nand it doesn't work and then it all of a\nsudden something happens and it we we we\nsee that it does work and that's a very\nrewarding moment and I think in this\ncase um once people saw Oh, I think this\nwas really even the the the reception of\nthat was a bit surprising because there\nis something about the real-time\ninteractive um capability that really\nsparks the imagination of oh I can\nactually walk into this environment\nmaybe generated environment and actually\nexperience it right so I think this was\na moment um that um later when I think\nabout it like we were kind of like um\nexcited about the the real-time nature\nof of the simulation and We really\nwanted to bring it to higher quality,\nmore general purpose uh simulations.\nSo Jack, I mean one of the um the\nmillion-dollar questions is, you know,\nlike even with a language model, it's\nit's stochcastically sampled, you know,\nwith this temperature parameter. Same\nthing here. I mean, with with Genie2,\nthe dynamics model is using this masked\ngit and it was it was run iteratively.\nAnd how do you square the circle between\nlike a stochastic\num neural network and yet it has\nconsistency? Right? So I look over here,\nI look back, I look there again, the\nthing is back. Like isn't it a bit weird\nthat a sub symbolic stochastic model can\ngive us apparently consistent like solid\nmaps of the world?\nThat's a really good question. Um I\nthink probably similar to language\nmodels that there are some fundamental\nthings about the world that you want to\nremain consistent. So with a language\nmodel um I think even though as you said\nthey can be like sarcastic models if\nthere is things that stated as facts in\ntheir in their context they will still\nprobably recall them correctly right\nwhereas new things are where they maybe\nhave more degrees of freedom to to\nchange things like that. So I'd imagine\nin a world like a genie generated world\nif you were to move around then maybe\nnew things would be have some degree of\nof um stocasticity to them right but\nthen once they've been seen once then\nthey should be consistent from that\npoint forward because the model knows\nwhen to use this stocasticity um and\nthis is kind of an emergent property\nfrom um the scale that we train at.\nYes. And we we'll save the emergent\ndiscussion. I was I was just telling the\nguys about my conversation with David\nKrakow the other day but maybe we won't\ngo there. Um but um the other really\ninteresting thing is so you know you\nsaid David Har you know 2018 with Schmid\nHuba the world models thing and um uh\nshow me in the presentation you defined\na world model as essentially being able\nto simulate the dynamics of something\nright if if a world model simulates the\ndynamics of of a system how could you\nfor example how could you measure that\nso I think it's very hard to exactly\nmeasure the quality of world models in\ngeneral and I I think when it comes\nespecially to visual um to visual\ngeneration if it's image models and and\nyou know generative models in in general\nvery difficult to to measure their\nquality because it's somewhat of fit is\nvery subjective right so I think for\nLLMs actually we're in a better place\nbecause we can measure their performance\nfirst of course there is perplexity just\na next token prediction problem but\nlater on we actually care about how they\noperate for for the task that we care\nabout right so we measure for example\ndownstream like performance on various\ntasks. But for when it comes to world\nmodels and today we focus mostly on the\nvisual aspect, right? So it's it's\nimportant to highlight that the world is\nmore than just visuals, right? Um but\nbut again uh for Genie we're focusing\nmore on that um because a lot is\ncaptured in the visual interaction um of\nthe world. Um so measuring how well a\nmodel is is doing really depends on the\ncontext and and also on how we want to\nuse it later. I think that's something\nwe have to keep in mind when we evaluate\nthings um in models. So we have in mind\none particular application that we think\nis really key and that's to be able to\nactually uh train and and let a AI\nagents interact with simulation\nenvironments. And I think that's\nsomething that um you know I'm coming\nmore from this kind like a kind like\nmaybe\nsimulation background but not so much\nfrom the um maybe training agents in\nsimulation environments that's not so\nmuch was wasn't my my original\nbackground but you know I think through\nthe interaction with you know other\npeople in deep mind that are kind like\nexploring that for a long time I was\nreally um you know over the last few\nyears I kind of came more and more to\nrealize how much potential there is in\nin that that Because if we really think\nabout it, AI would be limited by the\nability to perform physical experience\nexperiments, right? Because imagine that\nyou want to develop a new drug or a new\num collect treatment. Um you cannot\nreally do it in uh in the real world if\nit takes months for every step in the\nway right and the same we can think\nabout you know if we want to learn how\nto assemble something then again if I\nhave to train the robot in um in the\nreal world it might take very long. So\nthat's why the simulation of the real\nworld is really key and that's what we\nhope we kind like push a bit farther\nwith G3.\nYes. Very exciting. I spoke to a startup\nrecently and they they sketched out this\nfuture where we'll have um essentially\num a model platform where people doing\nrobotics can download the policies you\nknow so I'm I'm in a factory and I need\na policy for doing this particular\nthing. But of course, you know, they\nimagined that it's so scarce, it's so\ndifficult to get real world data that,\nyou know, there would be a marketplace\nand everyone would train their own\npolicies and they would like sell it to\nother people on the market. This is a\nslightly different vision. You're saying\nthat now we have a a world foundation\nmodel and essentially I could say, well,\nin this situation, I need to have a\nrobot policy for doing this particular\nthing. So, I can just spin off a job. I\ncan create the policy and and away we\ngo. So, is is that roughly correct?\nI think that is kind of the vision that\nwe have. So I think in robotics in\nparticular there's a lot of focus on\ndeploying robots in somewhat constrained\nsettings right so it might be for\nexample in someone's apartment that's\nvery staged right um almost as stage as\na podcast recording you know got like\nall these support staff watching around\nthis robot achie a achieve one goal\nright and from a control um perspective\nit might be very impressive but in terms\nof the stoasticity of the world that\nit's in it's very limited right um and\nif we look at simulated simulation\nenvironments they might accurately model\nphysics but they definitely don't model\nthings like weather or other agents or\nanimals or these kinds of things right\num whereas a model like Genie 3 because\nit's it has world knowledge that world\nknowledge extends beyond physics\nactually to also the behavior of other\nagents and as we showed you in that\nexample at the beginning with the the\nworld events that we can also inject\nright you can actually prompt to have\nyou know another agent crosses in front\nof you or like you know we had a herd of\ndeer run down the ski slope or something\nlike that and I think these are the kind\nof things that for robots to be deployed\nat large scale in the real world. The\nreal world is fundamentally populated by\npeople uh and other agents and this is\nsomething that we can gain from training\non this general purpose world model uh\nthat we just have no other approach I\nthink to scalably get this data in a\nsafe way as well right because the\nsafety is a critical element of this\nthat we can simulate things uh in a\nrealistic way without having to actually\num deploy agents in the real world.\nYes. And that was a very important\ndetail. So you can you can put a prompt\nevent in and you gave me an example of\nthere there's a skier going down the\nslope and then here's a guy with a\nGemini t-shirt. And I guess what I'm\nthinking about here is if we did train\nthese robot policies, we would need to\ndo probably some kind of curriculum\nlearning and some kind of diversity. So,\nyou know, we would start off with a\nsimple environment and then we'd add the\nguy with the Gemini t-shirt and then\nthere'd be a car coming along and maybe\nin reality there would be some kind of\nmeta process, you know, creating some\ngradient of complexity and, you know,\ndiversifying environments. I love that\nKen Stanley paper, you know, the poet\npaper doing something like that. But is\nis is that like a fairly reasonable\nintuition?\nSo I think it's still early to say\nexactly how u word models um like\ngenifree will be actually used for AI\nresearch. I think we can only kind of\ndirectionally say um we I in general I\nthink we we still um we see it in also\nin other generative\nmodels that there are some capabilities\nthat we actually discover right and we\ndon't necessarily know that they're\nthere and then through the interaction\ndevelopment we're actually seeing them\nemerge for example you know we we just\nrecent like a few days ago we then we we\nkind like shared that um you can write\nlike some text on a on a photo and\nprovide it to val and it just like it\nreads the the the text and it follows\nalso the spatial instructions right and\nI think that's that's for example\nsomething that we didn't necessarily\nexplicitly train them all to do but it\nit's capable of doing and I think here\nas well the capabilities of Genie free\nuh that we're exploring are still uh we\nstill discover new things and I think\nthat's something that we hope um that by\num first by having more like you know\ntesters and external testers that we\nalready shared some some uh uh like we\nbasically previewed the model to and\ngive us feedback. So we hope that\nthrough this kind like engagement with\nthe community um and we can better see\nhow those models will be useful um and\nthat's something that I expect to take\nsome time um as we we basically try and\nunderstand the best application. You\nknow, I'm a huge fan of open-endedness\nfor for example, and um certainly at the\nmoment when we prompt models, if we're\nquite generic, you know, in in in what\nwe put in the prompt, then we tend to\nget quite simplistic answers. So, you\nknow, a lot of um people doing computer\ngraphics when they prompt image models,\nthey they have so much specificity and\nthey deliberately take it, you know, on\nonto the tail of the distribution so\nthey get something that's novel and\ninteresting and so on. and and the real\nworld just always produces a sequence of\nartifacts which are novel and\ninteresting, you know, like you get\nrandom NPCs walk onto the onto the\nscreen and cars go in and so on. And is\nmy intuition correct that that at the\nmoment um as good as it is with with um\nwith Genie 3, you you tend to get quite\na specific scene and you don't have like\nrandom kind of planes flying over and\nsort of, you know, just random things\nhappening.\nYeah. So that's a really a good\nintuition, right? So it definitely is\nthe case that um the model is very it's\nvery aligned with the text prompt that\nthat it's given. So therefore there is a\nlot of emphasis placed on the quality of\nthe text prompt to describe the scene.\nUm but I I actually wouldn't see that as\nan limitation. I would see as a\nstrength, right? So firstly it means\nthat there's actually a lot of human\nscale still involved to create really\ncool worlds, right? And you see some of\nthe examples we showed you. We have some\nvery talented people that can do amazing\nthings with these models, right? and and\nthere is actually a lot of value add\nthere to do that right so it actually uh\nis a tool that can really amplify\nalready creative humans um in new ways\nand I'm definitely not the best at doing\nthis right and I can tell you that it\ndoes it is really impressive when\nsomeone is able to do that but on the on\nthe flip side from the the the agent\nperspective as well right so when we're\ntalking about designing environments for\nagents um and and you referenced poet\nwhich was um for me like poet and dwell\nmodels were the two papers that I just\nthought were eventually on a collision\ncourse right in That's basically why I\nstarted my research career. Um, and I\nthink poet was fundamentally limited\nbecause of the the environment in coding\nbeing an eight dimensional vector. Um,\nbut also the fact that there was no\nreally any notion of interestingness as\nwell, right? And in your recent\ninterview with with Jeff, right, he's\nobviously talked about how this problem\nis largely now solved with foundation\nmodels, right? So these foundation\nmodels can not only define what's\ninteresting based on standing on the\nshoulders of of human knowledge, right?\nBut they can also steer the generation\nof worlds in things like Omni Epic um um\nto do this. And that's in that case it's\ndone through code. But here we have text\nas a substrate as well. So in theory\nthis these kind of open-ended algorithms\nthat use language could actually be\nquite strong places to have these kind\nof like notions of interestingness and\nagents steer tasks through that space as\nwell. Yeah, I think I think this is the\nfundamental thing because certainly with\num creative models at the moment. Um\nlike weirdly counterintuitively you need\nmore skill to make it do something\ninteresting than you did before. Like\nthe average creative process now for\nsomeone designing a thumbnail on YouTube\nis they they mix together you know they\nmight use the contact model, they might\nuse an upscaler, they might then like\nyou know use another image generator\nmodel. you get this huge kind of\ncompositional tree of operations that\nhappen and it's very very highly skilled\nbecause a lot of the kind of um\nstructure for constraining the\ngeneration of these models still comes\nfrom our own abstract understanding of\nthe world and this is kind of what\nKenneth Stanley was saying in a sense he\nwas saying that we have this\nunderstanding of the world which is\nconstrained by things like symmetry and\nyou know various different rules and we\nthen kind of we we hint to the model we\nwe constrain the model in the prompt\nusing those things. Would the models\never be able to do that without the\nhumans needing to prompt them? So I I\nthink what's interesting is that\neventually what we find like what humans\nfind interesting and worth um you know\nmaybe watching or investigating or\nresearching it's it's eventually being\ndefined by people and and I think in the\ncase for example if it's video uh if\nit's it's video generation for example\nthen we see that people go and and find\nways that maybe we weren't like they use\nthe tool that we put in front of them to\ngenerate new things. So, for example, we\nhave people try to cut like we have the\nASMR videos of people cutting, you know,\nfruits made of glass, right? Which it's\nnot something you can do in the real\nworld. And the novelty comes from from\nthe prompt basically. Um, I think that's\nwhat you're alluding to. Um, so I think\nin in the case of we're still in a\nsimilar place I would say when we think\nabout the world models because you have\nto provide the description of the world\nthat you want to to maybe walk in into\nan experience but some elements are not\nlike would kind like emerge from and we\nwill be inferred from the prompt that\nyou provide. Right? So you can maybe\nwrite a very short prompt but still the\nworld will have much more richness. So I\nthink there's a question of where does\nthis rich richness coming from and I\nthink where it's different maybe levels\nof of this the the ability of models um\nto to bring this rich richness into your\nexperience but I think it's kind of like\nover time we see that it becomes more\nand more higher and higher and little uh\ninformation provided by users can\nactually generate very rich videos or\nexperiences. So I would say it's not\nlike a it's a bit of a evolving answer.\nI would say like over time I expect that\nmore inputs to the model or you can\nthink about it like the the the person\nis providing a seed and from that seed\nwe can maybe generate more like more\nelaborate\ndescriptions and finally an experience.\nSo I don't think of it as like a\none-step um process but more of like a\nseries of of um creative steps or each\none of them can be can happen by can be\ndone by a person or by an AI model and\ntogether they generate maybe something\nnew.\nYeah. And and that's what we're seeing\nplay out on Twitter that you know\nbecause the creative process is like you\nknow generate discriminate generate\ndiscriminate and we mimemetically share\nall of the prompts that work and that's\nwhy we've just created this beautiful\nfogyny of creative artifacts that are\nexploring the you know the the space of\nof these models which is beautiful and\nI'm thinking about the future. I mean I\nknow you probably can't speculate about\nthis but this could be the next YouTube\nit could be a new form of virtual\nreality. You know, in philosophy,\nthere's this thing called the experience\nmachine where you you plug yourself into\nthis better than life matrix simulation,\nand no one wants to leave the experience\nmachine because it's better than real\nlife. But we could co-create something\nlike that, right? We could we could have\nit could be on a on a phone or a virtual\nheadset, and we could create these\nworlds and portals between the worlds\nand it would just be a neverending\nsimulation.\nYeah. So, it's a that's a great\nquestion. So I mean going back a few\nsteps I think another really inspiring\nsort of thought experiment in this space\nbefore the generative models really\nbecame capable was something like\npickreeder right and so in that case it\nwas a very simple idea right it was just\nyou know evolving some some im images\nbasically um and some quite surprisingly\ncreative things emerged from that\nexperiment that I don't think um many\npeople would have expected right so you\nhad these beautiful um beautiful images\nbasically emerging uh to use the word\nemerging again emerging emerging from\njust evolving um evolving um user\npreferences over time, right? And we\ndefinitely see modern analogies of this\nlike you described with you know social\nmedia platforms sharing prompts and\npeople generating ideas and then it\nemerges in different ways or goes in\ndifferent ways like um like the VO ones\nwith people generating standup for\nexample and then suddenly there's tons\nof exciting content in that space and I\nthink that it's definitely fair to say\nthat what we've done with Genie 3 is\ncreate another form another platform or\ntype of model where this kind of\ncreativity could happen and it could\nalso lead to some unexpected exciting\nthings. Um, but I don't think we can\nspeculate too much at this point exactly\nwhat those will be. Um, other than say\nthat it it should be interesting and\nhumans will likely do cool things with\nit.\nYes, I I was discussing with Kenneth the\nother day whether because he's a big fan\nof um, you know, neuro evolution and I\nthink he's leaning towards the the evol\nyou know like creating an algorithm that\nrepresents evolution in of itself as\nbeing the way to, you know, explore\ninteresting fogynies. And for me,\npigreeder was like a kind of um\nsupervised human imitation learning. So\nit was almost like a reflection of the\nconstraints and the cognition that we\nhave. And I lean externalist a little\nbit. So I I think that a lot of se\nsemantics is about this embodied\nphysical interaction with the world and\nthat you know just via osmosis perhaps\ngets represented in our brains. But do\ndo you have a position on that? you\nknow, do do you think that just just\npure neural networks simulating the\nworld could could understand the world\nin the same way? So maybe first to like\nabout kind like the immersion or the the\nmaybe potentially using you know this\nkind of kind of models for actually you\nknow being immersed in it like I think\nthis is we're still very far like I\nthink it's really I said before that I\nthink the visual aspects are pretty much\nprimary right we're generating pixels\nand you know with free we also added\naudio but our embodied existence is so\nmuch more than that and I think\nsometimes we you know it's it's getting\nlost Right? Like because eventually we\nlike as as people we feel a lot we walk\naround. We have other senses. We we have\nthis sense of like where I am right now\nand and and I think that and and of\ncourse the physical interaction which is\nalso applicable to robots, right? So\nthere's still a large gap between where\nwe are right now and where and building\nyou know real full simulation of the\nworld that can actually provide all of\nthe information to an embodied agent. So\nI think there is definitely a gap there.\nUm that is interesting but it does show\nthat we're still very far you know in\nthose in that regard. Um and but I think\nas as Jack said basically building those\nkind of like experiences we do we do see\npeople uh try to com to build\nexperiences together and kind like\nexplore worlds together and I think\nthat's a very interesting uh direction\nfor us. Yes.\nYes. Yes. Um yeah so many things to talk\nabout there. I mean I suppose one one\nimportant step is this multi- aent\nsimulation thing right so quite a few\npeople have spoken about this certainly\nDavid Krakow he said that a lot of um\nyou know emergent intelligence is about\ncoarse graining when you have these\nsystems that can you know through a\nvariety of tricks accumulate information\nover time. So um you know eventually\nlike we developed a nervous system and\nculture and language and that allowed us\nto commun you know to accumulate\ninformation sort of like you know um\ntransgressing the the the hardware the\nDNA um evolution speed. So it's\nevolution at light speed and Max Bennett\nspoke about that in his um brief history\nof intelligence how you know like a lot\nof the evolution of the brain in culture\nwas about the propagation of information\nwithout needing to have direct physical\nexperience. So we can implicitly share\nsimulations with each other. So um you\nknow when we start to build these\nmulti-agent simulations do you think\nthat similar things might emerge where\nyou know um almost irrespective of the\nlife span of an individual agent that\nthe system could accumulate information\nand develop forms of agency and dynamics\nthat like simple systems couldn't.\nUh that's a really good question. And so\nI think the way I would see it from the\nstandpoint of Genie 3 where it is right\nnow is that it's sort of um it's a\nmulti-agent world, but that's only\ncontrollable in a in a in a single agent\nsetting, right? So a lot of the multi-\naentness about the world is sort of\nbaked into the the simulation around\nyou. Um they're almost like additional\ncharacters in the world rather than\nbeing like controllable agents. You can\ncontrol them if you wanted to with the\nworld events, right? So you could\nactually um control what the other\nagents are doing but otherwise it's\nalways kind of implicit in the weights.\nUm and what you see is that there is\nsome sort of like natural behavior of\nthem. So if you walk through a c crowd\npeople will move out the way for\ninstance. Um if you're if you if you\ncreate a driving world then when you\ndrive around the other cars move in a\nsensible fashion. Um, and go to go back\nto your actual question, I think you're\nsaying almost the system can almost like\nbootstrap from itself and learn um to to\nlearn to learn um sort of across the\ndifferent agents in the system. I think\nthe way I would see it right now is more\nthat the the model's sort of knowledge\nof of human behaviors can distill into\nthe the egocentric agent. Uh and that's\nactually something quite powerful that\nwe haven't really got with any other\nsimulation tool, right? Because if the\nother agents are moving around sort of\nin a way that we do, um then I think it\nmight even be a way of of our embodied\nagents learning things like theory of\nmind because they know for instance that\nif you're in a if you're in a if you're\nin a world walking around and say you go\nto cross a street, you sort of check the\ncues of the of the drivers for example,\nmaybe there's not a crosswalk um and you\nneed to know when to stop. You can see\nthat they're slowing down. So that's\nwhen you would go and the other agents\nshould be simulated in that fashion. So\nactually you can learn these kind of\ncues that you can't really learn any\nother way other than being deployed in\nthe real world and that obviously has\nsafety um safety risks and probably\nwouldn't be an advisable thing to do\nwith an agent that's learning from its\nown experience. Um so what we think with\nthis kind of model is that agents can\nreally learn these sort of social cues\nthings like theory of mind how to\noperate within humanike other agents but\nit's not the case that the model itself\nis then learning back from the agent\nthat's collecting experience. That might\nbe a future step, but not something\nwe've really considered in this work\nyet.\nYeah, it's fascinating. I mean, Shomy,\nwhat what do you think about that? I\nmean, certainly we we use tools, you\nknow, we have sex and GPS's and\ncomputers and calculators and all these\ndifferent things and I mean, do you do\nyou think about the the locus of\nintelligence being in our brains or do\nyou think if we built rich multi- aent\nsystems or or maybe even if we look at\nhumans and LLMs now like where do you\nthink of the locus of intelligence in\nthat system being? So I think there is a\num different types of intelligence\neventually and um you know as we make\nprogress towards understanding\nintelligence and building intelligence\nwe end up building like initially\nseparate models that can can kind like\naccomplish different tasks along\ndifferent dimensions of intelligence. So\nas I said before I think there is like\nin a way if you really think about it\nlike generating and simulating a world\nis not necessarily something that a\nperson can do right we're not actually\nthat's something that some people say\nokay we have a world model but\ndefinitely we don't have the same word\nmodel or an ability of like ve or genie\nfree right we cannot really simulate\nlike if you tell me a sequence of events\nI won't output pixels right I can maybe\nimagine at the lower level of detail\nwhat would happen if for for example you\nwould get up or if something happens in\nthe environment right and I can plan\naccordingly so I think there is like\nit's not completely parallel like we\ncannot just say that okay those models\nare exactly how we operate but I think\nwhat we do see is that some capabilities\nthat we wouldn't you know a few years\nago if you would come to me and say you\nknow where we'll be able to generate\nvideos um from text I would say okay I\nyou know it doesn't make sense to me\nlike uh I don't think that's going to\nhappen in few years but it did happen\nand other things that people thought\ngoing to happen way way before like\nmaybe self-driving cars now we have much\nbetter progress towards that but so they\ndidn't happen as fast as people thought\nso I think in this case um different\ntypes of intelligence made progress in\ndifferent ways and what I'm really\ninterested in is seeing how those types\nof intelligence can work together for\nexample if we have a model that can uh\nsimulate the world in a just in a in a\ndifferent level than was possible before\nand we have other models for example\nGemini that is able to maybe reason\nabout the world in a different maybe\nless visual way when we bring them\ntogether for example what would happen\nwould be we'll be able to um and so like\nthe the examples that we've demonstrated\nof the SEMA agent that is interacting\nwith Genie right those are two separate\nmodels trained completely separately but\nthen when they are put together they can\naccomplish maybe a new thing so I I'm\nreally excited about that yeah\nyeah that's amazing there's also this\nnotion that um so it's 720p\num you Genie3 was just creating these\nthese immersive and I use the word\nimmersive in intentionally because as a\nvideo editor I know that it's all about\nit's a bit of an illusion. So you are\ntrying to create a creative artifact\nthat is just beyond the predictive\nhorizon of the consumer and then they\nsuspend disbelief. Right? And so so in\nin a sense like we are cognitively\nbounded as observers. We see the world\nmacroscopically. We see chairs. we don't\nsee particles, you know, and so the\nworld can have descriptions at different\nlevels. And when when you interact with\num Genie 3, do do you see it kind of\nlike traversing the levels? Like if you\nzoom in on something, does it have a\ndifferent description or or is it\nlimited in some way? How do you think\nabout that?\nUm, so in some of the examples we\nshowed, there's the one where you're\ncontrolling this drone in the like by a\nlake and there's some trees and it's\nlike very beautiful scenery. And you do\nnotice in that one that when you focus\num your view in different areas, it\ndefinitely hones in on detail more. And\nso I think the model kind of learns that\num sometimes you don't need all this\ndetail, right? And actually it should\nfocus its efforts um sort of with the\nfocus of the agent. And I think this\ncomes a bit from I mean our emphasis on\non this with this model is to have an\nagentcentric sort of egocentric often um\nbut also can do third person but a model\nthat really feels like it's it's your\nview of the world right rather than um\nas a contrast to VO videos right which\nare more much more sort of like like\ncinematic in quality right the whole\nvideo is very high quality whereas G3\noften it does feel much like a your own\npersonal view in the world which I think\nis is quite a different experience,\nright? And it does have this different\nlevels of detail to it, right? Um\nin terms of things like um more more\nabstract representations, I think we're\nstill kind of exploring it to be honest.\nUm but it definitely has a slightly\ndifferent feel to it. Um because\nespecially for the first person view\nthat you get often when you're\nexperiencing it.\nHow do you think about that? I mean do\ndo you do you I mean it's so difficult\nfor us to know how these inscrutable\nmodels work but do you intuit it that\nit's simulating the world at multiple\nlevels of resolution? So yeah, it's a\nreally interesting uh question and way\nof thinking about it because you know\nwhen I first saw video models and in the\nsimulation of for example fluid dynamics\nand other you know aspects of reality I\nwas like how is it even possible to do\nthat in so little time or compute right\ncompared to comparing to actually\nrunning the entire simulation. Um so I\nthink that's first a surprising aspect\nof these models um and but it does come\nwith some limitations. I think what what\nwe basically see is that the models\nsomehow find ways to simulate um as you\nsaid like in a way that looks good,\nlooks reasonably realistic. Um I think\nwe see it with video models but as they\nget better those um kind of like\napproximations become even better. Um\nand maybe you know that's that's a good\nkind of like opportunity to think about\nthe difference in when we simulate the\nenvironment in an interactive way it\nbecomes much harder because if for\nexample you want to um just spill water\nlike a video of someone spilling water\non on some surface right so if the model\nis a video model it can just think about\nlike try it and generate the entire\nvideo end to end uh past and future can\nbe modified at the same time and\neventually you get some video that maybe\nlooks real. But with Genie free, we have\nbecause it's an interactive model, then\nthe user or the agent that controls it\ncan decide to intervene. They can maybe\nlook from different angle and we have to\ncreate the entire simulation frame by\nframe in a causal way and that makes the\nproblem much harder uh for the model.\nbasically you it cannot change the past\nright once the past happened you cannot\nchange it like in the real world right\nand and I think that's where uh we hope\nto see better physical\nkind like simulation but it also makes\nit much ch more challenging so at the\nthe your question about different levels\nfor example of of reality would it work\nif I just zoom in and look at the\nmolecules right and whe like I think\nthis is this highlights the amount of\ncomputation that actually happens in in\nthe real world that if we actually had\nto simulate it completely that's\nprobably be impossible but some of\nmodels find ways to um to approximate it\nto a certain degree that looks\nreasonable to to the observer which is\nus basically. Yeah.\nSo another interesting thing is um we we\ndo this thing called thinking and and we\nknow that neural networks they they they\nare sort of roughly computationally\nlimited. So they they they can be\ntrained to do a certain amount of\ncomputation in a certain amount of time\nand and that means that we can you know\nwe can do lots of things but there might\nbe certain types of things like for\nexample if you simulated someone solving\na Rubik's cube you might find that for\nwhatever reason it just doesn't have\nenough computation to do that thing. So\nwould there be an opportunity to create\na variable computation version where for\ndoing certain types of things it could\nthink more about it?\nYeah that's a really interesting\nquestion. I think some folks in the team\nwere also talking about this. So for\nexample um if in the in the future if\nyou wanted to be able to even write code\ninside the model um at that certain\npoint um maybe that requires some\ndifferent approaches right and I guess\nthen we already do have models that can\nwrite very good code I think um quite\nwidely available now um we also have\nmodels that can you know win a gold\nmedal the IMO for example right and um\nmaybe eventually you want to be able to\ndo this inside the simulation um because\nthat might be the next level is to be\nable to really develop embodied agents\nthat can blend these two different like\ntasks of physical tasks and thinking\nbased tasks. Um and so yeah, at a\ncertain point I think we will probably\nneed to kind of cross that gap. Um but\nfor now I think we probably focus much\nmore on the the visual uh quality and\nmore like physical simulation rather\nthan um the sort of like um more math\nand code type problems which typically\nhave more thinking style models. Um but\nI definitely think it is is an\ninteresting question. And I think it's\nalso um something where the model\ndefinitely has this like physical\nknowledge in it, but I don't know if the\nmodel itself could describe it. It\nprobably just has it implicitly in the\nweights. So another agent could actually\nlearn probably about about um the\nphysical world from the model, but the\nmodel doesn't necessarily know it and it\ncan't tell you it, but it just sort of\nimplicitly has that in the weights\nsomewhere. Uh and there's this kind of\ninteresting duality in a sense. uh which\ngoes back to the agent environment thing\nwhich I think is much more of our belief\nis that right now this is kind of nice\nsetup to have models that can focus on\ndifferent strengths and simulating the\nfuture versus thinking and understanding\nthe present.\nYes. What's your philosophy on that on\nthat show? Because um in in a way you've\nbuilt something which is even higher\nresolution than a language model. So in\nin in principle all the things a\nlanguage model could do as you were just\nsaying Jack could kind of emerge from a\nmodel like this. So is your philosophy\nto kind of build a massive model that\ndoes everything?\nI'm typically thinking about this more\nfrom a practical point of view. I think\nu there is definitely this kind like a\npuristic approach or we should have just\none model to do everything. But I think\nwhen when you know a lot of of the\nchallenges with modern machine learning\ncomes from actually building uh there is\na lot of engineering and um you know\nsoftware and hardware design that\nactually you know to build those things\nright uh train them and run inference\nand I think when we actually hand we we\ntry to actually uh design those systems\nthere are a lot of constraints and those\nconstraints basic basically um kind of\nlike impose on us some some ways in\nwhich we can we have to prioritize what\nwe want the model to do. Um I think\nespecially for Genifree when we're\nbringing the real time uh capability\nright real time is basically means that\nwe have to generate frames very fast\nright um multiple times per second for\nthe person or agent that interacts with\nit to to feel like this is actually you\nknow I can they can move around and and\nand feel the responsiveness of the of\nthe mall. Um so that sets some\nconstraints of on what the model um how\nmuch capacity we actually have. So I am\nI think when it comes to to the point of\nto basically to your question can we\nhave um one model to encompass all of\nthe aspects of intelligence we discussed\nbefore I think it boils down to the to\nwhat what are the set of requirements\nthat we have. If we don't care about\nreal-time interaction maybe we can do\nthat. um if uh we don't care about\nthings like how expensive it is to run\nbut ultimately we're trying to build\nmodels that are not just some you know\nthey don't end up just being as a\ntheoretical uh um kind of exercise we\nhope to actually bring them uh like\nother models to for for people to use\nand for to advance actual applications\nand I think that's where we kind of have\nto make those decisions and ultimately\nwe we pick the the the type of\ncapabilities we want to uh emphasize\nvery cool and 20 second answer, Jack. Is\nthere a sim toreal gap?\nUm, well, it depends how you define it.\nUm, I think that there's currently sim\nto real is actually a bit of a conflated\nterm. It's more sim to lab what people\ncurrently do. I think sim to really real\ncan only really be achieved with a\nphotorealistic world simulation tool\nlike Genie 3.\nYeah. Yeah. So, so you think this is\nactually a big step in the direction of\nI think it's the only way to solve it to\nactually get in the real world where\nthere's people and other agents in\ngeneral moving around rather than just a\nvery constrained label-like situation\nwhich has real world physics but nothing\nelse that's real.\nAmazing guys. This has been an absolute\nhonor. Thank you so much for coming on.\nAnd for folks at home, if you're\ndeveloping on Unreal Engine, might be\ntime to, you know, Yeah. Anyway, cheers.",
  "transcript_chars": 61508,
  "ingested_at": "2026-05-12T00:43:04.085651+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 56914,
    "like_count": 1737,
    "channel_id": "UCMLtBahI5DMrt0NPvDSoIRQ",
    "categories": [
      "Science & Technology"
    ],
    "tags": [
      "genie 3 google deepmind",
      "deepmind",
      "machine learning",
      "artificial intelligence"
    ]
  }
}