{
  "video_id": "vFIHrJuwVTc",
  "channel_slug": "deeplearningai",
  "channel_handle": "DeepLearningAI",
  "title": "AI Dev 26 x SF | Paige Bailey: Research to Reality",
  "duration_seconds": 743,
  "url": "https://www.youtube.com/watch?v=vFIHrJuwVTc",
  "upload_date": "20260520",
  "transcript": "Greetings everyone. My name is Paige.\nI'm the engineering lead for our\ndeveloper relations team at Google\nDeepMind and I have the pleasure here\ntoday to talk about a little of what\nwe're currently working on.\nUm, this is a lot of slides. I'm not a\nsuper fan of slides. So, if you want to\nlearn more about a kind of hands-on how\nyou can interact with the APIs in AI\nStudio, um, there will be a session\nlater this afternoon where we will be\ndoing a dedicated workshop to talk\nthrough some of these things.\nUm, I want to start with the mission of\nGoogle DeepMind.\nAnd it's really to build AI responsibly\nwith the intent to help humanity. And\nthis is everything from using AI to\nsolve science, um, to building a kind of\ncures for diseases, to also interacting\nin the physical world. And as a result,\nthis requires a lot of multimodal\nunderstanding, a lot of multimodal, uh,\nsort of real-world interactions. And\nwe'll see a couple of examples why\nGemini and our DeepMind models are\nreally accomplished at this later on in\nthe presentation.\nSo, to start, Gemini 3 is natively\nmultimodal. It's the latest in our\nfamily of models. And that means that it\ncan understand video, images, audio,\ntext, and code, and all of the above all\nat once. But it can also output multiple\nmodalities. So, if anybody has had a\nchance to see Nano Banana 2 or Nano\nBanana Pro, um, that's built on our\nGemini model series,\num, and allows you to create images, to\ncreate images and text interleaved, and\nto also edit images.\nWe also have our new Gemini 3.1 live\nmodel which gives you the ability to\noutput audio tokens natively. So, you\ncan have a conversation with the model,\num, you can share your screen, you can\nshare a video input, um, and this really\ncomes into play for things like robotics\nand augmented reality, which we'll see\nin a second.\nSo Pro is our kind of largest\nin the series of models. Flash is our\ngeneral workforce model. Most of the\nproducts at Google use Gemini 3 Flash in\nproduction.\nGemini 3.1 Flash light is very small,\nvery performant, very lightweight. And\nthen we also have our Nano model series,\nwhich is based on Gemma 4\nand our Gemma open model family, which\nallows you to even run models locally on\ndevice,\nwhich is quite cool.\nGemma 4 is our new open model family.\nHow many folks have had a chance to\nexperiment with Gemma 4? Amazing. So if\nyou haven't, I strongly suggest going to\ntake a look.\nIt's downloadable. You can sort of grab\nthe model from Hugging Face or similar.\nIt comes in four sizes. 2 billion\nparameters, which is small enough to fit\non a mobile device. 4 billion\nparameters, which is small enough to run\nlocally on your laptop. A mixture of\nexperts implementation that's 26 billion\nparameters and also a dense model that's\n31 billion parameters.\nBut really Gemma 4 is kind of punching\nabove its weight in in terms of model\nperformance versus size.\nYou can use it for a variety of things.\nEverything from vision understanding, so\nbeing able to analyze images, being able\nto analyze video and audio.\nAnd you can also even incorporate it\ninto your businesses\nsince it has a new Apache 2.0 license.\nAnd the team is very, very excited about\nthis and can't wait to see how people\nfine-tune the models and use them just\nout of the box for everything from\nyou know, on-device understanding for\nmobile phones to really, really detailed\ncancer research\nin hospitals.\nWhen you look at it compared to our\nGemini 3 model family, it's also doing\nquite well.\nYou can see\nkind of a few of the Anthropic models\navailable here today like Cloud 4\nSonnet. And the where the the kind of\nGemma models perform in compared to\nthose.\nRobotics is one of my most favorite\nareas for Gemini to Gemini to kind of\nhave a presence at Google DeepMind and\nthe Google DeepMind robotics projects\nare some of the most exciting that I've\never seen in my career.\nWe're able to use the Gemini models\nnatively.\nSo as an example, you can say something\nlike, \"Please go make me chicken Caesar\nsalad.\" Or please go clean up that\nspill. Or please go grab that blue ball.\nAnd what the model does since it's able\nto understand the natural world is kind\nof stitch together each one of the steps\nrequired in order to accomplish the\ntask. And then invoke an on-device model\nor just kind of\nGemini out of the box in order to\naccomplish each one of those steps. The\nmodels that you see here are available\non our Mountain View campus. So if you\nif you came by the the robotics lab, you\nwould be able to see some of these in\naction.\nAnd then we also have a partnership with\nStanford for their Pepper family of\nmodels,\nwhich is completely 3D printable. You\ncan download all of the designs and\nbuild it yourself. And it's running\nusing a Raspberry Pi using Gemini for\nboth the live capability of being able\nto talk to the model and direct it\ntowards actions, but also vision\nunderstanding.\nAnd happy to to share pointers or a link\nlater this afternoon as well.\nFor augmented reality, we've been\nbuilding more and more capabilities into\nour own glasses, but also uh but also\nmaking it available for things like Meta\nRay-Ban as well as any other\num any other smart glasses that are able\nto ping a REST API. Um you can see a\ncouple of examples here with\nintegrations with Google Maps. So, being\nable to do things like give live\ndirections. As you're walking along, it\ncan take in the geolocation\num from either your phone uh or\nsomething similar, and then give you\ndirections based on what it sees coming\nin through the glasses feed. Um these\nare another couple of examples for how\nyou can use AI Studio to generate apps\nthat are augmented reality apps. Um\neverything from again these kind of live\ndirections to giving you insights into\nwhat you're seeing as you're walking\ndown a busy street. Um to helping you\npractice basketball um or understand\nphysics to also building games that\nallow you to to create these unique\nexperiences that are only possible\nthrough extended reality and augmented\nreality. And one of the nice things\nabout the Gemini family of models is\nthat it's able to kind of double-check\nand verify um the information that it\nsees as kind of like a a quality check\nbefore before moving on to the next step\nof the model.\nWe also really, really love that using\nit for everything from helping me find\nitems that I might have left on a desk\nto being able to dynamically explain\nuh dynamically explain things that you\nmight see on a screen. So, whether it's\nmath equations, whether it's kind of a\nphysical system that you're observing in\nreal time. Um Gemini is able to to\nincorporate that into its responses. And\nwith our new Gemma model family,\num we're able to stitch together\nspeech-to-text,\nthe Gemma model and text to speech\ncompletely on device.\nUm so, if you don't want to send your\ndata external to the phone, um you're\nable to do all of that work just locally\non device uh using our open model\nfamilies.\nWe've also been investing in real-time\nspeech translation. So, as you're having\na conversation, whether it would be in\nGoogle Meet or in real-time, being able\nto speak in one language and then hear\nanother one and your own native language\num is pretty powerful. This is currently\nonly available via the API, um but is\nsomething that we announced at I/O last\nyear and have been really excited to see\nwhat people build with it long-term. If\nyou've read any of the Douglas Adams\nHitchhiker's Guide to the Galaxy books,\nit's kind of like having a Babel fish in\nyour pocket um and being able to to talk\nto anyone in whatever language you\nprefer.\nFor world models, we've also been\ninvesting pretty significantly using a\ncomposition of models. Um Nano Banana\nfor image editing, VEO for video\ngeneration, and then a model harness\nwhich also incorporates Gemini for help\nwith prompting and for design of the\nsystem, um as well as some of the code\ngeneration. And as a result, we have\nsomething new called Genie 3. Um Genie 3\ngives you the ability to describe just\nin natural language a scene that you'd\nlike to explore. Um anything from uh\nkind of create a world uh with volcanoes\nand kind of uh sparkly bunnies that I\ncan walk around um to what it would it\nbe like to have a man on a jet ski in\nLondon, kind of mobilizing around King's\nCross Station, to a corgi walking down a\nrainbow. Um and each one of these worlds\nis completely playable.\nEach uh frame in the in the sequences\nthat you see is generated dynamically.\nThere's no physics engine involved. And\nthe models are able to create these\nreally really unique video experiences\num for about a minute at a time for the\nones that we make available externally.\nAnd again, if you're curious about any\nof the things that I'm showing, make\nsure to come to the workshop later today\nand you can learn how to build it\nyourself or to interact with it\nyourself.\nAnd the last thing that I want to call\nout is something called Google\nanti-gravity. Um this is a partnership\nthat we've been doing with one of the\nteams at DeepMind uh under Varun Mohan,\nwhich is uh\nkind of building the next generation of\nan IDE. It's agent-first. It's what we\nuse at Google internally. I'm sure you\nall saw that Sundar just recently\nmentioned that over 75% of the code that\ngets checked in each week at Google is\ngenerated by AI.\nUm we see a ton of utilization for\nthings like agents at every hour of the\nday. Most of the people on my teams are\nmanaging fleets of agents at DeepMind.\nUh and a big chunk of this reason is\nbecause of anti-gravity. Um so if you're\ncurious about this, if you'd like to try\nit out yourself, again, the workshop is\nlater this afternoon. Um you can use a\nvariety of models within anti-gravity,\neverything from the Gemini model family\nto the Entropic family of models. Um and\nit's great for both kind of real-time\ninteractions, making to code bases, as\nwell as deploying agents in the agent\nmanager.\nUm so a lot to see.\nAnd I also want to say\nuh that there's never been a better time\nto be a founder. Um so if you uh if you\nneed uh support to grow faster, more\ncost-effectively,\num make sure to scan this QR code. We\nhave\na program at Google Cloud for startups\nwhere we allocate up to $200,000 or even\n$350,000\nto AI startups over 2 years to help you\nget everything that you would need to\nbuild. And this is everything from Cloud\nRun credits to credits for GCS storage\nto credits for our Gemini models.\nUm so if you're if you've ever wanted to\nstart a startup, there's absolutely\nnever been a better time, especially for\nsmaller teams of one to two people or up\nto five people. Um there's a lot to do.\nUh and with that, uh just want to say\nthank you. Go build. Um if you remember\nnothing from today, uh just go to ai.dev\nand you should find pointers to\neverything that I shared.",
  "transcript_chars": 10712,
  "ingested_at": "2026-05-21T19:15:44.687289+00:00",
  "source": "retry-no-transcript",
  "yt_meta": {
    "view_count": 195,
    "like_count": 6,
    "channel_id": "UCcIXc5mJsHVYTZR1maL5l9w"
  }
}