{
  "video_id": "bH2nP-aCFjk",
  "channel_slug": "openai",
  "channel_handle": "openai",
  "title": "Inside image generation’s Renaissance moment — the OpenAI Podcast Ep. 19",
  "duration_seconds": 1763.0,
  "url": "https://www.youtube.com/watch?v=bH2nP-aCFjk",
  "upload_date": "",
  "transcript": "Hello, I'm Andrew Mayne, and this is the\nOpenAI Podcast.\nOn today's episode, we're talking about\nImages 2.0 with researcher Kenji\nHata and product lead Adele Li.\nThey'll discuss why the new model\nrepresents such a major leap forward, the\nevaluations that mattered most during\ndevelopment, and what people are\ncreating with it now that it's widely\navailable.\nIf DALL-E was the Stone Ages, ImageGen 2.0\nis the Renaissance.\nIt's not only great artistically and\naesthetically,\nbut it also incorporates science, art,\narchitecture, all in one image.\nWe looked at it and we're like, all right,\nthis is better than ImageGen 1.\nAdele, tell me a little bit about how you\nbecame a product manager here.\nSo I joined OpenAI a little over two years\nago.\nAnd before OpenAI, I was an investor my\nentire career.\nOh, wow.\nSo I was in private equity and spent three\nyears at Redpoint\nVentures investing in AI and software\ncompanies.\nAnd when I first joined OpenAI, it was for\na completely different role.\nI was thinking about how do we build out\nour data and compute\ninfrastructure. And over time, made my way\nover to the product side.\nAnd for the last six months, have been\nworking on ImageGen.\nIt's interesting how you style yourself\ngoing from one role, then finding\nyourself into this space here, which is\nkind of cool to think about the idea\nthat you have this sort of ability to be\nuseful in different ways.\nAbsolutely. And I think the role of a\nproduct manager is just to do the job\nthat needs to be done, no matter what it\nis.\nAnd for ImageGen in particular, it's been\nreally awesome to flex a lot of\ndifferent muscles when it comes to\nbuilding products, working with\nresearchers like Kenji, but also thinking\nabout what is the gap in the\nmarket today that we want to fill and what\nis the opportunity that we want\nto grasp here. It's not the same market\nthat it was a year ago when we\nfirst released ImageGen 1.0.\nNow it's a very different landscape.\nThere are multiple image generation makers\nout there.\nAnd ChatGPT is a very different company\nand product itself too.\nAnd so really thinking about the evolution\nof ImageGen and its role within\nChatGPT has been really, really exciting\nto me.\nKenji, how did you end up working on\nimages?\nActually, like when I first started at\nOpenAI, I also started about two\nyears ago. I was working on like some\nrandom audio project initially.\nIt was my first project. And then at the\ntime, I just found my way just\nworking on helping them work on ImageGen\n1.0 prior to the launch.\nAnd so gradually I moved more and more\nonto the project and then I became\nfull time on it, basically.\nWhat has the reception been like right now\nfor\nthe model?\nIn the last two weeks since we launched\nthe model, usage is up\nmore than 50%. More than 1.5 billion\nimages are generated every week on\nChatGPT. And we've seen viral trends\nemerge across the world.\nAll the way from trends in Asia for color\nanalysis and stickers to US where\ncrayon and scribble are going viral.\nbut also a lot of people exploring\nemergent use cases.\nAnd I think it shows the dynamic range of\nthe model, but also how people are\nable to visually grasp the advancement of\nthe model almost immediately.\nI think the visual communication reaction\nthat we've seen from our users for\nthem to say, hey, this is the best,\nhighest fidelity, highest quality in a\nstatic model that we've seen has been\nreally awesome.\nThis felt like a really big shift, almost\nworthy of me not even being images\ntoo, but almost like just a new paradigm\nbecause just the capabilities are\nthrough the roof. What made that possible?\nWhen we started working on this project, I\nthink we sat down and we\ndiscussed what is the step change of\ncapability and use cases that we wanted\nto build towards.\nAnd we believe that image generation has\nthe ability to do so much more than\nwhat it does today. you could distill\nevery single output or visual content\nthat you see today into an image.\nAnd so that was the mandate that we sought\nout to improve.\nAnd with this 2.0 model, we've improved on\nvarious different dimensions.\nOne is text rendering. The ability for\ntext on a page is so much better\nfidelity. The language and words actually\nmake sense in their actual words.\nThe second of all is multilingual.\nSo we've really focused on making this\nmodel work in various different\nlanguages. And we're already seeing that\npeople across the world in Asia and\nEurope are really resonating with these\nadvancements.\nThe third is photorealism. I think we\nreally saw a lot of feedback from our\nprevious models that the output wasn't\nvery realistic or altered their face\nor their bodies. And so one of our\nmandates was how do we actually make the\nimage feel like more like yourself?\nAnd so all the things that you think that\nthe model knows, it does because\nit has imbued the knowledge of the world\ninto its conscience and is able to\nvisually communicate that back to you as a\nuser.\nAnd so putting that all together, I think\nwe really get a state-of-the-art\nimage generation model that is the best\naesthetic model out there on the\nmarket right now. That really represents a\nnew paradigm for image\ngeneration, which is a huge part of, I\nthink, AI progress at large.\nthat we have an opportunity to work on\nhere.\nWe often listen back, listen to feedback\non social media too.\nSo we kind of just take all these things\nand basically are just aware of it\nand try to make sure that they're\nmitigated or completely fixed in some\ncases in the next iteration.\nWhat kind of use cases are you seeing?\nWhat are you seeing people do with this\nnow?\nI think one that's particularly\nclose to like the research team as a\ngeneral is like infographics, text.\nI think text in images is like so much\nbetter nowadays.\nSo I think it just opens up a lot more\nproductive use cases.\nAnd from the research side, we kind of\nthink image generation used to\nalways be about fun and maybe unproductive\nthings.\nBut now we're really seeing steps forward\ninto productivity and image\ngeneration for any type of use case that\nyou can imagine it for.\nSo you mentioned text. I remember the\nearly models, no disrespect to\nchimpanzees, but getting to dispel like\nOpenAI even looked like a chimp did\nit. And then now I'm looking at pages of\ntext and finely detailed stuff.\nAnd I know that as models get smarter,\nvariable binding, the ability to put\nthings next to each other improves.\nBut this was just a big improvement.\nYeah. But I don't think it's completely\nunexpected.\nI think you see a lot of growth in\nbetween.\nWell, first you see between DALL-E 3 and\nGPT Image 1.\nIf you ask for a grid of random objects,\nyou go from maybe\nlike 5 to 8 in DALL-E 3 to maybe around 16\nin Images 1.\nAnd then with 1.5, we went to about 25 to\n36 consistently.\nAnd I think now we could probably do over\n100.\nI think this is like a test that we might\ndo internally.\nIt's just we just ask GPT, give me a list\nof 100 random objects.\nAnd then we just send that to our image\ngenerator and see how many are\ncorrect. And usually, you know, it'll get\nalmost all 100 correct.\nAnd that's, but you see the constant\ngrowth over time.\nSo I don't think it's like completely\nunexpected. It's just a steady pace.\nThat was a test I used to use for like the\nreally old models back with\nlike Ada, Babbage, and Curie, like list\n100 science fiction books.\nAnd then some of them would get, by the\ntime it got to like 22, we'd\njust start repeating stuff. So we realize\nthe model reached the end of it.\nSo we've seen stuff too, like 360, 360\ndegree panoramas.\nHow did that happen?\nYeah, that really came from the emerging\ncapability of\nthe model, which is the ability to render\nimages in any aspect ratio.\nWe discovered that people were generating\nreally long, amazing panoramics,\nskinny bookmarks as well. And one of the\ncool capabilities of the model is\nthat not only were you able to generate\nimages in this panoramic aspect\nratio, but you could also render images in\nthe style of 360.\nAnd we saw that it was really fun to\nactually view these images in a 360\nworld itself. And so that was a really fun\nfeature that we ended up adding\ninto the product and it's available on\nChatGPT on web and mobile right now.\nFirst thing I did was I made a version of\ndogs playing poker.\nPut that in there so you could sit there\nlike you're one of the dogs looking\naround in there, which was not something I\nexpected, but it's fun.\nYeah. I mean, it's really awesome to see\nhow people are exploring new use\ncases and fun things that they're creating\nwith the model, even far beyond\nwhat we expected users to be using it for.\nI think when we were designing the model,\nwe were really deliberate in\nunderstanding what people really wanted to\nsee from image generation.\nThere was a lot of latent demand in image\ngeneration.\nPeople were mostly using it for personal\nuse cases, but we definitely saw a\nlot of inklings of people wanting to push\nthe model in certain directions\nthat the model wasn't good at. So text\nrendering was definitely one of those\ndimensions that we really wanted to\nimprove on.\nMultilingual is another. And I think world\nunderstanding generally is so\nmuch better in this model. And that\ntypically means that now people online\nare sharing a bunch of examples of them\ncreating image done for all\ndifferent kinds of use cases that we\ndidn't even think existed out there.\nSo I think the model's understanding of\naesthetic beauty across\nmultiple different outputs, whether that\nit's like a fun meme, an image for\na five-year-old versus a professional\nconsulting deck.\nThe expansion of opportunity and outputs\nhas been amazing to see in this\nlatest model. It's funny too how one of\nthe things that was trending\nwas taking popular images or photos of\npeople and then having the model make\nlike kind of janky looking Microsoft Paint\nversions of that.\nYes.\nAnd did you think that was something you\nwould see was that people are\ngoing to use this incredibly capable tool\nto then go make, you know, these\nsilly looking things?\nYeah, it's funny because it takes a lot of\nintelligence to actually create something\nthat is imperfect.\nThat's what I tell people all the time.\nYeah.\nAnd it's definitely very interesting in\nthe viral trends that we're seeing\nonline right now. One thing that I think\npeople are really striving for is\nauthenticity, imperfection, nostalgia.\nWe're seeing that in the Microsoft Paint\nprompt, crayons, all different\nkinds of generations that people are\ncreating.\nAnd that really feels like the theme of\nconsumers is they want to interact\nwith AI in a very authentic, imperfect\nway.\nThey want to show their imperfections and\nuse AI to help make them look\ngood, but also show a more fun and goofy\nside of themselves.\nAnd I think that's self-expression via AI\nis something that we're really\nexcited about. And, you know, I think it's\nreally part of our mission as a\ncompany to make it easier for people to\nlearn more and distribute that\nintelligence, but also letting them\nexpress a version of themselves that\nmaybe wasn't possible before.\nKenji, was there a moment with this model\nwhere you're saying to yourself,\nwow, I think this is ready to go?\nYou know, as it's training, we take a\ncheckpoint and then we just sample\nfrom it, right? And just see, okay, how\ngood is this thing?\nAnd I think we just sampled a model, an\nimage, and we looked at it and\nwe're like, all right, this is better than\nImageGen 1.\nWe were just like, okay. I remember\nwatching the iteration of one of the\nearly versions of DALL-E and how at first\nit was sort of the wispy sort of\nweird sort of the tendril sort of thing\nand talking to one of the\nresearchers like, is that going to go\naway?\nIt's like, I think two, probably two runs\naway from that.\nAnd then just like that, the ability to\npredict that was amazing to me.\nAnd all of a sudden everything got crisp\nand clear. Yeah.\nAnd then also like looking at, you know,\nyears ago I'd played with like, you\nknow, GANS and like doing those things.\nYou have to squint and say, I think it's a\npickup truck or something like\nthat. Yeah. So it's interesting what you\nsee is you say, okay, this just all\nof a sudden got much better. Yeah.\nI mean, it was just very obvious. You just\ntake the early checkpoint.\nYou just sample an image from it.\nAnd then you just sample an image from,\nyou know, ImageGen 1.\nAnd you just look at the two and you're\njust, there's just.\nWhy do I like this garbage? I forgot what\nthe image was.\nIt might have just been like a picture of\nlike a woman on the seaside.\nLike, you know, overlooking a seaside.\nWe just looked at it and we're like, all\nright.\nIt was like no question. Yeah. That was\nthe big jump was the photorealism of\ngoing from something that looked that was\nmore of a glossy, idealized\nmagazine cover to something that looked\nlike a really good photograph.\nSo help me understand, like besides just\nmore compute, how did this happen?\nHow did you get a model that's much better\nand also that doesn't take an\nhour to generate an image?\nThe times are still, I remember in the\nDALL-E\ndays, like we would literally have to, you\nknow, tell us what you want.\nAnd then an hour later, it'd be on\nInstagram to now these things are in\nChatGPT and it's faster. How is it getting\nboth more intelligent and you're\nmaintaining the same speeds?\nI think we learned a lot in each release,\nlike between 1 and 1.5, now 2.\nAnd so we take each of the learnings that\nwe've made and we've, you know,\nlike, for example, speed, right?\nYou know, one of the things is like, oh,\ncan we make the model more token\nefficient or something like that?\nAnd we did a lot of work to make it\nproduce very good images with less\ntokens. I think the post-training for this\nmodel was very interesting in the\nsense that we really had to think about\nnot only does the model understand\nworld knowledge and how things look in\nscience, concepts, math, etc.\nin an image, but also what is the taste\nthat will resonate with users?\nWhat makes the model or output beautiful?\nHow do you make it look realistic?\nThese are all questions that we had to\ngrapple with when we were\npost-training this model. Because I think\nthat one of the things that was\nreally important for us was that this\nmodel was the strongest aesthetic\nmodel out there right now, which means\nthat it has more creativity in\nvarious different outputs, no matter what\nthat output is, if it's a\nprofessional output or a personal output.\nAnd so that range of training and the\nrange of use case, I think, made\ntraining this model a very interesting\nproblem.\nDo you have any personal favorite\nbenchmark tests you like to do, things you\nsay, I want to see it make an image of\nthis?\nI have an eval that I call the me, me, me\neval.\nOkay. It's essentially 100 photos of\nmyself and my friends and my family.\nAnd I put everyone in goofy positions.\nI have about a card or a birthday for\nevery single person.\nAnd I think it's a really great eval in\nthe sense that you only know the\npeople around your faces the best.\nYou also want to create funny things with\nthe model and do things that are\nrelevant. And so one thing for me as a\nproduct manager that I'm testing is\nnot only is the raw capability of the\nmodel really great, but also does\nChatGPT understand what I want in that\ncontext?\nYou know, ChatGPT remembers, you know,\nthat I have a brother, that I have a\nmom and dad and what they like to do.\nAnd so does the model accurately know how\nto insert pieces of\npersonalization in the moments that matter\nin the images?\nThese are things that I'm testing for. How\nabout you?\nBesides the grid one I mentioned earlier,\nthat's probably the one I've used\nthe most. For a while, I think Divya and I\nwere doing a lot about\nphotorealism. We're trying real hard to\npush on that.\nJust basically, I know Divya's favorite\none was like a woman holding a jug\nof orange juice. I don't know if you've\nseen this.\nThere's like so many images of a woman\nholding a jug of orange juice.\nWell, I actually feel like the researchers\nhad a more standard set of images\nthan they like to lead on. Yeah, and you\nget like the standard.\nCan it do somebody writing with their left\nhand or watching on their right\nhand and a clock showing this?\nI think the big leap of the images, like\nprobably 1 or 1.5, was like a half full\nglass of wine.\nThe wine glass folded the rim. Yeah,\nfolded the rim.\nYeah, exactly. And there were ways I was\nable to prompt it to do it, but\nit was really – had to get a really\ndescriptive, like, you know, red liquid\ninside this. This one is so fun to prompt.\nThere was a thing people said, oh, can it\ndo, like, you know, can it do,\nlike, pixel accurate pixel image style\nart?\nAnd somebody was like, no, it can't.\nAnd when I hear that, I'm like, okay,\nlet's try.\nAnd I found out if I gave it, like, a 64\nby 64 grid and I\nsaid, go draw the art in there.\nIt did. It just was able to put art into\nthere.\nAnd that was amazing to see those kinds of\nresults.\nAnd that's the promptability of this is\ninsane.\nHow do you plan for that? Does it just\nhappen?\nYou're like, oh, wow, this is better\nunderstanding this?\nPeople come to ImageGen with very vague\nprompts.\nYeah. Make it better. Make me look better.\nYou know, make me cuter. All these things\nare really vague.\nAnd I think it's really the job of the\nmodel and the harness to distill that\ninto actually what users want.\nAnd I think that's a personality of the\nmodel that we've trained over time\nthat we've really harnessed the power for.\nAnd honestly, I think it also yields a lot\nof really surprising results that\npeople may not expect. And that surprise\nis just part of the fun of using\nImageGen. I've seen like two kinds of\nprompting sort of emerge.\nAnd I remember back with DALL-E, I thought\nlike, oh, I'm a prompt engineer.\nI'll be great at this. Like, I'll be\nreally good at this.\nAnd I'd make a raccoon in space and be\nlike, feel proud.\nAnd I'd see an artist, somebody who wasn't\na prompt engineer, somebody who\nactually came from that world.\nAnd I'd watch them use their language.\nAnd they were doing amazing things. Yeah.\nAnd that seems like that's still holding\ntrue.\nDefinitely. I mean, we work with a group\nof artists very closely when we\ndevelop this model. And we're very\ninspired by artists, designers,\nmarketers, all these different professions\nthat I think have a different way\nof approaching their profession.\nAnd one of the things that was very\nimportant for us is we wanted to take\nthe inspiration as well as the best\npractices for those professions and\ndistill that into the way that people\ninteract with the model.\nAnd so that's something we've deliberately\ntried to focus on.\nOne hack that I've seen work really well\nis the ability to upload\ninspiration or context into the model.\nAnd the model has an incredible ability to\ntake the spirit of that context\nand translate it into the output.\nBut it's interesting because I think that\na lot of people worry that, oh, I\njust push in a button, I get something\nbeautiful.\nAnd each model, that gets better.\nIt's easier, as you said, to not have to\nput a lot of effort into it.\nBut when people do put effort into it,\nthey are getting even more amazing\nresults. And it seems like actually that\nif you're artistically inclined,\nyou're getting even greater control\nbecause now, like you said, it\nunderstands more about what you're talking\nabout when you talk about depth\nof field and these other things or\nwhatever you're trying to do.\nAnd as you mentioned, it was exciting to\nsee with earlier models, artists\nwho said, oh, I gave it my originals and\nit gave me these variations and I\nknow which one works.\nAnd just seeing that as this real creative\namplifier.\nYeah, definitely. I think having creative\ndirection or taste or judgment and\nbring that to the model is the best way to\npush it further.\nI think one thing about this model that\nI'm really excited about is how it\nexpands the creative outlet for people.\nI think the ability to create multiple\ndifferent styles or types or\nvariations has never been easier than with\nthis ImageGen model.\nAnd I think it's also understanding of\ndifferent contexts, like the way that\nit's able to shift what it's like to be\ngenerating an architectural diagram\nall the way to the aesthetics of a\nchildren's book.\nThe ability for it to move so seamlessly\nacross these vectors has been\nreally awesome. The ability to do great\ninfographics and diagrams is very\npowerful. What kind of feedback have you\nbeen getting from people in\nresearch and education?\nWe actually have an internal alpha channel\nwhere we\ntest our models.\nAnd in that, there's like a sub-channel\ndedicated specifically towards\neducators of any level, like elementary\nschool students all the way up to\ngraduate level. One of the coolest things\nI saw was there was a biology\nprofessor and he put like these graduate\nlevel textbook rendering pages of\nthings I had no clue about. And he said it\nwas perfectly accurate.\nI think the ability for this model to\ndistill very complex topics into\nsomething that is really easy to\nunderstand within an image is one of its\nstrongest capabilities.\nAnd we've seen this with students, with\nteachers who are using ImageGen to\nlearn different concepts, to also help\nthem create study guides, to help\nalso create personalized content.\nI think personalized learning is a huge\ntrend that we're very passionate\nabout. And I think the ImageGen model\nhelps you as a teacher create\nsomething that every kid can understand in\ntheir own language and their own\npreference. And that is something that\nwe're really excited about.\nWe're thinking about this in the context\nof also how do we bring more of the\nelements of ImageGen into ChatGPT at large\nso that when people are trying to\nlearn concepts, we're teaching them with\nImageGen.\nI remember when I was in school and kind\nof prior to a lot of kind\nof multimedia blowing up, posters were a\nbig thing, classroom posters\nexplaining stuff. This really reminded me\nof how powerful an infographic can\nbe because it allows you to bring as much\nattention as you want to it.\nAnd you can spend the time looking at it\nand seeing it, and you can put\na lot more detail into it. I think one\nreally awesome visual shift that I've\nseen with ImageGen is that now in internal\npresentations, over 50% of the\nslides are created with ImageGen.\nWow. And that permeation of communication\nvia images is so powerful when\nyou're trying to explain your concepts or\nillustrate what you mean.\nAnd I think infographics and the text\nrendering capability, as well as the\ncomposition of the text on the page, is\nincredibly powerful with this model.\nThe model's understanding of not only what\nto say, but how to present it is\na superpower. And I'm really excited about\nfuture explorations of this,\nwhere we can think about how do we make\nthis even better?\nHow do we improve the composition, the\ndifferent kinds of outputs, and also\nmake it editable in the product?\nThese are directions that we're really\nexcited about.\nHow do you see the progression of this?\nThis is great, but typically anytime I\ntalk to somebody to open an eye about\nwhat they're working on, they're like,\nyeah, this is good, but.\nI think we're still super early in\nexploring all the different use cases\nthat people are really trying to push the\nmodel with.\nAnd so one of the things that we're really\nexcited about is what is that\nnext stage for ImageGen, which is to\ncreate the creative agent.\nUltimately, the agent that can work\nalongside you, be your creative\nassistant, and really understand how you\nwork, what your preferences are,\nwhat is the output that you want to get\nto, and build the product and model\necosystem that helps users kind of have a\npersonal interior designer,\npersonal architect, personal wedding\nplanner, et cetera, all in one image.\nI'll tell you another thing that was kind\nof amazing was like, all right,\nbooks. And so every now and then I have a\nbook come out, I've got to\nchange my social media headers. And I just\nwent and I said, oh, find my book\ncover and create an appropriate size\nsocial media header that I can put on X\nor Facebook or whatever. I'm like, well,\nlet's see.\nFirst shot. First shot. Right aspect\nratio.\nEverything. We basically did that from the\nstart or trained the models to be\ngood at that from the start. I remember I\nworked on the initial de-risk of\nbasically it could do any aspect ratio\nthat you ask.\nYeah. Yeah, you can now really just easily\nspecify the outcome that you\nwant. Yeah. Like in the case of yourself,\nyou're like, I want promotional\nmaterial. I don't have an idea. I didn't\nspecify exactly what I wanted.\nBut the model was able to do the research\nand then give it to you in\nthe style and aspect ratio that was\nrelevant to you.\nAnd that's super powerful. We're already\nseeing this.\nyou know you're an author I've talked to\nreal estate agents who are using\nImageGen to help them create listings for\ntheir apartments or stage their\nlistings YouTube creators have talked to\nme about using ImageGen for their\nthumbnails and promotional content I've\ntalked to top artists who want to\nuse ImageGen to connect with their fans\nand I think the ability for all\ndifferent kinds of professions to start to\nuse ImageGen to help them with\nvisual creation is super powerful\nespecially if you're working in a visual\nand a creative industry, ImageGen is such\na hack in your professional\ntoolkit. I think it has to be a part of\neveryone's everyday workflow in the\nfuture. This does feel like the, I think\nit feels like the first time where\nanything I can reasonably come up with, it\ndoes a pretty good job of it.\nWe think it's a new paradigm for image\ngeneration altogether.\nLike if, you know, we set this in the\nlaunch video, if DALL-E was the stone\nages, ImageGen 2.0 is the Renaissance.\nYeah. And I think that is so true because\nthe model, It's not only great\nartistically and aesthetically, but it\nalso incorporates science, art,\narchitecture, all in one image together.\nAnd I think that composition and knowledge\nthat the model has just means\nthat the outputs are so much more\ntrustworthy, are more powerful, and enable\nso many more use cases. I think that\nImageGen and Codex is also an amazing\nintersection of the capabilities that\nwe're setting out to create with both\nImageGen as well as coding agents.\nSo many people are using ImageGen as a\nfirst step to designing a new website\nor creating a new app. And I think that\nintersection of having a really\nstrong aesthetic model, which is image\ngeneration, in combination with\nstrong coding abilities, means that now\nyou're able to zero shot really\namazing apps from scratch with both of\nthese tools.\nYeah, I asked it in Codex. I said, I took\nmy website.\nI said, could you make me, I had the\nImageGen.\nCould you create me some different\nconcepts for it?\nAnd I did these contact sheets. I asked\nfor contact sheets to that.\nGive me like four images there. And I\nsaid, oh, the one on the upper right.\nCan you go make that?\nAnd I watched Codex go make that, which\nwas like, this\nfeels like magic. And then they've\nimplemented as part of pets.\nAnd so like if you're using Codex and you\nsay, hey, I want to have like\nI have like I love Raven. So I have like a\nRaven.\nI said, can you make a Raven? And then I\nwatched it pull up the ImageGen\ntool and iterate and make the sprites for\nit.\nYeah. Sprite sheets are going viral.\nYeah. Same with game design.\nPeople are loving. using ImageGen to help\nthem create new worlds.\nAny hints on how to do better sprite\nsheets?\nI mean, I've tried to make, you know, GIFs\ninternally.\nAnd I think if I just use, like, the\nthinking mode or Codex, and you\nbasically just ask it to generate one\ninitial sprite, it's really good.\nAnd then you can just say, can you make\nthe rest?\nThe consistency across multi-images has\nbeen amazing.\nWe've seen a lot of people try creating\n10-page comic books with consistent\nstorylines, multi-page slides.\nI think that consistency of characters and\naesthetics\nis completely unique to this model.\nThat was an example too, where there were\na lot of workflows out there for\nworking with image models that were kind\nof janky, but you had to figure out\nhow to do. And it's great now because I\ncan do stuff where I can create\ncharacters and say, make a character sheet\nwith the different poses and\nstuff and just go feed it back in and say,\nokay, now doing this, now doing\nthat, now doing that. And that's just such\na – often, sometimes what we need\nis obviously a smarter model, but context\nlength did so much for ChatGPT,\ndid so much for coding. And with an image\nmodel, it's able to reliably\nreference these references.\nIt's incredibly capable. Yeah, for sure.\nAnd we're still trying to improve that as\nwell.\nIt's not perfect today. We're really\ntrying to develop this visual creation\nlayer for people because every single\nperson you have an aesthetic or\npersonal style or preference.\nAnd we're really trying to imbue that into\nthe product that we're building\nso that people can get to the output that\nthey're wanting easier and faster\nwith ImageGen. Any parting prompt tips for\npeople?\nWell, one of the things I would suggest\npeople try is ImageGen thinking.\nSo if you navigate to the thinking or pro\nmodels, we have a more powerful\nversion of ImageGen in that experience.\nAnd in that model, you actually are able\nto search the web, analyze files,\nleverage tools under the hood, which then\nyields a better quality and higher\ncomposition photo.\nAnd the suggestion that I have for\nprompting that experience is be\nopen-ended. I think the model will go and\ndo the exploration itself to\nunderstand and try to reason and find\ninformation that matters.\nAnd I also think giving it a sense of an\naesthetic is also super helpful.\nUsing and grounding that in a style has\nbeen really fruitful for a great\nresult. Good one. Good one.\nI think just being very particular about\nthe\nstyle or like what you like in general.\nLike for me, I like minimalist\ninfographics.\nSometimes I think the model can be a\nlittle dense.\nAnd so I just maybe I'm just a simplistic\nkind of guy so I just like\nvery very clean a very clean look so I\nlike that.\nAdele, Kenji, thank you very much.",
  "transcript_chars": 30042,
  "ingested_at": "2026-05-15T10:37:51.308000+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 6079,
    "like_count": 258,
    "channel_id": "UCXZCJLdBC09xxGZ6gcdrc6A",
    "categories": [
      "Science & Technology"
    ],
    "tags": []
  }
}