{
  "video_id": "dC_3ys349bU",
  "transcript": "Welcome back to another episode of Decoded. Today I'm sitting down with YC visiting partner Francois Chaubard to talk about one of the most important topics in AI today, Diffusion. Francois has been doing computer vision since 2012 when he started in Fei-Fei Li's lab, and after a decade running Focal Systems, he's currently back at Stanford finishing his PhD working on diffusion-based world models for AGI. We're going to break down what Diffusion is, how it's evolved over the past decade, and how it's used today. Francois, danke für being here. Thank you for having me. Well, we just got back from NeurIPS. We just spent a lot of time talking to researchers and thinking about all the newest models out there. I think we saw Diffusion pop up over and over and newer versions of this type of approaches that are not autoregressive LLMs. And so I wanted to talk to you about those today. So first, why don't we start by defining what is Diffusion? Diffusion is a very fundamental machine learning framework that allows you to learn any P data, Any probability of data for any domain, as long as you have the data. So you're trying to learn some data distribution. That's right. Now in a sense, all LLMs or all machine learning models are about learning data distributions. How does Diffusion in particular, what stance does it take or what approach does it take to being able to learn distribution? Yeah, I mean, I think you can use Diffusion to always do that. The thing where it stands out in particular is mapping from high dimensions to high dimensions, especially in low data regimes. So, say I only have 30 images of Gary, which I actually have some code that we're going to walk through. Cool. I only have 30 images of Gary, and again, we're in this 1,000 by 1,000 by 3 dimensional space, and I want to map to another 3 million dimensional space with only 30 training examples, and I can still do it. And it's pretty powerful in that way. Okay, cool. So you have this ability to use relatively small amounts of data compared to the dimensionality to learn a pData. That's right. Was circuits appropriately park her, eineliche Ankurse Forceотze. Wir 같이 passate nogle monter positions an einen. Wir haben einen Autodud describing. Wir müssen in den Gerissen stellen, das vậy. Wir haben diese Portal um, die neue V okay? Die neue Maxi herauszultäente. Es ist halt was ich Training davonverteit pour Takeimal. Und so dann wir flipen und dann wir versuchen, die Model zu reverseen, das ist es. Okay, cool. So es ist es ein Noiser und ein Denoiser. Und das Denoiser ist die Model, die du enden trainieren. Exactly, ja. Du hast es, die TeacherForce, und dann gibt es Noised-up Images, und dann haben sie es lernen, Intermediate Representation zu getehen, zu Pdata. Cool, nice. Und was ist Diffusion für heute? Was sind die App die es in den Wissen, die es in den Wissen? Es ist wirklich sehr surprising, wie das process ist. Ich denke, die 2015 Joshua Söld-Dixen-Papier war auf CIFAR-10, die ist nur images. Ich denke, es hat sich roots in Images, aber es ist viel mehr spürzend als Images. As you've gesehen, DeepMind hat sich die Nobel Prize für diese exacte Procedure auf Protein Folding. You can drive cars with this, with the diffusion policy paper, which is like an insane result. You can predict the weather. There's really no limit to the things that this can do. Yeah, it's pretty incredible to see. I mean, you have these image and video generation models that seem to be really advancing over the last few years. Stable diffusion is the one that I think many people have heard of, and then newer versions of it seem to be using this as well. And then, yeah, in the world of life sciences that my company was in, too, Ich denke, wir sehen, dass die neue Generation von Life Sciences und AI-Companyen sind, sind sehr viel investiert in diese sette Technologie. Es gibt eine Modell called DiffDoc, das funktioniert sehr gut für die Predicte, Small-Molecule Binding-to-Proteins. Und dann ja, AlphaFold, die Neues AlphaFold-Version, benutzen Diffusionen ziemlich viel. Es ist wirklich cool, dass die gleiche Core-Piece-Technologie apply zu so many different domains. Ja, ja. Das Klasse Modell hat sich über die Jahre über die Jahre, und da ist ein ganzes Buches, man könnte das Buches, so man könnte das Buches, so man könnte das Buches, aber vielleicht an der High-Level, we can try to trace out a few of the key innovations that happened, starting with the paper you already mentioned, that now led to the newest versions of these models. So how would you map those out? Like what was the first kind of turn of the crank from this very high-level diffusion process you outlined? What was the first version of that that started to work? Yeah. So I think the 2015 original Jasha paper put up all the key pieces, all the key components of modern diffusion. So now we're just playing with different things. So the scheduler, how do we add noise, at what weight? That's a whole part that we can discuss. What's the loss function? Should the deep learning model condition upon x of t predict the actual data, x of t minus 1? Or should it predict the error that was just added to it? Or should it predict the velocity, which is the error divided by the time? Should it predict the velocity of the start and the end? That's called flow matching. There's all these different plays on what the loss function is. So in all of those, the idea is still to do denoising. Yes. But the objective for each of them is somewhat different from each other. And they're all pretty closely related, whether it's basically a delta between two things, or the previous step, or the first step. How do these all actually come together? But these are a series of papers that happened one after another? Yeah, I think we just kind of hill climbed on this Faroche inception distance metric Das ist ein kooky, weird measure zu sehen, wie gut ein Image ist. Aber wir haben einfach getrennt, und besser an es, mit diesen kleinen Tricks. Es ist so, dass es sich nicht so schwer ist. Vielleicht predicting die Errung ist eigentlich eher. Und dann predicting die Velocität ist sogar eher eher than that. Und dann predicting die Global Errung across die ganze Diffusion-Schedule ist sogar eher eher than that. Und dann einfach finden die Errung und easier ways to basically sample from noise to data. And here when you say easier, was the ease largely driven by it was mathematically simpler or it was easier to implement and engineer or simpler to reason about? Or what got easier, really? It actually is that too, but I didn't mean it that way. What I actually meant was it's easier for the model to learn. But it is also, and we'll go through some coding examples, the math actually got easier. Ja sicher. Und soCIale Trusson für mehr deTurn, Thanksgiving´sstand das nochmal zu TI Bye at denn es, das Prime geht eins alan mit SIe Quirkung Ich denke wir haben mit Unets und das war die predominant Architektur Wir haben nicht wirklich die Architektur sehr viel aber dann haben wir diese Diffusion Transformers und diese Cross Mechanism und so ja wir einfach immer wieder an zu reduzieren FID. Interesting. Should wir in die Code examples? Let's do it. Ich habe über 1, 2, 3, 4, 5, 6, 7 Ich habe einen kleinen Beispiel von Gary, mit verschiedenen Levels of Success. Aber all die Structures sind die gleiche. So, die Jasha Paper, die Non-Equilibrium Thermodynamics Paper, hier sind die Gery, hier sind die Gery. Das ist das, was ich finde online. So, das sind die Gery, die Sie in der 1,000-1,000-1,000-1,000-1,000-1,000-1,000. Ja, ich denke, das sind 64x64. Ja, das ist wirklich ein sehr, sehr, sehr, sehr, sehr. Dann habe ich die neue Ebene. Ich habe das neue Ebene. Das war dann einfache, dass es einfacher. Ich habe das neue Ebene. Es ist das gleich. Ich habe das Diffusion-Schedule. Und das ist wahrscheinlich eine der wichtigsten, die Teil der Diffusion ist schwierig, dass die Möglichkeit ist, die die Noise-Schedule ist, die eigentlich die Hörste zu verstehen. Ich habe es versucht, zu verstehen. Und so, wenn du hier sehen, die Nr.de ist, die Nr.de ist, die Struktur. Und dann ist es, dass es, dass es Rande Static ist. Und wir wollen, dass es hier, hier, hier und hier, hier, hier, hier, hier, hier. Und so, die Interesse ist, und das ist, Joshua, wirklich, Er hat sich in den Momenten alles, was wir brauchen. Und es war ein paar kleine Tweaks, und er hat sich nicht zu sehen. Das ist, das ist die Parts, die wir uns nicht mehr sehen. Und wenn Sie hier sehen, die Noise Schedule. Es würde mir so, dass ich eine Linier Interpolation habe, zwischen dem Image und dem Noise. Ich würde mit 1 und 0, 1, die die Image und 0, die die Noise. Und dann wieder einmal. Und dann wieder einmal. Und dann wieder einmal. Aber wenn man das, dann ist das Massive Unstable. Weil die Incentanzierung der Error ist sehr klein in der Beginn. Wenn man an der Imagen hat. Auf der Relativ Basis. Und dann in der Ende, zu einem komplett Noiz, man muss noch ein paar Error sein. So, if you're a model and you're just looking at this little chunk of the noise schedule, then you have to handle a lot of error in one step. And on this side of the schedule, you need to handle such small amounts of error. And what you actually want is a relatively constant amount of error being introduced every single time step. Right. And the cumulative sum of all that error actually ends up looking like this curve here. That's the pink curve. Yeah. So they call this a beta schedule. Beta is the diffusion rate, the rate of diffusion that I'm doing while I'm rolling this thing out from time zero to time T, capital T. And so you can see here the beta schedule. So we usually have some beta min to beta max, and then one minus that is the alpha. And you can think about the beta as how much noise I'm adding at every time step. And you can think about the alpha as how much signal is being retained. Ja, sie wird es nicht mehr. Und dann die Terme das wirklich interessiert ist die Alpha Bar. Und diese sind die Weights, die sind. Und es hat diese 1-Sigmoid-sigmoid-höhung. Aber das ist die Noise Schedule. Und dann, wenn du das richtig, das hier, dann alles einfach hier funktioniert. Und dann, ich train die Model, und dann, wir können eigentlich... So, was der Training Objective wieder? So, du hast das Noise, und das Training Objective war, was zu tun, was zu tun, was zu tun? In this case, it's to minimize the KL divergence between the real distribution and the distribution that I'm learning. And so I won't go through the code for this one because it's a little bit hairier, but you can kind of see the result on these generated images after 100 diffusion steps at inference time. And you can see that the Farisheh Inception distance is 222, which is like extremely high today. Like modern day would be like maybe like 8 or 10 or something. Und was interessant ist, du hast du schon gesagt, dass du es ein vieles zu tun hat, dass es vieles zu tun hat, dass KL-diversion-basiert ist. Ich glaube, dass in diesen weiteren Modeln, es wird es deutlich simpler. Ich habe mich nicht nur noch das, weil ich glaube, dass es eine interessante Kontrast zu draw zwischen diesen beiden. Ja, so die nächste one, ich würde gerne so zeigen, ist Flow Matching, das ist so schön und einfach. Und das war aus Meta, Jeren Lippmann, Er hat sich gesagt, wir brauchen viel mehr davon. Was wir müssen, wenn du das, wenn du das, die Noisen process als, ich starte von Daten, ich sammle eine Vektor, und ich gehe in diese Richtung. Und dann wieder, ich gehe in diese Richtung. Und dann wieder, ich gehe in diese Richtung. Und dann hier, in die Noise. Und dann, du hast du das Ding, die gleiche auf die gleiche Richtung. Und du hast das sehr, was das eine Art Art. Das ist eigentlich sehr viel zu tun. Wir haben alle für ChatGPT oder Midjourney zu machen. Und es dauert ein paar Jahre. Was es ist 1,000 calls zu den Modell, wieder und wieder, weiter zu den Punkt. Und dann ist es so, dass wir die Stilin des Parfums haben, aber es ist eine Stilin des Parfums. Und so das ist was das so cool, zu mir, ist, dass sie gesagt, All of that intermediary results, there is a velocity, a global velocity, between the noise and the data. And it's just this direction. It's just this straight line. And I don't care where you are, go in that line. Wherever you are, you're over here, go in that line and teach it to go in that line. And that's what flow matching does. Let's see that in the code. I bet it's like five lines of code. It really is quite simple. And so it's pretty cool. So here you go. Das ist die Möglichkeit, die 10-15 Lines-A-Code derzeit ist die größte Maschinen-Learning-Procedure. So, ich habe einen Daten, ein Image von Gary. Ich habe einen Isotropik Gaussian-Noise, dass ich eine Sample von habe. Es ist ein Time dass ich in der Diffusion bin Und ich habe XT die Noised das ist zwischen Extremely Noises So that touches mich nur die Unsere akt patrons in den Volt einer Coolります I return that back to my training loop, which is the shortest amount of code training loop I've ever written, which is five lines of code. I have my batch. I have some time. I sample from that function I just explained before. And then I have my prediction from the model. I feed it in this some element, some noise up image, somewhere between lots of noise and little noise, x of t, let's call it. And I just want it to predict the velocity that I want to go. And this is also really powerful because here you have model abstracted, but that model can be any model. That's right. So you can put in whatever the relevant model is for your distribution, whether that's a protein model for proteins or if it's an LLM for text or an image-based model for images. That is a very clean abstraction as long as you can then predict this velocity and then move in that direction. That's right. This code here has nothing to do with images. It could be weather data. Es könnte ein Stock Market Daten, es könnte ein Trajektorys von Robotics und Tele-OPS-Setup, es könnte ein Protein, es könnte DNA, es geht nicht wirklich. Es ist alles die gleiche same Code. Und dann also, wir haben noch nicht über die Architecture. So, dieses Modell hier könnte es etwas, was es zu sein. Es könnte ein RNN, es könnte ein Unet, die es, die Traditionell ist. Modern же sie sie Drehen Ich verstehe Das ist ein sehr interessantes Result, in dass, especially this, um, I think we often assume, as models have gotten more sophisticated, that they become less accessible for people to understand. But this is quite literally 10 lines of code. Right. That explains essentially all of the most important kind of mathematical and fundamental foundations of the models that we all see as generating, basically, like, magical AI results on our phones. Of course, there's lots of engineering, how you scale them up. Right. That model could be a 100 billion parameter Also, es ist ein bisschen schwierig, aber es ist ein bisschen schwierig. Aber es ist ein bisschen schwierig. Aber es ist ein bisschen schwierig. Aber es ist ein bisschen schwierig. Aber es ist ein bisschen schwierig. Das ist richtig. Ja, und so, es gibt ein paar Tangenen Fälle zu Diffusion und alle haben eine andere Interpretation von was eigentlich passiert. Aber es ist alles die ganze Zeit. Und die meisten Menschen lernen Diffusion eigentlich get quite confused. Ich habe mir gesagt, dass ich diese Problematik-Grafik-Modelle habe, dass ich das Problematik-Grafik-Modelle habe. Und was wir eigentlich ist, dass wir eine Markov-Modelle machen ist, dass wir diese Markovian-Modelle machen. Okay, fine, aber es ist nur... Es ist nur... Noise minus Data. Und du solltest das erstellen. Und dann, wenn du das aus dem Physiksperspektiv und da all diese StatMech-Modelle haben, dass das Interpretation ist. Ich denke, es wird ein bisschen schwierig. Und die Stochastic Differential Equation, die denken, dass das ein SDE ist. Und ich denke, das ist alles gut. Aber in der Zwischenzeit ist es eigentlich ganz einfach. Das ist wirklich sehr gut. Cool. So, wenn wir hier hin, Sie sehen, dass es das wirklich die Vellocität ist. Ihr Ziel ist es, dass die Vellocität die Vellocität. Ihr Ziel ist, dass die Vellocität die Vellocität. Und die Vellocität. Das ist es. And that's super stable and it's really clean. And then at test time, for the physics people, this is like a Euler step kind of thing that you're doing where you call the model a bunch of times and you iteratively refine. So back to the hill climbing that we were talking about, I'll grab some random noise here, X, and I just do, and I call, basically reverse that noising process. Es ist denerer in my diffusion schedule, if I change that at test time, it doesn't work. And so you can't, like, oh, I want it even better, so I'll call it even more. That doesn't, you can't, I've tried it, it doesn't work. Yeah, there's various tricks people try there, but yeah. Yeah, and so, like, there's games played that is actually quite exciting. All the expense... But sorry, here you're saying that's not relevant, right? Because in this type of model, you don't have this time dependency. Well, so you do. So at this time, if you change, for example, the number of steps, if you double it, let's say that, Okay, so in that moment, you know, you kind of need to look at your time. Okay, so the idea of you do is going to be a little time for the second step and you expect to get even higher resolution images, it actually will just turn into white. Like, it actually just doesn't work at all. So you can't step beyond number of steps that was trained. That's an important detail. There are tricks that people are doing to try to compress that representation. So, like, if at train time, I train for 100 steps, Und so, wenn du mit 10-Steps trainst, wenn du mit X-Steps hast, du musst du mit X-Steps anstattest. Ich sehe. Du hast es über das Konzept der Squint Test. Warum hast du den Squint Test für einen Moment? Tell mich ein bisschen über das hier. Und dann, ich würde mich über die Diffusion Modelle in den Kontext der General Intelligence, die ich an. Jan LeCun hat dieses, wie, eine interessante Lekture. Er hat uns über die Erdbeutung und dass wir nicht brauchen, wir fassenen Wings, und wie das war ein Worte. Und zu diesem sagen, dass du 100% richtig bist. Aber wir haben zwei Wings. Wenn du die Worte broschen, und du schwindelst, und du bist ein Bird. Wir haben Helikopters, und wir haben Jets und die Röcke, wir haben da, endlich. Und so, es gibt viele Elemente in das Set-Auf-Things-Auf-Things-Auf-Flight. Und da gibt es verschiedene Pros und Cons. dass wir die Einwohner-Settung haben. Wir sind die Einwohner-Settung. Wir sind die Einwohner Wir sind die Einwohner Wir sind die Einwohner Und Und vielleicht LLMs kann man da sein Aber wenn ich und ich LLM Setup sehe ich sehe diese Monolithic Stacked Transformers, die gleiche Sache, Stacked, Stacked. Und es gibt drei Stages der Training. Wir haben das Pre-Train, SFT, Post-Train, und dann keine Learning-A-Tall, und dann das. Und es produziert sich genau eine Token-at-a-time. So iterative token. Iterative token in time, and it never goes backwards. And then you look at a brain, massive amounts of recursion. You have one learning procedure the whole time. You have these two lobes with a corpus callosum between them that's going back and forth like this. And we think, and then I definitely don't think in one token at a time. When I write code, I don't write one little character at a time. I never go backwards. And I kind of like am going backwards. I'm recursively improving. I'm going backwards again and again. I'm thinking in concepts. Es ist ein Dynamik-Prozess, das emitting die Konzepte und dann höher-level Konzepte und dann lower-level Manifestations. Ja, und ich bin sicher, dass das in den LLM ist, aber es ist almost, wie es stuck ist. Es kann nicht mehr in einem Schritt, sogar wenn es um einen Schritt ist, weil es die we trainieren. Es hat all das in den LLM, aber dann ist es, wie es ein Bottleneck, ultimately, ist es eine Token an a.m. Es ist eine Token an a.m. Und so, ich denke, dass das ist, wenn ich über Diffusion There's two main things that diffusion gives me. It doesn't get me all the way to pass my squint test, but it gives me two things that for sure the brain is doing. Number one, all of biology and nature leverages randomness. Randomness is good. And what is diffusion doing? It's leveraging randomness. If you give me data, I noise it up, and from that I can learn about the data. And can the brain add noise to input data? Absolutely. Absolutely. Neurons are massively random. Das ist die Lagnormal Distribution, Spiked Patterns und Dinge wie das. Und die andere ist das Emission eines einem Ding an einem Versus Thinking in Concepts und dann Decoding into ein Big Chunk of Text und Thought und Revisioning of the Previous Thoughts und Dinge wie das. Und so, ich denke Diffusion gibt mir beide Dinge, für sicher. People haben wahrscheinlich gehört of Stable Diffusion als eine sehr common application of this. Es ist eine Image Generation Model das war ziemlich widely available für die letzten Jahre. Was people may not be so aware of is all the other ways that Diffusion is used in the last few years in products that people are widely using. So what are some of the areas in which Diffusion is most widely accessible? Yeah, it's really any mapping from very high dimensional P data to very high dimensional action spaces or P data that you may want to map to. And so, I mean, yeah, of course everyone knows generating images because we've done mid-journey and things like that. und sogar noch mehr mit SORA und VIO und Flux und SD3 und Dinge wie das. Wir sind jetzt Videos, die sind nur nur Stable-Together, und VideoGen und ImageGen und Dinge wie das. Aber es gibt so viele mehr die wir sehen, das ist die die meisten mehr, in meinem Meinung nach, der neue applications ist. So, wenn du jetzt jetzt die Stenzen, Diffusion LLM war ein wichtiger Ja. Whether it's continuous diffusion LLMs or discrete diffusion LLMs. It's writing code now. It's creating proteins. DeepMind has won the Nobel Prize for that. There is robotic policies, this diffusion policy thing, which I think might actually be one of the biggest uses of it and will result in robotics actually working and Rosie the Robot actually working. Weather Forecasting, for the GenCast, is the most accurate weather forecasting system in the world. It's really anything. Even like I mentioned, Harrison working on the diffs, Diffusion for Failure Sampling. Just sampling for failures and bad things that could happen. We can do that as well. So a lot of the products where we see people actually using AI, especially for things other than just text-based chat, a lot of them are using Diffusion, especially on images, videos, increasingly now things like code and the life sciences. So, yeah, pretty wide berth of things. Yeah. In fact, I would say the only two holdouts right now where state-of-the-art is not Diffusion. Diffusion has eaten all of AI except two. AR, LLMs still are outperforming, and gameplay, and things like AlphaGo. And so MCTS is still state-of-the-art for those types of things. And so we haven't seen Diffusion really take a step in those two areas, but more research is needed. So, to bring the conversation to a head now, Wie sollte man sich über diese Forschung aneinanderfragen, wie als Forscher contributing zu den Fielen, oder als Founders looking to build a new product? Ja, ich würde denken, dass es vielleicht zwei zwei Kampen ist. Wenn du trainierst du, oder wenn du mit den Modeln in der Bühne in der Bühne in der Bühne in der Bühne. Wenn du in der Bühne in der Bühne in der Bühne in der Bühne in der Bühne in der Bühne, würde ich wirklich an Diffusion. Ich würde wirklich an die Diffusion interessieren. Du solltest du diese Procedure, Even if it's just to get a latent space that you can then train off of. And so there's no application in machine learning that I don't think you should be heavily looking at diffusion procedures as a fundamental piece of your training loop. In the case of people who are not training models, I would just update your prior on how good these things are getting. Now, this one of us che krank. Um, indem du Georg ist Licakt Reinhold questa benötigliche. Was immer auf der Premiere von OkavЗ cassiert像. Und nun, leaks Cole attending. Du gebildert das perfekte parties. Produktività thermπου es dirvorsisch.陣ze, äs agriculture. Fですね, Where doesn't people go? All these things are going to work and we're watching it happen. It may cost money and time and those kinds of things, but those are solvable things. Those are tractable problems that we can go solve. And also the core procedure of diffusion is getting better. That's another major... A lot simpler. A lot simpler and it's getting, like, we're just working better. And so skate to where the puck's going to go. Bet that Rose the Robot will work in people's homes. Bet that the protein folding is only going to get better and now we're going to apply Dasימene Freunden sind gut für Vizep 받고ne umuit übernatürliche Menschen mit свое deine Öknechanische Te Galaxy Möchung F leash Das ist Mit Important F alkalem Für Need Kat Woo-Woo.",
  "transcript_chars": 25325,
  "transcript_filled_at": "2026-05-23T14:30:55.084637+00:00",
  "transcript_filled_by": "groq"
}