{
  "video_id": "M-jTeBCEGHc",
  "channel_slug": "machinelearningstreettalk",
  "channel_handle": "machinelearningstreettalk",
  "title": "The Real Reason Huge AI Models Actually Work [Prof. Andrew Wilson]",
  "duration_seconds": 7429.0,
  "url": "https://www.youtube.com/watch?v=M-jTeBCEGHc",
  "upload_date": "",
  "transcript": "So when I say deep learning is not so\nmysterious or different, I'm not saying\nthat deep learning is not mysterious or\nnot different. I think it's actually\nboth. It's rather that the ways in which\npeople often think it's mysterious can\nbe relatively well understood both\nintuitively through a notion of soft\ninductive biases, but also formally in\nterms of rigorous generalization\nframeworks that have actually existed\nfor many decades. And a lot of these\nphenomena can also be reproduced using\nother model classes. I think deep\nlearning really is distinguished in its\nrelative universality. How broadly\napplicable it is relative to other model\nclasses. That doesn't mean that it's\nanywhere close to being completely\nuniversal, but it's sort of a movement\nin the direction of universality, a\nsignificant movement. It also does\nrepresentation learning incredibly\neffectively. It has properties of its\noptimization objective, its loss\nlandscape which are relatively different\nand surprising like mode connectivity.\nAnd so I think deep learning is\ncertainly different and mysterious but\noften not in the ways that that people\nmight might believe. So, I'm very\nexcited to be talking to you both even\nthough it can be challenging knowing\nthat there are so many people watching.\nBut I think there are so many\nfundamental misconceptions in the way\nthat people understand generalization\nand model construction and artificial\nintelligence. It's really important to\nhear a different perspective around, for\nexample, how it's completely fine to\nbuild a huge model that will also have a\nstronger bias for simple solutions. have\nmore of an aams razor-l like behavior\nthan even smaller models. And these\nsorts of perspectives actually help us\nunderstand phenomena that are often seen\nas very mysterious like double descent\nand benign overfitting and over\nparameterization and give us\na principled approach for thinking about\nhow we're going to build our own models\nfor whatever properties we're interested\nin.\nThere's this fundamental trade-off\nbetween bias and variance. And it feels\nlike you're saying you can have your\ncake and eat it and you can keep them in\nthe mixture and you still win. And that\njust goes against most people's\nintuition. So I think the bias variance\ntrade-off is an incredible misnomer.\nThere doesn't actually have to be a\ntrade-off.\nFolks, that interview with Andrew was\nabsolutely amazing. Keith came over to\nthe UK and we did it in my home studio\ntogether a few weeks ago. that's been on\nPatreon for a little while and I updated\nso much based on that interview. Andrew\nwas absolutely brilliant. So, I know\nyou're going to love it. But before we\nkick off, you've probably heard that\nhuman data is kind of the dirty secret\nof Silicon Valley. You know, human data\nis the reason why these AI models work\nso well because the open AIs and the\nanthropics, what they do is they hire\nhumans to do things like data set\ncuration and evaluation and post-\ntraining. And there is a ridiculous kind\nof uplift from using this human data.\nOur sponsor Prolific, what they want to\ndo is produce the first survey on how\nhuman data is being used in AI. And if\nyou volunteer and fill out this form for\nthem, they will give you first access to\nsee how you compare. So I'd really\nappreciate it if you did that. There's\nno personally identifiable information.\nLink is in the description. And we are\nalso sponsored by Twofer AI Labs. They\nare an incredible research lab based in\nZurich. They've just upgraded their\noffice. They've got an amazing new\noffice. They've hired 13 research\nengineers in the last year doing things\nlike reasoning and the ARK challenge.\nYou've probably seen some of the papers\nthat they've published on that. But they\nhave ambitions to build their own\nfoundation models from scratch. They've\ngot an amazing culture. And Benjamin\nKruier, the the director, is also very\ninterested in AI safety. So he's going\nthrough the Yudcowski book at the\nmoment. So if that seems like a fit for\nyou, please get in touch with Benjamin\nKruier. Go to twoflabs.ai\nor look in the description.\nAnd also MLST is sponsored by Cyber\nFund. Enjoy the show, folks. Well,\nAndrew, much of your work challenges\nconventional wisdom. Is that hard to do?\nIs there resistance in challenging\nstrongly held beliefs?\nSo, yes. I mean, this is what happens\nwhen you challenge conventional wisdom.\nBut in some sense, I think that we\nshould always be trying to do that\nbecause otherwise we're just preaching\nto the choir and then what's the point?\nIf no one if you're not changing\nanyone's beliefs about anything, then\nmaybe it doesn't make a difference. And\nso I think it's important to really try\nto understand what do a lot of people\nbelieve that might be wrong and then\njust unpack that. And it's also very\nexciting and fun. But it's challenging\nbecause of course the initial instinct\nwill be to resist whatever you're\nsaying. But then over time and if you\ntry hard enough and if you talk to\namazing communicators like you two then\nyou can start to have an influence. And\nI think that's really important because\nso much progress has been stalled, I\nthink, by just getting stuck on\nmisconceptions. Like once a certain\nnumber of people believe something, it's\nvery, very hard to change their minds,\nno matter what you say. And I think as a\nconsequence, we've been in all sorts of\nlocal minima in machine learning and AI\nresearch because we haven't been able to\nget unstuck from these erroneous\nbeliefs. And there's a whole roster of\nthings like this like the role of\nimplicit biases of stochastic\noptimization in generalization I think\nis significant but also significantly\noverstated. Um how we can have really\nlarge models that also generalize well\neven when there's a small number of data\npoints is something that is not very\nwell recognized. um and in fact I think\nis one of the primary drivers of scale\nbeing important for achieving good\ngeneralization. So not just flexibility\nthis simplicity bias that comes about\nthrough scale. I think another\nmisconception is this idea that\nwe should change our model depending on\nhow many data points we happen to have\navailable. And this is might even be the\nmost controversial one. The reason I\ndon't think we should is because we\nshould always honestly represent our\nbeliefs and our beliefs about the\nprocess that generated our data\ntypically shouldn't change depending on\nhow many data points we happen to have\naccess to. And you can actually\ndemonstrate that these principles work\nin practice. So you can have models that\nwill be very good when you have a small\nnumber of data points and also very good\nwhen you have a very large number of\ndata points. And so this relates to not\nnecessarily needing to have hard\nconstraints but instead combining\nexpressiveness with alam tracer.\nYeah. And you know you're you're only\nslowly starting to convince me to give\nup this 10,000 you know degree polomial\nis bad and and I'm only starting to\nchange because you value simplicity bias\ntowards simplicity as much as I do. And\nthe real key for me was understanding\nthat somehow scale\nhas a bias towards simplicity. And I\ndon't know why or where it comes from,\nbut I believe it. And it's weird.\nAnd can we just set this up? So in your\ntalk last year, your first slide, so you\nhad a bunch of students in the room and\nand you showed like a fairly sort of\nsimple linear correlated, you know,\nairline passenger data.\nAirline passenger data. Okay. It was it\nwas pretty linearly\nseasonality. and and you said here's\nthree models. One is basically y= mx\nplus c or something like that. Just a\njust a straight line. And I think was\nthe second one something like 10\nparameters and 10,000 parameters was the\nthird one. And almost everyone in the\nroom said they preferred one or two. And\nyou said at the end of this\nconversation, I'm going to convince you\nto prefer three, which was 10,000\nparameters.\nAnd I took a poll at the end and it did\nshift. And so that was promising. Yeah.\nI I sort of joke sometimes that if I\nhadn't met the airline passenger data\nset, I don't know really what my life\nwould be like now because it's it's\ndriven so much of my research. Um and\nit's just amazing how people are biased\ntowards choosing the linear function or\nthe cubic polomial even if in practice\nthey're not making that choice. Like on\nfor example, it's not uncommon to use a\nneural net with tens of millions of\nparameters to fit a training set with\ntens of thousands of data points. And\neven before deep learning was popular,\nwe were doing nonparametric statistics\nwhere we were working with models like\nGaussian processes that were inspired by\ntaking infinite limits of neural nets\nthat are more flexible than any neural\nnet you can fit in memory and other\npopular coariance functions as well like\nthe RBF kernel really are like saying I\nwant to use an infinite order\npolinomial. And so even in these kind of\nclassical statistical models we're\nimplicitly saying well if we're unhappy\nabout the third choice it's actually\nbecause it doesn't have enough\nparameters. We want infinitely many\nparameters, not just 10,000 parameters.\nAnd I think another way to say this is\nthat parameter counting is a very bad\nproxy for model complexity. Really, what\nwe care about is the properties of this\nsort of induced distribution over\nfunctions rather than just how many\nparameters the model happens to have.\nAnd so you can have a distribution over\nfunctions which is very flexible. It can\nrepresent many different solutions to a\ngiven problem. But it can also have very\nstrong preferences for certain types of\nsolutions over others. And strong\npreferences doesn't mean saying that\ncertain things are impossible\nnecessarily. They can just have epsilon\nprobability. And I think that that is\nreally meaningfully different than\nsaying, okay, we're we're going to have\na hard constraint and we're not going to\nrepresent those solutions. A because\nit's not an honest represent of our\nrepresentation of our beliefs to to have\nthose hard constraints and\nB because we see in practice that when\nwe do have these expressive models with\nsimplicity biases, they're much more\nadaptive. They're much more automatic.\nSo when you have a small data set, it\nsort of does the right thing. When you\nhave a large data set, it also does the\nright thing. And so you don't need as\nmuch human intervention. And arguably,\nthat's the definition of what really\nmachine learning is trying to achieve.\nIt's trying to build an intelligent\nsystem that doesn't require manual\nintervention. And so I think this is an\nimportant principle towards that goal.\nYou know, you you had a paper kind of\nremoving some of the mysteries of deep\nlearning, but I think it's it's still\nfair to say it's a bit mysterious where\nthe simplicity bias comes from at scale,\nisn't it?\nIt is. So there is some handwavy\nintuition around lost landscapes. And I\nthink this is borne out empirically. So\nwe can sort of understand for example\nthat the solutions that we're finding\nwhen we're building larger models are\nmore compressible and are flatter etc.\nBut this is really ongoing research and\nI think this is one of the most\nimportant questions to understand right\nnow is like why does scale rigorously\nspeaking produce a simplicity bias and\ncan we get that bias in a more elegant\nway than just building bigger models. I\nwas going to ask you, are you a theory?\nBecause I know you're a theory guy. Are\nyou a theory guy or an engineer?\nPresumably, you're both, but you know,\ndeep in your bones. Are you an engineer\nor are you a scientist?\nIt's really hard to choose. In some\nsense, both.\nI'm mostly driven by trying to\nunderstand things. And so, this can be\ndone in a variety of ways.\nA lot of our papers empirically try to\nunderstand model behavior. And so I feel\nlike this is a scientific approach to to\nmachine learning. And one thing that\nreally motivates me about this type of\napproach is whatever you learn will\nnever go obsolete. So quite often\nnewcomers to the field and even you know\nvery experienced researchers\nfeel distressed at the rapid pace in the\ndiscipline where you see methods getting\npublished at a conference and then\nbecoming obsolete within a month or when\nyou go to the conference everything\nyou're seeing is sort of you know been\nreplaced by some other algorithm and\nwonder okay well is there any point to\nme sort of investing myself\nsignificantly in building a model if I\nknow that it's not going to be used by\nanyone for any long period of time. And\nI think one way to address that is\nreally to try to combine what you're\ndoing with an understanding of why\nthings are working. So if you're\nbuilding a model that gets better\nperformance on some problem, there's a\nreason for that. And if you can\nunderstand the reason for that, that\nunderstanding will outlive that specific\nmodel and how widely it might be used.\nAnd I'm really hoping to do research\nthat will be relevant in hundreds of\nyears from now. And so I think these\nquestions around model selection for\nexample in AAM's razor people will never\nstop asking like hopefully they'll be\nable to go back and read not just my\nwork but like work that's been done in\nthis space and think okay this is useful\nto me in thinking about how to approach\nsome of these questions and so in this\nrespect I would say I'm a scientist and\nI try to combine classical theory with\nempiricism towards understanding model\nbehavior and I think if you really\nunderstand something hopefully it's\nsomething that you and demonstrate in\npractice. And so I also try to combine\nsome sort of practical demonstration\nwith a lot of the work that I do. Um,\nand this process also involves\nengineering. And sometimes understanding\nthose low-level engineering details\nbecomes really fascinating and leads to\nunexpected intuitions about the\nprinciples behind model construction. I\nthink this is something that's perhaps\nunderappreciated. Um, like quite often\nyou can have a great idea and whether it\nworks or not depends\nvery significantly on all the low-level\ndetails like numerical stability and\nother things like that. And when you get\nreally deep into those details,\nsometimes you can discover things at a\nhigher level that are also very\nsignificant and how we should think\nabout model construction and algorithm\ndesign.\nVery cool. Well, maybe um well, I was\njust going to share with you there's a a\nduality that Shannon pointed out that uh\nthat you you may like and you know it's\na duality between uh past future uh\nknowledge and control. He said um\nhe said uh we have no knowledge of the\nfuture but we can control it. We have\nknowledge of the past but we cannot\ncontrol it. And and I took that and\nrelated it to science and engineering.\nSo the way I look at the two sides of\nthat coin is scientists leverage control\nto gain knowledge. Engineers leverage\nknowledge to gain control.\nSo you can be both because the the new\nknowledge that you collect allows you to\nbetter control the environment, the\nfuture, the world, and collect even more\nlearning. Right.\nAbsolutely beautiful. Andrew, we haven't\neven introduced you yet. Can you can you\ncan you tell the audience about\nyourself?\nSo I'm Andrew Wilson. I'm a professor at\nthe Kuran Institute of Mathematical\nSciences and Center for Data Science at\nNew York University. My work focuses on\nhaving a prescription for how to build\nintelligent systems. What are the key\nprinciples involved in model\nconstruction? I think although the field\nhas made an extraordinary amount of\nempirical progress towards building more\nperformant machine learning systems,\nwe're still at early stages of\nunderstanding, you know, what principles\nshould we broadly embrace when we're\napproaching our own problems. And so\nthis involves work on understanding\ninductive biases. So what assumptions we\nshould be making. And so this relates to\nsymmetries like equavariances. Maybe\nwe're modeling molecules or rotation\nvariant. Images could be translation\nvariant. How do we represent those\ninvariances? How do we learn them\nautomatically? Um how do we discover\ninterpretable scientific structure in\nour data that tells us something maybe\nsurprising that we didn't know before\nthat will go beyond a particular\napplication? How do we represent\nuncertainty towards decision-m?\nArguably, a prediction uh that's just a\npoint estimate without any kind of error\nbars associated with it isn't really\nactionable in the real world. Um you\nknow, if you have a an autonomous car\nand it says there's a stop sign 5t ahead\nplus or - 10,000 ft, you can't really do\nanything with that information. But if\nit's plus or - 1 foot, then you can\nreally act on that information. And you\nknow observing that almost makes you\nparanoid like okay now I really need to\nrepresent uncertainty because if I don't\nhave that uncertainty then you know\nmachine learning can't meaningfully\nengage with the real world and basian\nmethods I think are a really great way\nof reasoning about uncertainty and so\nthat also forms a big part of my\nresearch program.\nAmazing.\nWell welcome to MLST. We have Dr. Dugar\nin the house. the first time we've met\nin person for we've been doing this for\neight years. Yeah.\nFive years on this channel, but we had\nthe previous channel as well. Yeah.\nAnd uh Keith came to my to my wedding on\non Friday. It's good to have you here,\nman.\nYeah. It's pretty crazy. We haven't met\nuntil now. It's It's going to be a\nblast.\nAbsolutely. Yeah, we've got some good\nstuff lined up. Um well, um I guess just\nto kick this off, Andrew, I was inspired\nby geometric deep learning. Um I\ninterviewed, you know, like Michael\nBronstein, Taka Cohen, Juan Bruner, Peta\nVilichovich. They had this geometric\ndeep learning blueprint. Our video on\nthat did half a million views. It was it\nwas an amazing uh video. And the basic\nhypothesis as as Jan um outlined the\nthree curses in um in machine learning,\nright? So there's the statistical curse\nwhich is that you only have so many data\npoints and the distance to those data\npoints, you know, or the density of\nthose data points is kind of cursed by\nthe dimensionality of your data. There's\nthe um the optimization curse which is\nthat you get stuck in these local\nminima. And there's the approximation\ncurse. And this is kind of where they\nwere driving to that you you have this\nfunction class and and you can be quite\nopinionated in how you structure that\nfunction class. But if you make it too\nsmall, you you incur approximation error\nwhere the actual test sample is, you\nknow, some epsilon distance from from\nthe approximation class. And they said\nthat all of these things are cursed. And\nI guess their prescription was this\nplatonistic idea that if we constrain\nthe models, so if we add bias to the\nmodels with these symmetries because the\ngenerating function of the universe is\nusing these symmetries anyway, then\nthere's no apparent approximation error\nin doing so. So why wouldn't we do it\nanyway? I think your your ideas are are\nkind of tangentially related to that.\nYou still think we should have biases,\nbut you you also think we can have our\ncake and eat it, so to speak. Mhm. So\nthis is a wonderful question and I've\nalso done a fair amount of work on\ngeometric deep learning particularly in\nterms of how we should represent\nequarian symmetries in scientific\ndomains and I agree that it's very\nappealing to say well if we know that\nsome constraint applies to our problem\nthen we should encode it in our model\nand I don't think that's particularly\ncontroversial. Well, there are some\nperhaps surprising results that in some\ninstances even when you know what the\nconstraint is, you can do as well or\nbetter when you don't represent that\nconstraint and this can do this can be\ndue to the dynamics of how you train\nyour model etc. Um but generally the way\nI would respond to this is um really to\nsay two things. The first is there\naren't that many instances in practice\nwhere we know exactly what constraints\nwe want to have even when it comes to\nthings like physical conservation laws.\nSo rarely are we modeling for instance\nclosed systems. You can have a dynamical\nsystem and maybe you have like a\npendulum with like wind in the air or\nwhatever and now you have some sort of\nviolation of of conservation of energy.\nAnd so you want to perhaps instead build\nmodels that are just biased towards\nthese constraints without being exactly\nconstrained. Secondly, when you do this,\nwhen you represent so-called approximate\nconstraints or you have soft\nconstraints, so you have a model which\nis very flexible, but it says, well, if\nwe can fit the data in a particular way,\nthen we want to do that. Quite often, it\nwill just collapse down onto those\nconstraints if it provides a consistent\nexplanation of what we observe. And so,\nif you're paying any kind of penalty for\ndeviating from those constraints and you\ncan perfectly explain your data with\nthose constraints, you'll just collapse\ndown onto that. And so I think as a\nprescription for model construction,\nit's often going to be fruitful to try\nto ex embrace expressiveness, but at the\nsame time have a simplicity bias which\ncan be formalized in terms of\ncompression. And this is really an\nhonest representation of our beliefs in\nmany cases. Like another way to describe\nmy philosophy for model construction is\njust honestly represent your beliefs.\nAnd we believe the real world is a\ncomplicated place. And if we combine\nthat belief with the idea that simple\nsolutions that are consistent with our\nobservations are more likely to be true,\nthen we can often see desirable behavior\nin quite a variety of different\nsettings. And so in terms of like having\ncake and eating it too, I think one of\nthe most surprising\nfindings that that we and and others\nhave had is that quite often you can\nincrease model expressiveness while\nsimultaneously increasing its biases. So\nlarger models are often more inclined\ntowards simple solutions. And there sort\nof demonstrations of this that have been\njust hiding in plain sight like double\ndescent. this idea that as you increase\nmodel flexibility, your generalization\nerror first gets lower. So your it\nimproves as you capture more structure\nin the data. It gets worse as you start\nto overfitit and then it gets better\nagain. And in that second descent,\ntypically all of the different models\nthat you're considering are fitting the\ntraining data perfectly. So the only\npossible way that larger models could be\ngeneralizing better is because they have\nsome other sort of bias like a\nsimplicity bias rather than being more\nexpressive. And so I think time and\nagain researchers express surprise at\nthe fact that they can have these\nmassive models, billion parameter plus\nmodels trained on relatively small data\nsets that aren't overfitting. But in\nfact actually if they've made the models\neven bigger they would be less likely to\noverfit. And so quite often\nexpressiveness and soft constraints can\nbe aligned.\nSo if parameters don't solve your\nproblem, you're not using enough of\nthem.\nWe always want more.\nPrediction with expert advice, we know\ntheoretically and and for me empirically\nthat by keeping all of these other\nexperts around in the mixture, you pay a\ncost for that.\nSo in my particular case, having the\nhistorical experts in the mixture, every\nsingle prediction, they still had some\nweight. you have to give them some\nepsilon weight\nbecause otherwise they would die and\nthat actually harms your performance. So\nthe question is, is it better for you\nwhen the new regime comes to pay the\ncost of learning the regime versus the\nswitching cost of bringing the old\nregime back? And that's kind of the same\nwith any ensemble or any um set of\nrestrictions on biases really that\nthere's this fundamental trade-off\nbetween bias and variance. And it feels\nlike you're saying you can have your\ncake and eat it and you can keep them in\nthe mixture and you still win. And that\njust goes against most people's\nintuition. So I think the bias variance\ntrade-off is an incredible misnomer.\nThere doesn't actually have to be a\ntrade-off. So the idea I guess behind\nthe classical bias variance trade-off is\nthat your generalization error can be\ncompartmentalized in these two terms. So\num bias sort of how well uh you uh are\nfitting the data essentially uh and\nvariance like how your fits vary\ndepending on uh uh like if you sample\ndifferent points from this distribution\nthat you're trying to model. Um and it's\ntrue that sometimes if you naively build\nlike a really large polomial for example\nyou can have low bias and high variance\nwhereas if you build a small polinomial\nmaybe you have low variance and high\nbias. Um, however, approaches like\nensembling are actually a good way of\ngetting low bias and low variance. And\nit turns out actually building large\nneural nets are another way of getting\nboth low bias and low variance, you\nactually have flexibility combined with\na simplicity bias. Uh, and this is\nwhat's leading to good generalization.\nAnd it's sort of another perspective on\ndouble descent. I think here's here's\nthe way I'll put the question is uh and\nI ran into this as a practitioner you\nknow back in the day so maybe I was just\nstuck in the hump of having like too\nmany but not enough parameters you know\nto to get to the double descent phase\nI'm not sure but I mean you know the\nwhat I experience is that you know\nhaving parameters in a model even if\nthey're very very small because some you\nknow I put in some term in the objective\nfunction that forced them to be small is\nnot the same thing as actually the\nsimpler model that just didn't have them\nat\nRight. Like I mean like we kind of you\nknow we'll talk about marginalization\nand kind of the basian basian\nperspective on that. So like overfitting\ncan be a real problem. So for example\nyou brought up you know conservation.\nIt's like well if I'm doing a model and\nI don't enforce conservation of energy\nand then as a result I end up with some\nsmall parameters that cause like a\nlittle bit of feedback and increasing\nyou know energy every single time a\nrobot you know takes some action that\ncan cause it to like spin out of control\nright where I actually did need it to\nconserve energy and not have that that\npositive feedback. So I guess I guess\nmaybe we're still struggling with you\nknow overfitting can be a problem. It's\na real problem. It's a known problem.\nHow do we know if we're overfitting in a\nbad way? We maybe haven't don't have\nenough parameters. We're stuck in kind\nof the the area before we got to double\ndescent or like how do you in practice\navoid the actual consequences of bad\noverfitting,\nright? So overfitting is real absolutely\nbut the conventional wisdom about how we\nshould approach it I think is\nfundamentally misguided. So, and this is\nrooted in things like the bias variance\ntrade-off like let's constrain our\nhypothesis space so that we can't have a\nbad fit to the data that will make bad\npredictions and so on. Whereas instead,\nI think if we just embrace the honest\nbelief that there are many possible\nsolutions even if they're not probable\nfor any given problem combined with this\nsort of simplicity bias, we won't tend\nto overfit. And interestingly,\nthe prescription is almost the opposite\nof what people think it it perhaps\nshould be in principle. like build a\nsmaller model is usually the the princ\nthe the sort of the the\nprescription for avoiding overfitting\nbut\nor to enforce simplicity.\nExactly. Yeah. Whereas in fact as we\nbuild bigger models we often actually\nstart to alleviate overfitting and and\ndouble descent is just a great example\nof this because that first ascent uh is\nis from overfitting the data but then it\ngets alleviated as we start to make our\nmodels bigger and bigger\nby some phenomenon that we still don't\nreally fully understand like this this\nthis ability of simplicity to start to\ncome back into the picture as you just\nmake it even bigger.\nRight? So in that second descent, the\nmodels are typically fitting the\ntraining data perfectly. And so the loss\nis not really the decisive factor\nanymore between which models we're\nselecting. It's something else. And so\nthose other biases start to dominate.\nAnd so this is actually really\nimportant, I think, when people talk\nabout phenomena like flatness. So\nflatness is this idea that if you\nperturb your parameters that you can\nstill get a relatively low value of the\nloss. And there are all sorts of debates\nabout like the the role of flatness and\nhow relevant it ought to be and\nunderstanding generalization etc. But I\nthink what a lot of these discussions\nmiss is it's just one of many properties\nthat control generalization. And so if I\nhad to choose between a model that has\nvery high loss but it's very flat versus\nlike a very flat solution versus a model\nthat that finds a solution that's low\nloss that's that's um relatively sharp.\nI would almost certainly choose the low\nloss solution. And so when you're in\nthat kind of second descent regime, you\nnow are controlling for the value of the\nloss and it's just the flatness of the\nsolutions that's increasing. And so, uh,\nthat's sort of one way of knowing maybe\nlike what side of the the the the curve\nthat you're on and thinking about how\nbig should you make your model. But I\nwould also just say make your model\nalways as big as possible. Um, just try\nto combine what you're doing with some\nsort of simplicity or a compression\nbias. And there's a question of like how\nyou do that, but I think there are good\num, good ways that we have of thinking\nabout how to do that.\nTrick question. Is predictive power the\nsame as understanding? I agree with the\nidea that representation matters.\nRepresentation meaning sort of how\nyou're solving the problem even if\nyou're getting the same performance in a\nparticular application. But the reason\nit matters is because\ndifferent representations that are\nachieving the same performance might\ngive you different performance than on\ndifferent problems. And so if we're\ntrying to build more general agents, we\nwant to understand what sorts of\nrepresentations are going to provide a\nbetter general description of the real\nworld. And so that means we want to\navoid things like shortcut learning and\nso on if it's just going to lead to good\npredictions in some contrived problem\nand not really in the real world. And so\nthere's this question of like can we\nunderstand what sorts of distribution\nshifts we might typically encounter and\ncan we build methods that have\nbroadly more robustness to a variety of\ndifferent types of realistic\ndistribution shifts. And I think this\nconnects to things like no free lunch\nsort of thinking. So the no free lunch\ntheorems say that every model is equally\ngood in expectation over all problems\nsampled uniformly from a distribution\nover all problems. And there are other\nno freelance theorems that say no single\nlearner can be good on all problems. The\nissue I think with these theorems is not\ntheir mathematical validity. What\nthey're saying is correct under the\nassumptions they're making, but rather\nthat the assumptions they're making are\nnot a good description of the real\nworld. So the real world is a small\ncorner of all possible data sets. It's\nnot drawn uniformly from a distribution\nof all possible problems. If we were to\ndo that, we would mostly just get noise.\nAnd so the question then is to what\nextent is the structure across real\nworld problems shared and at what level\nof abstraction can we represent that\nshared structure. And so my contention\nis that the distribution over real world\ndata is biased towards local mgraph\ncomplexity and so are some of the the\nmodels that we started to to develop.\nYeah. But can you give an example of\nwhere it was hard to confront a\nmisconception, why it mattered and what\nthe process involved?\nThere had been this sort of approach\nwhere you would take some approximate\nbasian inference procedure and pit it\nagainst deep ensembles as the non-basian\nalternative. And um normally I actually\ndon't care that much about what's being\ncalled basian or not. Like it's same\nwith like intelligence, what is int like\nwhether something's a good\nrepresentation all these things. Let's\njust connect this to whatever problem\nwe're trying to solve. But this was\nactually kind of problematic because the\ntakeaway seemed to be that if deep\nensembles were working better than some\nlelass approximation or some MCMC\nprocedure then the answer is to be\nnon-basian to be less basian than than\nwe have been historically. And that's\nturned out to be exactly the wrong sort\nof directionality for how we should\nthink about model construction because\nfor a given computational budget those\ndeep ensemble procedures were actually\ndoing a much better job of uh\napproximating the posterior basian\npredictive distribution. So doing\nmarginalization and so in fact the\nprescription should have been well we\nactually need to be more basian. And so\nlike this was sort of a frustrating\nthing but it was also very difficult to\num to approach because there had been so\nmany papers where people had just\nwritten basically that these deep\nensembles are the non-basian\nalternative. And so kind of coming out\nand saying well actually they're doing a\nbetter approximation of the basian ideal\nthan all these methods that are being\ncalled basian is sort of you know\nconfronting you know uh just sort of\nhundreds of papers in some sense at\nonce. Um but it felt like a very\nimportant thing to do. Um, and we had\nactually done it subtly in a lot of\npapers where there'd be some subsection\nof some paper that mentioned something\nlike this, but that was never really\ninternalized because it was never front\nand center. So I thought, okay, a blog\npost is the right way to do this. Um,\nand let's make it all about this. And\nuh, I think in the end actually it was\nthe blog post that changed people's\nminds. I I never saw a paper after that\nwhere people were making this\nseparation. Um, but also I think you\nknow it did it did sort of strike a\nnerve a little bit.\nYeah. But how should we approach model\nconstruction? How can we embrace\nexpressiveness without overfitting?\nAnd so when I say I want to embrace\nexpressiveness, uh there is some some\nsubtlety associated with that idea. And\nthat basically means that it in some\ncases maybe we're wanting to represent\nlots of solutions but we're assigning\nthem almost zero probability but not\nzero probability. So they're possible\nbut not plausible solutions in our view.\nAnd then if the data is telling us\nsomething that like well actually we\nreally should be paying attention to\ncertain type of structure that maybe\nwould surprise us the model can actually\nrespond to that. And if that structure\nisn't actually there then your model\nisn't going to perform a lot worse than\nthe model that is exactly constrained in\nthose ways. In terms of how we should\napproach model construction in general,\nif you have a soft bias, a gentle\nencouragement towards certain types of\nconstraints over others, quite often you\ncan do as well as the perfectly\nconstrained models. And the reason is\nyou're paying some sort of penalty, even\nif it's small, for deviating from that\nconstraint. So if you can fit the data\nperfectly with the constraint, you'll\njust collapse down onto that model. And\nwe noticed this in a work we had called\nresidual pathway prior for soft\nequavarian constraints. And so this was\na basian mechanism essentially to have a\ndistribution over solutions which would\nbe concentrated in some way around\ncertain types of equivariance\nconstraints. Just for the the sake of\nthe audience equariance is a\ngeneralization of invariance. Uh it\nbasically means if you have some\ntransformation t f of txals t of f ofx\nrather than f of txals f ofx um\nlike a cnn for example. So it it it\ncommutes in that case with the\ntranslation.\nExactly. Right? So if you translate the\nimage in some way the the pattern of\nactivations will translate in the same\nway across the different layers um\nrather than them just staying exactly\nthe same which would be invariance. Um\nso\nin this paper on residual pathway prior\nwe were interested in the strength of\nthis soft bias. So we basically had a\ndistribution over neural net parameters\nthat had a coariance matrix that lived\nin some equivarian subspace of our\nchoice plus some orthogonal complement.\nAnd so there'd be these waiting terms\nbasically that would represent you know\nhow strong do we want this bias to be\nfor equarians. And surprisingly to us at\nthe time it didn't matter very much like\nif as long as you had a very soft bias\nfor the the constraint it would often be\nas good as even a perfectly constrained\nmodel. basically for this reason that\nyou know you're still paying some\npenalty for deviating from the\nconstraint and so if it is a good\ndescription of the data you can fit it\nperfectly then you often collapse down\nonto that constraint although I wouldn't\nwant to dismiss the importance of trying\nto calibrate these biases it it can\nmatter in certain instances but in a lot\nof instances uh a very gentle bias is\nsufficient\nso maybe maybe just to put this into\nperhaps more familiar territory for\nother basians out there you know like\nthe assignment of prior just for\nparameters. It's always the goal to try\nand find a prior that's relatively\nignorant but encodes some very soft, you\nknow, type of constraint like maybe it's\na a scale invariant parameter or a loca\nor a scale invariant prior or location\nand variant, you know, kind of prior.\nAnd overall, a lot of times these priors\nare maybe worth like one data point or\ntwo data points, but they're small\nenough that they can be overridden quite\neasily by enough data. But even that\nsmall amount is enough to avoid sort of\nstupid answers like you know that that\nlike the chance of ahead is infinity or\nsomething like that. Is it kind of\nanalogous to that or\nI think that's a reasonable analogy. I\nwould also add that we can't get away\nfrom making assumptions. So even though\nI'm in favor of embracing expressiveness\nand having relatively soft biases for\ncertain types of solutions as opposed to\nhard constraints, machine learning means\nlearning by example. And we can't do\nthat without making assumptions. The\nquestion is just what assumptions should\nwe be making and at what level of\nabstraction? And perhaps it's enough in\na surprisingly large array of different\nproblems to embrace expressiveness in\ncombination with some sort of\nsimplicity. some AAM's razor bias that\ncan be formalized in terms of\ncompression.\nEmpirically, um, simple models work\nbetter.\nIronically, given the surprise people\noften express at this idea of wanting to\nalways embrace flexibility is that like\nbefore deep learning, the community had\nstarted to come on board with this\nnotion that we want arbitrarily flexible\nmodels. In fact, the class of models\nthat I was working on in my PhD,\nGaussian processes in machine learning\nwere um kind of inspired by this idea\nthat we want really really large neural\nnets. And Radford Neil, a statistician\nuh at Toronto at the time, working in\nJeff Hinton's group, was saying, okay,\nwe want to build models the size of a\nhouse and I'm a basian and I'm going to\nreally embrace expressiveness. So, I'm\ngoing to take an infinite limit of a\nneural net with an infinite number of\nhidden units and this is going to\nconverge actually to a Gaussian process\nusing a central limit theorem argument\nwith a particular type of coariance\nfunction. People thought, oh, well,\nthat's amazing. Let's just use Gaussian\nprocesses because they're so much more\nprincipled in a lot of other ways like\nthey're less sensitive to a bunch of\ndesign decisions, etc. they can be\nwritten in you know a very small number\nof lines of code and everyone anywhere\nin the world is going to get basically\nthe same answer etc and so people sort\nof moved in this direction of just\nembracing expressiveness but then once\nwe started working in these kernel\nformulations I think people perhaps\nstarted to forget that there was this\ndual space correspondence and we\nactually were working with models that\nwere more flexible than any neural net\nyou can fit in memory and finding we\nwere achieving very good generalization\nespecially on problems with a relatively\nsmall number of data points. So now to\nto get back to your question, um Radford\nNeil also had an interesting quote I\nbelieve in his PhD thesis that whenever\nyou have a simple model that performs\nwell, you can always build a more\ncomplicated model around it that will\nperform even better. And he gives this\nexample of handwritten character\nrecognition where you might have\nirregular writing styles, weird ink\nblotss on the stage on on the page. Uh\njust some sort of structure that you\nprobably haven't already accommodated in\nyour model that you can try to\naccommodate and you'll achieve better\nand better performance. So what I would\ndo in this situation you described is\nreally try to understand what's the\ninductive bias there that's leading to\ngood performance and how can I soften it\nin some way? How can I generalize this\nin a way that still honestly represents\nmy beliefs?\nSo, but there there seems to be\nsomething wrong with with this quote\nthat you can always build a more complex\nmodel because uh maybe you can always\nbuild a more complex model that has a\nhigher likelihood on the training data.\nBut overfitting is a real phenomenon.\nLike nothing in your work says that\noverfitting doesn't occur and can be\nharmful. So it's like there have to be\nsituations where if you diverge from the\nground truth model. I mean I'm sure I\ncould just build a system that has a\nground truth model simulate data and\nprovably show that a more complex\ninference doesn't generalize as well.\nSo that has to be true. So how do we\nknow in reality when we've stepped too\nfar?\nIt's a great question. I think we have\nto be careful about what we mean when we\nsay a complex model. So I think most\npeople would not consider Gaussian\nprocesses with an RBF coariance function\njust the standard coariance function\nkernel that's often used to be a complex\nmodel but it's highly expressive. So\nit's more expressive than any neural\nnetwork we can fit in memory. It just\nhas very very strong preferences for\ncertain types of solutions over others\nand this enables it to be\nextraordinarily data efficient. So one\nof the main use cases these days for\nGaussian processes is in something\ncalled basian optimization where you're\ntrying to maximize some sort of blackbox\nobjective. So it's not something you\nhave a closed form expression for like\nit could be generalization performance\nof a neural net as a function of some of\nits hyperparameters for instance or some\nreally costly physical simulation as a\nfunction of some parameters and you\nbasically want to query this objective\nas few times as possible in order to\nachieve a good result. Gaussian\nprocesses are an amazing surrogate model\nfor this objective and you use the\nuncertainty to do the exploration\nefficiently. So I think that you can\nhave expressive models that aren't\nnecessarily complex. Uh they still have\nvery strong simplicity biases, very\nstrong preferences for certain types of\nsolutions over others, but at the same\ntime they're representing a wide array\nof possible solutions to the problem.\nAnd I think that's how you can kind of\nreconcile what Radford Neil was saying\nwith what we see in practice around\nthings like overfitting. So I think\nthere's a common misconception that the\nexpressiveness of a model and its\ninductive biases are at odds with each\nother. uh the more expressive the model\nthe weaker its assumptions in some sense\nlike the the fewer inductive biases it\nhas the less data efficient it will be\netc. And what we found which has been\nquite exciting is that the larger you\nmake say big transformers actually the\nstronger its inductive biases like the\nmodels get both more expressive and they\nhave a stronger simplicity bias and so I\nthink you can\nexpand these two things together in some\nsense and this is how you can avoid say\noverfitting and other sorts of issues\nwith not achieving very good\ngeneralization and I think one of the\nthe clearest demonstration ations of\nthis is in a phenomenon called double\ndescent. Uh so double descent is this\nphenomenon where uh typically on the\nhorizontal axis you have the\nexpressiveness of the model the number\nof units for example in each layer of a\nresidual neural network and on the\nvertical axis you have generalization\nerror. And so initially generalization\nerror decreases as you increase the\nexpressiveness of the model and it's\nable to just fit the data better and\ncapture more structure. And then it\nstarts to decrease. it starts to go up\nand that corresponds to some sort of\noverfitting and then it decreases again\nand that's why it's called second double\ndescent because of that second descent\nand in that second descent all the\nmodels typically have about zero\ntraining loss. So the training loss just\nkeeps going down as you increase the\nexpressiveness of the model until\nroughly the number of parameters equals\nthe number of data points. Um, and what\nthat means is that the larger models in\nthat second descent cannot be\ngeneralizing better because they're more\nflexible. They're all fitting the\ntraining data perfectly. It has to be\nthat the larger models have some sort of\nbias which is enabling better\ngeneralization. And turns out that this\nis a simplicity bias, a compression bias\nthat we can measure.\nYeah, let's talk through that a little\nbit. So as I understand your thesis is\nthat there is some kind of generating\nfunction of the universe and in some\nsense it's it's quite simple. France\ntalks about this. He talks about the\nkaleidoscope effect that you know we\nhave the generating function and then it\ngets composed together in a myriad of\ndifferent ways and we see the\nkaleidoscope and intelligent people can\ndecompose the kaleidoscope back into the\noriginal generating function. and and\nyou've also said that natural data in\nparticular is is quite lowdimensional or\nquite simple and that neural networks\nprefer simple data and even that I want\nto take a slight issue of that because\nwhat I what I find is that neural\nnetworks in the early stages of training\nprefer simple data and when you continue\nto train them they they seem to\ncomplexify and complexify and learn more\nhigh frequency data. So how how does how\nhow do you think about that?\nIt's a great question. So it's a really\nimportant observation that neural nets\ntend to learn structure before they\nlearn noise and things like this.\nThere's a question\naround to what extent being able to fit\nnoise is hurting their generalization\ncapabilities. So there's this other\nphenomenon called benign overfitting\nwhere the model fits typically a mixture\nof signal and noise but the noise being\nfitting the noise doesn't significantly\ndegrade its generalization performance.\nAnd this is often seen as something\nthat's specific to deep learning. And in\nthe face of everything that we know\nabout generalization\nand that's partly because classical\nframeworks for trying to understand\ngeneralization like VC dimension and rat\nmacro complexity are essentially\nmeasuring a model's ability to fit\nnoise. However, there are other\ngeneralization frameworks like packbays\nand countable hypothesis bounds that we\nexplore which don't penalize an\nexpressive hypothesis space and instead\ntry to understand what sorts of soft\npreferences the model has for certain\nsolutions over others. And we've been\nable to achieve fairly tight bounds on\nthe generalization performance of these\nlarge models using something called a\nSolomonov prior. And so a solomonov\nprior says that we actually have a\nmaximally overp parameterized model. We\ncan represent every possible program on\na computer but we have exponentially\nstronger preferences for solutions with\nthat have what are called called lowcom\ncomplexity and so they're very\ncompressible. The komograph complexity\nis the the shortest possible program\nthat can generate our hypothesis.\nAnd the fact that we're able to get\nthese tight generalization bounds for\nthese very large models and in fact the\nmodels get sort of better uh uh sorry\nthat the generalization bounds get\nbetter as we make the models larger um\nsuggests that this is not a bad\ndescription of how these models are\nactually behaving. So um doing induction\nwith a Solomon of prior is called\nSolomon of induction. And so um it seems\nthat when we make these transformers for\ninstance very large we're combining this\nexpressiveness with this strong\npreference for local macro complexity\nsolutions. Another observation that's\nthat's been made is that models are\nbecoming increasingly general purpose.\nAnd so 20 or 30 years ago, the\ntypical prescription was to encode as\nmuch expert knowledge as possible into\nthe model you're constructing and tailor\nit very specifically to the problem that\nyou're considering because of results\nlike the no free lunch theorems that say\nthat every model is equally good in\nexpectation over all problems sampled\nuniformly from this distribution over\nall problems. Um there are several no\nfree lunch theorems. Another one says\nthat a single learner is not going to be\ngood on all problems. Now, these results\nare mathematically correct, but they\ndon't really correspond to the real\nworld data generating distribution. The\nreal world is a small corner of all\npossible data sets. It's not drawn\nuniformly from that distribution. If you\nwere to draw data sets from that\ndistribution, you would mostly get\nnoise. And so, I guess the question is,\nwell, what what is the real world data\ngenerating distribution really like? And\nit seems like there is a bias towards\ngenerating data with lowgrav complexity\nand our models share that bias and this\nis why we've seen increasingly general\nsystems. So we've moved from feature\nengineering basically hard coding\nstructure into our models uh to more\nmodality specific models and\narchitectures. So convolutional neural\nnets for vision, recurrent neural nets\nfor sequences, language etc.\num MLPS for tabular data and regression\nto transformers for almost everything.\nAnd so this isn't to say that\ntransformers have achieved general\nintelligence. Absolutely not. But\nthey're relatively speaking more general\nthan the predecessors. And we have seen\nthis kind of movement towards\nincreasingly general models. And our\ncontention is that this has been made\npossible by aligning with the real world\ndata generating distribution which seems\nto have a bias for lowcom complexity.\nAnd we had this paper on um no free\nlunch theorems\ninductive biases and come complexity.\nAnd I think one of the most surprising\nfindings in that paper was that um\nconvolutional neural nets which were\nclearly designed for image recognition.\nAnd so they have locality and\ntranslation echo variance and so on\nprovably have inductive biases for\ntabular data shaped as an image and the\nonly possible reason that could be the\ncase is because they both sort of share\nthis bias for lowcom complexity and that\nbias gets stronger as we make the model\nbigger.\nOkay. So I I have actually two questions\nabout this. Um and it's about the kgarov\ncomplexity. So I think if I heard you\ncorrectly on the one hand you're saying\nthat just stock neural network training\nof today so just transformers SGD\nbatchorm whatever people are doing seems\nto empirically\nexhibit bias towards lower kgav\ncomplexity models and this shows up by\ntheir generalization kind of falling\nwithin you know this bound you know that\nyou found and then you also mentioned\nkgav induction which I I believe would\nbe explicitly introducing, you know, a\nkind of penalty term or a risk term, an\nobjective part of the objective function\nthat has to do with Kgrov complexity and\nkind of maybe pushing the model a little\nbit further towards you know simple I\nthink you're talking about both. Um so\nmaybe if you could elaborate and also uh\nyou know where is this this bias towards\nsimplicity coming from like everybody\nknows the algorithms like which part of\nthe algorithm is is inducing this\nsimplicity bias is it because we're\nusing floatingoint numbers IE E or what\nlike where's it coming from\nthis is largely an open question\nalthough there are some intuitions and\nthis is something I'm really excited\nabout pursuing further in my research\nwhere does the simplicity bias\nespecially from scale originate there\nare some intuitions. So geometric\nintuitions around the loss landscapes\nfor the objectives that we use to train\nthese models. And so um when we're sort\nof minimizing training loss, we can try\nto geometrically understand the\nproperties of this landscape that we're\nminimizing. So it's been observed for\ninstance that flat solutions, meaning\nsolutions where you can perturb the\nparameters by some amount but retain low\ntraining loss, tend to generalize better\nthan sharp solutions that have the same\nvalue of the loss. And you could make a\ncompressibility argument for why that's\nthe case. Flat solutions don't need to\nbe represented with as much precision\nand so they're more compressible.\nAs you grow the size of these neural\nnets, the relative volume of these flat\nsolutions starts to exponentially\ndominate the volume of the sharp\nsolutions. And so you can imagine\nhoristically\nwhy is that? I mean do we know why or\njust it's just an empirical observation?\nYes. So we know why to some extent but\nit's our understanding is is somewhat\nhoristic. And so you could imagine for\ninstance having a region of the loss\nsurface with radius RA that's flat. So\nyou perturb your parameters within that\nradius and you have a low value of the\nloss and another region with RB that\nthat's that's you know or RB is much\nless than RA.\nNow as you grow the number of parameters\nin your model D RA to the D starts to\nreally dominate relative to RB to the D.\nM\nand we seem to observe this in practice.\nSo there was a result not from from my\ngroup um from Tom Goldstein's group at\nthe University of Maryland where they\nwere trying to understand\nthe role of the implicit biases of SGD\nand stochastic optimization in\ngeneralization.\nAnd so quite often uh as you perhaps\nhave alluded to SGD is often thought to\nbe a really integral component of\nachieving generalization in deep\nlearning. There's this idea that we have\nthese very complicated objectives that\nwe're minimizing that are very\nnon-convex etc. and SGD somehow saves\nus. That the implicit biases of SGD\nnavigate our procedure through some\nregion of the loss landscape that\nrepresents low loss solutions that do\ngeneralize rather than low loss\nsolutions that don't generalize. Well,\num well, it turns out you can actually\ndo full batch gradient descent and\nachieve pretty comparable generalization\nto what you would get if you're using\nSGD even if you don't try to make that\nimplicit regularization explicit in the\nloss. And there was another paper that\nshowed that if you even do guess and\ncheck, so you just randomly sample your\nsolution vector and then stop when your\nloss is below a certain threshold, the\ngeneralization will also be fairly\ncomparable to what you get if you use\nSGD or atom. And so that corresponds to\nthis geometric intuition. If you're just\nsort of throwing darts at the loss\nlandscape and then you stop as soon as\nyou have loss below a certain threshold,\nyou're much more likely to be in this\nregion of low loss and good\ngeneralization than low loss and bad\ngeneralization as you increase the\nnumber of parameters in the model. Now,\nthis is not an airtight argument. So,\nthere are lots of ways that you can\nincrease parameters in models and not\nreally influence the geometric\nproperties of the loss landscape.\nHowever, it does seem to correspond to\nwhat we observe empirically and um not\njust in terms of results like guess and\ncheck. If we look at something like\ndouble descent in that second descent,\nwe can measure something called the\neffect of dimensionality of these\nmodels. So that's the number of\nrelatively large values of the hessen\nwhich is sort of the number of sharp\ndirections essentially in the loss\nlandscape. And so that that decreases as\nas we as we make the model bigger.\nSo interesting. And to the second\nquestion about explicitly introducing\nyou know a penalty term for a koma grav\ncomplexity have you done much of that\nlooked into that is that useful or just\nnot necessary it's something I've been\nthinking about it's very hard to\noperationalize so you can\nevaluate these bounds by computing an\nupper bound on come complexity measuring\nthe compressed file size of your model\nafter training.\nMh.\nBut\nin order to\nuse that prior as some sort of\nregularizer, you would need to be\nconsidering sort of compression of some\nwhole set of different hypotheses, not\njust a single hypothesis that's found by\nthe model.\nAnd so there's an open question of how\nyou could try to operationalize that\nlike Solomon of induction is sort of an\nrepresents an idealized learning system.\nIt's not something that we can really do\nexactly in practice. It seems that\nneural nets are sort of approximating it\nand that's evidenced by these bounds and\nhow they're able to tightly characterize\ngeneralization behavior of these models.\nBut um it it's hard to sort of turn into\ninto some kind of regularizer. Uh I'm\nalso interested in how we can go beyond\nthings like grav complexity. So um come\ngraph complexity doesn't distinguish\nbetween incompressibility due to\nrandomness um so like noise for example\nin our data versus structural complexity\nand Scott Aronson actually had a really\ninteresting blog post related to this\nabout 10 or so years ago uh where he was\nimagining a system he had this kind of\nphysics analogy where you have coffee\nand cream and initially they're\nseparated liquids and you start to stir\nthem together and as you do this the\nentropy of this system is increasing.\nI've seen that one.\nThat was shared on our Discord recently.\nOh, very nice. Okay. Um and\njoin our Discord.\nI'd love to. Yeah. Um we actually have a\nYeah. Do I get ahead of myself? We have\nan idea of course for how we can do\nthis, but um so the the blog post uh\nsort of compares this to so the entropy\nof the system is increasing. The comra\ncomplexity is increasing. But the\nintuitive sophistication of that system\nis kind of non-monotonic. Initially it\nhas low entropy, low sophistication.\nthen sort of intermediate sophistication\nand entropy and then again sort of like\nsorry and then high entropy and kind of\nlow sophistication. And so you can think\nof this in machine in a machine learning\ncontext in terms of reasoning about the\nvalue of data. So if I sample data from\nlike a uniform random uniform\ndistribution that's going to be very\nincompressible. I'm going to need to\nmemorize it. It's sort of uncorrelated.\nThis could be useless for learning a\nrepresentation for training my model. I\ncould alternatively imagine some sort of\nsophisticated cellular automa problem,\nsome sort of game of life problem with\nvery sophisticated generalization rule\nor generation rules. Um that data\nactually could have an extraordinary\namount of value for learning a\nrepresentation. Um there was a paper\nthat looked at something briefly like\nthis called intelligence at the edge of\nchaos. And I think there are other\nresults like that that are coming out\nthat like you actually might want to\ntrain your models on this data with a\nlot of structural complexity even if you\nreally want you know in the end the\nmodel has some kind of aams razor bias\netc. Um, and so we've been thinking\nabout kind of measures of information\nthat might compartmentalize structural\ncomplexity and random complexity. And\nthis will help us reason better about\nthe value of data and developing priors\nsort of that are like Solomon prior but\nmight actually be more directly\naddressing the type of incompressibility\nthat we're interested in.\nYeah. I mean, so you said you said so\nmany interesting things. I mean, first\nof all, um, to Keith's point, you were\nsaying that this isn't a penalty term\nyet. Your hypothesis is that neural\nnetworks implicitly do this kind of um\ncompression which might be correlated or\nrelated to this chromograph complexity.\nSo many folks just kind of um they\nequate intelligence with compression and\nwhen I spoke with David Krakow he took\numbrage with that. He said you know\ncompression is a component of of\nintelligence but there are so many other\nthings going on as well and certainly\nwhen we look at things like the arc\nchallenge um there are so many possible\nsolutions. So a naive heristic of just\nselecting the simplest program isn't\nalways the best thing to do. There are\nmany possible selections you could make\nand you know so so you you you um\ndemonstrated this upper bound which used\nthis complexity term and in a sense\nthat's saying that um it could be no\nworse than this rather than it could be\nbut it could actually be so much better.\nAnd our empirical experience of deep\nlearning models is that they they seem\nyou know we call it shortcut learning\nbasically you know they they seem to\nhave found some superficial generalizing\nthing which does all the things you said\nwhen you mentioned the no free lunch\ntheorem. So you could take a CNN and you\ncould use it on tabular data you could\ntake a transformer and you could use it\non audio data. So it's almost like what\nwe've seen is that we've we've hit this\nfor want of a better word local minimum\nand it feels like we need something more\nto get to the real understanding of some\nof these problems. Does does that make\nsense?\nSo is compression intelligence? I think\nthis is a big debate right now. I think\nit is very closely associated with\nintelligence in a lot of ways. If we can\ncompress our data really effectively,\nthen in order to do that, we're\ndiscovering regularities that are\ntypically going to help enable\ngeneralization. And in some sense,\nphysical laws, for example, or a great\ncompressed representation of reality.\nAnd so I think that compression is\nreally intricately connected with what\nwe mean when we talk about building\nintelligent systems. There are instances\nin which this can go very wrong like in\nshortcut learning where for instance you\nmight have some spirious correlation\nlike maybe every time there's a blue\npixel in an image the label is a bird or\nsomething like this and so the model you\nknow forgets about the foreground and it\njust looks for some feature in the\nbackground in terms of generalizing on\nthat distribution that's actually not a\nbad idea that actually is the right\nthing to do so aams razor is really a\ngood principle but if we go out of\ndistribution so we see like a bird in\nyou know with a volcano or something\nbehind it or in a room um then you know\nit's not going to arrive at the right\nlabel. So there's this question of like\nare there other principles of induction\nthat will lead to greater robustness\nunder distribution shifts. I think in\ngeneral aam's razor still is the right\nthing to do. Like there are many cases\nwhere a compression\nwill actually not give you what you want\nmore broadly if you move beyond that\ndistribution. But in absence of\nadditional information,\nit seems like you can't really do better\nas a guess as to what's the right\nstrategy. And so I still think AAM's\nrazor is a very robust principle of\ninduction and perhaps one answer is just\nmore data. I think this will work well\nin some cases and not in others. Uh so\nwe have some work in progress on trying\nto build transformers for matrix\noperations. So this is actually kind of\nanalogous to the work that people have\ndone on transformers for\nmultiplication, addition, things like\nthis. So uh I guess you know these\nsystems LLM seem very impressively\neffective in some instances uh in being\nable to write and code and uh solve a\nvariety of different problems we didn't\nreally expect them to be able to solve.\nUh in other instances they're just\nshockingly bad. Um so you know how many\nRs are in strawberry like all reversing\nstrings counting and just basic addition\nand multiplication etc. And so there's\nthis debate I I think about whether we\nshould be giving them tools to do those\nthings. Like why not just give them a\ncalculator? I mean whatever they learn\num to do in terms of adding and\nmultiplying numbers is still not going\nto be as efficient as giving them access\nto to a calculator. Um so why don't we\njust do that? um or like you know is\nthere some sort of greater value in\nhaving to sort of learn a representation\nthat can do some of these things because\nmaybe even though we'll always want to\nuse a calculator if we're multiplying or\nadding numbers being able to do\nsomething like that reasonably well\nmight transfer into other settings that\nwe don't really anticipate and I'm\nprobably more in this category like I\njust think intellectually we should try\nto figure out how to do this without\ntools and so this work that we're doing\non transformers for matrix operations I\nthink is quite analogous so You could\nargue that addition and multiplication\nare just fundamental primitives for\ntrying to build intelligent systems.\nWe're not as good at it as a calculator\nis, but it might be important for us to\nhave some ability at being able to do\nthese things. The same could be said of\nmatrix operations. This is really the\nbackbone of all sorts of different\nlearning algorithms like Gaussian\nprocesses, which we discussed a little\ninvolve solving linear systems with a\ncoariance matrix, computing log\ndeterminance, etc. degenerative models\nlike normalizing flows involve log\ndeterminance um dimensionality reduction\nlike PCA etc involve other matrix\noperations and so they're just really\nubiquitous as kind of a primitive for\ntrying to build learning algorithms and\nintelligent systems and so I think if\ntransformers are ultimately going to\nbecome some sort of general intelligence\nthen they ought to be able to to be\ncompetent at these types of operations\nand so we were kind of representing then\nmatrices is as sequences of numbers and\nthen having the outputs be things like\nthe maximum iggon value or you know\nsolution to a linear system or whatever\nother operation we were considering that\nthe spectrum of igon values and we found\ninterestingly that when you\ntrain this approach on Gaussian random\nmatrices so you just sample every entry\nof the matrix from a standard normal\ndistribution. It will do fairly well at\nin distribution matrix operations. So\nother matrices sample from that\ndistribution that it hasn't seen before\nbut extraordinarily poorly even if you\ngo slightly outside of that\ndistribution. So give it an identity\nmatrix all ones on the diagonal zero\neverywhere else that'll have pretty low\ndensity has support under this Gaussian\ndistribution where matrices but pretty\nlow density under it hasn't seen\nsomething very much like that it will\njust completely fail. uh it won't do\nanything reasonable. And so we\nconsidered a number of interventions\nlike blooping sort of adaptive test time\ncomputation um enriching the training\ndistribution very significantly. So\nhaving sort of this like in some um uh\nspace of structured matrices that we\nwere sampling from so all sorts of\ndifferent matrix structures to toplets\nchronicer etc block diagonal low rank um\nand a variety of other interventions.\nAnd interestingly, we found once we had\ndone that,\nthis approach actually was able to\ngeneralize even to matrices that were\nout of distribution for this fairly rich\nsort of training set that we've created.\nAnd so it seemed to move more towards\nlearning an algorithm rather than just\ndoing statistical interpolation on the\ntraining data. And so this is I'm an\noptimist by nature. So I I am very happy\nabout the fact that we can use data to\nmove towards things like algorithm\ndiscovery, but now I remember why I\nstarted talking about this. So I think\nthere are instances where it's going to\nbe hard. So autonomous driving is is an\nexample where like you have outliers,\nbut they're different each time. So just\ntraining on the outliers isn't going to\nbe useful because they're going to be\nnew outliers that look very different\nfrom those outliers. And I just have the\nintuition that more data alone is not\nreally the answer to building robust\nautonomous driving systems.\nYeah. I just want to say um as far as\nwhy not just give them tools. I mean my\nanswer folks is because those have to be\nprogrammed and built by people. And the\nwhole point here is to allow machines to\ndo their own programming right machine\nlearning. And if they can't if they\ncan't even learn to do multiplication\nreliably and to generalize from decimal\nmultiplication to binary or hexodimal\nand nine digits to 36 digits, you know,\nwhat hope do we have that they're going\nto discover relativity or non-newtonian\nmechanics or any other kind of frontier\nfrontier things, right? I mean, isn't\nthat part of the goal here?\nAbsolutely. And I think being able to do\ncertain things well, we're finding might\nsurprisingly relate to doing other\nthings very well. And so I think we have\nseen this to a large extent with LLM. So\nwe had a paper where we just took a text\npre-trained LLM off the shelf and\napplied it to time series forecasting.\nSo we did everything naively. We weren't\neven really intending for this to be\nlike a proper project or a paper. We\nwere just curious, you know, if you just\ngive GPT like a sequence of numbers\nnaively encoded to strings and have it\nextrapolate the next sequence of string\ntokens, um, how would it compare to\npurpose-built time series forecasting\nprocedures? and it just worked way way\nbetter than we thought it could possibly\nwork. It didn't even really make sense.\nUm, and so we did sort of turn this into\na proper project. We did a little bit of\nwork on trying to improve the\ntokenization and think about uncertainty\nrepresentation etc. But most of that\npaper uh it's called large language\nmodels or zeroot time series forecasters\nwas focused around trying to understand\nhow this is even possible. Um, and in\nthe end it did start to feel a bit more\nlike maybe it's not just that you can do\nthis, maybe you should do it in some\ninstances. Like they did quite well on\non a variety of benchmarks. And to me,\nthis suggests with like a proper\ndedicated research effort on LLMs for\ntime series. You know, we could see\nthese systems actually working a lot\nbetter than the purpose-built models.\nAnd so what that shows is being able to\npredict the next words in sentences can\nactually transfer to being able to do\nother things like time series prediction\nreally well. And we had another paper on\nLMS for materials generation which was\nkind of similar. So we took a text\npre-trained LM off the shelf like a\nllama 2 model at the time and um we\nfine-tuned it in this case on atomistic\ndata represented as text. the locations\nof atoms and energies and things like\nthis. And the resulting system was able\nto generate inorganic crystals with\nfavorable properties better than these\npurpose-built approaches and even\nfoundation model approaches that had\nbeen trained um on that domain specific\ndata. And so one of the takeaways from\nthat project was like the textbased\npre-training was an indispensable\ncomponent in being able to achieve good\nresults on materials generation. And we\ntried to understand also in that paper\nwhy that was the case. We were\ncollaborating with some some chemists at\nfair in California and they were just\nvery curious about all MMS and\nfoundation models. So they were willing\nto sort of humor us a bit and help us\nsort of see what we could do but they\nwere just like skeptical throughout the\nproject until we we saw the results. And\nit's like well you can't deny the\nresults are great you know why is this\nhappening and part of it was that in\nbeing able to predict the next tokens in\nstrings you are learning principles of\ninduction like AAM's razor and so how do\nthose sort of manifest themselves in in\ncontext learning. So what it means is if\nor and in fine-tuning. So what it means\nis these models are going to be\npredisposed to discovering\ncompressible representations and that\nmeans for example salient symmetries. So\nthis was a problem where there was a\nrotation invariance and these models\nactually were very quick to learn these\nkinds of invariences because of the\ntextbased pre-training it. It sort of\ninstilled this principle. And so I think\nthis is also an example of how\ncompression can be a broadly applicable\nprinciple for induction. it's at sort of\nthe right level of abstraction that you\ncan start to see more relatively\nuniversal behavior. So for instance,\nthere are some theories on the success\nof foundation models that suggest that\ndifferent problems are just different\nprojections of some underlying reality.\nSo like platonic representation\nhypothesis is an example of this. And so\nyou can represent an image with pixels\nor with words and you know train the\nrespective models on those different\nmodalities and they learn similar\nrepresentations. I think that can be\ntrue in some instances but I also think\ndifferent problems are often truly quite\ndifferent from each other in terms of\ntheir low-level structure. So like if we\nhave molecules there's rotation\ninvariance maybe some other problem some\nimage recognition problem like character\nrecognition we might have translation\ninvariance doesn't matter if the two is\non the left of the screen or the right\nof the screen the label is still a two.\nThese are very different low-level\nfeature representations. So like the\narchitectures that you would typically\nuse um for each of those modalities and\nto respect those different types of\ninvariences would look very different.\nBut what they have in common with each\nother is they're both ways to compress\nthe respective problems that they're\nbeing applied to. So if you have a model\nthat has this compression bias, then it\ncan discover those salient symmetries.\nAnd we had this really surprising\nfinding that vision transformers\nactually can be more translation\nequariant than convolutional neural nets\nafter training. Which sounds impossible\nbecause connets by design are con are\ntranslation equivariant but they're not\nexactly translation equavariant because\nof aliasing artifacts and edge effects\nand things like this. And so this other\nmodel, this transformer with no explicit\nconstraint whatsoever, just a soft bias\nthat manifests itself increasingly at\nscale is able to discover a solution\nthat has lower equavariance error than\nthe convolutional neural net, which is\njust absolutely remarkable. So to come\nkind of full circle to your question, I\nthink that we're discovering more and\nmore that being able to solve certain\nproblems really well will translate in\nperhaps unexpected ways to being able to\nsolve other problems well. And so I\nthink especially when it comes to things\nlike addition and multiplication and\nmatrix operations like these are just\neven like obviously like fundamental\nprimitives for building intelligent\nsystems like even if we can use tools\nfor those things we want our\nrepresentations to be somewhat competent\nat them because it's going to be useful\nfor all sorts of other things we\nprobably haven't anticipated.\nSo AAM's razor has been mentioned\nmultiple times in in our in our in the\nlast few segments. I want to dive into\nthat bit because I love AAM's razor and\nI first came to really understand it\nwhen I became a basian and it's nice to\nhave a fellow you know arch basian uh\nand to talk with here because in basian\ninference if you will um AAM's razor has\na very explicit form like mathematical\nform and it's marginalization right it's\nsaying like if you have all these\nparameters around and you you compute\nthe average you do this integration over\nall these parameters that really that's\ntelling you like the probability of the\nmodel like given your you know full\nparameter space having been integrated\naway and that's really where you get you\nknow let's say a penalty for complexity\ncomes into play with with\nmarginalization. And it's like if you\nmake if you expand your model's\nflexibility or you make it more\ncomplicated without any gain in\ninference well then that kind of like\ncounts against you or without any gains\nof some weird simplicity in the form of\nyou know less curvature you know things\nlike that. And I was looking at some of\nyour talks which were really great from\nfrom five years ago. These this basian\ntutorials, basian deep learning\ntutorials you had, I think when you\nfirst when you first went to NYU. Really\nenjoyed them. Um and there was a lot of\ntalk about marginalization and the\nimportance of it and the intuitions that\nyou gained from that. Um, and I think\nbut more recently that's played less of\na role in the conversation or people\nhave just given up on trying to do the\nintegrals and they're just back to doing\nmaximum likelihood and maybe with some\nhacks and the objective function. So I'm\njust kind of curious, you know, as a\nbasian, your journey going from\nunderstanding like the beauty of\nmarginalization and the aams razor built\ninto basian inference versus what you do\nas a practitioner, what you see people\ndoing in practice, how it's playing out\nin the field or not. And maybe is there\na future in which marginalization and\ndoing these computational intractable\nintegrals or something could play a role\nin driving like even better, you know,\nmachine learning and inference. I'm so\nglad you asked. So there's almost\nnothing I like more than talking about\nbasian inference. And I'm so glad you\nsaid marginalization because I feel like\nthat's often overlooked. Like when\nsomeone says basian probably the word\nthat comes into your most people's mind\nis prior and is the prior good and how\ndo we know what would a good prior be\netc and they get worried about that but\nreally the prior is not the defining\nfeature of what it means to be basian.\nIt instead being basian means that you\nwant to represent the honest belief that\nyou have uncertainty over what solution\nis correct given a finite data sample.\nAnd so if I'm trying to estimate, you\nknow, the bias of a coin that I'm\nflipping, if I flip it once or twice,\nregardless of whether it comes up, you\nknow, heads twice or something like\nthis, I ought to have some uncertainty\nover the bias still. And that\nuncertainty is manifested through this\nprocedure called marginalization. So I\nthink we can think of regression to get\nsome intuition of this. So we could\nimagine like a bunch of different points\non a sort of y-x plot and maybe the\npoints look roughly like they're on a\nstraight line but not quite. We could\nalso imagine that there are many\ndifferent curves that will perfectly run\nthrough all of those points that all\nlook different from each other. Some of\nthem will be very wiggly, some of them\nwill be slowly moving, etc. Some of them\nwill basically just be like a straight\nline. And we wouldn't be able to say,\nwell, we know with 100% certainty it's\nthis curve that's the right description\nof our problem. But that's exactly what\nwe're doing, you know, 99 100 minus\nepsilon% of the time when we're training\nmodels in deep learning. That is not an\nhonest representation of our beliefs.\nAnd it's going to become a bigger and\nbigger problem the more expressive our\nmodel actually is because that means\nthere going to be many more different\nsettings of parameters that are\nconsistent with what we observe and\nwe're just betting everything on one of\nthem. And probability theory says no,\nthat's that's just wrong. And that's not\nwhat you should be doing like the sum\nand product rules of probability say you\nshould be doing marginalization. And so\nthat's all to say that basian\nmarginalization basically looking at all\npossible solutions that can be expressed\nby your model class weighted by their\nposterior probabilities\nis going to be most important when we\nhave a model that's very expressive has\na lot of parameters. So deep learning\nfor example especially relative to the\nnumber of data points that we're\nconsidering. And I think often people\nthink of it in in the opposite way that\nlike oh maybe basian methods are most\nrelevant in the realm of classical\nstatistics like if you're doing some\nlogistic regression and and you know\nmaybe you want to represent some\nuncertainty etc but it's not really\nreally suited for deep learning. It's\nreally the opposite. Um the challenge\nthen is how do we do this in a way\nthat's computationally tractable and I\nthink like with most things in life the\nanswer of course is nuance. So um it's\nnot that like we can either be fully\nbasian or not basian at all. Let's just\ntry to do what makes sense given the\nresources available to us. And so like\nif we're just doing standard training,\nyou can actually view that as a very\ncrude form of marginalization where\nyou're saying, okay, I'm representing\nthe the posterior as a point mass around\nthe most likely setting of parameters.\nAnd then we can say okay well maybe the\nposterior which is really like so the\nloss functions that we're minimizing are\nbasically negative log posteriors. Um\nmaybe the posterior looks nothing like a\npoint mass and it looks nothing like a\nGaussian. It's very multimodal. It's\nvery messy. But we can still do a better\njob of representing it with a Gaussian\nthan with a point mass. And so let's use\na Gaussian approximation. And indeed\nwhen we do that we often see better\ngeneralization because we're\nrepresenting all these other\ncomplimentary and compelling\nexplanations to our problem. And then we\ncan keep taking it from there. Okay,\nmaybe we can develop some MCMC procedure\nthat will explore the loss landscape and\ncapture something much richer than just\nunimodal Gaussian structure etc. And\nagain we see improvements in performance\nin doing that. And so I actually think\nthat basian methods have been an\nextraordinary success story in deep\nlearning and beyond.\nBut we don't hear as much about them now\nas we did maybe 10 years ago. And I\nthink there are a number of reasons for\nthat. So to some extent\nthey're a victim of their own success.\nSo I think there there was some\nlowhanging fruit in being able to do\nbetter approximate marginalization in\nneural nets around 2015 and there was\nreally a lot of progress between about\n2015 and 2020 in achieving increasingly\nbetter results uh that will also be kind\nof computationally e efficient. So there\nare procedures like one we developed\ncalled swag for instance which was based\non uh some insights we had into the the\nstructure of these loss landscapes that\nwill allow you to do basian\nmarginalization without really any\nsignificant additional cost on training.\nIt's a bit more expensive a test time\nand so on. It's also deep kernel\nlearning which uh basically just\nrequires a single forward pass through\nthe model and you get sort of some\nrepresentation of epismic certainty. Um\nand these procedures were adopted in\npractice. And so if I go like to some\nmore domain specific conference like\nsome workshop or conference on materials\nengineering and things like that, I'll\nsee lots and lots of talks using basian\noptimization, Gaussian processes, neural\nnetworks um with uh epistemic\nuncertainty representation etc. So this\nis really useful in practice but it's\nvery hard to kind of go beyond the\nuseful approximations I think we\ndeveloped without doing a significantly\num without making sort of a really\nsignificant you know 10ear kind of style\nmoonshot landing kind of investment in\nthose directions which I think is worth\nmaking. Um secondly there's the advent\nof LLMs and that did sort of change\nthings in practice. So like we went from\nmillion parameter models to billion\nparameter models and the role of\nepistemic uncertainty is also a little\nbit less clear I think in some instances\nwhen we're working with LLMs and so uh\nthis is just a real technical challenge\nand I think we're just sort of barely\nstarting to scratch the surface of how\nto think about it but the motivation is\nabsolutely there I think it's really\ngreater with these big models than it is\nfor virtually any other model class and\nin this sense I don't know how anyone\ncould not be a basian like How could\nanyone say I don't want to represent\nepistemic uncertainty? I don't like\nusing jargon. So there's like there's\naliatoric uncertainty which means\nirreducible uncertainty often associated\nwe have given a data set we have some\nnoise on that data set there's nothing\nwe can do about it. That's ali\nuncertainty. Epistemic uncertainty is\nuncertainty that's reducible with more\ninformation. And we believe that\nuncertainty is there. And there's even\nan argument that that might be the only\nuncertainty in the world, right? Like\nthings that seem intrinsically random,\nlike I could roll a dice and it might\nseem like, okay, there's a one six\nprobability that it'll land on any one\nof the sides. But if I had enough\ninformation, like if I could sort of\nexactly know the strength of the throw,\nthe wind in the air, the friction on the\ntable, and I had the right physical\nmodel, I should be able to predict\nexactly what side it's going to land on\nwith each throw. So that's just saying\nokay the more information we gather the\nless uncertainty we have even for this\nprocess that somehow seems intrinsically\nrandom and physicists modern physicists\nnow you know probably do believe that\nsome aspects of the universe are\nintrinsically random radioactive decay\nand things like this um but many others\nhave not historically like Einstein was\na determinist um and so he would have\nbelieved that the only uncertainty there\nis is epistemic uncertainty and I think\nfor practical purposes that's not an\nunreasonable belief um and so it's\nabsolutely crucial that we try to model\nit in some way to not do it is just to\ndo something that's mathematically\nincorrect and could come at a huge cost.\nThere's also a change I think that's\nhappened in this movement to LLMs in the\nway that people think about data. So it\nused to be the case that you kind of\nhave a fixed data set and then you throw\neverything you can at it to get the best\npossible performance. Maybe you have\nsome scientific problem and you just\nreally want to do well and you're\nlimited by computation and other things\nto some extent but it's not the case\nthat you have this sort of tradeoff to\nmanage for a given computational budget\nbetween the size of the data and the\nsize of the model and so that's\ndescribed by these scaling laws like\nchinchilla scaling laws and that is\nstarting to become the assumption in\nmainstream machine learning that you\nhave almost like an arbitrary amount of\ndata and you have a fixed computational\nbudget and you want the best possible\nperformance under that computational\nbudget. In that case, you know, rather\nthan trying to be more basian or\nsomething like this, it could actually\nmake sense just to use more data. Um,\nbut the but the other huge benefit\nthough is because you mentioned um, you\nknow, that for whatever reasons that\nwe're still trying to understand these\nlarge, you know, large neural networks,\ndeep neural networks have this kind of\nsimplicity bias towards simplicity,\nright? And marginalization is the\nultimate AAMS razor. You know, it's like\nAAM's guillotine or something that\nreally can force models, you know, force\nyou to select, let's say, the simplest\nmodel that's consistent with the data in\na in a sense. So, I think there's\npossibly more to gain there, like just,\nyou know, trying to understand in\ncomputationally tractable ways how to\npull in more of those effects of\nmarginalization like on inference. I\nmean, is that true?\nAbsolutely. Yes. So basian\nmarginalization has this automatic aams\nrazor bias and I actually started\nworking on basian deep learning\nironically when I saw a talk on\noptimization which might be viewed as\nalmost the opposite of being basian. So\nJorge Nosadol gave a talk at Cornell\nwhere I was initially a faculty and he\nwas presenting on small batch biases\nbasically of stochastic optimizers and\nhe had this figure where he was\ncontrasting flat optima with sharp\noptima and he had this horizontal\ntranslation of the the loss landscape\nand under that translation the flat\nminimum still had reasonably good test\nloss but the sharp minimum had very bad\ntest loss and so he was saying this is\nwhy we want to do stocastic optimization\nwith small batches because it's going to\nbe more likely to find these flat\nsolutions and in some sense like the\nother optimizers like his LBFGS and so\non were like too good they were like\nconverging to these sharp minima and if\nlike you just care about minimizing loss\nthen maybe that's not bad behavior but\ngeneralization depends more than on just\ngetting a low value of loss at least\nwith the loss functions that we're using\nand so I saw that figure and I thought\nwell this is actually really great\nmotivation to be basian because if\nyou're being basian you're not just\nbetting everything on one solution\nyou're integrating under sort of a\nflipped version of that curve. And so\nautomatically most of the volume will be\nin these flat regions. And that's so\nmuch more elegant than approaching this\nfrom the optimization perspective where\nyou have to say then okay I need to\nrigorously define what it means to be\nflat like maybe it's going to be like\nthe largest value of the hessen. Maybe\nit's going to be flat\ntensor or something like that.\nExactly. Yeah. Maybe flatness in random\ndirections um versus sharpest\ndirections. uh uh maybe it needs to be\nparameterization invariance like that's\nthat's been often a criticism of of\ndiscussions around flatness etc that\nit's most ways of measuring flatness\naren't parameterization invariant I have\nthoughts about that separately but\nanyway you you might then look at fisher\nmatrices or something okay so then once\nyou've decided on what it means to be\nflat and no one will completely agree\nwith you and most people will will\nstrongly disagree whatever you choose\nthen you have to decide like how much am\nI going to penalize sharpness no one\nwill really know what to do about that\neither and like the idea is you want to\nbuild a loss function is a better proxy\nfor generalization by accounting for\nthings like flatness, but it's just this\nreally messy rabbit hole. Whereas, if\nyou're just doing marginalization, this\nis all happening under the hood. You\ndon't need to worry about it. And so,\nthat's really elegant. And then there's\nAAM's razor in terms of model selection.\nAnd I I'd recommend that everyone check\nout chapter 28 of David Mai's\ninformation theory and and uh inference\nlearning algorithms book. Uh it's titled\nAAM's Razor. And I don't agree with some\nof the stuff in that chapter. And in\nfact, we wrote a paper about this on\nbasian model selection and marginal\nlikelihoods. But it's it's a very\nbeautiful description of automatic aams\nrazor. And he has these just\nextraordinary visualizations to try to\ndemonstrate how that's possible. So he\nhas like these also, you know, everyday\nexamples as well like that this idea\nthat maybe you have like a block behind\na tree, but if you had x-ray vision, it\nwould appear to be like two blocks of\nequal height and color or 10 blocks or\nsomething like this. But given that you\ndon't have X-ray vision and you don't\nthink this is a trick question, you'd be\npretty confident it's got to be one\nblock. And this seems like some\nmanifestation of AAM's razor and you\nmight try to rationalize this by saying,\nwell, it would be just a remarkable\ncoincidence to have two blocks standing\nnext to each other of equal height and\ncolor. But he argues that this is\nactually a quantifiable consequence of\ndoing basian marginalization. And so he\nbasically has this conceptualization\nwhere you have all possible data sets on\na horizontal axis and the probability of\ngenerating a particular data set under\nyour model on the vertical axis. And\nthat's the marginal likelihood, the\nprobability that you would generate your\ntraining data under your model prior.\nAnd so if you have the oneblock model,\nyou're not going to be able to generate\nvery many different data sets. So most\nof the data sets on the horizontal axis\nhave no support, no PD given M. But the\nones that it can generate, it's going to\nhave to give a lot of probability for\nbecause this is a proper normalizable\nprobability density. Similarly, if you\nhave the 10 block model, you can\ngenerate all sorts of different types of\nobservations, but you're going to have\nto sort of spread that mass more thinly\nbecause this is like a proper\nnormalizable probability density. And so\nfor a particular data set that's\nconsistent with both of these models,\nlike the block behind the tree example,\nthe oneb block model is actually going\nto have significantly more probability.\nAnd this is even not factoring into\naccount the idea that maybe we have some\nprior preference for simplicity and that\nthere's like we think that one block is\nmore likely than 10 block or something\nlike that. He's just like let's just\nforget about this ratio of prior odds\nand only consider this ratio of marginal\nlikelihoods. And so it's just a\nbeautiful demonstration of how basian\ninference automatically encapsulates a\nnotion of AAM's razor. This is just a\nfundamental question in science like if\nyou have arbitrarily many hypotheses\nthat are consistent with any number of\nobservations you can ever record how do\nyou choose between them what what what's\nthe principal way of approaching that\nproblem and the answer to some extent I\nthink is given by basian marginal\nlikelihoods there are um very subtle\nissues I think with some of the ways\nthat this can be done and so we wrote a\npaper all about that but largely\nspeaking it's um you know something that\nI think people should acquaint\nthemselves with because it it is sort of\ngetting at something very fundamental\nand it's has extraordinary practical\nvalue.\nWhat was that paper was that\nthe the one that we were sort of looking\nat. So we had a paper called basian\nmodel selection the marginal likelihood\nand generalization and so uh that paper\nis really trying to sell two sides of a\nstory. On the one hand, it's trying to\nconvince the readers that the marginal\nlikelihood is something quite\nextraordinary and that there's a reason\nwe should be interested in it because if\nit isn't, then you know, if you haven't\nheard of it, like why even read\ncriticisms of it if it's something that\ndoesn't matter. Um, the other part of\nthe paper is questioning whether it's\nanswering exactly the question we want\nto be asking when we're doing model\nselection towards trying to achieve the\nbest possible generalization. And so the\nquestion that the marginal likelihood\nanswers is what is the probability that\nmy prior generated the training data\nthat's different than\nwhat is the probability that my\nposterior after I've observed the data\nis going to lead to reasonable\npredictions.\nAnd so we can construct examples that\nreally illustrate this difference. Like\nyou could have a uniform prior over\nsolutions that you expect to be easily\nidentifiable from the data. And so the\nposterior will contract very\nsignificantly around something that will\nactually be quite reasonable and make\nreasonable predictions, but the marginal\nlikelihood will be really bad. And so we\nconstruct all sorts of examples where\nthere's actually a misalignment between\nthe marginaliz marginal likelihood and\ngeneralization for these reasons. Um you\ncould also overfit um even though you\nhave sort of some robustness to\noverfitting like if you're considering\narbitrarily many models you could just\nget super unlucky and have a model which\nis like a point mass prior around\nsomething that can only generate that\ntraining data set but isn't going to do\nanything reasonable. So the marginal\nlikelihood might prefer that but it's\nnot going to lead you to a model that's\ngoing to make good predictions. It is\nvery useful as sort of a horistic for\nmodel selection in many instances. It's\ngot incredible practical value in\nlearning things like hyperparameters\nthat control complexity in Gaussian\nprocesses. Um, and I think it often is\nthe right tool for scientific hypothesis\ntesting which is subtly different. And\nso you could this is actually a real\nhistorical example. So there was a\ndispute between statistitians. I think\nColumbia University thought that um\ngeneral relativity was not the\nexplanation for Mercury's irregular\norbit. Uh it's called Mercury's\nperihelion. There must be some hidden\nplanet or some orbital debris or\nsomething like that. And so basian\nstatistitians actually went ahead and\ncomputed the marginal likelihood\nassociated with general relativity\nexplaining Mercury's orbit versus these\nalternative hypotheses like um some\nmodification to Newtonian gravity etc.\nAnd because general relativity was so\nfalsifiable, like its predictions were\nso sharp and consistent with what we\nobserved, it had orders of magnitude,\ngreater marginal likelihood than\nsomething like a modification to\nNewtonian physics where you have to have\nsome distribution over the modification\nmight sort of enable you to explain what\nwe see, but it also is going to generate\nother data sets and so its mass is going\nto be spread more thinly. And so I think\nthat's actually a beautiful\ndemonstration of how the marginal\nlikelihood can be used for for\nscientific hypothesis testing.\nYeah. Are there any other heristics? I\nmean we've spoken about marginal\nlikelihood and and model complexity and\nso on and I'm just thinking of take the\ngame of life you know you have all of\nthese simple rules and what we do there\nis is we we do this you know this\nsequence of computations and it's\nirreducible as as Wolram would say and\nwe just look at the dynamics right I\nmean is is there something to that is it\nmy intuition is that just by looking at\nthe thing in isolation doesn't tell you\nsomething like actually running it you\nknow in in the real world over several\nsteps of computation that that's where\nthe information is about whether the\nmodel is good or not. Is is that is that\na fair intuition?\nThat's a great intuition and it's\nconnected with the marginal likelihood\nand um sequential coding and ideas and\ninformation information theory. So this\nis something we've been thinking about a\nlot actually in trying to go beyond\nkomograph complexity and very much\nrelated to your earlier question\nconnected to benign overfitting this\nobservation that models first tend to\nfit structure and then they start to fit\nnoise and so if we can think of some\nsort of like compute limited come\ncomplexity um maybe we could start to\nget some idea of how to do model\nselection and the marginal likelihood\nactually can be written as the uh so the\nmarginal likelihood is the probability\nof the data under your model condition\non your model And so you can use the\nchain rule of probability to write that\nas the product of P of DI given D less\nthan I essentially. Um and so you take\nthe log of that and it looks sort of\nlike the log under your sort of training\ncurve as you're observing more and more\ndata. And so these these things are all\nconnected together. Um and I think this\nis a very reasonable way of trying to\nunderstand how to do model selection\nproperly. Sort of thinking about\ncompressibility but also thinking about\num the dynamics of training. how like a\nmodel's representation evolves with\nincreases in computation.\nYes. And and David Ko I mean when he\ntalks about intelligence he says it's\nwas it inference adaptivity and\nrepresentation but when he was talking\nabout emergence he said it's about a\nfundamental reorganization in in the\nmicro substrate. So um you know take\nNabia Stokes for example it's a\nreorganization where this new higher\nlevel description now does you know is a\nbetter description than at the molecular\nlevel. And surely there must be a thing\nin training dynamics as well that you\nknow when if if there is some kind of\nemergent behavior um there would be a\nfundamental reorganization and then with\nsome emergentist optimization where you\nare looking at the macroscopic behavior\nyou would actually select that\nunderlying model which generated it.\nMhm.\nYeah. I mean I think that's that's that\ntype of reorganization is is what's\nhappening in and double descent and um\nprobably groing too, right? Where like\nis it that's an interesting\nYeah. Yeah. like you you know it starts\nto reorganize like it's it hasn't it's\nnot that it's found lower lower loss or\nanything it's just because of the\nsimplicity bias all the hyperplanes has\nstarted to adjust into like a simpler\nphase you know almost of of parameter\nspace\nyeah we didn't we speak to Dan Daniel\nRoberts about that he had that\ncriticality thing in during\nright right that was a similar idea\nthe deep learning the there's a we\ntalked to these physics guys who who you\nknow were trying to start if you will a\nphysics-based perspective perspective on\ntheory of deep learning and they they\nhad some really interesting points about\ncriticality and sort of that yeah I\nagree I think grocking is similar right\nit's like you you start forcing more and\nmore data and essentially because it\ncan't memorize anymore it's forced to\nreorganize into these simpler you know\nmore generalizable representations\nyeah more training is that\nwell more more training and and also\nit's related to the size of the model\ntoo right like is it\nis it becomes if if it's too small to\nreally memorize anymore. It has to\nreorganize.\nWell, that's a good question, too,\nbecause you had this amazing paper out\nand we we were skimming it earlier and\noh god, what was it? Um, it's not deep\nlearning is not so mysterious after all.\nNot so different after all.\nNot not so mysterious or different.\nYes. And and so so um you're talking\nabout this benign um overfitting and you\nknow double descent and over\nparameterization, but would you would\nyou lump potentially grocking in with\nthat as well?\nThat's a great question. And so in the\nintro at the end of the intro to that\npaper I say that I'm not talking about\ngroing or scaling laws much in this\npaper because these phenomena are often\nnot treated as particularly mysterious\nor distinct to to neural nets. Um but I\nwould say that they fit into this sort\nof suite of generalization phenomena\nthat we can try to understand using some\nof the same tools. And so the\ngeneralization bounds that I present in\nin that paper now are being used by us\nand others to try to understand the root\nof scaling laws. I mean they seem like\nremarkable laws of nature almost like if\nyou increase computation by a certain am\namount you can predictably improve\ngeneralization on a set of tasks but\nit's sort of this empirical law and so\nwe want to understand why and this\nrelates to why you know larger models\nmight have stronger simplicity biases\nand things like that and so definitely\nthe generalization frameworks that we\npresent in that paper that I present in\nthat paper can um be used to\nshed light on scaling law behavior. uh\ngrocking is not something I've thought\nabout too specifically, but it does seem\nto be the case that by training for\nlonger, the model is doing some\nreorganization that enables a more\ncompressible solution. And so it would\nbe very interesting. I'm sure people\nhave done this like measure like you\nknow the flatness of the solution and\nthings like that as you proceed through\ngroing. Um and I think this is actually\nan older like many things older than\npeople might realize. So like double\ndescent is thought of as like a modern\ndeep learning phenomenon but it was\nactually first presented in the 1980s.\nUm and so this has been sort of around\nfor a while and I remember like Ilia\nSatsgiver and others talking about\nbehavior that was very similar to groing\nwhere like the training loss isn't\nreally changing but training for longer\nactually leads to better generalization.\nAnd this relates a little bit um to a\nprocedure we had called stochcastic\nweight averaging. And so the idea there\nis you want to ramp up the learning rate\num to a relatively high constant\nlearning rate and then maintain a\nrunning average of the parameters as\nyou're sort of traversing this loss\nlandscape with SGD or atom or whatever\nelse. And what happens when you do that\nis you're sort of spinning around the\nperiphery of flat solutions. Um and by\ntaking an equal average you get to move\ninside that region and get a much sort\nof like flatter solution. um and that\nreliably led to better generalization\nand it's quite convenient because you\ncan just load up a pre-trained model and\nthen increase the learning rate do this\nfor you know a certain number of epochs\nand get sort of better generalization as\na consequence and so I think you know\ngrocking might be related to to that\nsort of behavior as well\nI think I think this ability to to move\num in the parameter space during\noptimization I personally think that's\nalso part of the explanation for double\ndescent right because it's like as you\nprovide\nmore and more flexibility and it can\nshift around a bit. So I'll give you\nI'll give you why I think that because\num you know you're familiar with integer\nprogramming. So like where you're trying\nto solve some system of equations where\nyou're looking for a solution only among\nintegers you know so you have whatever\nthousands of integers extremely\ndifficult combinatorial optimization\nproblem you know because what what am I\nsupposed to do try out like all the\ndifferent integers there's no smoothness\netc. So people figured out like really\nearly on, you know what, let's just\nexpand the parameter space to go from\nintegers to just floating point numbers\nand then just do a floating point\noptimization and when we get to the end\ncrystallize it to sort of the nearest\nyou know integer solution that's like\nyou know pretty turns out being pretty\ngood right and it's this ability to move\nlike within this numerical space\nsmoothly you know that allows it to find\nlike reasonable integer integer\nsolutions and And I think that's sort of\nwhat happens with double descent too,\nright?\nYou're preaching to the choir. So adding\nflexibility.\nYes. So we can have Yeah. flexibility\nand\nbiases. Yeah.\nRight. So you can move around a bit more\nand then almost stumble into or maybe\nit's you know by virtue of whatever the\nstructure or the landscape or SGD you\nknow whatever it is but just the ability\nto kind of and I think there's a paper\nabout these wormholes like sort of being\nable to wormhole to a better a better\nsolution you know because you have a\nmore complex space and the higher and\nhigher the dimensionality is the more\nlikely they're sort of a wormhole to a\nnear a nearby good solution and maybe\nit's related to this connected modes.\nExactly. It sounds a bit like mode\nconnectivity. So actually after I heard\nthat talk from Jorge Nosadol about\nflatness and its role in generalization\nand its connection to small batch\noptimization and thought well I want to\ndo basian deep learning. I had done a\nlot of basian ML and I had done some\ndeep learning but I hadn't thought about\nbasian deep learning until I saw that\nfigure in his talk. I thought well\nbefore we get into basian\nmarginalization let's try to understand\nthe geometric properties of these\nobjectives that we're minimizing first\nso that we can come up with like good\nposterior approximations and so in some\nsense this is a bit basian agnostic it\nshould be useful even if you're just\ndoing classical optimization and then we\nhad this you know the first thing that\nwe encountered was this discovery we\ncalled mode connectivity which shows\nthat if you retrain your neural net for\nexample with different initializations\nand find seemingly different modes in\nfact you can actually walk from one mode\nto the other in sub subspace without\nincreasing the training loss at all\nalong the way. And so before that it had\nbeen believed that the different\nsolutions that we would find for example\nby retraining our model uh were isolated\nfrom one another. So you walk in any\ndirection and you increase the loss a\nlot along the way. And there there were\nsome results from Goodfellow and others\nto suggest that intuition. Um we showed\nthat you can introduce very simple sort\nof parametric curves like a polygonal\nchain with just a single turn or a\nquadratic bezier curve. You can anchor\nthe end points with whatever solutions\nyou find in this procedure that finds\nyour two parameter vectors w1 hat w2 hat\nand then as you vary the the parameter\nof this curve t from 0 to one you walk\nfrom one to the other and then the idea\nis well how do you sort of learn this\ncurve and you can discover these by\nminimizing your loss uniformly in\nexpectation over the curve. So if you're\num doing classification, this looks sort\nof like a line integral of cross entropy\nloss normalized by arc length. And it's\nactually a pretty simple objective to\ntry to minimize because you could just\nsort of sample uniformly along the curve\nt and then take gradient steps with\nrespect to the parameters of the curve\ntheta that you're trying to learn. And\nyou can always do this and the larger\nyou make the model um the less of an arc\nlength they're kind of or sorry less of\na sort of bent turning that you need to\ndo the more it looks almost like a\nstraight line path between the two\nsolutions. And so what this showed is\nthat there were these regions within the\nloss landscape that were extraordinarily\nflat. So they all had sort of zero loss.\nAnd what was especially interesting\nabout them was that the different\nparameters in these mode connecting\ncurves actually led to models which made\nvery different predictions on the test\nset. So of course they're making the\nsame predictions on the training set to\nhave the same loss but on the test set\nthey were different representations. And\nso this meant that you could ensemble\nthem and get much better performance for\nexample just uniformly sampling on the\ncurves. And then this led to this\nstochastic weight averaging procedure\nwhere we were sort of thinking well how\ncan we sort of spin around these like\ncontiguous regions of flat solutions and\nfind something that's centered within\nthem that ended up being fairly\npractical. Yeah.\nBut even um I suppose also related to to\nthe groing question is where does the\ngradient come from? Because you just\nsaid well you know it's the same on the\ntraining data but on the test data it's\nactually behaving differently.\nRight? Where does the signal come from?\nlike what what happens when you continue\nto train a model which has ostensibly\nconverged\nbut it's it's behaving different on vow\nwell I mean I have a I have a thought on\nthat like I don't maybe maybe this\nhelpful but so I'm I'm big fans of\nBalstriero you know Randall Balstrio's\nkind of he helped me personally at least\nthrough through the work on the spline\ntheory of deep to understand them better\nand I think what's happening there is\nbecause if you if you agree with that\nand you kind of think of these splines\nas as essentially just\nunbelievably hyperdimensional honeycomb,\nright? And all these shared kind of\nhyperplanes that are activating. I think\nwhat's happening there is they're really\nslowly just moving. You know, this\nslopes are just slightly changing and\nthen they happen to hit a phase change\nwhere it's like, wow, now we can combine\nall these hyperplanes into a much like\nless lower complexity or, you know,\nsimpler, more parsimmonious kind of\ncombination. And that's that's what it\nis. Like I think that's what's\nhappening. It's just slowly slowly\nshifting and then snaps into place.\nRight.\nWell, yeah. And I don't know whether you\nsaw his ICML paper from last year, but\nhe he he had a spline interpretation of\ngroing and he said exactly that.\nBasically, you know, um during groing,\nthe splines just suddenly\nand slowly moving\nbecause because there are so many of\nthem overlapping. But what happens is it\nforms this honeycomb where they just\nkind of compress together.\nAnd I I should have said this to David\nKow because he said there's a\nreorganization in the micro substrate.\nAnd what's this if it's not a\nreorganization where the the you know\nthose little hyperplanes if you like the\nhoneycomb it just does that during\ndroing during groing. Yeah, it's\nfascinating. Grocking is not something I\nhave thought about too specifically but\nit seems to me that you probably are\nentering some region of the loss\nlandscape that doesn't have different\nloss but is providing a representation\nwith very different properties. And so\nthis does relate to procedures like\nstocastic weight averaging where again\nin the end the solution doesn't really\nhave a different loss than what you\nwould have gotten training in a standard\nway but it has properties that will lead\nto better generalization um better\ncompressibility etc. I still don't quite\nunderstand where the drive to more\ncompressible slashs simpler\nrepresentations is coming from just\nmechanistically like in the optimization\nprocess you know I mean I don't know\nmaybe it's just like you said maybe it\nrequires less floating point precision\nand so that that the optimize the jitter\nin the optimizer or something like that\nyou know I I just I'm really curious\nabout the actual mechanism you know that\nforces you there like if there is no\nchange in loss Mhm.\nOr maybe it's so tiny we just don't\nreally pay attention to it. I don't\nknow. But I'm super curious like what's\nreally driving it downhill, if you will,\nto a simpler solution.\nYou might sort of get to start to I I I\nmean this is really speculation and I\nhaven't thought a lot specifically about\ngroing, but it could be that you're just\nsort of edging your way inside because\nof gradient noise and so on towards a\nflatter solution as you continue to\ntrain.\nYeah. Yeah. Just some type of Yeah. It's\nreally fascinating. This sort of made me\nwonder actually after the mode\nconnectivity discovery like what what do\nthese loss landscapes really look like\nthe the path that we considered\ninitially were just one-dimensional\neventually we started thinking about\nmulti-dimensional loss volumes kind of\nand so we had a paper which was creating\na simplex we were basically adding\nvertices to the simplex and sampling\nuniformly within the resulting simplex\nand trying to sort of add vertices such\nthat when we did that we would sort of\nhave low loss and also maximize the\nvolume of the simplex. And so that sort\nof enabled us to find these sort of like\nmulti-dimensional loss surfaces. And\nthen we had this that all had sort of\nlow loss. Um and we had this picture I\nguess in the front of that that paper\nwhere we started with this sort of\noriginal understanding that all the\nlocal optima are kind of isolated to\nwormhole-like tunnels between the\ndifferent optima to this idea that maybe\neverything is just sort of connected\ntogether in some manifold that's\nembedded in this really highdimensional\nspace. And what seems like a sharp or a\nflat optima might actually just be like\nhow converged are you with inside that\nmanifold because if you're on the edge\nit'll look it'll look sharp because you\nmove in most directions and you increase\nthe loss quite a bit but if you move in\nspecific directions it's very flat and\nso I think we still don't have a full\nunderstanding of what that looks like\nand I think it has really fascinating\nimplications for the generalization\nbehavior. It's\nreally hard to think in a billion\ndimensions. I mean,\nyeah, it really is. Andrew, um, so let\nlet's sum up a little bit. Your\nphilosophy is we should have maximally\nflexible models with soft\nregularization.\nHow is that actionable to um\npractitioners? I mean, what what what\nbecause you also um gave many practical\nempirical examples, you know, um\nensembling models together and and you\nspoke about this residual pathway prior\nand so on.\nHelp us understand. Mhm. So my\nphilosophy is that we should honestly\nrepresent our beliefs in the way that we\ndo model construction. And our honest\nbeliefs are usually that the real world\nis a nuanced place and we're going to\nwant to have model expressiveness in\norder to represent that nuance. At the\nsame time, it can't just be flexibility.\nIt has to be flexibility combined with\nsome sort of simplicity bias. It should\nhave some kind of aams razor bias. It\nturns out perhaps surprisingly that\nmaking transformers and other types of\nneural net models larger often actually\nenhances rather than reduces a\nsimplicity bias. And this is in many\ninstances mostly what's been responsible\nfor the better generalization behavior\nof larger models. So we had this example\nwith double descent where in the second\ndescent all the models have basically\nzero training loss but the larger models\nare generalizing better. can't be\nbecause they're more flexible. It has to\nbe because they have some other bias,\nsome sort of simplicity bias. So, how do\nwe operationalize this? I think one way\nis well embrace expressiveness. So,\nchoose a model class that is going to be\nable to represent lots of different\nsolutions. In terms of the simplicity\nbias, well, we're finding empirically\nthat increasing model size can help with\nthat. And so if you're able to build a\nreally big model, you probably should\nboth for the expressiveness and the\nsimplicity bias.\nBut what if what if I'm a researcher who\nlike I just want I want the simplicity\nbias to go to 11, but I can't afford\nmore parameters. Is there something I\ncan tweak in my objective function to\njust push me a little bit more towards\nsimplicity without breaking things?\nYes. So there are a lot of things you\ncan do\nI think\nand we'll talk about them but the um the\nquestion of how can you more elegantly\nencode this compression bias beyond just\nmaking the model bigger is really a\nfascinating open research question. So\nmy contention and I might be wrong is\nthat\nin many cases a 7 billion parameter\nmodel is not doing better than a 1\nbillion parameter model primarily\nbecause it's more expressive. It's\nactually because of this simplicity\nbias.\nAnd perhaps things like knowledge\ndistillation can help make this\nargument. If it's possible to really\ndistill a large model into a much\nsmaller model, then it means that the\nsmall model has some setting of its\nparameters that can provide a good\napproximation to the large model. It's\njust not able to find those parameters\nwhen trained directly on the data. It\nneeds the help from the teacher model.\nAnd so the question is can we build the\n1 billion parameter model with some sort\nof like explicit bias that would have\notherwise come about through scale in\nthe 7 billion parameter model and find\nthose solutions itself. I don't think\nwe're anywhere close to being able to do\nthat. I think it's it's sort of a an\nopen research program. However, there\nare little things we can do that will\nhelp. So like\nthere are ways of intervening in the\ntraining procedures. So like the\nstochastic weight averaging we discussed\num that will help find a flatter more\ncompressible solution. There are um of\ncourse all sorts of regularizers that\nthat can be useful in in certain\ninstances. Um basian marginalization can\nbe very helpful in sort of encoding an\nautomatic simplicity bias in what we do.\nUm and so there are all sorts of\ninterventions that will help us with\nthis. But I think the dream is that\nmaybe we can embrace flexibility in you\nknow 15 20 years from now by building\nthese nonparametric models that like\nreally do have an infinite number of\nparameters and are more expressive than\nany model we're using right now but then\nthey have this sort of like more\nexplicit compression bias that is\ninterpretable and is getting us what\nbuilding huge models is inelegantly\ngiving us right now.\nThat's interesting. I mean the the other\nkind of interpretation that some folks\nat home might have based on what you've\nsaid is going back to Rich Sutton's\nbitter lesson right so he said that\ndesigning things is bad don't put your\nsymmetries in there don't don't kind of\ncreate these multifaceted systems just\nscale and lots of computation is the way\nforward because it seems like one\npotential interpretation of what you're\nsaying is that we should create hybrid\nsystems that make models more flexible\nby combining different modalities and\nwhat you're kind of saying is we should\njust have bigger models. So are you a\nSutton guy or or not?\nSo I think the bitter lesson is widely\nmisunderstood and incomplete.\nThe bitter lesson\ndescribes how over short time scales\nit's been appealing to try to encode\nour knowledge into our procedures in\norder to achieve good results. uh and\nthat this is also a cognitive bias that\npeople have. They like encoding\nexpertise and things like this and being\nthoughtful and elegant about problem\nsolving. But over longer time scales\nthan a typical research project,\ncomputation becomes cheaper. And so\nquite often more brute force seeming\napproaches based on search and learning\nend up working a lot better than\nelegantly encoding our priors and our\nconstraints etc.\nAnd there are a number of examples in\nthe essay like deep blue uh for playing\nchess in the late 90s uh speech\nrecognition where\nsome researchers were trying to model\nthings like physiology of the voice box\netc to try to get any possible advantage\nbut then they were beat out by\nstatistical procedures like hidden marov\nmodels and of course alpho and and other\nprocedures like this too. And so the\ntakeaway seems to have been because of\nhow computation is becoming cheaper over\nlonger time scales,\nquite often it's going to be more\npractical to try to build procedures\nthat are based on search plus learning\nas opposed to feature engineering.\nI think to a large extent this is true.\nBut what it doesn't say is that in order\nto learn, you need to make assumptions.\nAnd so machine learning, as we\ndiscussed, is learning by example. And\nwe can't do that without making\nassumptions. So if we go to like maybe\nthis coin toss example where we're\ntrying to estimate the bias of the coin,\nyou have to make some assumption like do\nbefore I start doing that experiment,\nare we assuming a uniform bi uniform\nbias? uh uh are we thinking it's looking\nsort of more like uh you know centered\naround 0.5 or something like that like\nbeing unbiased but it we have support\nfor other things. these assumptions are\ngoing to influence how we do induction.\nAnd so like we just can't get away from\nmaking assumptions. And so the question\nis just what assumptions should we make\nand to what extent can they be\nuniversal? And when it comes to things\nlike scaling laws that describe how we\ncan reliably improve performance with\nincreases in computation, if we're able\nto make better assumptions, we can\nactually change the scaling exponents\nwhich act which which mean that we'll\nget exponential improvements in\nperformance with increases in\ncomputation, which is just remarkable\nmotivation for trying to do this. And\nit's not just a contention that this\nmight be possible. We're actually\nstarting to see some evidence of this.\nAnd so in our own work, we've been very\ninterested in how we can produce\nstructured representations for linear\nlayers and neural nets. And so um this\nmight sound very abstract, but there's a\nvery well-known example of doing this.\nSo you can actually start with a fully\nconnected layer having every possible\nconnection between two between the nodes\nand two two layers. Um remove a bunch of\nthe connections and enforce parameter\nsharing and you have a convolutional\nlayer. And so mathematically what you've\ndone is you've replaced a dense matrix\nmultiply with a matrix that has sparity\nand locality and uh that's a toplit\nmatrix. And so there's this question of\nlike could you actually systematize this\nprocess of creating structured layers\ntowards better compute optimal\nefficiency? And if you were to do that,\nwhat sorts of principles should we be\nembracing? Should we embrace things like\nparameter sharing because convolution\nfor example has parameter sharing. um\nshould we embrace uh sparsity? Should we\nembrace other sort of features? And so\nwe introduced this sort of Einom\nformulation over structured matrices\nthat contains um all sorts of different\nstructures as special cases as well as\nall sorts of different novel structures\nand this taxonomy of interpretable\nhyperparameters that kind of control the\nproperties of these structures. So how\nfast is it to do a matrix multiply? Um\nwhat's the rank of the matrix? What how\nmuch parameter sharing there is? And\nwe'd have sort of the continuous values\nfor these parameters that would control\nthese things. And what we found is that\nactually in general towards compute\noptimal efficiency assuming we have you\nknow as much data as we ever need. Um\nparameter sharing was actually not a\ngreat principle. Um uh and that was a\nbit surprising. You also want to have\nfull rank structures and that move from\nsurprising to sort of disappointing.\nIt's like well can we go beyond what\nwe're doing with dense matrices then\nbecause that's what we're using now. Um\nand the answer is yes if you can get\nfaster multiplies. So we proposed a\nstructure called a block tensor train um\nwhich is related to another structure\ncalled a tensor train and another\nstructure called a monarch matrix. It's\nbasically a sum of monarch matrices. Um\nand um this sort of is full rank doesn't\nhave parameter sharing but you can do\nmultiply faster than you can with a\ndense matrix and this allows you to\nbuild wider layers for a given\ncomputational budget. And this this did\nhave sort of a meaningful effect on the\nscaling exponents. And so this is sort\nof at a proof of concept level. um and a\nlot of work would need to go into\nparallelized implementations and so on\nof these structures. Um but it showed\nthat it's possible and we're not the\nonly ones. So there are um other groups\nuh uh that have been looking at neural\ntangent kernel inspired ideas to\nunderstand different regimes of learning\nlike easy versus hard feature learning\nand so on and trying to think about well\nwhat principles are actually going to\nmodify these scaling exponents and uh\nother things like I think Albert Goo and\nothers have been looking at like\nchunking and transformers and the\ninductive biases that are you know come\nwith that and like whether we can change\nthese inductive biases towards better\nscaling exponents and so I guess in\nshort learning requires assumptions and\nso uh I don't think we can neatly\ncompartmentalize like elegant ideas from\nhaving successful learning combined with\ncomputation uh we really do need to\nthese things not only are not at odds\nwith it with one another they really\nstrongly go together\nyeah when we spoke I mean by the way out\nof all of those assumptions sparity\nseems like a really good one there's\nalways one thing where you think that\nseems like a really good one Daniel\nRoberts was saying you know in in many\neffective theories in physics. That's\nthat's one of the earliest assumptions\nis is sparity. But then you kind of get\nto other factors to consider like\ncomputational complexity and training\ntractability and and so on because in an\nideal world we we would make these\nthings sparse. What's stopping us?\nIt's a great question. So I'll say that\nI I think convolutions are a great idea.\nAnd if you're in this setting that I\ndescribed earlier where you have a fixed\ndata set and you want to do something\nreasonable in order to achieve the best\nperformance, this will often work really\nwell. And even if we're moving towards\nsoft inductive biases from hard\nconstraints, I would normally advocate\nfor having some sort of convolutional\ninductive bias, even if it's not a hard\nconstraint anymore. And that's what we\nwere doing with residual pathway prior.\nAnd there's been some interesting work\non convolutional VITs and things like\nthis to try to have these sorts of soft\nbiases for more efficient learning. Um,\nthe reason that we found in this\ninstance that parameter sharing didn't\nseem to be a good principle towards\nbetter compute optimal scaling is\nbecause we were in this setting where we\ncould have as much data as we ever\nneeded. And there wasn't much of a\ngeneralization gap. And so we basically\njust need to fit the data as efficiently\nas possible. And so this basically led\nto the principle of having as many\npossible parameters per flop. Um, and\nyou might sort of wonder then well can\nyou go beyond one parameter per flop?\nAnd I think things like sparse mixtures\nof experts actually allow you to do this\nbecause you have some sort of gating\nfunction which might be acting over um e\ndifferent um MLPS and you only have some\nsubset of them that are active. And so\nyou kind of divide your computation by K\nover E where K is the number of active\nefforts experts relative to a model with\nthe same number of parameters. And so we\nalso tried to push that principle a bit\nfurther in this work where rather than\nhaving the mixture of expert uh gating\nfunction operate over whole MLPS, it was\nactually operating over individual\nlinear layers within the MLPS and the\nattention projection matrices. And so\nyou would represent these layers as like\nsums of structured matrices and then you\nwould have this gating function acting\nover the rank index of this sum. And\nthis would allow like much finer grain\nrooting decisions across the experts and\ndid lead to sort of better efficiency\nfor a given amount of computation. But\nthe short reason I think is we found\nparameter sharing to not be helpful\nbecause you basically just want to\nreduce your loss as efficiently as\npossible and you have sort of as much\ndata as you could ever need and there\nisn't much of a generalization gap.\nAlso reminiscent of of VIT surprised\neveryone that that it did so well\ncompared to CNN's. But um final question\nand we were talking about this in the\ncar on the way over Andrew. Um the\nelephant in the room is that GPT5, you\nknow, we have all these huge\noverparameterized models and on the\nsurface they seem to be doing very well.\nThey're they're benchmaxing and they're\nthey're brilliant to use, but it feels\nthat there's something missing. I mean,\nI I think they're not intelligent. I\nbelieve you would agree with that\nstatement. What's missing and what's\nnext?\nSo, one of the things I'm most excited\nabout is developing AI systems that can\ndiscover new scientific theories at the\nlevel of general relativity or quantum\nmechanics. And we haven't even really\nscratched the surface in being able to\ndo this.\nIt's not even clear how datadriven that\nprocess would be, how much it would look\nlike symbolic logic if we were to think\nabout what Einstein did when he proposed\nrelativity and how we might want to\nwrite that down as an algorithm and\nautomate it on a computer. And so I\nwould love to see progress in this\ndirection. I think that some of the\nideas that we've discussed around\ncompressibility could play an important\nrole in how we think about\nselecting for scientific hypothesis.\nAlso, ideas around universality, like\nwhat sorts of assumptions might be more\nuniversal than others and at what level\nof abstraction.\nBut it's something that we haven't\nreally made progress on at all despite a\nnumber of very exciting research\nprojects in AI for science where neural\nnets for example are being used as\nblackbox function approximators in some\nsort of pipeline targeted at a very\nspecific application. And I think this\nis extraordinary work and it's really uh\na way in which machine learning is\nclearly doing a lot of good in the\nworld. But I think it's time to to try\nto go beyond that paradigm towards\nreally giving us new scientific insights\ninto the data that we didn't have\nbefore. And in fact, that's, you know,\nsomething that I'm I'm really more\nexcited about than anything else in in\nterms of how technology might develop in\nthe future. Like if I were to go a\nthousand years in the future, one of the\nfirst questions I would have is, well,\ndo we understand something about physics\nthat we didn't before? Do we understand\nhow the brain works? Um and this is sort\nof a conventional sort of like approach\nto science where the theory is really\nthe quantity of primary interest and the\napplications of course are important but\nthey're not primarily why like no\nindividual application is primarily why\nwe care about the theory. Like GPS would\nbreak within minutes if we weren't\naccounting for gravitational time\ndilation and general relativity. But\nEinstein wasn't thinking about GPS when\nhe proposed relativity. And if you have\nthe theory, then um you can sort of\nsuggest all sorts of applications that\notherwise wouldn't have been on the\nhorizon. And like we probably could\ntrain a neural network to correct for\ngravitational time dilation without\nunderstanding what's going on, but that\nwouldn't be nearly as exciting or useful\nas as having the theory of relativity.\nMhm. Yeah. Professor Wilson, thank you\nso much for joining us today. It's been\namazing.\nThanks so much. It's been a real\npleasure. If you want to learn more\nabout this work from Andrew, please\ncheck out the paper deep learning is not\nso mysterious. Uh also their paper\nbasian deep learning and a probabilistic\nperspective of generalization which also\ndiscusses basian principles in deep\nlearning. And also their recent paper\ncompute optimal LLMs provably generalize\nbetter with scale that has further\nanalysis on the source of the simplicity\nbias uh arising from scale. Cheers.",
  "transcript_chars": 132915,
  "ingested_at": "2026-05-12T00:42:33.011895+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 71549,
    "like_count": 1486,
    "channel_id": "UCMLtBahI5DMrt0NPvDSoIRQ",
    "categories": [
      "Science & Technology"
    ],
    "tags": []
  }
}