{
  "video_id": "3rWSvrFahIY",
  "channel_slug": "ycombinator",
  "channel_handle": "ycombinator",
  "title": "5 Papers That Show Where AI Research Is Heading Right Now",
  "duration_seconds": 4615.0,
  "url": "https://www.youtube.com/watch?v=3rWSvrFahIY",
  "upload_date": "",
  "transcript": "Thank you guys so much for coming. This\none will have much a much more applied\nbent based on the feedback. We have a\nbunch of really cool people that I'll\nintroduce in a second, but we're\ncovering AI for uh biology by my\nfavorite one of my favorite co-\nresearchers, Yas Beg. We have Luke um\nout of Tatsu's lab talking about\nselfplay, Alpha Zero style selfplay for\nLLMs. Super excited about that. Arnob\nwill be uh presenting he's a researcher\nat Giga on stream rag uh super different\napplication you know thinking about uh\nreal realtime voice uh agents uh Robert\nGeorge working on lean for science super\nexciting and then the AI token maxer\nhimself Luke Worthwine\ncool so I want to introduce some like\ncall for presentations you know maybe\ninspire some of my interest and maybe\ninspire some some of you guys to jump up\nand and ask for a presentation on this\nstuff. I think memory has been like the\nhot topic for at least the last year and\na half. There's been so many papers from\nmem zero to recursive language models\ncartridges out of uh our lab hnet, you\nknow, dynamic chunking stuff. There's so\nmany different ideas and so I'm\ndefinitely interested in that area. If\nyou guys want to present on that, I did\nthis Nome Brown podcast I think a couple\nweeks ago launched and like he's still\nof the view that this human generated\nsubspace H is still if we train on that\nwe can test time compute our way out of\nit and recursive self-improve out of it\nall the way to get to this F minus H.\nAnd I just like really struggle with\nthis and I really really don't see how\nit's probable. Not that it won't it's\nnot possible, but it's just not probable\nthat we'll sample all of that. So, I'm\nreally interested in that and that's\ndefinitely in Luke's. We were talking\nabout that a bunch and I think that\nbasically the left side is alpha go, the\nright side is alpha zero. And I think\nthat alpha zero unbiased by um humans\nmeandering is uh the way we'll get to\nmuch more intelligent systems, maybe\neven dare say agi. And the right way to\nsay this um would be um I want to be\nvery careful. If the full solution space\nf is f, training on known human\nsolutions will limit you to some typical\nset h despite any feasible amount of\ntest time compute or recursive self um\nimprovement. You won't feasibly sample f\nminus h. Um and especially all of it. If\nif it's infinite recursive self\nimprovement, infinite test compute\nmaybe, but we don't have infinite. So\nlife's a pom dp and this is a we're\nfinite horizon mdp.\nIntelligence per sample. I think this is\nlike the two major problems left in my\nopinion are intelligence per sample,\nintelligence per watt. And intelligence\nper sample. I always think about this\nlike as I get one new sample, I do this\ncontinuous learning. What is the right\nthing to do if I'm trying my goal is to\nmaximize performance condition upon that\nn. Most people's answer to this right\nnow in practice is ICL. And I actually\nplayed with this. As you increase the\nnumber of samples in ICL, it is not\nmonotonically improving in performance.\nAnd so it actually starts to bob and\nweave. It gets worse. It gets better\nsometimes. Um and then it hits a cliff,\nwhich is the context um uh length,\ncontext length that the model was\ntrained on, and it literally just stops.\nSo it clearly can't go on forever. It\ndoesn't monotonically improve. And I\nstarted playing around with Laura. I\nthink Laura at higher at lower ranks for\nuh lower amounts of sample size actually\ndoes impressively well. And then it has\nthis kind of arc. They both peter out\npretty quickly as you increase number of\nsamples all the way until you do SFT\ngroup O all the way at the end. And if\nyou look at this, it's kind of weird\nthat like you have this like ICL is the\noptimal thing to do in the beginning and\nthen training the whole thing. If I get\none new sample, I want to retrain with\nLaura at some rank on N plus1 samples\njust to get that little bump in\nperformance. And there's a different\noptimal thing to do all along the way as\nI stream and I get more and more\nsamples. And that's just not how we are.\nSo we're kind of like monotonically\nimproving. And the more chess games that\nMagnus Carlson plays, he just keeps\ngetting better. The, you know, 10,000\nhour rule, etc., etc., we just keep\ngetting better. And it's the same algo.\nAnd so I think there's just something\nreally different happening in us. And so\nthere's must exist some learning\nprocedure that is has a much higher\nintelligence per sample. And then\nintelligence for Watt out of my lab,\nIvonica and John who will hopefully come\nto the next one uh and give a talk on on\nthis. And I just think it's the right\nway to think about it. Arguing that\nhaving smaller models um are sometimes\nactually better from an intelligence per\nwatt perspective. Alternatives to back\nprop for those that know me, I'm very\nhot on this. Back to the brain um and\nhow we learn. There's very little\nevidence that the brain is taking the\ntranspose of the weight matrix and there\nmust exist some other learning\nprocedure. I'm highly interested in\nSPSA, but if there's alternatives that\nI'm not aware of, please like recommend\nthem. And I'm really interested in novel\nbreakthroughs. Yaso is one of my\nfavorite AI researchers, but he's mostly\nfocused on bio and he always um sends me\nbiopers and it's super interesting.\nWhether it's about how birds navigate\nthe world via iron in their liver\napparently that's how they they actually\nnavigate\ncrazy robotics, uh speech, other things\nlike that as well. And then of course\nunhinge founder hacks very interested in\nthat as well. Call for ideas on ways to\nmake the club better. So, better ways to\nmeet you if it's a lightning round if\nit's not. Um, some people talked about\nsome AI benchmarks that we could\nactually launch together. That'd be kind\nof fun. Club challenges to challenge\neach other. And then, uh, any open\nsource ideas that you want to hack on\ntogether for this club or something or\notherwise. It' be very interesting. All\nright, that's all I got. Thank you so\nmuch.\n>> Hi. Yeah. Uh, thanks France for that\nintroduction. We've been labmates now\nfor like two years, something like that.\nYeah. Um, I think France is a great\nexample of someone who brings very\ncreative and very out of distribution\nideas to our group all the time, even if\nI sometimes like have no idea where he\ngets them from. Um, but um, that being\nsaid, uh, he asked me to give a talk on\nsome bioai things and I thought why not?\nUm, so I'll be presenting on this paper\nthat came out just last week from Biohub\nfolks um, here in California, not too\nfar actually in the Bay Area. I think\nthey just moved to the city actually.\nUm, so I am a second year PhD student\nwith France, but I'm also co-advised by\nSteve Quake over at Stanford. Sort of\nanyone in biology probably will know\nSteve at least tangentially done a lot\nof work in bioengineering and all kinds\nof applications was director of the\nbiohub where this work came out um just\nuntil recently. So uh a lot of uh a lot\nof overlap but uh the high level pitch\nfor this work is that I know most of\nthis audience is probably more like AI\nML types. I want to talk a little bit\nabout how sort of a lot of these ideas\nfrom sort of that's motivating a lot of\nprogress in language modeling and AI\nvery broadly have been sort of recently\nbeen translating into biology with a\nfocus on this recent paper because I\nthink it really does a really excellent\njob of interrogating how scale which you\nknow at some level has been like the\nfundamental primitive in terms of\nassumptions that we as a community have\nin terms of how to make things better um\nhas actually been playing out for a lot\nof these biological problems\nparticularly protein biology. So um\nthere won't be like much bio in this\ntalk. I'll try to focus more on the ML,\nbut feel free to ask questions. Um so\nyeah, like I called this talk the uh\nbitter lesson comes from biology. The\nactual paper title is right below that.\nBut I mean just a quick refresher. I'm\nsure everyone in this specific audience\nprobably read Richard Sutton's famous\narticle. You know, basic premise here is\nthat like, you know, across the past 70\nyears of AI, methods that win are\nmethods that are general that sort of\nexploit really fundamentals of like\nscaling compute and data as opposed to\nmethods that sort of handgineer human\ndomain, human domain knowledge. And\nSutton always cites his work in or like\nyou know a lot of the work that was\nearly in the field. So alpha go alpha go\nzero that's sort of just inordinately\nscaled compute and then for a long time\nthey were far worse than sort of expert\nsystems until they eventually overtook\nand then exponentially improve past them\nright knowledge systems win at first but\nthen eventually sort of these like big\nlarge dumber models will like um you\nknow win in the long run this is sort of\na new goal in biology is to what extent\ncan we study like this is actually also\ntrue for a lot of these sort of\nbiological AI problems right um the bet\nsort of behind this paper and they do a\nreally exploring is well the same\npattern basically also saw protein\nbiology right um can we take like you\nknow a lot of these ideas in scaling law\nanalysis here is from the um uh famous\nyou know neural scaling laws paper and\nthen translate them for all these\nproblems that we care about for say\ndesigning a drug or like you know trying\nto understand how like a cell works\nright so on the left is something that\nwe trust it's like a language model we\nhave like this nice smooth log linear\nscaling laws that we see that moss like\npredictively falls as a function of\ncompute and data uh the question on the\nright is whether this curve will exist\nfor bio spec more generally but proteins\nspecifically in this paper you know sort\nof does our LLM recipe transfer or does\nbiology really out of distributional\ndomain relative to language sort of\nbreak it right that's the bet and like\nin this talk I'll basically chat about\nlike three like vignettes from this\npaper that sort of interrogate to what\nextent this is true\nso this slide is all the biology you'll\never need about a quarter of the slide\nat least for this presentation um let's\ntalk about proteins so in your body\nthere's like broadly three major classes\nof macro molecules. There's lipids,\ncarbohydrates, and proteins. Um, a\nprotein is just a string of amino acids,\na special type of biomolelecule. There's\n20 varieties of amino acids. If you put\nthem together into a sequence, you can\nhave like a virtually infinite number of\npossible molecules that then fold into.\nSo, you can think about it largely as\njust every single protein is this 20\nletter alphabet. And that string\nspecifically determines a unique 3D\nshape. And by virtue of the shape of\nthat protein, what job it does in your\ncell like presents catalyzes a reaction,\nkeeps pathogens out, etc. Um, The work\ndone in this paper is that their goal is\nto train um ESMC sort of their like\nthird major or fourth major iteration in\na long series of models of the group of\nevolutionary scale group originally at\nMeta then their own company now Biohub\nhave been training for a few years now\nwhere the cell is very similar to\nlanguage where we let's just take\nhundreds of millions of years of or sort\nof evolved sequences that we've sort of\ngone out and found across biology both\nin humans but also across bacteria and\nin our environment and just go train a\nbig mass language model on them, right?\nSo, I mean nowadays we mostly train NTP\nmodels, but the pitch here is that if I\ntake some protein represented as like\nyou know these strings of 20 tokens like\nhide a few and then can I train a really\nbig BERT style transformer to predict\nmass positions as a function of the\nother ones nearby. And the crucial part\nis that we never tell it anything about\nthe protein beyond just the sequence.\nRight? So all this guy's got access to\nis just this string and it's being asked\nto basically learn things about the\ngrammar of that protein as a function of\nwhich protein other amino acids tend to\nco-occur with right so like I think\nthere's like this old saying in natural\nlanguage processing it's like you'll\nknow a word by the company that it keeps\nand here the idea is that you'll know a\nprotein by amino acids it keeps and the\nbet is that if we do this at scale just\non the simple sequence task we will\neventually get all these sort of other\nproperties of protein say like structure\nthat we do care about sort of for\nAnd uh yeah like I said before there's\nbeen a lot of prior work on this like\nlargely from evolutionary scale but a\nfew other or a few other sort of groups\nworking largely on this bit. Uh this\ntable will be a map for the rest of the\ntalk. So like every row is a concept you\nprobably already know from natural\nlanguage and then analog onto the\nprotein context. So tokens become amino\nacids. The internet becomes sort of all\nevolution sequence databases all the\nproteins we can actually go out and\nmeasure. Mass token prediction stays as\nmass token prediction and sort of\nemerging capabilities. we talk about\nlanguage model having like become\nemerging structure and function within\nlike basically understanding of a\nprotein and then there's also sort of\nlike this really fun stuff at the\nbottom. So like recently there's been a\nlot of advancements in sort of these\ninterpretability toolkits from the mechi\nfolks you know things like sparse\nautoenccoders some of the earliest work\nbasically in really trying to\ninterrogate um using the toolkit the\nlanguage modeling community has built to\nunderstand language models now in a\nprotein language model setting. So I'm\ngoing to fill in the right hand column\nfor the rest of this talk with sort of\nevidence and three questions which is\nthat do these models learn with scale?\nUm can they basically substitute for a\nlot of these handbuilt features sort of\ndoes the bitter lesson hold and like\nwhat do these representations actually\nencode interpretably? So question one is\ndo scaling laws even hold in the protein\ncontext in the way that we see them in\nthe language context. First let me just\ntry talk a little bit what we measure\nwhen we're talking about sort of\nemergent properties sort of how do we\nactually like study the model see if\nit's learning anything right we need a\nproxy for does the model understand\nprotein structure for instance right um\nthe one the authors use in this paper is\nthat they look at the internal or they\nbasically take the model representations\nduring training and they use this to\npredict um long distance protein\ncontacts so the idea here is that\nproteins have a one-dimensional sequence\nbut they fold into complex\nthreedimensional shapes and if the model\nis sort of understood something complex\nabout the protein structure or something\nemerging about the protein structure. It\nshould be able to predict um contacts\nthat occur over long distances sort of\nnearby contacts are rather kind of\nobvious and this is like a really\nchallenging object for it to get just\nsort of denovo purely from sequences\nalone. They called this P at L right um\nsort of a long contact precision at some\ngiven length and it's just a clean\nunsupervised readout sort of structural\nknowledge inbuilt in the model that's\nlearned during this language modeling\nobjective. Um on the right I plot the\nperformance of this or the authors plot\nthe performance of this I should say\nagainst training compute for the for\nbasically this new model family the\nauthors have built recently called the\nESM cranberry at 300 million 600 million\n6 billion parameter scales. uh\ninterestingly and they had this fit line\nwhich is basically this predict compute\noptimality curve which they um estimated\njust from sort of lowend training runs.\nSo relatively low computational budget\nand they find it actually extrapolates\nvery cleanly to real model training runs\nmeaning so the answer is like do these\nmodels with scale and this data at least\nsuggests that the answer is like yes\nright like you do see this nice log\nlinear curve right if you keep investing\nmore and more compute you training more\nand more protein data with larger and\nlarger models um you see the same exact\nsame broad qualitative shape as the LM\nscaling or sort of LM setting and arrest\nretransfers cleanly meaning that without\nany kind of like predisposed part of the\nmodel that we've taught to look at\npurching structure even didn't get any\nprotein structures. It does a good job\nof sort of picking these out just from\nsequence co-occurrence patterns.\nUm there's like one interesting twist\nthough is that I said before there's\nbeen a lot of prior work from this group\nas well as others and trying to answer\nthese scale questions. So not the first\nones to look at this but previous models\num so the sort of the prior generation\nESM2 models shown in um purple here had\nactually not shown the same behavior.\nThey sort of hit this wall where they\nkept adding more parameters and they got\ndiminishing returns and you had this\nsort of flattening out the scaling\ncurve. this ESMC or ESM Cambrian model\nsort of the green line keeps climbing\nwith no plateau. And their fix for this\nwasn't really like they came up with\nlike a really clever inductive bias in\nthe architecture. Not to say there isn't\na lot of excellent engineering work in\nthis paper, but really it was just data\nscaling, right? They um had about 50\nmillion training samples in their\noriginal ESM2 paper and here they just\npushed that to 2.8 billion by pulling\nlargely in metagenomic data. So\nessentially amino acids or protein\nsequences that have been found from\nsequencing DNA actually out in like dirt\nand oceans and like human guts from like\norganisms that nobody has like really\never cultured or even has really really\nelucidated. And their conclusion is that\nmore data ends up being really important\nand keep getting sort of uh are\nbasically justifying the cost for\nincreasing compute. So it's like the\nprotein version of LM data wall\nconversation, right? Like except here in\nbiology, evolution has been generating\nthis train data for for four billion\nyears and not humans in like the past 30\nor so. And you know compared to tokens\nin natural language like I mean we've\nonly sampled like less than 1% of all\nknown protein sequence diversity and\nthat's like only currently at this\nmoment in time let alone like all of the\nsequence diversity evolution has sampled\nsince the beginning of life on Earth.\nUm sort of question two in this paper I\nthink is interesting is that um it's\nsort of the most bitter lesson part and\nthey really try to evaluate to what\nextent their paper can do or how well\ntheir model trained purely on mass\nlanguage modeling objectives can compete\nagainst a structure based model with\nsort of handtuned inductive components.\nSo I'm sure you're all familiar with\nAlphaold won the Nobel Prize a few years\nago was sort of a landmark moment in bio\nreally show that these computationals\nhave a lot of value in the biology\nsphere. Alfold is brilliant but its\npower comes from basically building\nhandput in or handbuilt inputs sort of a\nmanual feature curation called a\nmultiple sequence alignment or an MSA.\nSo to fold a protein it goes and finds\nhundreds of evolutionary cousins of that\nprotein and stacks them up. Um these\npatterns of sort of coariation across a\nfamily are essentially this encodesical\ninformation you need to do to be able to\nget structure. This is like a beautiful\ndomain engineering application and it's\nthe sort of like really good human\ncrafting objective bias that the bitter\nlessons at least claims should\neventually lose right think like hog\nfeatures in CD and compared to sort of\nthings we used to do before this is\nactually like far more bitter lesson\nthan like say building a whole physics\nsimulator for a protein but it's also\nreally slow to do this right we need to\nbuild this huge databases the sequence\nalignment it takes time right and it's\nabsent precisely where you often want it\nfor instance the antibody design task we\ncome back to at the end um ESM just\nthrow this away all it says is it just\ntakes input sequence and instead of an\nalignment and it just feeds in the\nmodel's representations as the input to\ntheir structure predictor and these are\njust like per residue embeddings. So\ntake your input sequence you get a set\nof like per amino acid just like some\nnumerical representation and we just\ntrain the specialized module to do\npredicting the like large protein\nstructure right so this folding network\nit's kind of like a projection into uh\n3D corded space so same target same\noutput no handbuilt features and the\nquestion becomes can this general model\nrepresentation like match the sort of\nspecialist model in getting that MSA\nvalue one interesting architectural note\nthough for sort of the more ML folks in\nthe crowd um the one there in their\nprojection networks for the part that\nconverts representation to structure.\nThere is actually one really interesting\nfeature that actually builds off some of\nthe work from our lab alum Dan Fu. Um\nand they have a actually a looped model,\nright? I mean there's a lot of\nexcitement about these recently for\nparameter sharing and I just think\nthey're cool algorithmically for a\nnumber of reasons.\nAnd this is gives them basically a lever\nby which they can scale inference time\ncompute, right? So essentially they have\na model that predicts structures and\nthey have a procedure by which\nrepresentations can be fed through a\nseries of layers and sort of refine\ntheir structure predictions without\nnecessarily retraining or any kind of\nfine tuning. This is like our test time\ncompute access and something like say\ndiffusion steps could be or like test\ntime sampling from LA lab. I'll keep\nthis in mind just for later results. And\nthe sort of like headline figure I would\nsay pointing out here is that um they\nbasically show that yeah their technique\nworks really well. Um just a quick\ndefinitions on the left we have this\nthing called DOCQ pass rate. This is\njust a metric for how good your\nstructure prediction was. Essentially\nit's a measure of the fraction of test\ncases where the predicted shape of two\nproteins stick that stick close together\nis close enough to be like really useful\nto realistic settings. And there are two\ngroups in each panel. One is for single\nsequence with no MSA and the other is um\na single sequence plus an optional MSA\nyou can also feed to the model or is\nrequired for competitor model. And when\nwe look at the sort of outcomes from\nthis, what we see is that for general\nprotein protein complexes, ESM fold 2,\ntheir sort of new projection model from\na single sequence with no MSA lands\nwithin about three points of alpha 3\nwhich does take these handcrafted\nfeatures. So we get near par without the\ncrutch and but if we look at the\nantibody applications which is on the\nright on the left hand side here right\num the modality but is like you know\nessentially behind like all modern NC or\nMAB based drugs or monol antibodies tons\nand tons of applications in human\nbiology and biotech. um we are actually\nwinning or are comparably winning or the\nauthors are comparably winning. So\nbroadly like single sequence ESM fold 2\ndoes actually build alpha fold 3 sort of\n50 versus 47 on this really specific\ndesign task that people really do care\nabout and biologically this makes a lot\nof sense compared to say like um other\nclasses of proteins the amount of sort\nof sequence variation that's been\nsampled in the space of all known\nantibodies relative to structure is\nconsiderably smaller considering their\nenormous diversity. So the headline\nisn't that MSAs are dead yet, right?\nIt's that handle features only help\nwhere it's abundant and basically where\ndrug designers really need it often does\ngo away. And this general method still\nbasically just save a lot from\npre-treating across all known\nrevolutionary contacts.\nSo we're not quite there yet. And one\nother thing though worth flagging is\nthat a second point says give it MSA. We\ncan also scale the amount of test time\ncompute. So how many loops we run in\nthis recursive model in order to prove\nperformance. And we do see um basically\nreturns on this meaning that like the\nbetter loss at least at inference time\nalso seems to broadly hold. And it's not\njust accurate just as an aside. This is\njust quick um it's also just much\nfaster. MSA construction just takes a\nlot of classical computational biology\ntime. So at least if throughput is your\nconcern or latency is your concern, you\ncan with this single representation like\nget quicker results. Though the wall\nclock times here are like you know well\nwithin like I would consider to be\npretty good to start off with.\nUm, and the last bit, I'll get through\nthis a little bit quicker, is just they\ndid a really interesting analysis of\nlike sort of mechanistic\ninterpretability, like what are these\nmodels actually learning and sort of can\nwe find features that are interpretable\nas humans in the same way that sort of\nlanguage modeling folks in the Mechai\ncommunity have found in language models\nlike anthropic has. Um here they sort of\njust apply the same tool or they borrow\na lot of the tools for like sparse\ncoding analysis here where they look at\nactivations from these models and try to\nsee if they can decouple them find these\nlike mono semantic activating directions\ninside their feature spaces and they ask\nis this also going to be a property in\nprotein models and their answer is\nlargely yes um right so from like pure\nfill-in-the-blank pre-training the\nmodel's latent space decomposes into\nclean features that correspond to real\nbiological concepts here the these\nconcepts have been annotated by LM\nagents And they're organized actually\nquite interestingly in a nice hierarchy.\nSo you have like features that\ncorrespond to say individual amino acids\nat the bottom then like structural\nmotifs then like whole protein domains,\nright? So look longer or larger portions\nof the individual protein molecule up to\nlike functional sites and whole protein\nroles, right? And none of this was\nsupervised, right? The model like\nlearned to organize its latent space\npurely just through MLM, which is like\ncrazy. Um\nI'll just with one example maybe to\nclose things out before I finish\neverything and I think I'm actually have\none more slide after this. Um this is a\ninstance of a feature activation that\ncorresponds to a really specific\nwell-known protein motif called the\nnucleophilic elbow. This is a type of\ncatalytic domain that's used in a lot of\nenzyme catalysis. It's really\ninteresting because it's evolved\nmultiple times in multiple different\nproteins unrelated to each other. So\nit's a it's a vitif biology keeps coming\nback to and the model has basically\nlearned to identify in the four quite\nstructurally diverse proteins from like\nboth evolutionary distance as well as\nthe rest of the protein. So it's like\nfound a consistently occurring motif in\nvery different backgrounds. So it's like\nit's basically learn to look at the\nright thing not just sort of memorizing\nlike you know broad similarly comparable\nsequences. It's like a deeper level of\nintuition.\nAnd if you look at the sort of the whole\nSE activation space, you can find like\nnice structures that sort of correspond\nto like various known aspects of\nbiology, right? This organization isn't\njust local, it scales to all of life,\nright? So they um ended up building\nactually a huge atlas of their pro with\ntheir model afterwards sort of just\nfolding and analyzing um millions of or\nup to I think seven billion proteins.\nThis is the largest atlas I think out\nthere in alpha protein structure\ndatabases, more than alpha folds even.\nAnd they've predicted like you know O of\na billion of these as I mentioned before\nand laid them out here by the\nrepresentations in SAPE space and you\nget like a really nice interesting like\nprotein space family map right you can\nfind that there's clear families that'll\ncross clear here are like for instance\ncrisper castine enzymes which if you're\nnot a biologist maybe you still probably\nhave heard of and really important for a\nlot of biotechnology applications it's\nkind of like a Google maps from proteins\nand it's produced all as a byproduct of\nthe model right like just naturally it's\nlike picked up evolutionary relationship\nas well as functional ones just denovo\nfor free which I think is like I don't\nknow if you're maybe not a protein nerd\nlike me I just think this is like\nutterly crazy right\num so like just to finish like does a\nbitter lesson scale to biology not\nperfectly yet I mean some of this\nanalysis still requires a lot of\nhandcrafted features and it's not fully\ncompetitive but we're getting very close\num but even if we just don't care about\none specific downstream the model just\nfrom a relatively quite small amount of\nor like a relatively quite simple\npre-training objective and a lot of data\nhas like learned an enormous amount of\nbio that we can reverse interrogate\nafter the fact um and just for record\nlike they found that data scaling does\nkeep improving. Um I want to just point\nout you know partially as a process like\nour partially just like try to convert a\nlot of smart people like we have in the\naudience there's lots of folks work on\nML a lot of applications software um\nbiology is a great place to work in ML\nbecause the models are still really\nyoung and the other thing is that the\ndata is increasing exponentially per\nyear and that rate of increase is also\ngoing up meaning that like we're not\ndata limited it's a great time to work\nin this space and we need a lot of these\ntools uh and any audience members\nwatching this on YouTube similar pitch\num just as one last thing um I didn't\nget talking into detail but the one\napplication they use for their models\nfor inverse design. So they actually\ndevelop a lot of potential protein drugs\nand they validate a lot of use at least\nin um wet lab settings to show that\nthese are potential like proteins that\nyou can design using this model purely\nin sequence space for the most part by\nthe way um with the exception of like\none structure head at the end um that\nbind various like known molecules that\nhave therapeutic effect right so for\ninstance uh this PDL1 binder is\nbasically the most or is like basically\na medication that is now the sort of big\nsuccess of amunotherapy it's helped\nplenty of patients with cancers in ways\nthat historically have never been able\nto tackle before, right? And developing\nmedications that sort of targeted this\nprotein was immensely challenging. And\nlike if we can basically reduce the\ncosts for developing such future drugs\nfor future targets, it would have\nenormous human impact. So like even if\nthe data scale doesn't sell you, then\nmaybe some of the human impact will. But\nbroadly speaking, it's a really exciting\ntime and it's wonderful to see that a\nlot of these lessons are at least\ntranslating and people are really making\nsteady progress.\nOkay, next we have Luke. um second year\nPhD out of uh Tatsu and Tangu's lab uh\nfresh from the UK. Then he went to\nHarvard CS uh worked on adversarial\nrobustness and now post-training\nselfplay and is directly uh uh in the\nspirit of this um alpha zero kind of\nmindset and so we've been chatting with\nthat about that a lot. All right, please\nwelcome Luke.\nOkay. Hi everyone. Um, yeah, I'm Luke.\nUm, I guess I'll be presenting on\nthis paper we put out uh a few months\nago called Scaling Selfplay with\nSelfguidance. I guess more generally,\nI'll be talking about selfplay for LMS.\nUm, this work was with some great\nco-authors, Caillou, Kan, and my two\nadvisers, Tatu and Tangu.\nOkay, so um, what does the current\ntraining stack look like for big LMS?\nTwo simple parts basically. We pre-train\nthe model on web text and then we\npostrain it. And interestingly recently\nthe post- trainining we've ended up\nspending you know a huge amount of\ncompute on doing large scale long\nreinforcement learning runs. And what\ndoes that reinforcement learning look\nlike? You collect a huge number of tasks\ncoding tasks maths tasks tasks\ninteracting with different bits of\nsoftware. And you just have the agent\ntake a bunch of actions in those\nenvironments. you get some reward back\nand we train the model on that data\nupwaiting the good rollouts down waiting\nthe bad rollouts and like I said the\ninteresting change that's happened is\nwe're now approaching the amount or even\nsurpassing that we're spending on\npre-training actually on this very long\nrunning RL post training and I've swept\nsome things under the rug that we do at\npost training as well like a bit of\ninstruction tuning and and uh alignment\nbut really most of the compute spent on\nthese long RL runs\nokay so we also know that as we increase\nthe number of uh RL tasks during post-\ntraining and we increase the amount of\ncompute we get better downstream\nperformance and I think this is best\nillustrated by this like really\nbeautiful plot from the composer 2\ntechnical report from cursor where what\nthey're both basically showing is they\nhave loads of RL tasks such that they\nonly ever the model only ever sees each\ntask once and so on the x-axis scaling\ntraining step is basically each training\nstep I'm putting in some compute and a\nnew RL task and what they show is nice\nsmooth line as you increase the amount\nof tasks and compute you put in, you get\nthis reliable improvement.\nAnd I guess they had this nice eval set\non the left, but they also have a\ndownstream benchmark on actual coding on\nthe right. And that's also like\nincreasing reliably in a really nice\nway. This recipe tells us great, just\ncollect more and more RL tasks, put them\nin a loop, and model going to keep\ngetting better and better. But\ngenerally, we're going to have to\ncollect these RL tasks by hand, which\nmight be a problem if you want to keep\non feeding. You'll notice log scale on\nthe x-axis there. And I guess there's\nanother problem where you might think\nthat um eventually we'd like the model\nto surpass any of the problems we can\ngive it.\nSo I guess the question that Cell Play\nasks is how can we automatically\ngenerate new RL task to the model, train\non those and repeat.\nOkay, so like I said in traditional RL,\nwe'll have a predefined task and we\ntrain the model on that predefined\nenvironment and task.\nBut in selfplay, we do something\nslightly different where the model does\ntwo things. It's going to generate RL\ntasks and it's going to attempt to solve\nthose tasks. And crucially, we train it\nto be better at both of these things.\nSo, we train it to be better at in\nvirtual commas, we'll go through what it\nmeans to be better to generate tasks and\nthen also to get high reward in those\ntasks.\nSo, how do we fit I guess some papers\nwe've likely seen from the past into\nthis description of selfplay? Because\nyou might be thinking, this doesn't look\nexactly like what I thought of when I\nread the alpha go alpha zero paper. So\nthose traditional works we'd call\nsymmetric selfplay.\nAnd in this case uh let's say in alpha\ngo how you train the model is you have\nthe go agent and then you have the rules\nof go and you have the go board. That's\ngreat but that is nonrl environment I\ncan interact with. Like I need an\nopponent to play against. And so this\ngenerate RL task part. They have an\nolder version of the agent take the role\nof the opponent. So in this case\ngenerating the task because I just put\nan older version of myself in there and\nnow I have a nice RL task. It's a go\nboard with an opponent.\nSo this would traditionally be called\nsymmetric selfplay because the model is\ntaking on the same ball twice, a go\nplayer.\nMore recently, however, uh in the LM\nspace, there's been the rise of\nasymmetric selfplay. This actually hails\nfrom a lot of older work on control\nproblems and things like this. But\nasymmetric selfplay, we instead more\ngenerally just have a model that I will\ncall in this talk a conjecturer that\nwill just generate entire RL tasks for\nthe solver to then operate in. The\nsolver is the equivalent of the agent\nhere. So the conjecturer might come up\nwith a coding problem and then come up\nwith a bunch of unit tests and then\nit'll go into that environment to do a\nbunch of rollouts, get reward and train\non that.\nGreat. So so why do why am I excited\nabout selfplay? Why do I think you\nshould be excited about selfplay? So I\nguess this first point is\nthe first point I have to go go in some\nsome depth. So so in principle nothing\nbounds learning. And what do I mean by\nthat? So if I take a bunch of\ndemonstrations from humans and I train a\nmodel on that, I think it's clear that\nthe model will never get better than\nthose demonstrations.\nSo the next step is okay, I'm going to\ncreate a bunch of environments out of\nthe model learning those environments.\nThat's regular RL. We have two problems\nthere. One, if you ace all of the\nenvironments, you'll never get any\nbetter. Or the second problem is if I\ncan't even get any reward in those\nenvironments, I will also never get any\nbetter. So selfplay on the other hand is\ngoing to say I'm going to keep on\ngenerating new learning signal with new\ntasks. learn it and just keep on\nimproving hopefully forever.\nAnd indeed, we saw this was the case\nwith two-player games like Go. It just\nkept on getting better beyond human uh\nperformance and kept improving. So, the\npromise for LM is I can take some I can\ntrain on a bunch of human data. I get to\nlike human level and then I can run\nloads of selfplay and go far beyond that\nand hopefully solve really interesting\nproblems with with our models.\nBut unfortunately, this is not how it\nworks.\nSo in practice if I run which we'll get\ninto this talk if I run selfplayer for a\nlong time it plateaus I the model stops\nimproving at some point which is the\nexact same that happens when you run RL\nlike as much as I'm trying to tell you\nthere's a bunch of secret source going\non like it doesn't actually play out.\nSo basically this paper we try to figure\nout like why is this happening and then\nlike do one step to solving the problem\nbut by no mean by no means completely\nsolve it. Okay. So to begin with we need\nto understand like the baseline LM\nselfplay algorithm pretty simple we're\ngoing to sample synthetic tasks from the\nconjecturer which is just our model\nconjecture and solver same model just\ngiven it two different names the model\nwill then the solver then attempts them\nand we verify the correctness using\nsome reward signal somehow like perhaps\nthe conjectur wrote unit tests for us to\ncheck and then we're going to update the\nsolver just on all the correct rollouts\nand then this is the key part the\nconjecturer gets updated ated on this\nreward which is zero if the prover if\nthe solver could not solve the problem\nand one minus the solver rate otherwise\nokay what is that actually doing that is\nbasically saying all the conjecturer\nmust do is produce problems that are\nhard for the solver model and I think in\nprinciple that makes a lot of sense the\nidea is if the conjecturer can ace this\nI will keep on giving you problems at\nthe frontier of your capability you will\nbe able to solve them and learn from\nthem and we'll just keep on expanding\nand expanding expanding and get better\nand better and\nOkay, so let's see how this recipe does.\nSo we take in our paper like 3,000\nformal math problems. So this is just uh\nin lean for you can write out the\nproblem statement in this coding\nlanguage in math. So you write out a\nmath problem in this coding language.\nYou can write the proof in the coding\nlanguage. Then you can automatically\nverify if it's correct. So we take 3,000\nproblems and we run like our best RL\nbaseline on it. And this is the amount\nof compute we put in here. And on the y\naxis we have how many problems you\nsolve. And you can see it plateaus out\nand we fit a law and it asmmptotes at\nlike 60%.\nAnd if we on the right hand side we're\ngonna say how much synthetic new task\ndid we generate. RL generates no\nsynthetic task. So by construction this\nstays at zero. Now I'm going to fill in\nthe vanilla selfplayer with that solver\nrate reward. And I'm not going to show\nthe left for now. We see as time goes on\nthe conjecture gets better and better at\nits job. It keeps on generating more and\nmore tasks on the frontier of the\nsolver's capabilities which seems really\ngood and yet these tasks are completely\nuseless. The cell play does no better\nthan regular RL.\nSo this is not very promising. So now we\nneed to understand why. And here what\nI'm visualizing or I'm literally showing\nyou is one of the problems the\nconjecture generates late on in\ntraining. And we don't really understand\nthis. I've highlighted in blue that the\nconclusion to the statement in lean. And\nif anyone's using that this is horrific.\nThis is an incredibly complicated,\noverly complex disaster of a statement.\nAnd so what is basically happening is we\nreward the conjecture for producing\ntricky problems. But the easiest way to\nproduce tricky problems is produce these\nbasically messy, artificially complex\nand elegant problems. It is the exact\nequivalent of if I wanted you to get\nlike 50% solve rate problem, I could\njust give you like a three-page long\nhigh school calculus problem and you\nwould make some little mistake\nsomewhere. But that was a completely\nuseless synthetic problem for like other\ntasks we care about in maths for\nexample.\nGreat. So how do we fix this in a minute\nbecause I've been talking too slowly.\nSo we've diagnosed this problem. Here is\nlike roughly at a high level how we try\nattempt to solve it. So there are two\nparts of our algorithm SGS self-guided\nselfplay. We're going to take the set of\nproblems we cannot solve the 3,000\nproblems and we do two things. one for\neach of those problems we cannot solve\nwe're going to get the conjecturer to\nproduce a related problem to it. So once\nyou prompt it to produce a synthetic\nproblem that is related. So this way\nwe're trying to ground the synthetic\ndata distribution in a distribution of\nproblems that we think is good at least.\nAnd next if you still just trained on\nthe solvent rate reward you would\neventually ignore this prior and still\nproduce that junk. So we're going to\nintroduce a new reward signal which is\nthe model takes on a third role and it\nwill literally judge it looks at\nsynthetic problem and the target problem\nit came from and decide if these two\nthings are actually related and not\noverly complex. So we call this third\ncomponent a guide.\nOkay. So the algorithm looks like this.\nIt's very similar. We for every target\nproblem we haven't solved we'll sample a\nconjecture uh from the conjecture that\nis related to it. We then will attempt\nto solve them. And then what changed\nhere is when we update the conjecturer,\nwe now have this dual reward. One, we\nstill want the problem to be tricky.\nThat is important. So we can get RL\nsignal on it for the solder. And we'll\nmultiply it by this guide score. Great.\nOkay. There's a bunch of kind of\nsubtleties we cover in the paper that I\nwill skip over because we don't have\nloads of time. If you want to talk about\nI'm going to say largeish scale RL\ninfra, the academic size. That's what I\nspent most of my time doing. So I would\nlike to talk about that, but there is\nnot time. So let's just look at the head\nheadline results here. Here is basically\nthe same type of plot. I've put the RL\nbaseline on here. Recall that like\nstandard selfplay is exactly in line\nwith that. We've also put parallel\nsampling down here just to show you that\nindeed RL at least gives you a boost.\nAnd I guess I wouldn't be here unless\nour method works better. So the method\ndoes work better. Um like ground how\nmuch better it's doing. We we we were\nusing a 7 billion parameter model here\nand this is it like 670B like big\nbrother and we spend eight times as much\ncompute doing the selfplay at this you\nyeah we do eight times compute the\nselfplay but we get like to the ability\nof that larger model at least it's pass\nup for ability so you spend a lot more\ncompute but we are able to get this like\nlittle 7B guy to do as well as the\nbigger model but very sadly you will\nnotice like this is not at 100%. So like\nthe work is is by far not done. The\nproblems itself plays like you would\njust ate all the problems here and so\nthere's a bunch of well there's lots of\nfuture work but luckily a PhD is very\nlong so I'll be able to work on that.\nUm but yeah that's the summary.\nAwesome. Thank you so much Luke Bailey.\nOkay uh next one we have Arnab Matei. Is\nthat the right way to say it? Matei. um\nwho is a researcher currently at Giga\none of YC's fastest growing companies I\nthink market cap is like 400 million 300\nmillion something like that now so\nreally fast growing YC company uh PhD\nUniversity of Washington focus on bandit\nlearning um yeah please let's tell us\nabout stream rag\n>> so um there's a paper by the group at\nmeta and I kind of chose chose this\npaper to kind of maybe highlight some of\nthe new emerging challenges that are\ncoming up especially in a voice AI kind\nof setup. Um my goal with this talk is\nnot more like about talking specific\ndetails about this paper but more like\nto highlight the good problems that they\nhave identified and I feel like there's\na lot of research that is to be done\nhere and it also kind of closely mirrors\nwhat at least I do in my production uh\nsetup like I look at these kind of\nproblems I do the research and then I\ntry to come up with a method that will\nwork probably in production. So yeah,\nlet's get started. This is a very\nclassical setup where uh you probably\nask a input question to an alm and it\ngives an output. And if you remember\nmaybe from 2023 maybe there was a lot of\nhallucinations but especially like say\naround citations and all but over time\nmaybe the hallucinations went down and a\nbig role\nwas uh rag like you kind of give the\ninput query to a rag system. it kind of\ngoes and figures out relevant\ninformation that needs to be provided to\nan LLM and then the LM probably gives\nyou an output which is hopefully not\nhallucinated.\nNow uh a lot of voice AI uh startups are\nalso coming up and\na natural expectation with the voice AI\nis that okay you're having like a\nconversation like oh you can ask oh\nwhat's the weather like and the agent\nwould reply like hey the weather\ncurrently is like 22°C\nand so on and maybe you can ask a\nfollow-up question so it's more like\nconversational in nature\nand so even here as well you would like\nthe output to We like there shouldn't be\nany hallucination. Especially in voice,\nwe care about this even more because\nfrom a human perspective, it's difficult\nto kind of actively catch hallucinations\nwhen you're listening to it compared to\nlike when you're reading it over text.\nSo one might ask, okay, what's the issue\nwith just using rag here? like can't you\njust take the input query take apply rag\ngive the relevant information to the\nvoice agent and get the output the issue\nis that rag would add a lot of latency\num like for example if I ask a voice\nagent some question and the voice agent\ntakes 10 seconds to reply that's not at\nall natural especially if you want to\nhave some sort of natural conversation\nso that's where this paper kind of looks\nat A very clever idea I would say uh\nwhich is like instead of like waiting\nfor the question to end and then\nactivate your rag pipeline you kind of\nstart analyzing the words that are being\nspoken by the user and somehow figure\nout a way to\nrun the rag system while the question is\nbeing spoken. Like for example uh like\nyou might ask like hey what's the\nweather today like I'm I want to decide\nbased on that whether I want to go out\nor not. The main question is in the\nfirst part. So the second part of your\nquestion might be irrelevant. So we want\num some sort of uh mechanism via which\nwe can figure out okay uh when to call\nthis rack system and appropriately get\nthe right uh information.\nSo this particular paper focuses on two\napproaches. Uh the first one is fairly\nsimple. Um so it's called fixed rag uh\nfixed interval streaming rag. So the\nidea is like you divide the audio into\ncertain blocks and after each block\narrives you can like run rag on every\nblock. So\nuh so after when the block B arrives you\nrun your rag get the results for the rag\nRB and you keep on doing till this uh\ntill the end probably. Now the question\nthe main question here is like which\nblock to consider because you ideally\ncannot like wait the entire goal was you\ncannot wait till the end and then run\nthe rag. So what do you do? So the main\nuh maybe idea here is that rack pipeline\nhas lot of mini components. So maybe\nsome of the components are like easy to\nrun or like more faster to run. So for\nexample uh you can kind of get some\ndocuments very quickly and you can say\nokay for the entire query what were the\ntop documents and for the intermediate\nquery what were the top documents and\nare they matching or not? This is just\none of the ideas which is from the\npaper. Uh and then based on that you can\ndecide okay should I go ahead with the\nintermediate query and uh just do the\nentire rack pipeline on that. Um so the\nthing I want to stress is not the method\nper se but the point that okay when you\nare getting this uh input in chunks at\nwhat point can you stop and say that\nokay like this chunk is like super\nrelevant for me. Uh so this is like an\nactive question I would say like how\nwould you do that? This uh paper does it\nin a very simple manner which is just to\nmaybe look at the initial path of the\nrack pipeline and if they kind of match\nlike if the end path matches the\nintermediate part then you go ahead with\nthe intermediate and let the full rack\npipeline complete. Another approach\ncould be like you you can probably\nfine-tune a model to kind of trigger on\nits own like when to call the rack\nbecause in the previous approach you\nwere calling rag on every single chunk.\nSo maybe that's computationally\nwasteful. So what you can probably do\nis when a particular chunk arrives you\ncan maybe fine-tune some model and ask\nit to decide whether\nuh this chunk uh is like in critical new\ninformation and you should generate a\nnew query or the query that you\ngenerated based on the past chunks are\ngood enough for you to just answer the\nquestion. And uh based on that you can\ngenerate the final audio.\nYeah. So in in the paper they kind of\ndescribe a post- training pipeline. What\nthey do is like they kind of uh for the\npartial uh spoken uh question they kind\nof generate some pseudo queries using\nsome LLM and then uh they run a rag on\nthat and they look at the retrieved\ndocuments\nand based on the retrieved documents\nthey kind of decide okay is this uh\npartial query like uh something new or\nis it already like we already have the\nuseful material.\nSo in this okay um in this paper\nessentially they are kind of basing\ntheir decision based on the retrieval\nquality of the partial question so far.\nThat's that would be my takeaway. But\nmaybe there are different ways in which\nyou can do this assessment. Maybe you\ncan look into the semantic of the\nquestion so far like is the partial\nquestion so far good enough for me to\nanswer this question just by looking at\nthe question. No no no need to do this\nentire lag pipeline. So there are my my\npoint is like you need not uh this need\nnot be the only way there might be so\nmany different ways and that's where the\nresearch maybe uh is required like while\na user is speaking their question how do\nwe like on like why instead of waiting\ntill the end how do we like figure out\nokay this part of that question is good\nenough for us to go and do the retrieval\nyeah so that that's what they do\nprobably I'll\nquickly give a glimpse of the results\nfrom the paper. Um\nthis paper is a year old. So they were\nlike yeah looking at some smaller open\nsource models. Um so they were they kind\nof considered the rag benchmark\nconverted into audio and uh showed that\nthe latency kind of decreases for the\nsynthetic data sets by 0.5 seconds and\nfor human data sets like human spoken\ndata sets by uh almost like 1.5 seconds\nand uh the accuracy\nuh comparison like uh if there was rag\nuh after the final query and streaming\nrags it kind of remains the same like um\nyeah so yeah so that's what the paper is\nabout so like the key takeaway is like\nthere are some interesting small\nproblems here like but if you can crack\nthe small problems it can lead to huge\ngains in the production yeah thank you\nokay next up we have Robert George. Um,\ncome on up. Uh, thirdyear PhD at\nCaltech.\n>> Yes.\n>> Okay. My brother got his PhD at Caltech.\n>> Um, and, uh, you work on AI for math and\nscience.\n>> Yeah.\n>> Um, and what are you going to tell us\nabout?\n>> I'm going to tell us about lean.\nBasically, Luke already told a little\nbit, but I want to go more in depth. So,\nI'm going to be talking about lean and\nwhat I think is this new era of verified\nintelligence. Um so let's get into it.\nSo again there's bunch of breakthroughs\nin the past like couple of weeks itself\nlike first I want to go back like two\nyears before you know like we said that\nIMO open and even deepine actually at\nthe 2024 IMO got the gold medal then you\nknow there's this very famous problem\nlist which is very famous right now\nwhere people are trying to kind of solve\nnew open unsolved odos problems and you\nknow you can see that it's keep on keep\non increasing with the new models from\nlike open AAI depend and all. Um just\ntwo weeks ago OpenAI claimed to solve\nanother big breakthrough 80-year-old\nOdosh problems. You know Terry Tower was\nhas this promotional video at OpenAI\nwhich he showed really well about these\nkind of things. And then last week\nDeepmind released something also solving\nbunch of new not only ODOS problems but\nproblems in like other different fields\nright but this paper is cool because\nthey also use some kind of formal\nverification in the loop. So I want to\nsay that you know we all took like high\nschool calculus we took undergrad\ncollege math courses and all this you\nknow informal math is very very flexible\nright um your your professor say\nsometimes you know proof by QED like you\nknow sometimes it's like proof by\nintimidation or something that right\nthere many of the steps are not fully\nwritten down but this is where I believe\nthat you know formal world is like you\nhave to be fully explicit right and I'll\ntalk and introduce the language lean\nagain before lean In past couple of\nhundred years, you know, people have\nbeen doing formal math a lot, but you\nknow, lean has just kind of this really\ngood design language just kind of taken\noff, right? So again, first thing is\nit's very easy to check if a proof is\ncorrect or not. You cannot fool this\ntheorem prover. Secondly, uh it's\nscalable. Again, there's bunch of issues\nover there, but I can talk more about\nthat soon. So before that I just want to\ngive you like a precursor. So people do\nknow about like um there was a thing\npreviously like in the 1990s even right\nnow actually 2020s and all this is\nthere's this thing called automatic\ntheorem provers which are basically like\nSMT solvers um they are basically um\nminimal effort from humans you know but\nthey're very limited expressivity in\nwhat type of mathematics they can encode\nin some sense right and on the far right\nhand side you can see interactive\ntheorem provers like lean rogue Isabel\nwhich are very have a much more stricter\nlike expressive logic system. So it's\nbased at least some of them are based on\ndependent type theory but much more\neffort from humans to kind of write down\nthese proofs right like if you're\ntalking about like 10 years people have\nbeen contributing to this very famous\nlibrary called math lab in lean um\nthere's a lot of human effort to kind of\npick premises and all this and again we\nall know how good LLMs are right now at\nkind of combining with these kind of\ntheorem provers to kind of do proof\nchecking for like research level math\nright and it's so much news that I you\nknow if If I go on Twitter right now, I\ncan open up a bunch of posts saying how\nmuch progress past couple of hour\nprobably in some sense right so first\nthing is I want to introduce why leen\nyou know Luke mentioned this formal very\nmessy language but I actually think it's\na very beautiful language again one can\nargue no but um it's a very fast\nlanguage again it's also people think of\nit as only a theorem prover but it's\nactually also a functional programming\nlanguage right you can use it as a\nprogramming language itself so it's\ncompile checking Um it's very good\nunified. So this is what like the proofs\nand programs. Um you can do like meta\nprogramming, you can do macros, custom\nautomation, you know, you can I've seen\npeople trying to even create like games\non with using lean, right? It's actually\nsuper cool. So lean has something called\nthe foreign face interface where you can\ndo like external library bindings like\nyou can do on the CUDA or something that\num I want to point out the math liy. I\nthink that is the coolest biggest\nformalized math library out there. um I\nforgot how many number of lines probably\nat least in a million or so but all of\nthese are really high quality math right\nfrom like say topology to algebraic\ngeometry and all this and again it's an\ninteractive theorem prover so you always\nhave to you know the human can sometimes\nbe in a loop but it is also a very\nscalable language because you know not\nonly frontier labs are pushing a lot of\nmoney into it and also the world is but\num there's more data being generated\neither through synthetic or like a lot\nof people like even myself I do manual\nformalizations. Um so just very short I\ndon't want to take time but this is how\na simple lean code looks like like in VS\ncode you have like an info goal view\nwhich shows like what are the current\nkind of sub goals. So goal is basically\nlike what are you trying to prove at\nthis step. So the first theorem is like\nyou're basically showing associivity of\naddition of like natural numbers right\nlike a plus b plus c is equal to a plus\nc plus b and each line in a proof is\nusually called like a tactic. So usually\nwhen people talk about like proof search\nthey mean like you know you can search\nover this kind of tactic space there are\nmethods where you do foolproof\ngeneration but you know these are the\ntwo different axis. So this is how lean\ncode looks. It's not as bad as it seems.\nIt's a steep learning curve. I think\nit's much better than even C++ in some\nsense like learning but um you at least\nget really um at least for me I get very\nhappy when I see oh I've fully proven\nthis theorem right there's no\nassumptions like I cannot like handwave\nor fool the lean kernel basically like\nyou have to be fully 100% sure um now I\nwant to talk about the formalization\nbreakthroughs right I talked about\ninformal but actually the first book was\nactually in 2020 Ilia and uh Stan was\nfrom open they released something called\nGPDF um this was first generative\nlanguage model for automatic theorem\nproving mini F2F is just like a Olympia\nlevel kind of competition but you see\nthe amount of progress like it's kind of\nexponential right like from open source\nmodels big players in China in the US\nCanada like across the entire world um\nlast year's IMO you know again deep mind\nclaimed to not have used lean I if you\nsee the open air solutions some kind of\nDSL of lean kind of stuff in the\nsolutions um but even Steve prover from\nChina also got the IMO gold And then\nobviously there's a bunch of like axi\nimprover there's harmonic AI like they\ngot recently in the pakam they got all\nthe 12 problems solved\nmost of the odos problems now when\npeople are saying they kind of um claim\nto have a solution using AI they also\nprove it um using like say Aristotle\nfrom harmonic um and then another\namazing work was kind of this fields\nmetal work from math inc and obviously\nthe Google Google deep mind stuff in\nsome sense right and again I love the\nfact that you know everyone's is talking\nabout math and all but you know for me\npersonally there's also these two other\nbubbles right like there's also code now\none can argue what is program\nverification as well you know bugs are\nreally expensive it's like a huge\ntrillion dollar industry wide coding is\nall of a sudden really great like\neveryone is generating but we I want\ncode that needs guarantees right I think\nthat's like something which I'm very\ninterested in and also AI for science\nmatters like there's uh repro uh\nreproducibility and all this kind of\nstuff so I want to go through this\nreally fast but um LMS can write code\nbut can they prove it's correct um you\nknow there's scale of generated code\nthere's that of bugs uh how can you kind\nof capture human intent and the\nverification language and again in short\nI want to talk about like program\nverification is like there's these three\nconcepts where humans actually always\nhave some kind of like specification\nabout like what they want their code to\ndo so a proof is basically saying that\nthe code kind of satisfies that\nspecification um there's this work which\nI introduced called bridge where you can\nuse this lean as a functions programming\nlanguage to kind of elicitate the llms\nto kind of prove this kind of code\nbetter. Um so I like this code from max\ntagm where they say that we should shift\nfrom actually wide coding to like very\ncoding right. Um verifiable coding will\nbe like definitely I think a much more\nbetter way. Um and you should contribute\nto CS lab. This is started from Clark\nBarry's group at Stanford. There's bunch\nof from deep mind and all. But if you\nwant to contribute to CS concepts and\nall, you should definitely contribute\nwith CSL. Um I want to go through\nquickly just about one last work about\nuh torch which I recently introduced.\nThis is the first unified framework for\nactually writing down neural networks in\nlean. So you have this kind of full like\npytor style like tensor system.\nEverything compiles down to a shared\nintermediate representation. You can\nkind of prove properties of specs like I\ncan show you some examples. You have\nlike verified floatingoint arithmetic.\nyou can kind of do even like neural\nnetwork verification like certified\nrobustness kind of stuff right and again\nthere's bunch of applications which I\nshow but I think one cool thing that\nI'll show this and the next slide is\nthat you know you can show that the\nflash attention is equal to like at\nleast in the spec level is equal to the\nuh normal standard attention right again\nwe don't worry about like IO and all\nthis processing also you can a very\nstandard fact is like the attention\nmechanism is permutation in if you don't\nhave position like curtains so I\nactually kind of trained a GP2 style\nlike Karpathi's thing in torching itself\nfully natively in lean right and you can\nkind of prove properties about it and\nall this um one thing I think I can end\nwith this slide is that thinking machine\nlab last year released something about\num this kind of non-determinism even\nwhen you have like temperature zero um\nwhen you put it into your LM inference I\nactually kind of formalize this whole\nsystem in torch lean all the way down to\nlike almost a GPU kind of like small\ncuda level kernel verification because\nthe whole goal in this blog was saying\nthat the tiny floatingoint arithmetics\ncan flip the final argmax in the kind of\nthe batch thing. So again there's a blog\nyou can check it out on my website but\nuh I was very very cool that you can\nkind of do real life software\nverification\num in some sense and uh again there's a\nbunch of different slides I have but I\nkind of want to end on this note just\nfor the sake of time but you know I see\na future where uh science like even code\ncan be formally verified through a lot\nof building blocks which people are\nputting a lot of effort in and this is\none of the examples that I think is like\nmy fuse matter like kind of contribution\nto the ammo wall in some sense.\n>> All right, great job. Okay,\nfor our last presentation, it's going to\nbe the antithesis of lean\nand token maxing to the max. Um, very\nexcited uh to introduce Luke Orthwine,\nhis close friend. um we're friends in in\nuh in Woodside together. Um and did his\nuh CS degree at Harvard, then ran growth\nat WeChatad from 2012 till 2015. Uh\nwhich is why we call him the lion of\nHong Kong.\nUm and now has been running his startup\nchannel AI and is probably the most\nunhinged technical CEO that I know. So\n>> thank you Francois.\nUm yeah so the the idea behind this talk\nis sort of um what we uh at channel have\ndone to try to take the the best\nadvantage of sort of rethinking how you\nshould do software engineering in this\nworld of agentic programming assistance\ncla etc. Um and really you know the the\nways in which uh I think many\nassumptions about what good programming\nis are now sort of the opposite of what\nyou should be doing. Uh and these are\nsort of what we have have worked through\nourselves and found very useful and\nwanted to share with all you guys to\ngive some context. channel AI. We're a\nconsumer entertainment uh AI business.\nUh and we're really focused on the\nproblem of automating as much as\npossible of not just software\ndevelopment but content development. How\ndo you really create like an endto-end\nsystem uh that is pure AI that uh gets\npeople to pay you money uh and stay\nengaged etc. Uh we've had pretty solid\nsuccess with that so far. Um, and it's\ninspired us to think in our own\nworkflows, how can we just sort of max\nthis and and be as far ahead of the\ncurve as possible. Um, and chess is an\nimperfect analogy to what programming\nused to be like, but I think the ways\nthat it uh is useful is like yeah, maybe\nprogramming before you wanted to be very\nlinear. You wanted to predict the\nfuture. You wanted to design very\nthoughtfully systems that would be like\nrobust and work well uh and and be\ncorrect. Um, and even if you're trying\nto do something sloppily, it's still\nlike a single threaded process where you\nonly are worrying at a given moment\nabout what's in front of you. Um, and to\nme, I'm a big fan of real-time strategy\ngames using Agentic systems. Feels\nexactly like playing real-time strategy\ngames to me. Uh, and there are a lot of\nproperties of those games that are very\ndifferent from chess. Um, one thing and\nespecially if you look at like highle\nplay uh there is no single aspect that\nyou can do perfectly and like succeed.\nYou have to be balancing many different\nthings at once. You have to always have\nyour economy running, your production\nrunning, your units doing something\nproductive. You need to be engaging. And\nso this notion of like how do you\nmaximally parallelize both what your\nsystems are doing but also your\nattention so that you are adding the\ncorrective\nuh feedback that's necessary as you\nlearn new things as the map is exposed\nall this kind of stuff. Um anyway this\nto me feels like exactly what like\ncoding with agents is like um and this\nwhat we'll talk about. Um so in terms of\nlike tools we've built just to like\nground this in a very simple thing. This\nis the LW stuff is just like our linear\nwork trees. Um, a lot of people early on\nstarted using realizing how useful git\nwork trees are when you do coding\ndevelopment. Having separate uh I assume\neverybody kind of knows where they are,\nbut in case not like you know it was\nfine to have one repo on your machine\nwhen you were the only one doing\ndevelopment. Now you need to have like\nlots and lots of repos on your machine\nall doing development in parallel. Uh\nall compiling separately and like not\nstepping on each other's toes. Um and so\nthe combination of like uh using work\ntrees, using task management software,\nuh having the actual work itself be\nportable, um which is what the team bit\ncomes in, and then like sticking in\nautonomous agents, one or many different\nones on a given workflow. Um the way\nthat we basically ship stuff, the way I\nship stuff, uh is I have an orchestrator\nagent that's run by Claude usually, but\ncould be codeex 2. Uh, I try to have as\nminimal a number of keystrokes as\npossible to go from like here's an idea\nof something that needs to be fixed to\nwork being started on it because I can\ncourse correct that work later. Think\nlike grabbing a unit and just like\nclicking across the map and you'll come\nback later to like make it work\neffectively. Um, status tracking,\nwatching your mini map, it's the RTS\nequivalent uh from the orchestrator of\nall the different uh spawned workers\nthat you have working. Uh, and then all\nthose workers being instructed basically\nto try to go as far as they can, really\nput like a really low premium on their\ntime and effort and a high premium on\nyours. So even if they're going to be\nwrong, even if they're going to need to\nbe corrected later, it's better for them\nto push as far as they can before they\nask for feedback. Uh, so that you can\njust have a lot of them running in\nparallel, even if it's wasteful from\nlike a per per token standpoint. it's\nlike saving you a lot of time or letting\nyou do more things at once. Anyway, so\nthey try and take everything all the way\nto a PR uh not just a PR but also like a\nsummary that's well I'll get into that\nlater anyway. So uh uh and then like how\ndo you take each the results of every\nworker who completes something and like\nfeed it back into the system so that the\nsystem learns and becomes better again\nlike without the human having to type a\nlot of things or doing minimal work so\nthey can do a lot of these things at\nonce. Uh and then other pieces like how\ndo you tag in other teammates? we'll\nalso get into. Um, but anyway, this is\nvery much like an RTS where you're like\nproducing units, trying to move them\naround, trying to constantly adapt to\nstuff, but also with really high\nvisibility, not just like spawning 20\nagents and like hoping that you'll, you\nknow, solve this problem for me, make no\nmistakes, and it'll just work in the\nend, cuz that doesn't actually happen in\nproduction.\nUm, so like some general guidelines or\nor or practices uh that that that we use\nthat I use uh at least um but but that\nwe've we've uh spread through our team\nis like trying to run almost everything\nincluding scripts that you run because\nsometimes scripts are a lot better and\nsave on context space than than just\nlike doing everything by the LLM\nobviously but running everything from\nthe cloud instances always like never\ntyping anything outside of it if you can\navoid it. Uh having this portability\nbecause a lot of times you start work on\na ticket, you start work on something\nand actually the reason you're stuck on\nit is cuz someone else on your team or\neven maybe another machine. Maybe you're\nrunning it locally on your computer and\nthen you're like, \"Oh like I got\nto go home now, but I want this to run\novernight and I make it really easy to\nmove it elsewhere uh and let other\npeople pick it up. Uh maybe it needs\nmore compute to do something. Whatever.\nIt needs more memory.\" Um, and uh, and\nthen also just like always running in\ndangerously skip permissions mode like\nwhenever possible. Uh, if you can't be\nrunning in dangerously skip permissions\nmode, do what you need to do to like\nmake a sandbox so you can, but if you're\nhaving to give feedback at any regular\npace, like you're going to go really\nslow. Uh, and then like so what yeah,\nwhat do the workers do? As I mentioned\nbefore, they're always trying to go to\nPR. Uh, they are not rigorously adhering\nto like the given spec you do. they're\ntrying to learn and adapt to it as they\ngo because your specs will be wrong. Uh,\nand it's okay for them to make\nassumptions because you can correct them\nuh as you catch them. Um, and then, you\nknow, for like, for example, front-end\ndevelopment doing every everything is\nlike pre-baked into the worker spawn. So\nboot the local dev server, run tests\nyourself on it, have it ready and\nwaiting so that the human can just come\nand open a browser tab pointing to the\nright port and they can just test the\nthing as quickly as possible. Minimizing\nthe number of human steps that need to\nbe taken and like clicks to just move\nsomething forward to the next step uh\nstep. Um and also just like lots of\nthings baked in that are like what are\nthings that we know really reliably? the\nagent's going to be bad about how do we\nlearn about those things, bake them in,\nput them in uh to not just like the\ncloud MD file, but also like broader\nreaching graphs that you have of MD\nfiles uh which I'll get to later uh to\nmake those things\nless of a problem. So, for example, one\nof like the really obvious things that\nClaude is super bad at today is\npredicting how long it'll take to do\nsomething. If you ask it like how long\nis it going to take to solve this\nproblem be like a maybe like two weeks\nof like you know one engineer's work and\nin practice it takes like one prompt and\nit can do it in 20 minutes cuz it's\ntrained on what it would have taken a\nhuman to do those things that's all it's\nlike basis for training data the these\nsystems haven't been around long enough\nfor that to be updated and I think\nthey'll like always be behind anyway so\nyou can take all these things and be\nlike no no never trust yourself in these\nways uh and uh and then also like you\npeople think a lot and a lot of times\nit's kind of true that like the code is\nthe source of truth but the code is\noften like a really expensive source of\ntruth for the agents to pull context out\nof and it's actually really cheap\nespecially when you have all the context\nloaded in memory to like aggressively\ndocument things in a way that benefit\nfuture agents. So uh not just like\nwriting comments in the code but also\nstructured linked uh um sort of wiki\nstyle knowledge knowledgebased files\nthat will make future agents have an\neasy time um basically take advantage of\nthe context as much as you can uh and\nalso helps the visibility of humans and\nand audit auditability of what you do.\nUh, so macro by default, micro win it\ncounts is another RTS principle. Like\nyou can't win a game of RT uh like RTS\ngame usually if you're just really good\nat moving your individual units because\nif you didn't make any units, you're\njust going to lose. Uh so yes, it's\nimportant to like deep dive and tunnel\nvision into certain things that are\nreally critical. Some tickets for sure\ntake a long time, but anytime you're\nlike tunnel visioned into something, you\nshould always be thinking, how do I\nspawn as many other little things that\ndon't take my cognitive bandwidth as\nmuch and just like move those things\nforward? Um, so that always you're\nbasically like maxing out your cognitive\ncapacity. Um, and again, like things can\nwait. You can come back to them like 3\ndays later. It's not that expensive and\nyou can just ask Claude like remind me\nwhat the hell I was doing with this\nthing. All this stuff is really cheap.\nwhat's expensive but doesn't feel\nexpensive is like not doing these things\nat the same time. Um anyway, so macro\nnecessary, micro useful, but you can win\nhonestly in RTS games and I think in a\nlot of things, including in programming,\nif you just macro enough, if you just do\nenough things, you'll kind of uh\nstupidly adjust your way towards\nsomething that's good if you're just\nalways really quickly identifying\nproblems and solving them. Um and yeah,\nthis is gets back to like the high\nvisibility thing. So, one of the things\nthat I really like about you like how I\nset things up is it's not like a lot of\nagents that are kind of tucked away and\nthat you have to like dig in hard to\nactually read what their ongoing stream\nis and what they're actually doing like\nlike in an RTS game like you click\nbuttons to immediately jump to different\nkey points in the map so you can always\nbe auditing stuff and always like catch\nit and correct it quickly if it's a\ncritical thing. That true I that too I\nfind is like super useful in\nprogramming. Uh because again like\nthey're going to make mistakes all the\ntime. They're going to like go in wrong\ndirections and you definitely save time\nand value if you catch them early, fix\nthem, course correct. Uh so you should\nbe kind of like looking around between\nyour different agents, monitoring them\nwhile you are also trying to have as\nmany as you can. Um another thing to\nthis point that I personally like a lot\nuh and is like a big thing in RTS games\nis audio. So, like the only way that you\ncan manage a big army across the whole\nmap is to have lots of audio cues where\nit's like your base is under attack or\nyou know this guy's moving or whatever\nthing is happening. You don't have to be\nlooking at you can hear and it's like\nokay I need to put my attention to\nthis thing and you know based on like a\nlot of variety these audio cues that you\ncan learn and they're good like\npneumatic devices. Uh what's important?\nWhat do I need to act on right away?\nWhat don't I? So, like the way I run my\npersonal setup is I actually have all of\nmy individual agent uh like T-M sessions\nmapped to different Warcraft and\nStarcraft units uh that are colorcoded\nand themed based on the type of ticket\nit is. And then they play the actual\nsound effects from Warcraft and\nStarcraft units. So, I immediately know\nand like visually identify. I don't even\nhave to read like this tab needs my\nattention. This thing's going on.\nAnyway, like to me it just seems like a\nnatural way of like take advantage of\nthese and and again like Cludes made all\nthese things for me really quickly as\nlike a side ticket that I was working on\nover time while I worked on eight other\nthings. So it's like why not do these\nthings and these these devices pneumatic\ndevices uh or or whatever like cues for\npeople are really optimized in gaming\nand they like know what good sound\ndesign is to like be memorable and\notherwise catch your attention in\ndifferent ways. Um, yeah, and like cult\nuse of color, icons, anything that's\njust like quicker to read and process\nbecause I actually do think like these\nthings matter a lot, especially if\nyou're trying to uh really aggressively\nget a lot of stuff done and the sky is\nkind of the limit in how you can do that\nstuff. Another thing we built internally\nis like an APM tracker. Uh, and I'll\njust show quickly here. Um,\nso and this this is Warcraft 3, which is\nlike one of the lower APM requiring\nprofessional RTS games, but this is what\nit looks like to actually play this game\nwell uh at the at the top level. And one\nof the things that you'll notice is like\nno APM is not the uh the thing that like\nif you max it, you're the best player in\nthe world, but nobody is good who\ndoesn't have high APM. And so you can\njust kind of take that as a mental\nrubric like if I'm like thinking\nand like typing slowly and like am I if\nthis was a competition, would I really\nbe the best? Like do I really need to\ntake that much time in everything I'm\ndoing? and how much can I just take like\nlots of little micro decisions and you\nknow fall toward the right uh the right\ngoal or toward making things better. Um\nanyway, so this is just something like\nwe we you know each of us run like\npersonally on our computers and keep an\neye on and it's just like just keep\ntrack of like are things moving and this\nthis APM is not like clicks you have\nbecause I don't think that's like a\ngreat tracker for for for agent use. We\nuse tool you tool calls. It's like how\nmany tool calls are your agents doing\nper minute? Uh this minute, this five\nminutes, this hour, this day, this seven\ndays, like how do you max all those\nthings and have high numbers. Um and\nagain, it's like it's it's one metric\namong many, but it's how are you\nactually being really productive or are\nyou really doing the most you could be\ndoing if you have a low APM? Uh probably\nnot. So otherwise like things probably a\nlot of people know uh easy way to to use\ntokens more effectively is just like do\na lot of things in parallel do different\nthings with the same agent do different\nagents in parallel it will uh invariably\nlike for complex tasks usually give you\na better outcome than if you did it by\nyourself and just like in an RTS like\nyou should be spending your resources\nyou should never have your claude tokens\nlike sitting unused that's really\ninefficient economy like use them all\nevery hour period that you man. Um,\nknowledge base. This is like a really\nbig thing that that for us I think is\nstill like somewhat early on. But, uh,\nthis whole presentation I made and\nstarted the exact same way that, uh, I'm\njust describing how I do tickets, which\nis I went to Claude, I took what France\nasked me to talk about, I pasted it in,\nI said, \"Look at our knowledge base and\nhow we do stuff.\" And put together a\nPowerPoint presentation based on the\nphilosophies embedded in there and like\nwhat I've told you before. and he didn't\nlike oneshot it, but it's like\nI maybe did like 15 edits to it, you\nknow, and and got to this presentation.\nUh, and then I refed it all back into\nthe knowledge base and said like learn\neverything that I've said and all the\nthe the advice I've given and\ncorrections I've given and like make\nthose better instilled in the knowledge\nbase. And this knowledge base is\nbasically just because like linked docs\nare much faster diverse by LLMs. And so\nuh and you can encode everything\nincluding business knowledge and indeed\nlike Claude and and Codex are really\ngood at coming up with features and\nstuff if they have enough knowledge\nabout your business. Uh so trying to\nbuild this up in an automated way is\nsuper useful. People come up with their\nown tickets. Uh because if you have\nsomething you could do everybody you\nshould just like do it. Everybody should\nbe full stack all the time. Uh be\nreactive. Uh and uh even if agent does\nit way worse than you or slower than\nyou, it's still better to have it do it.\nAnd uh it's easy to change things when\nthey're screwed up. Satisficing is a\nword from economics is like do things\nsatisfi like enough but not perfect. Uh\nreally really key principle for like\neverything. Uh mix different ticket\nsizes at the same time. Uh you know in\nlike we we've three and a halfx our\noutput uh PRs per engineer per month. uh\nboth because LM have made ourselves\nbetter, but like when we like really\nadopted this stuff broadly with everyone\non the team this last month, we grew\nanother 60% in our PRs per engineer per\nmonth. So like you're not going to get a\nlot smarter, but the thing you can train\non yourself is like how do I act like\npeople who are good at doing these kinds\nof things really well like RTS pro\nplayers? What does it look like to be\nlike optimal in this and how can I learn\nthe methods of doing it just like\nprogram like an RTS pro? Thank you.\nOkay, I think that's all we have. Um,\nnow I think Vikica, we have cookies, ice\ncream, and popsicles and mochi donuts.\nOkay, what is a mochi donut? It's\ndelicious. Okay. Um, yeah. So, thank you\nguys so much for coming. It was a lot of\nfun. Uh, I will send out a feedback\nform. Please review it and give me give\nme back your thoughts. um think about\nthose uh call for presentations and\ncalls for ideas. If you guys have ideas,\nlet's let's definitely hear them. Um and\nuh looking for more papers coming up\nprobably in in two weeks. I think we're\nalready fully slated. Um and then\nbasically the first one in July uh you\nknow, we're looking to fill out as well.\nSo if you wanted to present, please let\nme know. That's all I got. Thank you\neveryone.",
  "transcript_chars": 83128,
  "ingested_at": "2026-06-17T04:31:56.146307+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": null,
    "like_count": null,
    "channel_id": null,
    "categories": null,
    "tags": null
  }
}