{
  "video_id": "XCLODgdCmKA",
  "channel_slug": "dwarkeshpatel",
  "channel_handle": "dwarkeshpatel",
  "title": "Evolution designed us to die fast; we can change that — Jacob Kimmel",
  "duration_seconds": 6321.0,
  "url": "https://www.youtube.com/watch?v=XCLODgdCmKA",
  "upload_date": "",
  "transcript": "You always have to start by asking\nyourself, did evolution spend a lot of\ntime optimizing this? If yes, my job is\ngoing to be insanely hard. If no,\npotentially there are some lowhanging\nfruit. And so I think that puts human\naging and longevity really in this\ncategory of problem in which it should\nbe relatively speaking easy to try and\nintervene and provide health. We have a\ngene called TRIM 5 alpha. Trim 5 alpha\nonce protected against an HIV like\npathogen. It's currently protecting\nagainst a virus which no longer exists\nand you can edit it back to actually\nrestrict HIV dramatically. You can\nreprogram a cell's type and a cell's age\nsimultaneously just by turning on four\ngenes. Out of the 20,000 genes in the\ngenome, the tens of millions of\nbiomolelecular interactions, just four\ngenes is enough. That's a shocking fact.\nToday, I have the pleasure of chatting\nwith Jacob Kimmel, who is president and\nco-founder of New Limit, where they\nepigenetically reprogram cells to their\nyounger states. Jacob, thanks so much\nfor coming on the podcast.\nThanks so much for having me. Looking\nforward to the conversation.\nAll right, first question. What's the\nfirst principles argument for why\nevolution just like discards us so\neasily? Look, I know evolution cares\nabout our kids, but if we have longer,\nhealthier lifespans, we can have more\nkids, right? Or we can care for them\nlonger, we can care for our grandkids.\nSo, is there some plyotropic effect that\nanti-aging medicine would have which\nactually selects against you staying\nyoung for longer?\nYeah. So, I think there there are a\ncouple different ways one can tackle\nthis. One is you have to think about\nwhat's the selective pressure that would\nmake one live longer and encode for\nhigher health over longer durations. Do\nyou have that selective pressure\npresent? There's another which is are\nthere any anti- selective pressures that\nare actually pushing against that. And\nthere's a third piece of this which is\nsomething like the constraints of your\noptimizer. If we think about the genome\nas a set of parameters and the optimizer\nis natural selection, then you've got\nsome constraints on how that actually\nworks. You can only do so many mutations\nat a time. You have to kind of spend\nyour steps that update your genome in\ncertain ways. So tackling those from a\nfew different directions like what would\nthe positive possible selection be as\nyou highlighted it might be something\nlike well if I'm able to extend the\nlifespan of an individual they can have\nmore children they can care for those\nchildren more effectively that genome\nshould propagate more more readily into\nthe population and so one of the\nchallenges then if you're trying to\nthink back in sort of a thought\nexperiment style of evolution of of uh\nevolutionary simulation here would be\nwhat were the conditions under which a\nperson would actually live long enough\nfor that phenotype to be selected for\nand How often would that occur? And so\nthis brings us back to some very\nhypothetical questions. Things like what\nwas the baseline hazard rate during the\nmajority of human and private evolution?\nThe hazard rate is simply what is the\nlikelihood you're going to die on any\ngiven day? And that integrates\neverything. That's like diseases from\naging. That's getting eaten by a tiger.\nThat's falling off a cliff. That's like\nscraping your foot on a rock and getting\nan infection and dying from that. And so\nfrom the best evidence we have, the\nbaseline hazard rate was very, very\nhigh. And so even absent aging, you're\nunlikely to actually reach those outer\nlimits of possible health where aging is\none of the main limitations. And so the\nnumber of individuals in the population\nthat are going to make it later in that\nlifespan where using some of your\nevolutionary updates to try and actually\npush your lifespan upward is relatively\nlimited. So the amount of gradient\nsignal flowing back to the genome then\nis not as high as one might intuitively\nthink.\nRight? By the way, just on that, often\npeople who are trying to forecast AI\nwill discuss basically how hard did\nevolution try to optimize for\nintelligence and what were the the\nthings which optimizing for intelligence\nwould have prevented evolution from\nselecting for at the same time which\nwould make it so that even if\nintelligence were a relatively easy\nthing to build in this universe,\nit would have taken evolution so long to\nget a human level intelligence. And\npotentially if it was if intelligence\nwould be really easy then it might imply\nthat you know we're going to get to\nsuper intelligence and you know Jupiter\nlevel intelligence etc etc the sky's is\nthe limit. So one argument you know like\nbirth canal sizes etc or the fact that\nyou know we had to spend most of our\nresources on the immune system. But what\nyou just sent that at is actually an\nindependent argument that if you have\nthis high hazard rate, that would imply\nthat you can't be a kid for too long\nbecause you got to, you know, the kids\ndie all the time and you got to become\nan adult so that you can have your kids.\nYeah. You got to contribute resources\nback to the group. You can't just be a\nfreeloader. You need to get calories, go\nout in the jungle, get some berries.\nLike if you're just um if you're just\nhanging out learning stuff for 50 years,\nyou're just going to die before you get\nto have kids yourself. So obviously\nhumans have bigger brains than other\nprimates. We also have longer\nadolescences which help us make use\npotentially of the extra capacity our\nbrain gives us. But if you made the\nadolescence too big uh then you would\njust die before you get to have kids.\nAnd if that's going to happen anyways,\nwhat's the point of making the brain\nbigger? Aka, you know, maybe maybe\nintelligence is easier than we think and\nthere's a bunch of contingent reasons\nevolution didn't turn as hard on this\nvariable as it could have.\nI entirely agree with that particular\nthesis. You know, I think in biology in\ngeneral, when you're trying to engineer\na given property, be it being healthier\nlonger, be it making something more\nintelligent, and this is true even at\nthe micro level of trying to engineer a\nsystem to manufacture a protein at high\nhigh efficiency. You always have to\nstart by asking yourself, did evolution\nspend a lot of time optimizing this? If\nyes, my job is going to be insanely\nhard. If no, potentially there are some\nlowhanging fruit. And so, I think this\nis a good argument for why potentially\nintelligence wasn't strongly selected\nfor. And I think actually the lifespan\nargument plays back into intelligence to\na degree. You start ask\nfor instance in some hypothetical\nuniverse my fluid intelligence lasts\nmuch longer if the number of people who\nare reaching something like 65 is very\nsmall in a population you're not\nnecessarily going to select for alals\nthat lead to fluid intelligence\npreservation late into life. This is\nactually part of my own pet hypothesis\naround some of the interesting\nphenomenology in when discoveries are\nmade throughout lifespans. So there are\nsome famous results where for instance\nI'm going to get the exact age a little\nbit wrong but in mathematics most great\ndiscoveries happen roughly before 30.\nWhy should that be true? That doesn't\nmake sense. You can sort of put down a\nbunch of societal reasons for it. Oh\nmaybe you sort of you know become stay\nin your ways. Your teachers have you\nknow caused you to restrict your\nthinking by that point. But really\nthat's true across centuries. That's\ntrue across many different unique\ncultures around the world. That's true\nin both cultures from the east and\ncultures from the west. That seems\nunlikely to me. I think a much simpler\nexplanation is that for whatever reason,\nour fluid intelligence is roughly\nmaximized at the time where the\npopulation size during human evolution\nwas maximal. If you had to pick an age\nat which fluid intelligence was selected\nmost strongly for, it's probably around\n25 or 30. That's probably about the age\nof the adults in the large populations\nthat were being selected for during most\nof evolution. And so I think there's a\nlot of reason here to think that\nactually there's interplay between many\nfeatures of modern humans and how long\nwe were living and how that dictates\nsome of the features that occur that\nrise and fall throughout our lives.\nSo in one way this is actually a very\ninteresting RL problem, right? It's a\nlong horizon RL problem 20-year horizon\nlength and then there's a scalar value\nof how many kids you have. I guess\nthat's survive etc. Um, and if you know\nhow hard or I I don't know, but if\nyou've heard from your friends about how\nhard um RL is on these models for just\nvery intermediate goals that last an\nhour or a couple hours, it's actually\nsurprising that any signal propagates\nacross a 20 20 year horizon. Um, by the\nway, on the point about fluid\nintelligence speaking, so not only is it\nthe case that in many fields uh\nachievement peaks before 30, in many\ncases if you look at the greatest\nscientists ever, they had many of their\ngreatest achievements in a single year.\nSo um\nyeah the annual mirror bliss yeah\nexactly exactly yeah non what is it\noptics gravity calculus uh 21\ndo you know the Alexander von Halt story\nno\nso Alexander von Halt's one of the most\nfamous scientists in history is kind of\nforgotten now but he had this one\nexpedition to South America where he\nclimbed Mount Chimberazzo at a time when\nvery few Europeans had done that and so\nhe was able to observe various\necological layers that were repeated\nacross latitudes and across altitudes\nand it caused him to formulate an\nunderstanding of how selection was\noperating\nplants at different layers in the\necosystem. And that one expedition was\nthe basis of his entire career. And so\nwhen you see something named Hmel, just\nto give you a sense of how famous this\nguy is, it's usually Alexander von H.\nIt's not like this is some like massive\nprosperous German family name that just\nhappens to be really common. It's this\none guy. Um and so really it was like\nthis singular year in which he conceived\na lot of our modern understanding of\nbotany and selective pressure.\nInteresting. So that's one out of three\ncomponents of uh the evolutionary story.\nYeah. Yeah. So then the next piece of\nthe evolutionary story is like is there\nanything selecting against longevity?\nLike okay, let's just pretend everything\nI said was wrong. Can I still make an\nargument that maybe evolution hasn't\nmaximally optimized for our longevity?\nOne argument that comes up and I'll\ncaveat and say I don't know how strong\nsome of the mathematical models that\npeople put together here are. You can\nfind people using the same idea to argue\nfor and against. But there's this notion\nwhat's called kin selection that if you\nsort of take a selfish gene view of the\nworld that really this is the genome\noptimizing for the genome's propagation.\nit's not trying to optimize for any one\nindividual. Then actually optimizing for\nlongevity is a pretty tricky problem\nbecause you have this nasty\nregularization term which is that if\nyou're able to make a member of the\npopulation live longer, but you don't\nalso counteract the decrease in their\nfitness over time, meaning you maybe\nextend maximum lifespan, but you haven't\ntotally eliminated aging, then the\nnumber of net calories contributed to\nthe genome as a function of that\nperson's marginal year and their own\ncalorie consumption is less than if you\nwere to allow that individual to die and\nactually have two 20-year-olds, for\ninstance, that that sort of follow\nbehind And so there is a notion by which\na population being laden demographically\nwith many aged individuals even if they\ndid have ficundity persisting out some\nperiod later in life is actually net\nnegative for the genome's proliferation\nand that really a genome should optimize\nfor turnover and population size at max\nfitness.\nI love this idea of aging as a length\nregularizer. So I mean people might be\nfamiliar with the idea that when\ncompanies are training models they'll do\nhave a regularizer for you can do chain\nof thought but don't make the chain of\nthought too long and then you're saying\nlike how many calories you consume over\nthe course of your life is that one such\nregularizer\nthat's interesting okay and then the\nthird point was\nthe third piece is basically\noptimization constraints so I think this\nis where another ML analogy is helpful\nwhich is something like well actually a\ntwo-layer neural network is technically\na universal approximator but we can ever\nactually fit them in such a way. And why\ndoes that occur? People will wave their\nhands, but it basically comes down to we\ndon't really know how to optimize them\neven if you can prove out in a formal\nsense that they are universal\napproximators.\nAnd so I think we have similar\noptimization challenges with our genome\nas the parameters and evolution as the\noptimization algorithm.\nAnd one of those is that your mutation\nrate basically bounds the step size you\ncan take. So if you imagine that at each\ngeneration you get some number of\ninputs, you can select for some number\nof alals. Well, the max number of\nvariations in the genome is set by your\nmutation rate. If you dial your mutation\nrate up too high, you probably get a\nbunch of cancers. So, you're select\nyou're selected against. If you have it\ntoo low, you can't really adapt to\nanything. So, you end up with this happy\nmedium, but that limits your total step\nsize. And then the number of variants\nyou can screen in parallel is basically\nlimited by your population size. And so,\nfor most of evolution, there are lots of\nforces constraining population size as\nwell. One of the dominant source of\nselection on the genome is really\nprevention of infectious disease. And it\nseems like when you study the history of\nearly modern man, infectious disease is\nactually what shaped a lot of our\npopulation demographics. And so there's\na lot of pressure pushing for those step\nsizes, those updates to the genome\nreally to be optimizing for protection\nagainst infectious disease rather than\nother things. And so even if you imagine\nthat maybe the arguments on the former\nand the the first and the second of\nthese possible, you know, positive\nselection being absent for longevity and\npotentially some negative selection\nexisting, you could, I think, construct\na reasonable argument for why humans\ndon't live forever, why, you know, the\ngenome hasn't optimized for that simply\nbased on these optimization constraints.\nYou have to imagine not only that the\npositive selection is there and the\nnegative selection is absent, but that\nwhen you think about sort of the\nweighted loss term of all the things the\ngenome is optimizing for, that the\nweight on longevity is high enough to\nmatter. And so even if you imagine it's\nthere, if you simply imagine that the\nlambdas are dialed toward infectious\ndisease resilience more effectively,\nthen you can construct an argument for\nyourself. And so I think really when you\nstart to ask why don't we live forever?\nWhy didn't evolution solve this? You\nactually have to think about an\nincredibly contingent scenario where\nboth the positive selection is there,\nthe negative selection is absent, and\nyou have a lot of our evolutionary\npressure going toward longevity to solve\nthis incredibly hard problem in order to\nconstruct the counterfactual in which\nlongevity is selected for and does arise\nin modern man and in which we are\noptimal. And so I think that puts human\naging and longevity and health really in\nthis category of problem in which\nevolution has not optimized for it. Ergo\nit should be relatively speaking\nrelative to a problem evolution had\nworked on easy to try and intervene and\nprovide health. And I think in many ways\nthe existence of modern medicines which\nare incredibly simplistic. We are\ntargeting a single gene in the genome\nand turning it off everywhere at the\nsame time. And yet the fact that these\nprovide massive benefit to individuals\nis another sort of uh positive emission\nor piece of evidence. Antibiotics are an\neven more clear case of that because\nhere is something that evolution\nactually cares a lot about right. So it\nfeels like antibiotics should\nwhy didn't humans evolve their own\nantibiotics. Yeah, it's actually an\nexcellent question that I haven't heard\nposed before. Um so we think about where\ndo antibiotics come from? To your point,\nwe could synthesize them. They're just\nmetabolites largely of other other\nbacteria and fungi. You think about the\nstory of penicellin. What happens?\nAlexander Fleming finds some fungi\ngrowing on a dish and the fungi secrete\nthis penicellin antibiotic compound and\nso there's no bacteria growing near the\nfungi. and he says he has this light\nbulb moment of oh my gosh they're\nprobably making something that kills\nbacteria. There's no primmaasia reason\nthat you couldn't imagine encoding an\nantibiotic cassette into a mamleian\ngenome. I think part of the challenge\nthat you run into is that you're always\nan evolutionary competition. There's\nthis notion of what's called the red\nqueen hypothesis. It's an illusion to\nthe story in Louiswis Carol's Through\nthe Looking Glass where the red queen is\nrunning really fast just to stay in\nplace. So when you look at sort of\npathogen host interactions or\ncompetition between bacteria and fungi\nthat are all trying to compete for the\nsame niche, what you find is they're\nevolving very rapidly in competition\nwith one another. It's an arms race.\nEvery time a bacteria evolves a new\nevasion mechanism, the fungus that\noccupies the niche will evolve some new\nantibiotic.\nAnd so part of why that there is this\ncompetitiveness between the two is they\nboth have very large population sizes in\nterms of number of genomes per unit\nresource they're consuming. There are\ntrillions of bacteria in a drop of water\nthat you might pick up. So there's\ntrillions of copies of the genome,\nmassive analog parallel computation, and\nthen at the same time, they can tolerate\nreally high mutation rates because\nthey're proarotic. They don't have\nmultiple cells. So if one cell manages\nto mutate too much and it isn't viable\nor it grows too fast, it doesn't really\ncompromise the population in the whole\ngenome. Whereas for mezzoans like you\nand I, if even one of our cells has too\nmany mutations, it might turn into a\ncancer and eventually kill off the\norganism. So basically what I'm getting\nat, and this is a long-winded way of\ngetting there, is that bacteria and\nother types of microorganisms are very\nwell adapted to building these complex\nmetabolic cascades that are necessary to\nmake something like antibiotics. And\nthey are necessary, it's necessary to\nmaintain that same mutation rate and\npopulation size in order to maintain the\ncompetition. Even if our human genome\nstumbled into making an antibiotic, most\npathogens probably would have mutated\naround it pretty quickly. And actually\nthat should imply that there's\nthrough evolutionary history millions of\nquote unquote naive antibiotics\nwhich could have acted as antibiotics\nbut now basically all the bacteria have\nevolved around it. Do we see evidence of\nthese like historical antibiotics that\nsome fungi came up with and the bacteria\nevolved around and all there's evidence\nfor remnant in their DNA?\nI'm I'm going a bit beyond my own\nknowledge here. So I want to say my\nstrong hypothesis would be yes. I can't\npoint to direct evidence today. There\nare some examples of this where for\ninstance bacteria that uh fight off\nviruses that infect them, bacteria\nphasages have things like crisper\nsystems and you can actually go and look\nat the spacers, the individual guide\nsequences that tell the crisper system\nwhich genome do you go, where do you cut\nand you find some of these guides that\nare very ancient. It seems like this\nbacterial genome might not have\nencountered that particular pathogen for\nquite a while. And so you can actually\nget sort of an evolutionary history of\nwhat was the warfare like? what were the\nvarious conflicts throughout this\ngenomic history just by looking at those\nsequences. In mammals where I do know a\nbit better, we do have examples of this\nwhere there is this co-evolution of\npathogen and host. Imagine you have some\nantiathogen gene A fighting off some\nvirus X. Well, you then actually update.\nSo now you have virus X prime and\nantipathogen gene A prime. Now virus X\nprime goes away, but actually virus X\nstill exists and we've lost our ability\nto fight it. Those examples really do\nhappen. And so there's a prominent one\nin the human genome. So we have a gene\ncalled trim 5 alpha and it actually\nbinds a indogenous retrovirus that is no\nlonger present but was at one point\nactually resurrected by a bunch of\nresearchers and it was demonstrated that\nit is the case. We have this indogenous\ngene which basically fits around the\ncapsit of the virus like a baseball and\na glove and prevents it from infecting.\nAnd it turns out if you look at the\nevolutionary history of that gene and\nyou trace it back through monkeys you\ncan actually find that a previous\niteration inhibited SIV which is the\ncousin of HIV in humans. And so old\nworld monkeys actually can't get SIV\nwhereas new world monkeys can and humans\ncan obviously. And so it seems like what\nhappened and you can you can actually\nmake a few mutations in trim 5 alpha and\nfigure find that this is true is that\ntrim 5 alpha once protected against an\nHIV like pathogen in the primate genomes\nand then there was this challenge from\nthis massive indogenous retrovirus and\nit was so bad that the genome lost the\nability to fight off these HIV like\nviruses in order to restrict this\nindogenous retrovirus. And you can see\nit because that retrovirus integrates\ninto our genome. There are like latent\ncopies like the, you know, half bodies\nof this virus all throughout our DNA\ncode.\nAnd then this particular retrovirus went\nextinct. Reasons unknown. No, no one\nknows why. But we didn't like re-update\nthat piece of our host defense machinery\nto fight off HIV again. And so we're in\na situation where you can go in and take\nhuman cells and make just a couple edits\nin that trim 5 alpha gene and it's\ncurrently protecting against a virus\nwhich no longer exists and you can edit\nit back to actually restrict HIV\ndramatically. So there are plenty of\nexamples. So you could imagine the same\nthing for antibiotics. We're like, hey,\nthis particular, you know, defense\nmechanism went away because the pathogen\nevolved its own defense to it. Well, the\npathogen might have lost that defense\nlong ago. And if you could sort of\nextract that historical antibiotic, that\nhistorical antifungal, potentially, it\nactually has efficacy.\nIsn't the mutation rate per base pair\nper generation like one in a billion or\nsomething?\nIt's quite low. So you're saying that in\nour genomes we can find some extended\nsequence which encodes how how to bind\nspecifically to the kind of virus that\nSIV is and the amount of evolutionary\nsignal you would need in order to have a\nmultiple base pair sequence. So each\nnucleotide consecutively would have to\nmutate in order to finally get the\nsequence that binds to SIV.\nThat seems almost implausible that you\ncould like I mean I guess evolution\nworks so like we can come up with new\ngenes, right? But like how how would I\nthat even work?\nI think I think a great explanation for\nunderstanding a lot of evolution and how\nyou're able to actually adapt to new\nenvironments, new pathogens is that gene\nduplication is possible. And this\nexplains a whole lot. If you look at\nmost genes in the genome, they actually\narise at least at some point in\nevolution from a duplication event. So\nthat means you've got gene A, it's\ndoing, you know, it's performing some\njob and then some new environmental\nconcern comes along. Maybe it's like a\nlack of particular to source of\nnutrient. Maybe it's a pathogen\nchallenging you. And maybe gene A, if it\nwere to dedicate all of its energies, so\nto speak, you were to mutate it to solve\nthis new problem, could be adapted with\na minimal number of mutations, but then\nyou lose its original function. So we\nhave this nice feature of the genome,\nwhich is it can just copy and paste. And\nso occasionally what'll happen in\nevolution is you get a copy paste event.\nNow I've got two copies of gene A and I\ncan preserve my original function in the\noriginal copy. And then this new copy\ncan actually mutate pretty freely\nbecause it doesn't have a strong\nselective pressure on it. So most\nmutations might be null. I've got two\ncopies of the gene. I can have lots of\nmutations in it accumulate. Nothing bad\nreally happens because I've got my\nbackup copy, my original. And so you can\nend up with drift. So you're saying that\neven though the per base pair mutation\nrate might be one in a billion, if\nyou've got a 100 copies of a gene, then\nthe sort of like mutation rate on a gene\nor on a low Hamming distance um sequence\nto the one you're aiming for might\nactually be quite high and so you can\nactually get the target sequence.\nIt's not that the base rate goes up.\nIt's not like DNA polymerase is, you\nknow, more erroneous or that you're just\nlike doubling it. It's not like, oh\nwell, I've got two copies. That is true,\nbut I don't think it's the main\nmechanism. The ma one of the main\nmechanisms that just makes it difficult\nfor evolution to solve a problem is that\nif a mutation breaks a gene or somewhere\nalong the path of edits, imagine there\nare three edits that take a host defense\ngene from restricting SIV to restricting\nthis new nasty PT endogenous retrovirus.\nWell, if one edit just breaks the gene,\ntwo edits just breaks the gene, three\nedits fixes it, it's really hard for\nevolution to find a path whereby you're\nactually able to make those first two\nedits because they're net negative and\nnet negative for fitness. And so you\nneed some really weird contagion\ncircumstances. So through duplication,\nyou can create a scenario where those\nfirst two edits are totally tolerated.\nThey have like no effect on fitness.\nYou've got your backup copy. It's doing\nits job. And so even though the mutation\nrate is low, some of these edits\nactually aren't that large. I'm going to\nforget the number of edits, for\ninstance, in trim 5 alpha for this\nparticular phenomenon we're talking\nabout for memory, but it's in like the\ntens. It's not it's not that you need\nmassive kilobase scale rearrangements.\nIt's actually a fairly small number of\nedits. And basically you can just align\nthe sequence of this gene in new world\nversus old world monkeys and then for\nhumans and you find there's a very high\ndegree of conservation\nconceptually. Is there um some\nphoggenetic tree of gene families where\nyou've got the transposons and you've\ngot like the gene itself but then you've\ngot like the descendant genes which are\nlike low heming distance. Um I don't\nknow is there like some conceptual way\nin which they're categorized? Okay.\nYou can arrange genes in the human\ngenome by homology to one another. And\nwhat you find is even in our current\ngenome, even without having the full\nhistorical record, there are many many\ngenes which are likely resulting from\nduplication events. One like trivial way\nthat you can check this for yourself is\nlike just go look at the names of genes.\nAnd very often you'll see something\nwhere it's like gene one, gene 2, gene\nthree, or you know, type one, type two,\ntype three. And if you then go look at\nthe sequences, sometimes those names\narise from like they were discovered in\na common pathway and they have nothing\nto do with each other. A lot of the time\nit's because the sequences are actually\nquite darn similar. And really what\nprobably happened is they evolved\nthrough a duplication event and then\nmaybe did some swapping with some other\ngenes and you you ended up with these\nquite similar quite homologous genes\nthat now have specialized functions. So\nit's like when evolution has a new\nproblem to solve. It doesn't have to\nstart from scratch. It starts from like\nwhat was the last copy of the parameters\nfor encoding a gene that is getting\nclose to solving this. Okay, let's do a\ncopy paste on that and then iterate and\nfine-tune on those parameters as opposed\nto having to start with like abinio some\nrandom stretch of sequence somewhere in\nthe genome has to become a gene.\nInteresting man. This is fascinating.\nOkay. Um, back to aging. You'll cancel\nyour evening plans. I've got so many\nquestions for you and I keep going. Um,\nso the second reason you gave which was\nthat\nthere's selective pressure against\npeople who get old but\nstill keep living but they're like\nslightly less fit.\nThey're sub-optimal from a calorie input\nperspective. Yeah, the number of\ncalories they can gather for the\npopulation.\nThat's how people love thinking about\ntheir grandpas, you know,\noptimal optimal provider right there.\nUm, anyways, so a concern you might have\nabout the effects of longevity\ntreatments on your own body is that you\nwill fix some part of the aging process\nbut not the whole thing. Mhm.\nIt seems like you're saying that you\nactually think this is the default way\nin which an anti-aging procedure would\nwork because that's the reason Evolution\ndidn't optimize it for it. It's just\nthat like we're only fixing half of the\naging process and not the whole thing.\nWhereas sometimes I hear longevity\nproponents be like no, you know, we'll\nget the whole thing. There's like going\nto be a source that explains all of\naging and we'll get it. Um whereas your\nevolutionary argument for why evolution\ndidn't optimize against aging relies on\nthe fact that aging actually is not\nmonocausal uh and evolution didn't\nbother to just fix one cause of aging.\nYeah, I I think that's correct. I I\ndon't think that there is a single\nmonocausal explanation for aging. I\nthink there are layers of molecular\nregulation that explain a lot. For\ninstance, I have dedicated my career now\nto working on epigenetics and trying to\nchange which genes cells use because I\nthink that explains a lot of it. But\nit's not that there is like some\nupstream bad gene X and all we have to\ndo is turn that off and suddenly aging\nis solved. And so I think the most\nlikely outcome is that when we\neventually develop medicines that\nprolong health in each of us, it's not\ngoing to fix everything all at once.\nThere's not going to be a singular magic\npill, but rather you're going to have\nmedicines that add multiple healthy\nyears to your life, years you can't\notherwise get back, but it's not going\nto fix everything at the same time. You\nare still going to experience for the\nfirst medicine some amount of decline\nover time. And this gives you an example\nof if you think about evolution as a\nmedicine maker in this sort of\nanthropomorphic context, why might not\nhave been selected for immediately. What\nwould the AI foundation model for\ntrading and finance look like? It would\nhave to be what LLM NLP or what the\nvirtual cell is for biology. And it\nwould have to integrate every single\nkind of information from around the\nworld from order books to geopolitics.\nNow, think about how insane this\ntraining objective is. Here's this\nconstantly changing RL environment with\ninput data that's incredibly easy to\noverfitit to where you're pitted against\nextremely sophisticated agents who are\nlearning from your behavior and plotting\nagainst it. Obviously, there's very few\nthings in the world that are as complex\nas global capital allocation. It's a\nsystem that reflects billions of live\ndecisions in real time. Now, as you\nmight imagine, trading an AI to do all\nof this is a compute intensive task.\nThat's why Hudson River Trading\ncontinually upgrades its massive\nin-house cluster with fresh racks of\nbrand new B200s being installed as we\nspeak and more on the way. HRT executes\nabout 15% of all US trading equities\nvolume and researchers there get\ncompensated for the massive website that\nthey create. If the newest researcher on\nthe team improves an HRT model, their\ncontributions are recognized and\nrewarded right away regardless of their\ntenure. If you want to work on high\nstakes, unsolved problems unconstrained\nby your GPU budget, check out HRT at\nhusten rivertrading.com/4ache.\nAll right, back to Jacob. All right, so\num uh evolution didn't select for aging.\nWhat are you doing? What's your approach\nat New Limit that you think is um is\nlikely to find the true cause of aging?\nYeah, so we're working on something\ncalled epigenetic reprogramming, which\nvery broadly is using genes called\ntranscription factors. I like to think\nabout these as sort of the orchestra\nconductors of the genome. They don't\nperform many functions directly\nthemselves, but they bind specific\npieces of DNA and then they tell which\ngenes to turn on, which genes to turn\noff. They eventually put chemical marks\non top of DNA on some proteins that DNA\nsurrounds. And this is one of the\nanswers, this particular layer of\nregulation called the epiggenome. It's\nthe answer to this fundamental\nbiological question of how do all my\ncells have the same genome but\nultimately do very different things.\nYour eyeball and your kidney have the\nsame code and yet they're performing\ndifferent functions. And that may sound\na little bit simplistic, but ultimately\nI think it's kind of a profound\nrealization. And so that epigenetic code\nis really what's important for cells to\ndefine their functions. That's what's\ntelling them which genes to evoke from\nyour genome. What is has now become\nrelatively apparent is that the\nepigenome can degrade with age. It\nchanges. The particular marks that tell\nyour cells which genes to use can shift\nas you get older. This means that cells\naren't able to use the right genetic\nprograms at the right times to respond\nto their environment. You're then more\nsusceptible to disease. You have a less\nless resilience to many insults that you\nmight experience. And our hope is that\nby remodeling the epiggenome back toward\nthe state it was in when you were young,\nright after development, that you'll be\nable to actually address myriad\ndifferent diseases whose one of strong\ncontributing factors is that cells are\nless functional than when you were at an\nearlier point in your life. So, we're\ngoing after this by trying to find\ncombinations of these transcription\nfactors that are able to actually\nremodel the epiggenome so that they can\nbind to just the right places in the DNA\nand then shift the chemical marks back\ntoward that state when you were a young\nindividual. M\nif you're just making these broad uh\nchanges to a cell state uh through these\ntranscription factors which have many\neffects, are there other aspects of a\ncell state that are likely to get\nmodified at the same time in a way that\nwould be deletarious or would it be um a\nsort of straightforward effect on cell\nstate?\nOh, how I wish it were straightforward.\nUm no, it's very likely the each of\nthese transcription factors binds\nhundreds to thousands of places in the\ngenome. And one way of thinking about it\nis if you imagine the genome is sort of\nthe base components of cell function,\nthen these transcription factors are\nkind of like the basis set in linear\nalgebra. It's different combinations and\ndifferent weights of each of the genes.\nAnd so most of them are targeting pretty\nbroad programs. And there are no\nguarantees that aging actually involves\nmoving perfectly along any of the\nvectors in this particular basis set.\nAnd so it's probably going to be a\nlittle tricky to figure out a\ncombination that actually takes you\nbackward. There's again no guarantees\nfrom evolution that it's just a simple\nreset. And so it's actually a critical\npart of the process that we run through\nas we try and discover these medicinal\ncombinations of transcription factors we\ncan turn on is to ensure that they not\nonly are making an age cell revert to a\nyounger state. We measure that a couple\ndifferent ways. One is simply measuring\nwhich genes those cells are using. They\nuse different genes as they get older.\nYou can measure that just by sequencing\nall of the mRNAs which are really the\nexpressed form of the genes being\nutilized in the genome at a given time.\nYou see that age cells use different\ngenes. Can I revert them back to a\nyounger state? colloally we call this,\nyou know, a looks like assay. Can I make\nan old cell look like a young one based\non the genes it's using? And maybe more\nimportantly, we go down and drill to the\nfunctional level and we measure can I\nactually make an age cell perform its\nfunctions, its object roles within the\nbody the same way a young cell would.\nAnd these are the really critical things\nyou care about for treating diseases.\nCan I make a hypatocite, a liver cell in\nGreek, um, function better in your\nliver, so it's able to process\nmetabolites like the foods you eat, how\nit's able to process toxins like alcohol\nand caffeine. um can I make a T- cell\nrespond to pathogens and other antigens\nthat are presented within your body? So\nthese are the ways in which we measure\nage and so we need to ensure that not\nonly does the combination of TFS that we\nfind actually have positive effects\nalong those axes but we then want to\nalso measure any potential detrimental\neffects that that emerge.\nSo there are canonical examples where\nyou can seemingly reverse the age of a\ncell for instance at the level of a\ntranscriptto but simultaneously you\nmight be changing that cell's type or\nidentity. So Shinyamanaka, a scientist\nto woman Obel in 2012 for some work he\ndid in about 2007 discovered that you\ncould just take four transcription\nfactors and actually just by turning on\nthese four genes turn an adult cell all\nthe way back into a young embryionic\nstem cell. So it's a pretty emer amazing\nexistence proof that shows that you can\nreprogram a cell's type and a cell's age\nsimultaneously just by turning on four\ngenes. Out of the 20,000 genes in the\ngenome, the tens of millions of\nbiomolelecular interactions, just four\ngenes is enough. That's a shocking fact.\nAnd so we actually have known for many\nyears now that you can reprogram the age\nof a cell. The challenge is that\nsimultaneously you're doing a bunch of\nother stuff as you alluded to. You're\nchanging its type and that might be\npathological. If you did that in the\nbody, it would probably cause a type of\ntumor called a terteratoma. So we\nmeasure not only at the level of the\ngenes a cell is using. Do you still look\nlike the right type of cell? Are you\nstill apatite? Are you still a T- cell?\nIf not, that's probably pathological.\nBut you can also use that same\ninformation to check for a number of\nother pathologies that might develop.\nDid I make this T- cell\nhyperinflammatory in a way that would be\nbad? Did I make this liver cell uh\npotentially neoplastic proliferate too\nmuch even when the organism is healthy\nand undamaged? And you can check for\neach of those at the level of gene\nexpression programs. And then likewise\nfunctionally before you put these\nmolecules in a human, you actually just\nfunctionally check in an animal. You\nmake an itemized list of the possible\nrisks you might run into. Here are the\nways it might be toxic. Here are the\nways it might cause cancer. Are we able\nto measure deter deterministically and\nand empirically that that doesn't\nactually occur? Okay, this is a dumb\nquestion, but it will help me understand\nwhy an AI model is necessary to do any\nof this work. So, you mentioned the\nYamanaka factors. From my understanding,\nthe way he identified these four\ntranscription factors was that he found\nthe 24 um transcription factors that\nassociated uh that have high expression\nin embryionic cells and then he just\nturned them all on in a sematic cell.\nBasically, he systematically removed\nfrom this set until he found like the\nminimal set that still induces a cell to\nbecome a stem cell and that just like uh\ndoesn't require any fancy AI models etc.\nWhy can't we do the same things for the\ntranscription factors that are\nassociated with younger cells or express\nmore in younger cells as opposed to\nolder cells and then keep eliminating\nfrom them until we find the ones that\nare necessary to just make a cell young.\nI wish it were so easy. Um you're\nentirely right. You know, Shinyamanaka\nwas able to do this with a relatively\nsmall team with relatively few resources\nand achieve this remarkable feat. So,\nit's entirely worth asking, why can't a\nsimilar procedure work for arbitrary\nproblems in reprogramming cell state,\nwhether it be trying to make an age cell\nact like a young one, disease cell act\nlike a healthy one, why can't you just\ntake 24 transcription factors and\nrandomly sort through them?\nSo, there were two features of Shinyo's\nproblem that I think make it amendable\nto that sort of interrogation that\naren't present for many other types of\nproblems. And this is why he's such a\nremarkable scientist. Most of science is\nproblem selection. You don't actually\nget better at pipetting or running\nexperiments after a certain age, but you\ndo get better at picking what to do. And\nand he's amazing at this. So the first\nfeature is that measuring your success\ncriterion is trivial in the particular\ncase he was investigating. He's starting\nwith sematic cells that in this case\nwere a type of fiberblast, which\nliterally is defined as cells that stick\nto glass and grow in a dish when you\ngrind up a tissue. So it's like sounds\nfancy, but it's a very very simplistic\nthing. So he's starting with fiberblast.\nYou can look at them under a microscope\nand you can see their fiberblast just\nbased on how they look. And then the\ncells he's reprogramming toward are\nembryionic stem cells. So these are tiny\ncells. They're mostly nucleus. They grow\nreally, really fast. They look\ndifferent. They detach from a dish. They\ngrow up into a 3D structure. And they\nexpress some genes that will just never\nbe turned on in a fiberblast by\ndefinition. So actually how he ran the\nexperiment was he just set up a simple\nreporter system. So he took a gene that\nshould never be on in a fiberblast,\nshould only be on in the embryo, and he\nput a little reporter behind it so that\nthese cells would actually turn blue\nwhen you dumped a chemical on them. And\nthen he ran this experiment in many,\nmany dishes with, you know, millions\nupon millions of cells. The second\nreally key feature of the problem is\nthis notion that those cells he's\nconverting into amplify. They divide and\ngrow really quickly. So in order for you\nto find a successful combination, you\ndon't actually need it to be efficient\nalmost at all. The original efficiency\nYamanaka published the number of cells\nin the dish that convert from sematic to\nan induced puropotent state back into a\nstem cell is something like a basis\npoint or a tenth of a basis point. So\nlike 01.001%.\nIf these cells were not growing and they\nwere not proliferating like MAD, you\nprobably would never be able to detect\nthat you had actually found anything\nsuccessful. It's only because success is\neasy to measure once you have it. and\neven being successful in very rare\ncases, one in a million amplifies that\nand you can detect it that this I think\nwas amendable to to his particular\napproach. So in practice what he would\ndo is dump these factors or this group\nof 24 minus some number eventually\nwhittling it down to four, he would dump\nthese onto a group of cells and over the\ncourse of about 30 days, just a few\ncells in that dish, like a countable\nnumber on your fingers would actually\nreprogram, but they would proliferate\nlike mad. They form these big what we\ncall colonies because it's like a single\ncell that just proliferates and forms a\nbunch of copies of itself. They form\nthese colonies. You can see with your\neyeballs by holding the dish up to the\nlight and looking for opaque like opaque\nlittle dots on the bottom. You don't\nneed any fancy instruments. And then you\ncould stain them with this particular\nstain and they would turn blue based on\nthe genetic reporter he had. So now we\nlook at those key features of the\nproblem and we pick any other problem\nwe're interested in. I'm interested in\naging. So that's the one I'm going to\npick for explanation. How difficult is\nit to measure the probability or the\nlikelihood of success or whether you've\nachieved success for cell age? Well, it\nturns out age is much more complicated\nin terms of discriminating function than\nactually just comparing two types of\ncells. An old liver cell and a young\nliver cell primascia actually look\npretty darn similar. It's actually quite\nnuanced the ways in which they're\ndistinct. And so there isn't a simple\ntrivial system where you just like label\nyour one favorite gene or you can\njust give the young cells cancer.\nThey'll grow, you know, you want to see\nthem.\nJust just make the old ones cancer and\nthen they'll grow. Yeah, Dorash, you've\nsolved it for me. Um, so the there's no\ntrivial way that you can tell whether or\nnot you've succeeded. You actually need\na pretty complex molecular measurement.\nAnd so for us, a real key enabling\ntechnology, and I don't think our\napproach would really have been possible\nuntil it's emerged, was something called\nsingle cell genomics. So you now take a\ncell, rip it open, sequence all the\nmRNAs it's using. And so at the level of\nindividual cells, you can actually\nmeasure every gene that they're using at\na given time and get this really\ncomplete picture of a cell's state,\neverything it's doing, lots of mutual\ninformation to other features. And from\nthat profile, you can train something\nlike a model that discriminates young\nand aged cells with really high\nperformance. It turns out there's no one\ngene that actually has that same\ncharacteristic. So unlike in Yamanaka's\ncase where a single gene on or off is\nlike an amazing binary classifier, you\ndon't have that same feature of easy\ndetection of success in aging. The\nsecond feature is as you highlighted, we\ncan't just turn these into cancer cells.\nSuccess doesn't amplify. And so in some\nways the bar for a medicine is higher\nthan what Yamanaka achieved in his\nlaboratory discovery. You can't just\nhave 0001% success and then wait for the\ncells to grow a whole bunch in order to\ntreat a patient's disease or you know\nmake their liver younger, make their\nimmune system younger, make their\nendothelium younger. You need to\nactually have it be fairly efficient\nacross many cells at a time. And so\nbecause of this, we don't have the same\nluxury Yamanaka did of taking a\nrelatively small number of factors and\nfinding a a success case within there\nthat was pretty low efficiency. We\nactually need to search a much broader\nportion of TF space in order to be\nsuccessful. And when you start playing\nthat game and you think, okay, how many\nTFs are there? Somewhere between a,000\nand 2,000. Depends on exactly where you\ndraw the line. And devi developmental\nbiologists love to argue about this over\nbeer, but let's call it 2,000 for now.\nAnd you want to choose some combination.\nLet's say you guess it's like somewhere\nbetween 1 and six factors might be\nrequired. The number of possible\ncombinations is about 10 to the 16. So\nif you do any like math on the back of a\nnapkin, in order to just screen through\nall of those, you would need to do many\norders of magnitude more single cell\nsequencing than the entire world has\ndone to date cumulatively across all\nexperiments. And so it's just not\ntractable to do exhaustively. And so\nthat's where actually having models that\ncan predict the effect of these\ninterventions comes in. If I can do a\nsparse sampling, I can test a large\nnumber of these combinations and I can\nstart to learn the relationship of what\na given transcription factor is going to\ndo to an age cell. Is it going to make\nit look younger? Is it going to preserve\nthe same type? I can learn that across\ncombinations. I can start to learn their\ninteraction terms. Now I can use those\nmodels to actually predict in silico for\nall the combinations I haven't seen\nwhich are most likely to give me the\nstate I want. And you can actually treat\nthat as a generative problem and start\nsampling and asking which of these\ncombinations is most likely to take my\ncell to some target destination in state\nspace. In our case, I want to take an\nold cell to a young state. But you could\nimagine some arbitrary mappings as well.\nAnd so I think as you get to these more\ncomplex problems that don't have the\nsame features that Sha benefited from,\nwhich were the ability again to measure\nsuccess really easily, you can see it\nwith your bare eyes, you don't even need\na microscope, and two amplification. As\nyou get into these more challenging\nproblems, you're going to need to be\nable to search a larger fraction of the\nspace to to hit that higher bar.\nSo we can think of these transcription\nfactors as these basis directions and\nyou get like a little bit of this thing,\na little bit of that thing, and some\ncombination. And evolution has designed\nthese transcription factors to is that\nyour claim that they're to have the\nrelatively modular uh self-contained\neffects that work in predictable ways\nwith other transcription factors. And so\num we can use that same handle to uh to\nour own ends.\nYeah. Yeah. That that would be very much\nmy contention. And one piece of evidence\nfor this is that's the way the\ndevelopment works. You know, it's kind\nof a crazy thing to think about, but you\nand I were both just like a single cell\nand then we were a bag of undifferiated\ncells that were all exactly alike. And\nthen somehow we became humans with\nhundreds of different cell types all\ndoing very different things. And when\nyou look at how development specifies\nthose unique fates of cells, it is\nthrough groups of these transcription\nfactors that each identify a unique\ntype. And in many cases actually the\ngroups of transcription factors, the\nsets that specify very different fates\nare actually pretty similar to one\nanother. And so evolution has optimized\nfor being able to just swap one TF in or\nswap one TF out of a combination and get\npretty different effects. And so you\nhave this sort of like local change\nleading to local change in sequence or\nor gene set space leading to a pretty\nlarge global change in output. And then\nlikewise many of these TFS again are\nduplicated in the genome. And because\nmutations are going to be random and\nthey're inherently small changes at the\nlevel of sequence at a given time,\nevolution needs a substrate where in\norder to function effectively, these\nsmall changes can give you relatively\nlarge changes in phenotype. Otherwise,\nit would just take a very long time\nacross evolutionary history for enough\nmutations to accumulate in some\nduplicated copy of the gene for you to\nevolve a new TF that does something\ninteresting. And so, I think we're\nactually in most cases in biology due to\nthat evolutionary constraint. Small\nedits need to lead to meaningful\nphenotypic changes in a relatively\nfavorable regime for generic\ngradient-like optimizers. You know, it\nwould be uh maybe a little bit\noverstating to say evolution is like\nusing the gradient, but there is a\nsystem kind of like uh if you've heard\nof evolution strategies where uh\nbasically the way you optimize\nparameters is you can't take a gradient\non your loss. So you make a bunch of\ncopies of your parameters, you randomly\nmodify them and then you compute a\ngradient on your parameters against your\nloss and so you can take a gradient in\nthat space. That's kind of how I imagine\nevolution is working. And so you need\nlots of those little edits to actually\nlead you in to have meaningful step\nsizes in terms of the ultimate output\nthat you have.\nInteresting. uh you're just like\ndesigning the Laura that goes on top of\nuh\nyeah yeah in a way and to think like you\nknow why would transcription factors\nmaybe this is getting a little bit too\ngigabrained about it but like why is the\ngenome even have transcription factors\nlike what's the point why not just have\nevery time you want a new cell type you\nlike engineer some new cassette of genes\nor some new totally denovo set of\npromoters or something like this I think\none possible explanation for their\nexistence rather than just an\nappreciation for for their for their\npresence is that while having\ntranscription factors allows a very\nsmall number of base pair edits at the\nsubstrate of the genome to lead to very\nlarge phenotypic differences. If I break\na transcription factor, I can delete a\nwhole cell type in the body. If I\nretarget a transcription factor to\ndifferent genes, I can dramatically\nchange when cells respond and have, you\nknow, hundreds of their downstream\naector genes change their behavior in\nresponse to the environment. And so it\nputs you in this regime where\ntranscription factors are a really nice\nsubstrate to manipulate as targets for\nmedicines. In some ways, they might be\nlike evolution's levers upon the broader\narchitecture of the genome. And so, by\npulling on those same levers that\nevolution has gifted us, there are\nprobably many useful things we can\nengender upon biology.\nYeah, you're you're sort of hinting that\nuh if we analogize it to some code base,\nwe're going to find a couple lines that\nare like commented out that's like\ndaging, you know, and then like unhyen\non parenthesis. And\nI don't know about that, but if I can\ngive you can I'll give you like a real\ncringe analogy that sometimes I deploy,\nbut it requires a very special audience.\nI think you'll probably be the one who\nfits into it.\nYou're you're flattering our uh\nlisteners.\nOnly cringe listeners will appreciate\nit, but your audience will love this.\nI don't know about your audience, but\nbut you will. Um you can kind of think\nabout it like, you know, if you think\nabout how attention works, like queries,\nkeys, values.\nTFs are kind of like the queries, the\ngenome sequences they bind to are kind\nof like the keys. Genes are kind of like\nthe values. And it turns out that\nstructure then allows you to very\nefficiently in terms of editing space.\nyou can change just one of those\nembedding vectors in this case one of\nthose sequences and get dramatically\ndifferent performances or title total\noutputs and so I do think it's kind of\ninteresting how these structures recur\nthroughout biology you know in the same\nway that the attention mechanism seems\nto exist in some neural structures I\nthink it's kind of interesting that you\ncan very easily see how that same sort\nof querying and information storage\nmight exist in the\ninteresting yeah a previous guest and a\nmutual friend Trenton Burkin has a had\ndid a paper in grad school about how uh\nthe brain implements\nattention.\nYeah. Or Eddie Chang has found like\npositional encodings probably exist in\nhumans using neuropixels. If you haven't\nread these papers, oh yeah, he so he\nimplants these neuropixel probes into\nindividuals and then he's able to, you\nknow, talk to them, look at them as they\nread sentences. And what he finds is\nthat there seem to be certain\nrepresentations which function as a\npositional encoding across sentences. So\nthey fire at a certain frequency and it\njust increases as the the sentence goes\non and then like resets. And so it seems\nexactly like what what we do when we\ntrain large language models where you've\ngot some sign\nfunction. The way we're going to learn\nhow the brain works is just like trying\nto first principles engineer\nintelligence and AI and it just like\nhappens to be the case that each one of\nthese things has a neural correlate.\nGemini CLI just oneshotted an automated\nproducer for me in one hour. Basically,\nI wanted this interface where I could\njust paste in a raw episode transcript\nand then get suggestions for Twitter\nclips and titles and descriptions and\nsome other copy all of which\ncumulatively takes me about half a day\nto write. Honestly, it was just\nextremely good. I described the app I\nwanted and then asked Gemini to talk\nthrough how I would go about\nimplementing it. It walked through its\nplans. It asked me for input where I\nhadn't been sufficiently clear. And\nafter we ironed out all the details,\nGemini just literally oneshotted the\nfull working application with fully\nfunctional backend logic. Making this\napp literally took 10 minutes, including\ninstalling CLI. Then I spent 15 minutes\nfine-tuning the UI, messing around. And\nby the way, this process did not involve\nme actually editing or even looking at\nany of the code. I would just tell\nGemini how I wanted things moved around\nand the whole UI would change as Gemini\nrewrote the files. Despite building and\nthen fine-tuning an entire working\napplication, the session context didn't\neven get 10% exhausted. This is just a\nsuper easy and fast way to turn your\nideas into useful applications. You can\ncheck out Gemini CLI on GitHub to get\nstarted. All right, back to Jacob. If\nyou're right that transcription factors\nare the modality evolution has used to\nhave complex phenotypic effects optimize\nfor different things. Two-part question.\nOne, why haven't pathogens which have a\nstrong interest in having complex\nphenotypic effects on your body also\nutilized the um transcription factors as\nthe way to you over and steal your\nresources? And two, we've been trying to\ndesign drugs for centuries. Why aren't\nall the big drugs, the top selling\ndrugs, ones that just um modulate\ntranscription factors?\nYeah. Yeah. Why don't we have a million\nof these pills? Okay, I'll try and take\nthose in stride. And they're pretty\ndifferent answers. First answer is there\nactually are pathogens that that utilize\ntranscription factors as part of their\nlife cycle. So like a famous example of\nthis is HIV. HIV encodes a protein\ncalled TAT. And TAT actually activates\nNFCAPPA B. And so HIV, sorry to back up\na little bit, is a retrovirus. Starts\nout as RNA, turns itself into DNA,\nshoves itself into the genome of your\nCD4CT cells. And so then it needs this\nornate machinery to actually control\nwhen does it make more HIV and when does\nit go latent so it can hide and your\nimmune system can't clear it out. And\nthis is why HIV is so pernicious is you\ncan kill every single cell in the body\nthat's actively making HIV with like a\nreally good drug, but then a few of them\nthat have like lingered and hunkered\ndown just turn back on. And so people\ncall this the latent reservoir. Same\nwith HEP B, right?\nWell, HEP B and HEPS can both do this\nsort of like latent sort of behavior.\nUm, and so HIV is probably the most\npernicious of these. And one way it does\nit is this gene called TAT actually\ninteracts with NFCappa B. NFCappa B is a\nmaster transcription factor within\nimmune cells. Typically, if I'm going to\nlike horribly reduce what it does, and\nsome immunologists can can crucify me\nlater, it like increases the\ninflammatory response of most cells.\nThey become more likely to attack given\npathogens around them on the margin. Um,\nand so it'll turn on an FCAPA B activity\nand then uses that to drive its its own\ntranscription and its own life cycle.\nAnd so it I can't remember quite all the\ndetails now exactly of how it works, but\npart of this circuitry is what allows it\nto in some subset of cells where some of\nthat upstream transcription factor\nmachinery in the host might be\ndeactivated, it goes latent. And so as\nlong as the population of cells it's\ninfecting always has a few that are like\nturning off the transcription factors\nupstream that drive its own\ntranscription, then HIV is able to\npersist in this latent reservoir within\nhuman cells. So that's just one example\noffhand. Then there are a number of\nother pathogens and and unfortunately I\ndon't have quite as much molecular\ndetail on some of these, but they will\ninterface with other parts of the cell\nthat eventually result in transcription\nfactor transllocation to the nucleus and\nthen transcription factors being active.\nThis actually segus a little bit to your\nsecond question on why aren't there more\nmedicines targeting TFS? In a way, I\nthink many of our medicines ultimately\ndownstream are leading to changes in TF\nactivity, but we haven't been able to\ndirectly target them due to their\nphysical location within cells. And so\nwe go several layers upstream. If you\nthink about how a cell works in sensing\nits environment, it has many receptors\non the surface. It has the ability to\nsense mechanical tension and things like\nthis. And ultimately most of what these\nsignaling pathways lead to is to tell\nthe cell use some different genes than\nyou're using right now. That's often\nwhat's occurring. And so that ultimately\nleads to transcription factors being\nsome of the final aectors in these\nsignaling cascades. So a lot of the\ndrugs we have that for instance inhibit\na particular cytoine that might bind a\nreceptor or they block that receptor\ndirectly or maybe they hit a certain\nsignaling pathway. Ultimately, the way\nthat they're exerting their effect is\nthen downstream of that signaling\npathway, some transcription factor is\neither being turned on or not turned on,\nand you're using different genes in the\ncell. And so, we're kind of taking these\nlike crazy bank shots because we can't\nhit the TFS directly. So, that sort of\nbegs the question like why can't you\njust go after the TF directly?\nTraditionally, we use what are called\nsmall molecule drugs where they're\ndefined just by their size. The reason\nthey have to be small is they need to be\nsmall enough to wiggle through the\nmembrane of a cell and get inside. And\nthen you run into a challenge, which is\nif you want to actually stick a small\nmolecule between two proteins that have\na pretty big interface, meaning like\nthey've got big swaths on the side of\nthem that all, you know, sort of line up\nand and uh form a synapse with one\nanother, then you would need a big\nmolecule in order to inhibit that. And\nit turns out that TF's binding DNA is a\npretty darn big surface. And so small\nmolecules aren't great at disrupting\nthat and certainly even worse at\nactivating it. So small molecules can\nget all the way into the nucleus, but\nthey can't do much once they're there.\nThey're just too small. And then the\nother classic modalities we have are\nrecombinant proteins. We make a protein\nlike a hormone in a big vat. We grow it\nin some Chinese hamster ovary cells. We\nextract it. We inject it into you. This\nis how for instance like human insulin\nworks that we make today. Or you make\nantibodies. Antibodies produced by the\nimmune system. These run around and find\nproteins that have a particular\nsequence. They bind to it and often they\njust like stop it from working by\nglomming a big thing onto the side. So\nthose are too big to get through the\ncell membrane. So then they can't\nactually get to a TF or do anything\ndirectly. So we take these bank shots.\nSo what changes that today and why I\nthink it's pretty exciting is we now\nhave new nucleic acid and genetic\nmedicines where you can for instance\ndeliver RNAs to a cell that can get\nthrough using tricks like lipid\nnanoparticles. You wrap them in a fat\nbubble looks kind of like a cell\nmembrane. It can fuse with a cell put\nthe mRNAs in the cytool you can make a\ncopy of a transcription factor there and\nthen it transllocates the nucleus the\nsame way a natural one would and exerts\nits effect. And likewise, there are\nother ways to do this using things like\nviral vectors. But I think we've only\nvery recently actually gotten the tools\nwe need to start addressing\ntranscription factors as first class\ntargets rather than treating them as\nlike uh maybe some ancillary third order\nthing that's going to happen.\nInteresting. So the drugs we have can't\ntarget them, but your claim is that a\nlot of drugs actually do work by binding\nto the things we actually can target and\nthose having some effect on\ntranscription factors. So this brings us\nto questions about delivery which is the\nnext thing I want to ask you. You\nmentioned lipid nanoparticles. This is\nwhat the co vaccines were made of. The\nultimate question if we're going to work\non deaging is how do we make every\nsingle cell in the body? Even if you\nidentify what is the right transcription\nfactor to deage a cell and even if\nthey're shared across cell types or you\nfigure out the right one for every\nsingle cell type, how do you get it to\nevery single cell in the body? Um, yeah,\nhow do we how do\nhow do you deliver stuff? How do you get\nthem in there? So, I think there are\nmany ways one could imagine solving it.\nI'll sort of like narrow the scope of\nthe problem to saying I think delivering\nnucleic acid is a pretty good first\norder primitive. Ultimately, the\ngenome's nucleic acids. The RNAs that\ncome out of it are nucleic acids. So, if\nyou can get nucleic acid into a cell,\nyou can drug pretty much anything in the\ngenome effectively. So, you can reduce\nthis problem to asking how do I get\nnucleic acids wherever I want them to\nany cell type very specifically. So\ntoday there are two main modalities that\npeople use, both of which have some\ndownsides. The first one that we've\ntouched on already is lipid\nnanoparticles. These are basically fat\nbubbles and by default they get taken up\nby tissues which take up fat like the\nliver. Um and they can be used sort of\nlike Trojan horses. They can release\nsome arbitrary nucleic acid usually RNA\nmaybe encoding your favorite genes in\nour case transcription factors into the\ncell types of interest. You can play\nwith the fats and you can also tie stuff\nonto the outside of the fat like you can\nattach a part of an antibbody for\nexample to make it go to different cell\ntypes in the body and I think the field\nis making a lot of progress on being\nable to target various different cell\ntypes with lipid nanop particles. So\neven if nothing else worked for the next\nseveral decades I think companies like\nours would have more than enough\nproblems to solve and with the cells\nthat we can actually target. Another\nprominent way people go after this is\nusing viral vectors. The basic idea\nbeing viruses had a lot of evolutionary\nhistory and very large population sizes.\nThey've evolved to get into our cells.\nMaybe we can learn something from them.\nEven better Trojan horses. So, one type\nof virus people use a lot, it's called\nan AAV. Those AAVs um carry DNA genomes\nand so you can get genes, whole genes\ninto cells that they've got some\npackaging sizes. You can think of it\nkind of like a very small delivery\ntruck, so you can't put everything you\nwant into it. They can go to certain\ncell types as well. And then on top of\njust where do you actually get the\nnucleic acid to begin with, you can\nengineer the sequences a bit. And that\nbasically allows you to add like a a\nnotgate on it. You can make it turn off\nthe nucleic acid in certain cell types,\nbut you're never going to use the\nsequence engineering to get nucleic acid\ninto cells where it didn't get delivered\nin the first place. So you can sort of\nstart broad with your delivery vector\nand then use sequence to narrow down to\nmake it more specific, but not the other\nway around.\nSo I think both of those methods are\nsuper promising. Again, if nothing else\nemerged for decades, we'd still have\ntons and tons of problems as a\ntherapeutic development community to\nsolve even using just those. I do think\nI have one sort of very controversial\nopinion which, you know, people can\nroast me for later.\nYou have just one.\nYou're trying to solve aging. I think\nyou have only one.\nI have many controversial opinions. One\nof them is that I think both of these\nprobably in the limit will not be the\nway that we're delivering medicines in\nthe year 2100.\nUm, if you think about viral vectors, no\nmatter what, there always going to be\nsome amount some amount of immunogenic.\num you're always going to have your\nimmune system trying to fight them off.\nYou can play tricks, you can try and\ncloak them, etc., etc., but they're\nalways going to have some toxicity risk.\nThey also don't go everywhere. It's not\nthat we have examples of like a single\nviral species that infects every cell\ntype in the body and we just need to\nengineer it to make it safe. It's we\nwould have to also engineer the virus to\ngo to new cell types. So, there's some\nlimitations there. LMPS likewise have\nsome problems. They can go to tons of\ncell types. That's what largely we're\nworking on. We're super excited about\nit. But, there are some physical\nconstraints. They just have a certain\nsize and they have to get from your\nbloodstream out of your bloodstream\ntoward a given target cell and they have\nto not fuse into any of the other cells\nalong the way. So there's a whole gamut\nthey have to run. Ultimately, I think\nwe're probably going to have to solve\ndelivery the way that our own genome\nsolved delivery. So we have this same\nproblem that arose during evolution,\nwhich is how do I patrol the body, find\narbitrary signals in the environment,\nand then deliver some important cargo\nthere when some set of events happens.\nhow do I you know find a specific place\nand only near those cell types release\nmy cargo and really the the problem was\nsolved by the immune system. So we have\ncell types in our body T- cells and B\ncells which are effectively engineered\nby evolution to run around invaginate\nwhatever tissues they need to. They can\nclimb almost anywhere in the bodies.\nThere's nowhere they can't get access\nalmost and then once they sense a\nparticular set of signals and they've\ngot a very ornate circuitry to do this.\nThey run basically an ANDgate logic.\nThey can release a specified payload.\nAnd right now, the way our genome sets\nthem up, the payload they release is\nlargely either uh enzymes that will kill\nsome cell that they're targeting or kill\nsome pathogen or some signal flares that\ncall in other parts of the immune system\nto do the same thing. So, that's super\ncool. But you can think about it as a\nmodular system that evolution's already\ngifted us. We've got some signal and\nenvironmental recognition systems. So,\nwe can find particular areas of the body\nthat we want to find and then some sort\nof payload delivery system. I can\ndeliver some arbitrary set of things.\nAnd I imagine if we were to like rip von\nWinkle ourselves into 2100 and wake up,\nthe way we will be delivering these\nnucleic acid payloads is actually by\nengineering cells to do it to perform\nthis very ornate function. Those cells\nmight actually live with you. You\nprobably will get engrafted with them\nand they might persist with you for many\nyears. They deliver the medicine only\nwhen the environment within your body\nactually dictates that you need it. And\nso you'll actually won't be seeing a\nphysician every time this medicine is\nactive. Rather, you'll have a more\nornate responsive circuit. The other\nexciting thing about cells is that\nthey're big and they have big genomes.\nAnd so you actually have a large pallet\nto encode complex infrastructure and\ncomplex circuitry. So you don't need to\nlimit yourself to like the very small\nRNAs you can get in that might encode a\ngene or two or in our case a few\ntranscription factors. You don't have to\nlimit yourself to this tiny AAV genome\nthat's only a few kilobases. You've got\nbillions of base pairs to play with in\nterms of encoding all your logic. So I\nthink that's ultimately how delivery\nwill get solved. We've got many many\nstepping stones along the way. But if I\ncould like clone myself and work on an\neven riskier endeavor, that's probably\nwhat I would do.\nThis is actually in in I mean in a way\nwe treat cancer this way with CARTT\ntherapy, right? We take the tea cells\nout and then we tell them go find a\ncancer with this receptor and kill it.\nBut is the reason that works is that the\ncancer cells we're trying to target are\nalso free floating in the blood. And is\nthat what it targets? Basically, um\ncould this deliver to literally every\nsingle cell in the body? Not literally\nevery single cell. I'll like asterisk it\nthere. So, example, tea cells don't go\ninto your brain. You don't have they\ncan, but it's generally a pathology when\nthey get in there. So, it's not like\nliterally every cell, but almost every\ncell in your body is surveiled by the\nimmune system. So, there are very very\nfew what we call immune privileged\ncompartments in your body. It's things\nlike the joints of your knees and your\nshoulders, your eyeball, and your brain\nbasically. Um, there might be a couple\nothers. I think the ear probably falls\ninto that category. A funny way of\nthinking about this is all the gene\ntherapy people using viruses they want\nto deliver to the immune privilege\ncompartments because their their drugs\nare immunogenic and they're limited to a\nvery very small set of diseases. So in a\nway it's like the shadow of all the\ndiseases you can't address with viruses\nis what you can address with cells and\ngiven the complimentarity between them\nit's like okay you can probably cover\nthe entire body. Um, and so they can't\nliterally go everywhere. But I think\nyour analogy to to the the carti work is\nis very apt as well, where you can think\nabout of that two component system. I've\ngot some detection mechanism for the\nenvironment I want to sense to perform\nsome function and then I have some sort\nof payload that I deliver. Cartis\nengineer the first of those and leave\nthe second exactly the same as the\nimmune system does. So they engineer go\nrecognize this other antigen that you\nwouldn't usually target some protein on\nthe surface of a cell for instance and\nthen deliver the payload you would\nusually deliver if it was infected by a\nvirus or if you saw that it was foreign\nin some way whereas cancer cells usually\ndon't actually look that foreign. Most\nof their genes are the same genes that\nare in your normal genome and that's why\nit's hard for the immune system.\nInteresting. You know, it's funny that\nwhenever we're trying to cure infectious\ndiseases, we just have to deal with,\nviruses have been evolving for\nbillions of years with our, you know,\noldest common ancestor and they know\nexactly what they're doing and it's so\nhard. Um, and then whenever we're trying\nto do something else, we're like,\nthe immune system has been evolving for\nbillions of years and it knows what it's\ndoing and how do we get past it?\nYeah.\nYeah. The red the red queen race is like\nquite sophisticated and if you want to\njust like throw a new tool into biology,\nyou somehow have to get around one side\nof that equation,\nright? given the fact that it's um some\nmixture somewhere between impossible and\nvery far away uh and it's necessary for\nfull curing of aging. Um does that mean\nthat in the short run in the next few\ndecades we'll have some parts of our\nbody which will have these amazing\ntherapies and then other parts which\nwill just be stuck the way they are. So\nyou mentioned hpatocytes are some of the\ncells that you're able to uh actually\nstudy and or deliver to and these are\nliver cells. So you're saying look I can\nget drunk as much as I want and it's not\ngoing to have uh an impact on my longr\nrun liver health because then you'll\njust inject me with this therapy but for\nthe rest of my body it's going to age as\nnormal. Um what is the implication of\nthe fact that the delivery seems to be\nlagging much behind your understanding\nat some point your understanding of\naging? Yeah, ju just to give give the\ndelivery folks credit, they're currently\nahead. There are currently no\nreprogramming medicines for aging and\nthere are medicines that deliver nucleic\nacids. So like they're still winning the\nrace against us right now, but but to\nyour point, I hope the lines cross. I\nhope we over out compete them.\nUm, so I do think actually even if you\nwere able to only target some subsets of\ncells, it's not that you would see like\nthis strange Frankensteinian benefit in\nhealth in some aspects and and lack of\nbenefit entirely in others. I think what\nwe found across the history of medicine\nis that actually the body's an\nincredibly interconnected complex\nsystem. And if you're able to rescue\nfunction even in one cell type in one\ntissue, you often have knock-on benefits\nin many places that you didn't initially\nanticipate. One way we can get examples\nof this is is through transplant\nexperiments. So both in bone marrow and\nin liver, for example, we have fairly\ncommon transplant procedures that occur\nin humans. And so we can compare old\nhumans who get livers from young people\nor old people and in a way ask a pretty\ncontrolled question. What occurs as a\nfunction of just having a young liver?\nIs it that for example you can eat a lot\nof fatty food and drink a lot and be\nfine or is it that actually you see\nbroader benefits? And the latter seems\nto be true. They have reduced risk of\nseveral other diseases and overall\nbetter survival as a function of having\na younger liver than they do for an\nolder one. suggesting that actually\nbecause these tissues are so\ninterconnected many of these organs like\nthe liver like your atapost tissue or\nendocrine organs they're also sending\nout signals to many other places in your\nbody helping coordinate your health\nacross multiple tissue systems even just\none tissue can benefit other feature\nother tissue systems in your body at the\nsame time hs are another example where\nthere are many circumstances where one\nof the and I'll summarize this is uh\nmostly examples taken from a wonderful\nbook by Frederick Applebomb who uh\ntrained with Don Thomas the physician\nwho invented human bone marrow\ntransplants. Um there are many\ncircumstances where patients got a bone\nmarrow transplant and actually cured\nanother disease they had as a result\nmaybe unanticipated where it's even just\nthe replacement of this one special cell\ntype HSC's has knock-on effects\nthroughout the body. You know there were\nsymptoms of these diseases that\npresented in myriad ways throughout\ntheir system but ultimately its root\ncause was even just a single cell. Um\nthere are counter examples as well where\nyou can go into animals and break even\njust one gene in one specific subset of\nT- cells. You can break a gene in their\nthat encodes for a transcription factor\nin their mitochondria called Trem and\nyou actually dramatically shorten the\nlifespan of mice. One gene in one\nspecial type of sea cells can give you\nthat type of pathology. And so the it\nsort of implies the inverse may also\nexist.\nIs this related to why has so many\ndownstream positive effects that seem\neven not totally related to its um\neffects just on making you leaner? Yeah,\nI think it's one one example because\nit's a it is a hormone and your\nendocrine system coordinates a lot of\nthe complex interplay between your\ntissues. I don't think the story is\nfully written yet on exactly why GLP-1\nand GIP one, you know, broadly\nincredinatic medicines like Ompic have\nso many knock-on benefits, but I think\nthey're a great example of this\nphenomenon. If someone told you, I'm\ngoing to find a single molecule and I'm\ngoing to drug it and it's not only going\nto have benefits for weight loss, but\nalso for cardiovascular disease, also\npossibly for addictive behavior and\nmaybe even preventing ner degeneration.\nYou would have told them they were\ncrazy. And yet, just by acting on the\nsmall number of cells in your body which\nare receiving this signal, the interplay\nand the communication between those\ncells and the rest of your body seems to\nhave many of these knock-on benefits.\nSo, it's just one existence proof. Very\nsmall numbers of cells in your body can\nhave health benefits everywhere. Um, and\nso even if cellular delivery does not\nemerge by 2100, as I imagine it will,\nthen I still think that you're going to\nhave the ability to add decades of\nhealthy life to individuals by\nreprogramming the age of individual cell\ntypes and individual tissues.\nInteresting. How big will the payload\nhave to be?\nUh, how many transcription factors?\nYeah,\nI think just a countable number. I think\nsome of those that we found today that\nhave efficacy are, you know, somewhere\nbetween one and five and that that's a\nsmall enough number that you can\nencapsulate it in current mRNA\nmedicines. Um so already in the clinic\ntoday there are medicines that deliver\nmany different genes as RNA. Um so there\nare medicines where for instance it's a\nvaccine as a combination of flu and\nCOVID proteins and they're delivering 20\ndifferent unique transcripts all at the\nsame time. And so when you think about\nthat already as a medicine that's being\ninjected into people in trials, the idea\nof delivering just a few transcription\nfactors is seemingly quotient. And so\nthankfully I don't think we'll be\nlimited by the size of the payloads that\none can deliver. One other really cool\nthing about transcription factors is\nthat the indogenous biology is very\nfavorable for drug development. The\nexpression level of transcription\nfactors in your genome relative to other\ngenes is incredibly low. So if you just\nlook at like the rankordered list of\nwhat are the most frequently expressed\ngenes in the genome by the count of how\nmany mRNAs are in the cell,\ntranscription factors are like near the\nbottom. And that means you don't\nactually need to get that many copies of\na transcription factor into a cell in\norder to have benefits. And so what\nwe've seen so far and what I imagine\nwill continue to play out is that even\nfairly low doses of these medicines,\nwhich are well within the realm of of\nwhat folks have been taking for now more\nthan a decade, um are are able to induce\nreally strong efficacy. And so we're\nhopeful that not only will the actual\nsize of the payload in terms of like\nnumber of base pairs not be limiting,\nbut the dose shouldn't be limiting\neither.\nAnd is it would it have to be a chronic\ntreatment or could it just be a onetime\ndose?\nIn principle, it could be one time. I\nthink that would be an overstatement for\ntoday, but I can sort of talk you\nthrough the evidence from like the first\nprinciples back to the reality of like\nwhat's the hardest thing we have in\nhand.\nSo epigenetic reprogramming is basically\nhow the cell types in our bodies right\nnow are able to adopt the identities\nthat they have. And the existence proof\nthat those epigenetic reprogramming\nevents can last decades is that my\ntongue doesn't spontaneously turn into a\nkidney. Um so these epigenetic marks can\npersist for decades throughout a human\nlife or you know hundreds of years if\nyou want to take the example of a\nbowhead whale which uses the same the\nsame mechanism. Um and we also know that\nwith very targeted edits other groups\nhave done this folks like Luke Gilbert\nnow at the Arc Institute who I think of\nas like one of the the great unsung\nscientists of our time um have been able\nto make a targeted edit in a single\nlocus and then show that you can\nactually make cells divide 400 plus\ntimes over multiple years in an\nincubator in the lab. So imagine like a\nhot house where you're just trying as\nhard as you can to break this mark down\nand it can actually persist for many\nyears. Um other companies have actually\nnow dosed some editors similar to the\nones that that Luke developed in his lab\nin monkeys and shown they last at least\na couple years. So in principle the\nupper bound here is really long. You you\ncould potentially have one dose and it\nlasts a very long time. You know\npotentially decades as long as it took\nyou to age the first time maybe. We\ndon't have data like that today. So I\ndon't want to overstate. Um, we do have\ndata that these positive effects can\nlast several weeks after a dose. And so\nyou could imagine even without many\nleaps of faith up toward this upper\nbound limit of what's possible just from\nthe data we have in hand now that you\ncould get doses, you know, every month,\nevery few months, and actually have\nreally dramatic benefits that persist\nover time rather than needing, for\ninstance, to get an IV every day, which\nmight not be tractable. So, we've got\n1600 transcription factors in the human\ngenome. Is it worth looking at nonhuman\nTFs and seeing what effects they might\nhave or are they unlikely to be the\nright search base?\nI think it's less likely. I think you\nhave a prior that evolution has given\nyou a reasonable basis set for\nnavigating the states that human cells\nmight want to occupy. And in our case,\nwe know that the state we're trying to\naccess is encoded by some combination of\nthese TFS. It it does arise in\ndevelopment. Obviously, we're trying to\nmake an old cell look young, not look\nlike some Frankenstein cell that's never\nbeen seen before.\nThat said, we don't have any guarantees\nthat the way aging progresses is by\nfollowing the same basis set of these\ntranscription factor programs in the\ngenome that are encoded during\ndevelopment. So, I don't think it's\nunreasonable to ask, would your eventual\nideal reprogramming medicine necessarily\nbe a composition of the natural TFS or\nwould it encode include something like\nTFS from other organisms as you posit or\neven entirely synthetic transcription\nfactors as well? things like supers\nsocks. Supers socks is a particular\npublication from Sergey I mispronounced\nhis last name Vichenko where they\nmutated the socks tube gene and they\nmade more efficient iPSC-c\nreprogramming. So they could take\nsematic cells and turn them into pur\npotentotent stem cells more effectively\nthan you could with just the the\ncanonical yamaka factors which are oct4\nand mick. IPS-C reprogramming never\nhappens in nature. So there's no reason\nto necessarily believe that the natural\nTFS are optimal. And so even really\nsimple optimizations like just\nmutagenizing one of the four Yamanaka\nfactors we already know about or\nswapping some domains between a few TFs\nseem to improve things dramatically. So\nI think that's a pretty good signal that\nactually there's a lot of gradient to\nclimb here and the potentially for us\nthe endstate products we're developing\nin 2100 are more like synthetic genes\nthat have never existed rather than just\ncompositions of the natural set. What\nabout the effects of aging which are um\nokay so I don't know your skin starts to\nsag because of the effects of gravity\nover the course of decades. Is that a\ncellular process? How would how would\nsome cellular therapy deal with that?\nThe best evidence is that it's probably\nnot cellular. So the reason your skin\nsags is there's a protein in your skin\ncalled elastin which does exactly what\nyou'd think it would based on the name.\nIt kind of keeps your skin elasticy like\na waistband and holds it to your face.\nSo you have these big polymer\npolymerized fibers of elastin in your\nface. And as far as we understand it,\nyou only polymerize it and form a long\nfiber during development. And then the\nrest of your life, you make the\nindividual units of the polymer, but\nthey for reasons no one, as far as I can\ntell, understands, they they fail to\npolymerize. And you can't like make new\nlong cords to hold your skin up to your\nface. So I think the eventual solution\nfor something like that is likely that\nyou need to program cells to states that\nare extra physiological. There might not\nbe a cell in your body. It's not just\nlike a young skin cell from a\n20-year-old is better at making these\nfibers. As far as we can tell, they\ndon't. Um, but you could probably\nprogram a cell to be able to\nreinvigorate that polymerization process\nto run along the fiber and repair it in\nplaces where it's damaged. Obviously,\nthese things get made during\ndevelopment. So, it's totally physically\nfeasible for this to occur. Maybe\nthere's even a developmental state which\nwould be sufficient to achieve this. I\ndon't think anyone knows, but that would\nbe the kind of state that one might have\nto engineer denovo even if our genome\ndoesn't necessarily encode for it\nexplicitly.\nInteresting. Okay. What is Arum's law?\nIram's law is a funny portman who\ncreated by a friend of mine Jack Scanell\nwhere he inverted the notion of Moor's\nlaw which is the doubling of compute\ndensity on uh silicon chips every few\nyears. So Moore's law has graciously\ngiven us massive increases in compute\nperformance over several decades. And\nEim's law is the inverse of that because\nin bioarma what we're actually seeing is\nthat there's a very consistent decrease\nin the number of new molecular entities.\nSo new medicines that were able to\ninvent per billion dollars invested and\nthis trend actually starts way back in\nthe 1950s and persists through many\ndifferent technological transitions\nalong the way. So it seems to be an\nincredibly consistent feature of trying\nto make new medicines. Hm. So, um, in a\nweird way, Aram's law is actually very\nsimilar to the scaling laws you have in\nML where you have this very consistent\nlogarithmic relationship of you throw in\nmore inputs and you get consistently\ndiminishing outputs. Um, the difference\nof course is that this trend in ML has\nbeen um used to raise exponentially more\ninvestment um and to drive more hype\ntowards AI. Whereas in biotech um you\nknow modular new limits new round it it\nhas driven down valuations driven down\nexcitement and energy with AI at least\nyou can sort of internalize the extra\ncost and the extra benefits because\nthere's this general purpose model\nyou're training so this year you spent\n$100 million training a model next year\na billion dollars the year after that 10\nbillion but it's one general purpose\nmodel unlike we made money on this drug\nand now we're going to use that money to\ninvest in 10 different drugs in 10\ndifferent bespoke ways. Okay. Anyways, I\nwas gearing up to ask you what would a\ngeneral purpose platform where even if\nyou had diminishing returns, at least\nyou can have this sort of like less\nbespoke way of designing drugs look like\nfor biotech.\nOkay, I'm going to slightly dodge your\nquestion first to maybe analyze\nsomething really interesting that you\nhighlighted which is you have these two\nphenomena again ML scaling and then\nscaling in terms of the cost for new\ndrug discovery. Why is it that the\npatterns of investment have been so\ndifferent? I think there are probably\ntwo key features that might explain this\ndifference. one is that the returns to\nthe scaled output in the case of ML\nactually are expected to increase super\nexponentially if you actually reach AGI\nit's going to be a much larger value\nthan just even a few logs back on the\nperformance curve that that people are\nfollowing whereas in the life sciences\nthus far each of those products we're\ngenerating further and further out on\nthe e- room slot curve as time moves\nforward haven't necessarily scaled in\ntheir potential revenue and their\npotential returns quite so much and so\nyou're seeing these increased costs not\ncounterbalanced by increased ROI The\nother piece of it that you highlighted\nis that unlike building a general model\nwhere potentially by making larger\ninvestments, you're going to be able to\nsolve a broader addressable market,\nmoving from solving very narrow tasks to\neventually replacing large fractions of\nwhite collar intelligence.\nIn biotech, when you're traditionally\nable to develop a medicine in a given\nindication, I was able to treat disease\nX. It doesn't necessarily engender you\nto be able to then treat disease Y more\nreadily. Typically where these firms,\nbiotech firms in general have been able\nto develop unique expertise is on making\nmolecules to target particular genes. So\nI'm really good at making a molecule\nthat intervenes on gene X or gene Y. And\nit turns out that the ability to make\nthose molecules more rapidly isn't\nactually reducing the largest risk in\nthe process. And so this means that the\nability to go from one or two outputs\none year to then going to four the next\nis much more limited. And so this brings\nus then to the question of what would\nthe general model be in biology? And I\nthink it kind of reduces down to how do\nyou actually imbue those two properties\nthat create the ML scaling law curve of\nhope and you know bring those over to\nbiology so that you can take the rooms\nlaw curve and potentially give it the\nsame sort of of potential beneficial\nspin. So I think there are a few\ndifferent versions of this you could\nimagine but I'll address the first\npoint. How do you get to a place where\nyou're actually able to generate more\nrevenue per medicine so that potentially\nthe outputs you're generating are more\nvaluable even if each output might cost\na bit more. Traditionally, when we've\ndeveloped medicines, we go after fairly\nnarrow indications, meaning diseases\nthat fairly small numbers of of people\nget. And that's actually increased in\nterms of the narrow scope of what\nmedicines are addressing as we've gone\nforward in time. And so, this is sort of\nan ironic situation where we've gone\nfrom addressing pretty broad categories\nof disease like infectious disease to\nnarrower and narrower genetically\ndefined diseases that have small patient\npopulations because these only affect a\nfew people. If you think about the value\nfunction of a medicine is, you know, how\nmany years of healthy life does it give\nhow many people? If how many people is\npretty small, it just really bounds the\namount of value you're able to generate.\nSo, you need to then be able to find\nmedicines that treat most people. All of\nus will one day get sick and die. So,\narguably the TAM for any really\nsuccessful medicine could be everybody\non planet Earth. Um, so we need to find\na way to be able to route toward\nmedicines that address these very large\npopulations. The second piece then is\nhow do we actually build models that\nenable us to take the success in one\nmedicine we've developed and lead that\nto an increased probability of success\non the next medicine. Traditionally, we\nhaven't been able to do that. Maybe\nyou're better at making an antibbody for\ngene Y because you made one for gene X 5\nyears ago, but it turns out making an\nantibbody isn't really the hard part of\ndrug discovery. Figuring out what to\nmake an antibbody to target is the hard\nthing about drug discovery. What gene do\nI intervene upon in order to actually\ntreat a disease in a given patient? Most\nof the time we just don't know. And so\nthat's why even if a given drug firm\nbecomes very good at making antibodies\nto gene X, they have a successful\napproval. When they then go to treat\ndisease Y, they don't necessarily know\nwhat gene to go after. And most of the\nrisk is not in how do I make an\nantibbody to treat my particular target.\nIt's in figuring out what to target in\nthe first place.\nI'm not sure how to understand uh this\nclaim that we don't, you know, we know\nhow to engage with the right hook. We\njust don't know what that hook is\nsupposed to do in the body. I don't know\nif that's the way you'd describe it.\nWith another claim that I've seen that,\nyou know, with small molecules, we have\nthis Goldilocks problem where they had\nto be small enough to uh percolate\nthrough the body and through cell walls\netc., but big enough to interfere with\nlike protein for interactions that\ntranscription factors might have or\nsomething. So there it seems like\ngetting the hook is the big problem.\nYeah. In this particular case, if we we\nbound ourselves to we must use small\nmolecules as our modality, then there\nare lots of targets which are very\ndifficult to drug. There are many other\nmodalities by which you can drug some of\nthese genes. And I would say I don't\nhave a formal way of explaining this.\nthat if you were to write out a list of\nwell-known targets that many many folks\nwould agree are the correct genes to go\nafter and to try and inhibit or activate\nin order to treat a given set of\ndiseases and the only reason we don't\nhave medicines is that we can't figure\nout a trick in order to be able to drug\nthem. It's a fairly small list. It would\nprobably fit on a single page. Whereas\nthe number of possible indications that\none could go after and the number of\npossible genes that one could intervene\nupon, especially when you consider their\ncombinations, is astronomical. I think,\nyou know, the experiment you could run\nhere is if you lock 10 really smart drug\ndevelopers in a room and you tell them\nto write down some incredibly high\nconviction target disease pairs where\nthey're sure if they modulate this\nbiology, these patients are going to\nbenefit and all they need is some\nmolecular hook as as you put it in order\nto do this. It's a relatively short\nlist. What you're not going to get is\nanything approximating the paniply of\nhuman pathologies that develop. And you\ncan actually look for this. There are\nsome existence proofs you can look for\nout in the universe. Which is to say, if\nthe only problem was that we didn't have\nthe ability to drug something using\ncurrent therapeutics that we can put in\nhumans, we should still be able to treat\nit in the best animal models of that\ndisease because we can use things like\ntransgenic systems. You can go in and\nyou can engineer the genome of that\nanimal. And so this gives you all sorts\nof superpowers that you don't have in\npatience but allow you to for instance\nturn on arbitrarily complex groups of\ngenes in arbitrarily specific or broad\ngroups of cells in the organism at any\ntime you want at any dose you want in\nthe animal and for the majority of\npathologies we just don't have many of\nthose examples.\nOkay. So then what's the do you what is\nthe answer to what is the general\npurpose um\nthe general purpose model like how\nevery marginal discovery increases the\nodds you make the next discovery or\nsomething like that.\nSo there are multiple ways one might\napproach this problem. the most common\ntoday. Um, this is often what people are\ndescribing when they talk about a\nvirtual cell. This is sort of a a very\nnebulous idea, sometimes numinous if\nyou'll let me describe it in that way as\nwell. Um, but I think most concretely\nwhat most people are trying to do is\nmeasure some number of molecules or some\nsort of uh perceived emissions like the\nmorphology of a cell and then perturb it\nmany times, turn some genes on, turn\nsome genes off and measure how that\nmolecular morphological state changes.\nThe notion is that there's a lot of\nmutual information in biology. So if I\nmeasure something like most commonly all\nthe genes the cell is using at a given\nmoment which you can get by RNA\nsequencing that I get a decent enough\npicture of most of the other complexity\ngoing on. And so that I can for instance\ntake a bunch of healthy cells and a\nbunch of cells that are in a diseased or\nage state and I'm able then to compare\nthose profiles and say okay my disease\ncells use these genes, my healthy cells\nuse these. Are there anti-interventions\nI can find that I'm able to do\nexperimentally in the lab that shift one\ntoward the other? And then the hope\nwould be because you're not never going\nto be able to scan combinatorily all the\npossible groups of genes just to to make\nthat concrete there. I'm just going to\nbe be round with it, but there's\nsomething like 20,000 genes in the\ngenome. You can then choose, you know,\nhowever many genes in your combination\nyou want. It's not crazy to think of\nhundreds at a time. That's what\ntranscription factors control. That's\nhow development works. So the number of\npossible combinations is truly\nastronomical. You just can't test it\nall. So the hope would be that by doing\nsome sparse sampling of those pairs,\nyour inputs are here's what the cell\nlooked like beforehand. Here's the\nparticular genes I perturbed. You have\nsome measurement then of the state that\nthe cell resulted in. So here's which\ngenes went up. Here's what here's which\nwent down. And then you can start to ask\nonce I've trained a model to predict\nfrom the perturbations to the output on\nthe cell state. What would happen for\nsome arbitrary combinations of genes?\nAnd now in silic possible things that\none might do and potentially discover\ntargets that take my disease cells back\nto something like healthy cells. So\nthat's another version of what would a\nall-encompassing model look like where\nyou actually have compounding returns in\ndrug discovery.\nRight. And you basically described um\none of the models you guys are working\non at New Limit. You you're training\nthis model based on this data where\nyou're taking the entire transcriptto\nand just labeling it based on how old\nthat cell actually is. If you've got all\nthis data you're collecting on how\ndifferent perturbations are having\ndifferent phenotypic effects on a cell,\nwhy only record the uh like whether that\neffect correlates with more or less\naging? Why can't you also label it with\nall the other um effects that we might\neventually care about and eventually get\nthe full virtual cell? Because that's\nthe that's the more general purpose um\nmodel, right? that not just the one that\npredicts whether a cell looks old or\nnot.\nYeah, absolutely. So I think what we\nactually do both today. So we can train\nthese models where basically the inputs\nare a notion of what that cell looked\nlike at the starting place. Here's what\na generic old cell looked like and then\nrepresentations of the transcription\nfactors themselves. We derive those from\nprotein foundation models. They're\nlanguage models basically trained on\nprotein sequences. Turns out that gives\nyou a really good base level\nunderstanding of biology. So the model\nis kind of starting from a pretty smart\nplace. And then you can predict a number\nof different targets from some learned\nembedding. The same way you can have\nmultiple heads on a language model. And\nso one of those for us is actually just\npredicting every gene the cell is\nexpressing. Can I just recapitulate the\nentire state and guess what effect these\ntranscription factors will have on every\ngiven gene. And you can think about that\nas like an objective rather than a value\njudgment on the cell. I'm not asking\nwhether or not I want this particular\ntranscriptto. I'm just asking what it\nwill look like.\nAnd then we also have something more\nlike, you know, value judgments. I\nbelieve that that transcriptto looks\nlike a younger cell and I I'm going to\nselect on that and train a head to\npredict it where I can d noiseise across\ngenes and then select for younger cells.\nBut you could do that for arbitrary\nnumbers of additional heads. What are\nsome other states you might want? Do I\nwant to polarize tea cells to a less\ninflammatory state in somebody with an\nautoimmune disease? Do I want to make\nliver cells more functional in a patient\nwho is suffering from certain types of\nmetabolic syndrome? Be that, you know,\nmaybe even orthogonal to the way that\nthey age. Do I want to go in and change\nthe way a neuron is functioning to a\ndifferent state to treat a particular\ntype of neuro degenerative disease?\nThese are all questions you can ask.\nThey're not the ones we're going after,\nbut that that is the more general\nbroader vision.\nThis is so similar to in LLMs. You have\nfirst imitation learning with\npre-training that builds a general\npurpose uh representation of the world\nand then you do RL about a particular\nobjective in math or coding or whatever\nthat you care about and you are\ndescribing an extremely similar\nprocedure where first you just learn to\npredict perturbations in genes to uh\nbroad effects on the cell and that's\nlike that's the sort of pre-training\njust like learn how cells work\nand then there's another afterward layer\nof these like value judgments of okay\nwell how would how would we have to\nperturb it to have effect X which\nactually seems very similar to how do we\nget the base model to answer this math\nproblem or answer this coding problem. I\ndon't know. Um I don't know if people\nusually put it this way, but it actually\njust seems like an extremely extremely I\nmean that makes me more optimistic on\nthis because like LLMs work, right? And\nRL works.\nYeah, they do. Um yeah, I think the\nconceptual analogy is very apt. You\nknow, we don't actually use RL at the\nmoment, so I don't want to overstate the\nlevel of sophistication we've got, but I\nthink the general problem reduces down\nin a similar way. And so you can think\nabout, you know, your your earlier\nquestion of what does the general model\nlook like that enables you to actually\nhave compounding returns in drug\ndiscovery? Well, you might have\nsomething like this base model which as\nyou said just predicts this object\nfunction of how are these perturbations\nhitting these targets going to change\nwhich genes are turned on and off in\nthis cell.\nThen there's an entirely other task\nwhich is well which genes do you want to\nturn on and off and what state do I want\nthe cell to adopt? Our lens on that is\nthat across many different diseases\npeople will have age is one of the\nstrongest predictors of how they're\ngoing to progress whether that disease\narises. And so in many many\ncircumstances you have evidence in\nhumans where you can say ah if I could\nmake the cell younger maybe that's not a\nperfect fix but that's going to\ndramatically benefit not only patients\nwho have a diagnosed disease but it\nmight actually help most of us stay\nhealthier longer even subcl clinically\nbefore anyone would formally say that\nwe're sick. Now that's another more\ngeneral function. The same way that in\nLLMs you might have to create these\nparticular RLVF environments. You need\nto have places where you can, you know,\nstate a value function of the particular\ntask that you're trying to optimize for.\nIn drug discovery, you would then need\nto know well what are the cell states I\nwant to engineer for? That's kind of the\nnext generation of what a target might\nbe. Beyond just which genes do I want to\nmove up and down and which gene\nperturbations do I put in, you then need\nto know what cell state am I engineering\nfor? What do I want this to be?\nA bunch of labelers in Nigeria like\nclicking different pictures of cells\nlike, \"Oh, this one looks young. This\none looks old. This one looks really\ngreat.\" I love that one. Potentially.\nPotentially. It's more like\ndevelopmental biologists locked in a\nroom as my friend Cole Trapnel would\nsay.\nUm, it seems like what you're describing\nseems quite similar to perturb. And\nwe've had perturb I don't know when it\nwas done. Uh, what year was it?\nThere were three papers almost\nsimultaneously in 2016.\nOkay. So, almost a decade. Um, I don't\nknow. We're still waiting I guess for\nthe big breakthrough supposed to cause\nuh and this is the same procedure. So\nwhy why is this going to have an effect\nthat\nwhy has this taken so long? Yeah. Yeah.\nGood questions. Um so the original\nprocedure you know was created by a\nbunch of brilliant folks. There's a\ngroup in Idamit's lab at the Visman\ninstitute a vivv's lab at the broad\nwhere trade dixit a friend of mine\nhelped work on this and then Jonathan\nWeisman's lab at UCSF where Brit Adamson\ndid a lot of the early work. They all\nconstructed this this idea where you can\ngo in and you label a perturbation that\nyou're delivering to a cell. So this is\ntypically a transgenic perturbation,\nmeaning you're integrating some new gene\ninto the genome of a cell and that turns\nanother gene on or off. They used\ncrisper, but there's lots of ways to do\nit and the concept's pretty general. And\nthen you attach on that new trans gene,\nthat new gene you put into the genome of\nthe cell, some barcode that you can read\nout by DNA sequencing. So now when you\nrip the cells open, you're able to not\nonly measure every gene they're using,\nbut you also sequence these barcodes and\nyou know which genes you turned on and\nwhich are off. So you can then start to\nask questions like, well, I've turned on\ngenes A and C. What did it do to the\nrest of the cell? So that's the general\npremise of the technology. And so it's\nuseful to just set that up because it\nexplains why this didn't all happen\nearlier.\nOne, the actual readout ripping the\ncells open and sequencing them used to\nbe pretty bad and it used to be really\nexpensive and it's gotten much better\nover time. So the metric people often\nthink about here is like cost per cell\nto sequence. It used to be measured in\ndollars and now it's measured in cents\nand down to the fractions of cents\nbecause that cost curve has made it uh\nhas improved dramatically. The cost of\nsequencing has likewise come down. So\neven beyond the actual reagents\nnecessary to rip the cell open and turn\nits mRNAs into DNA that are ready for\nthe sequencer, now the sequencer is\ncheaper. The other piece is actually\ngetting these genes in and then figuring\nout which ones are there started out\npretty bad. So when we started with this\ntechnology, it was a beautiful proof of\nconcept, but I don't think anyone would\ntell you it was 100% ready for prime\ntime. When you sequence a cell, only\nabout 50% of the time could you even\ntell which perturbation you put in.\nSometimes you just like wouldn't detect\nthe barcode and you'd have to throw the\ncell away or you detect the wrong\nbarcode and now you've like mislabeled\nyour data point. So this might sound\nlike a trivial sort of technical piece,\nbut imagine you're running this\nexperiment the old fashioned way where\nyou test different groups of genes in\ndifferent test tubes on a bench. Now\nimagine you hired someone who every\nother tube labels it wrong. So when you\nthen collect data from your experiment,\nyou basically have no idea what happened\nbecause you just like randomized all\nyour data labels. You wouldn't do much\nscience and you wouldn't get very far\nthat way. So a lot of those technologies\nhave improved to the point where you had\na number of processes which are pretty\ninefficient and you multiplied a lot of\nthese things together and ended up with\nlike a very small outcome of successful\ncells you could actually sequence.\nThey've all improved to the degree where\nnow you can actually operate at scale.\nAnd then groups like ours have had to do\na bunch of work in order to actually\nenable combinatorial perturbations,\nturning on more than just one gene at a\ntime, which it turns out is much much\nharder for the same reason we're just\nalluding to. Imagine you're having\ntrouble figuring out which one gene you\nput in the cell and turned on or off.\nNow imagine you have to do that five\ntimes correctly in a row. Well, if you\nstart out with the original sort of\nperformance of like you could detect\nroughly 50% of them, then the fraction\nof cells that would be correctly labeled\nis like one over two to the n where n is\nthe number of genes you're trying to\ndetect. And very quickly it's like more\nof your data is mislabeled than\nlabelled. So there's lots of technical\nreasons like this that have gotten\nworked out over time. And so only now\nare we really able to scale up where\nwe're able to run experiments that are\nin the millions of cells in just a\nsingle day at for instance a small\ncompany like New Limit. There was a\npoint even just six or seven years ago\nwhere the companies that made these\nreagents were publishing the very first\nmillion cell data set just as a proof of\nconcept and only they could do it as the\nconstructors of the technology. And now\ntwo scientists in our labs can generate\nthat in an afternoon. If it actually is\nthe case that the um this is actually\nvery similar to the way LLM um dynamics\nwork then the once this technology is\nmature and you get the GPT3 equivalent\nof the virtual cell what you would\nexpect to happen is there's many\ndifferent companies that have um you\nknow are doing these cheap uh perturb\nlike experiments and building their own\nvirtual cells or at least a couple um\nand then they're like leasing this out\nto other people who then have their own\nideas about well we want to see if we\ncan come up with the labels for this\nparticular thing we care about and test\nfor that. What seems like happening\nright now is at least at New Limit you\nwere like we know the end ca end use\ncase we're going after it would be like\nif um cursor or whatever is like we're\ngonna in like 2018 is like we're going\nto build our own LLM um from scratch so\nthat we can enable our application\nrather than some foundation model\ncompany being like we don't care what\nyou use it for we're going to build this\num does that make sense like uh it seems\nlike you're combining two different\nlayers of the stack um and it's just\nbecause nobody else is doing\nthe other layer and so you're you're\njust doing both of them. I don't know\nwhat to extend this analogy maps on, but\nyeah. Yeah. Maybe to play with the\nanalogy a bit, imagine that, you know,\nyou think about New Limit as an LLM\ncompany. If if I'm going to put us in\nthe shoes of Cursor, which Oh, oh, so I\nwish. Um, imagine we're trying to in\n2018 create Cursor Tab, but we're not\ntrying to create a full LLM,\nright?\nI'm not I don't know enough about the\nunderlying mechanics to know if that\nwould have been feasible, but it's a\nmuch more feasible problem than trying\nto create like their most recent cursor\nagent or compete with like modern cloud\ncode, right? I think that's roughly the\nequivalent where the problem we're\nbreaking off is a subset of the more\ngeneral virtual cell problem. We're\ntrying to predict what do groups of\ntranscription factors do to the age of\nvery specific types of cells. We only\nwork on a few cell types at New because\nthose are the only cell types or some of\nthe only cell types today we believe we\ncan get really effective delivery of\nmedicines.\nAnd so we think they're just more\nimportant because we can act on them\ntoday. If we solve the problem of what\nTFS to use, we can make a medicine\npretty quickly. So in a way we're\ncarving out a region of this massive\nparameter space and saying if we can\nlearn the distribution of effects even\njust in this small region it's going to\nbe really effective for us and we can\nmake really amazing products unlike the\nworld has ever seen. Um and over time we\ncan expand to this more general corpus\nof predicting every possible gene\nperturbation in every possible cell\ntype. And so I think that's maybe the\nway the analogy maps on. But it is true\nthat we are vertically integrating here.\nWe're generating our own data in a way\nthat's proprietary. We think we have a\nmuch much larger data set for this\nparticular regime than the rest of the\nworld combined and that enables us to\nbuild what we think are the best models\nand in many cases what we found is that\nunlike with LLMs where a lot of the data\nthat was necessary to build these was\nsort of uh a common good. It was\nproduced as a function of the internet\nshared across everyone it's pretty\ncommon across all the domains everyone\nwants to use it for. This biological\ndata is still in its infancy. It's like\nimagine we're in like the early 1980s\nand we are just now thinking about\ntrying to create some of the first web\npages. That's kind of the era we're in.\nAnd so we're going after generating some\nof our own data in this very niche\ncircumstance like building the very high\nquality corpus, the Wikipedia that you\nmight train your, you know, overly\nanalogized now LLM on. Um, and then\nbuilding the first products based on\nthat and then expanding from there. And\nso we think that's necessary because of\nwhere we are today. there isn't this\ninternet-like equivalent of data that\neveryone can go out and and reap rewards\nfrom.\nInteresting. And then um this is more\nmore a question about the broader pharma\nindustry rather than just new limit\nwhich is that in the future how are\npeople going to make money if you have\nyou know with the GLPs we've got\npeptides from China that are just a gray\nmarket that people can easily consume\nand presumably with these future AI\nmodels even if you have a patent on a\nmolecule maybe finding an isomorphic\nmolecule or an isomeorphic treatment is\nrelatively easy if you do come up with\nthese crazy treatments and if pharma in\ngeneral was able to come up with these\ncrazy treatments, will they be able to\nmake money?\nMhm. The gray market piece I'll maybe\nput aside and say, you know, that's sort\nof a um IP enforcement at a geostrategic\nlevel that I'm maybe not qualified to\nspeak to. But I do think it comes down\nto to IP enforcement effectively. Um, I\nthink for that gray market piece,\nanother another reason that sort of the\ntraditional pharmaceutical industry, I\nthink, will still continue to reap the\nmajority of rewards here is that most of\nthe payment in the United States, which\nprovides most of the revenue for drug\ndiscovery in the world, goes through a\npayment system that is not just direct\nconsumer. It goes through payers. And so\nif you have the opportunity to either\nlike order a sketchy vial off of some\nwebsite from some company in Shenzhen or\nyou can go through your doctor and get a\nprescription with a relatively low copay\nfor tresepide the real thing. I think\nmost patients will go for trespide. I\nthink you and I probably live in a millu\nof people who are much more comfortable\nwith ordering the vials from Shenzhen\nthan most people might be. Um but I I\ndon't consider that to be a tremendous\nconcern at large. I do think the the\nbroader point of if you have medicines\nwith very long-term durability, how do\nyou reimburse them or if just the\nbenefits are very long-term and you know\nsort of acrue in the out years a\nchallenge we have in the US system is\nthat the average person churns insurers\nevery 3 to four years that number\nfluctuates around but that's the right\norder of magnitude and that means that\nif for instance you had a medicine which\ndramatically reduced the cost of all\nother healthcare incidents but it\nhappened exactly 5 years after you got\ndosed with it no insurers technically\neconomically incentivized to cover that.\nAnd so I think there are a couple models\nhere that can make sense. One is\nsomething called pay for performance\nwhere rather than reimbursing all of the\ncost of the drug upfront, you actually\nreimburse it over time. So say you get a\nmedicine that just makes you generically\nhealthier and you can, you know, measure\nthe reduced rates of heart attack and\nreduced rates of obesity and various\nother things and you get this one dose\nand it lasts for 10 years. each year you\nwould pay something like a tenth of the\ncost of the medicine contingent on the\nidea that it was actually still working\nfor you and you had some way of\nmeasuring that. So that's a big a big\nchallenge in this industry is like how\nwould you demonstrate that any one of\nthese medicines is still working for the\npatient. In the few examples we have\ntoday, these are things like gene\ntherapies where you can just like\nmeasure the expression of the gene and\nlike okay the drug is still there. But\nit gets more complicated when you have\nsome of these sort of longer term net\nbenefits. And the idea would be that\nthen each insurer is incentivized to\njust pay for the time of coverage that\nyou're on their plan. And we already\nhave a framework for this in you know\npost affordable care act in the US where\nyou know pre-existing conditions no\nlonger really exist. So patients are\nable to freely move between payers and\nyou could sort of treat the presence of\none of these therapeutics lowering this\npatient's overall healthcare costs the\nsame way we treat a a pre-existing\ncondition. I think this is something\nthat the system is still overall\nfiguring out. So what I'm saying here is\none hypothesis about what the future\nmight look like, but I think there are\nalternative clever approaches people\nmight think about for reimbursement. I\nalso think over time we're going to move\nmore toward a direct to consumer model\nfor many of these medicines which\npreserve and promote health rather than\njust fixing disease. You're seeing what\nI think are really some of the most\ninnovative examples of this right now\nfrom Lily around the incredibly\ndirect. So for the first time, rather\nthan going to a pharmacy which interacts\nwith a PBM which interacts with your\nprimary care physician, now you can get\na prescription from your doctor, go\nstraight to Lily, the source of the good\nstuff, and you're able to order\nhigh-quality drug from them when not\ninvolve, you know, some intermediary\ncompounder in the middle that might not\neven make your molecules properly. Um,\nand I think as these medicines develop\nthat have actual consumer demand because\nyou feel it in your daily life, you're\nactually seeing a benefit from it. It's\nnot just something that your physician\nis trying to get you to take. that that\nmodel will start to dominate and that\nmeans that this sort of like payment\novertime for some of these long-term\nbenefits might be able to be abstracted\naway from our current payer system where\nit turns every few years and now a sort\nof like payment overtime plan the same\nway we finance other large purchases in\nlife seems very feasible.\nThe the reason I'm interested in this is\nthat healthcare is already 20% of GDP. I\nthink it's grown like notable\npercentages in the last few years. It's\nlike this is a fraction that is quickly\ngrowing. Um, and most of this I I I\nshould have the numbers I should have\nlooked the numbers up, but the\noverwhelming majority of this is going\nto administering treatments that have\nalready been invented. Um, which is\ngood, but nowhere near as good as\nspending this enormous sum of resources\ntowards coming up with new treatments\nthat in the future will improve the\nlives of people that will that will have\nthese ailments. Um, I mean, one question\nis just how do we make it so that more\nlike we're going to spend 20% of GDP on\nhealthcare, it should at least go\ntowards like coming up with new\ntreatments rather than just like paying\nnurses and doctors to keep administering\nstuff that kind of works now. And two,\nif the cost of drugs ends up being at\nleast from the perspective of the payer\nends up being you need a doctor to give\nyou some scan and before I can write you\num a prescription and then they need to\nadminister it and they need to make sure\nthat you're doing okay, etc., etc., Then\neven if for you to manufacture this uh\ntherapy might cost um you know tens of\ndollars per patient for the health care\nsystem overall it might be tens of\nthousands of dollars per patient. I\nactually I'm curious if you agree with\nthose orders of magnitude. I\nI think that's correct. So I think the\nstat is something like drugs are roughly\n7% of healthcare spend. I I could be a\nlittle bit wrong on that but the the oo\nis right.\nRight. So basically even if we invent\nde-agging technology or especially if we\ninvent aging technology how how should\nwe think about the way it will net out\nin the fraction of GDP that we have to\nspend on healthcare will that increase\nbecause now people just had to go\neverybody's lining up at the doctor's\noffice to get a prescription and um you\ngot to go into clinic every week or will\nthat decrease because the other\ndownstream ailments from aging aren't\ncoming about.\nI I think the latter is much more likely\nto be the case. So just as like some\nquick heruristics part of the reason I\nthink there are many reasons that\nhealthcare costs so much in the US. One\nof them is something like bomb's cost\ndisease which is you know very unrelated\nto pharmaceutical discoveries but you\nknow is something that we will have to\nsolve in the system. Part of it's like\nthe disintermediation of the actual\ncustomer and the actual provider and and\nthese are things that biotech probably\nisn't going to be able to solve as an\nindustry alone. That's probably a larger\na larger economic problem. But when you\nthink about how will this affect sort of\nthe total amount of healthcare that will\nneed to be delivered, if you have more\nof these what I like to think of as sort\nof like medicines for everyone,\nmedicines that keep you healthier longer\nrather than medicines that only fix a\nproblem once you're already very sick, I\nthink you actually avoid a lot of the\ntypes of administration costs, not just\nadministration like admins at hospitals,\nbut the cost of administering existing\nmedicines and therapies to you going\ndown. One stat on why I think that's\ntrue. Something like a third of all\nMedicare costs are spent in the final\nyear of life, which is shocking when you\nrealize that the average person on\nMedicare is I don't know the exact\nnumber, but probably a decade plus\ncovered by it. And so there's an\nincredible concentration of the actual\nexpenses once someone is already\nterribly sick.\nSo helping prevent you from ever having\nto access the intensive healthare\nsystem, meaning something like an\ninpatient hospital visit. If you can\nprevent even just a couple of those\nvisits over a long period of someone's\nlife with a medicine like an increment,\nlike a reprogramming medicine that keeps\nyour liver, your immune system younger,\nI think on net that actually starts to\ndrive health care spend down because\nyou're you're sort of shifting some of\nthat burden from the administration\nsystem to the pharmaceutical system. And\nthe pharmaceutical system is the only\npiece of healthare where technology has\nmade us more efficient. As drugs go\ngeneric, actually the cost of\nadministering a given unit of healthare\nis going down. And the grand social\ncontract is that they eventually go\ngeneric. That's the way our current IP\nsystem works. So I think you know if you\nwere to get the question of like when\nwould you like to be born as a patient.\nYou always want to be born as close to\ntoday as possible because for a given\nunit in terms of pharmaceuticals for a\ngiven dollar unit of expense, you can\naccess more pharmaceutical technology\ntoday than has ever been possible in\nhistory. Even as healthcare costs\neverywhere else in the system have shot\nup. And so pharmaceuticals are the one\nplace where because of the mechanism of\nthings going generic and the fact that\nour old medicines continue to work and\npersist over time, you're actually able\nto get more benefit per dollar.\nOkay, final question. Um, so pharma is\nspending billions of dollars per new\ndrug it comes up with and surely they\nhave noticed that the lack of some\ngeneral platform or some general model\nhas made it more and more expensive and\ndifficult to come up with new drugs. and\nyou say perturbed seek has existed since\n2016 and you as far as you can tell you\nhave the most amount of that kind of\ndata which would could feed into a\ngeneral purpose model. So um what is\nwhat is like the traditional farm\nindustry on the other coast up to if you\nif I went to the the head of R D at Eli\nLilly or Fiser or something do they\nthink that this is like they have some\ndifferent idea of the platform that\nneeds to be built or they're like no\nwe're all in on the bespoke game uh\nbespoke for each drug.\nYeah I so I'll just correct one thing to\nmake sure I'm not overstating. We have\nway more data for a particular the\nlimited sub problem we're tackling which\nis overexpressing TFS in combinations. I\nthink we have way more data than anyone\non full stop there. But even more\nspecifically, I feel very very confident\nwe have more data than anyone looking at\ntrying to reprogram a cell's age. And so\nthat's where we're we're way larger than\nthan the rest of the world. When we\nthink about just general single cell\nperturbation data, various flavors, then\nI think there are other groups which\nwhich have very large data sets as well.\nWe're still differentiated because we do\neverything in human cells with the right\nnumber of chromosomes. Whereas it's very\ncommon to do things in like cancer cell\nlines which have 200 chromosomes. So\nlike is that human? I don't know.\nDepends on depends on how you actually\nquantify these things. Um so then if\nyou're going to go ask the leaders of of\nsome of the traditional pharmaceutical\nfirms like are you trying to build a\ngeneral model? I think some of them have\nin-house like AI innovation teams that\nare working on this. They're really\nsmart people there. But I think it's a\ngeneral trend. I think you can think\nabout some of the modern pharma a bit\nlike venture capital firms where they've\nover time externalized a lot of their\nR D and so they often have divisions of\nexternal innovation which you can kind\nof think of as like the corp dev version\nof venture capital. They work with the\nbiotech ecosystem to have a number of\nsmaller nimble firms explore really\npioneer ideas like the types of things\nwe're working on and then eventually\npartner with them once they have assets\nthat are later downstream. And so I\nthink the industry has sort of\nbifurcated where smaller biotechs like\nours take on most of the early\ndiscovery. The stat I'm going to get a\nlittle bit wrong from memory, but it's\nsomething like 70% of molecules approved\nin a given year come from originally\nsmall biotechs rather than large pharmas\neven though you look at the actual like\ndollars of R D spend on the balance\nsheet and it's like largely in big\npharma.\nAnother level intermediation.\nAnother disintermediation and and part\nof the reason for that that difference\nin cost is they're running most of the\ntrials. most people partner with farmer\nto run trials where a lot of the costs\nare incurred. So it's not just that like\noh all large farmers are horribly\ninefficient or anything like that.\nUm and so I think some of them would\ntell you like these ideas are really\nexciting. We have an external innovation\ndepartment if we don't have one\ninternally or we were collaborating with\na startup that's doing something\nsimilar. And so you can kind of think of\nthe market structure like you have a\nbunch of biotechs which are kind of like\nthe startups in your ecosystem and then\nthey're working with something like an\noligopsin of pharmas where it's like a\nlimited number of buyers for this\nparticular type of product which is a\ntherapeutic asset that is ready for a\nphase 1 phase 2 trial and so there's a\nvery liquid market for the phase 1 phase\n2 assets and that's the point at which\nthese partnerships can can come to\nfruition and so I think that's what a\nlot of those leaders would say now some\nof them by by contrast uh for instance\nro bought back in 2013\nR D is currently run by a viv forge, one\nof the scientists I admire most in the\nworld who is like a thousand times\nsmarter than me. Um, and you know, she's\none of the people who invented this\ntechnology and has a big group doing\nthis sort of work there. So, it's not\nlike every pharma takes that view, but I\nthink that's sort of a general trend.\nInteresting. Um, full disclosure, I am a\nsmall uh angel investor in New Libert\nnow, but that that did not influence the\ndecision to have Jacob on. This is super\nfascinating. Thanks so much for coming\non the podcast.\nAwesome. Thanks for cash. I hope you\nenjoyed this episode. If you did, the\nmost helpful thing you can do is just\nshare it with other people who you think\nmight enjoy it. Send it to your friends,\nyour group chats, Twitter, wherever\nelse. Just let the word go forth. Other\nthan that, super helpful if you can\nsubscribe on YouTube, and leave a\nfivestar review on Apple Podcast and\nSpotify. Check out the sponsors in the\ndescription below. If you want to\nsponsor a future episode, go to\ndwarcash.com/advertise.\n[Music]\nThank you for tuning in. I'll see you on\nthe next one.",
  "transcript_chars": 130317,
  "ingested_at": "2026-05-15T04:57:13.565381+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": null,
    "like_count": null,
    "channel_id": null,
    "categories": null,
    "tags": null
  }
}