{
  "video_id": "l2hro8DemsM",
  "channel_slug": "statquest",
  "channel_handle": "statquest",
  "title": "Human Stories in AI: Fabio Urbina",
  "duration_seconds": 2131.0,
  "url": "https://www.youtube.com/watch?v=l2hro8DemsM",
  "upload_date": "",
  "transcript": "hello I'm Josh starmer and welcome to\nhuman stories and AI with stack Quest\nand lightning AI in this series we'll\nhear about the career journeys of\npassionate AI experts from their humble\nbeginnings to conquered challenges will\nbe inspired by the realworld experiences\nof professionals thriving in the ever\nevolving AI\nlandscape human stories and AI is\nbrought to you by lightning AI code\ntogether prototype type train and deploy\nAI web apps all from your browser with\nzero setup personally I love lightning\nAI because it makes it super easy to use\nand learn from the stat Quest coding\ntutorials just go to the web page click\non the Run button and Bam you get code\nthat you can play with without\ndownloading anything or installing any\npackages today we have special guest\nFabio urbina an associate director at\ncollaboration\nPharmaceuticals Fabio combines\ncomputational tools and machine learning\nwith classical small molecule molecular\nand cell biology techniques to address\npreviously difficult to probe scientific\nproblems specifically Fabio finds\nsolutions to drug Discovery with machine\nlearning so without further Ado Fabio\ncan you tell us about your journey to\nwhere you are right now at\ncollaborations Pharmaceuticals how did\nthis all start okay if we go as far back\nas possible so uh my great-grandfather\nback no I'm kiding far um no it's a\ngreat question so um yeah I guess as a\nkid I was always really interested in\nscience it was something I was always\nwanted to do something that I really\nenjoyed you know I'm sure many people uh\nLike Me grew up watching Nature\nDocumentaries and science documentaries\nand really enjoyed the process of\nlearning and so since I was little I\nessentially wanted to go do something in\nthe Sciences\num you know didn't know what that meant\nexactly as a kid you know you really\ndon't know what you're looking for but\nsomething in that realm and so uh\nanother sort of component to my I guess\nupbringing is that I was really into\ncomputers um so you know I remember we\nget our first computer as one of the\napples you know Green text Oregon Trail\nand floppy disc it was great and so that\nthat sort of became one of my hobbies\nessentially on the side as I was going\nthrough school and whatnot so in college\num I decided to get a biolog a Bachelor\nof Science in biology and a minor\ncomputer science and so I sort of just\ndabbled in um biology and that's really\nwhere a lot of my interest lies I\nremember taking my first like cell\nbiology class and thought it was like\nthe coolest thing ever and so sort of\nwhere I decided to to keep my focus in\nthat\narea and um yeah so essentially you know\ngot a Bachelor's of Science and biology\nand uh there's a really interesting\ninternship program that I remember\nseeing uh on a door of one of our\nbuildings and it was to come up to um\nMassachusetts General Hospital and do\nlike a couple of years of lab tech\nresearch once I graduated in a um lab\nthat worked in a really rare disease\ncalled familial disautonomia which I'm\nsure most people have not heard of\nsomething like 400 people in the world\nhave ever been affected by it um so it's\na very very rare disease and so not\nknowing what exactly I wanted to do once\nI graduated college which I think a lot\nof people find themselves in that that\nsort of interum what to do um I decided\nto go there and become a lab tick for a\ncouple years and that's when I got\nreally invested in essentially rare\ndiseases and uh early stage drug\ndiscovery which is is what I focused on\nwhile in that lab so we were essentially\ntrying to find a drug of any sort that\ncould potentially be a therapeutic for\nthis rare disease um this famili dis\nanomia and so that experience made me\nrealize that I really wanted to kind of\ndive head first into\nbiology and really into the cell biology\nof things and really into the the\nnitty-gritty of of what makes cells kind\nof tick and that sounds a little bit\nquite different from my current area of\nresearch but it'll come back around in a\nminute here and so um I joined a lab at\nUNCC Chapel Hill which was the lab of\nStephanie gupton and there I worked on\num neuron cell biology and neuron\ndevelopment and essentially how do\nneurons grow and make the connections\nthey do in your brain from you know\nindividual Blobs of cells to this very\nvery complex\nstructure and one thing I wanted to\nbring to that experience is some of my\ncomputational background so throughout\nmy time\nof working in Biology one thing I did\nwas I always kind of weaved in some\ncomputer science or computation or\nstatistical sort of analysis kind of\ngrouped in to the actual experimental\nwork and so that sort of is what I did\nover the course of my PhD is um applied\ncomputational image analysis to a lot of\nthe cell biology Imaging we were\ndoing so once I finished my PhD I I kind\nof realized I didn't really want to go\ninto Academia\nit wasn't really what I was that wasn't\nreally the career trajectory that I\nwanted to go so I wasn't quite sure what\nI wanted to do so you you hear there's a\nlittle bit of luck in this for sure when\nit comes to finding out you know where\nyou go in life but I started looking for\ninternships and Industry positions and\nso essentially started looking around\nfor um you know posts for internships\nthat looked really interesting and by\nbiotech um companies sort of in the\narea and um so good was this while you\nwere still a student or um or is this\nafter You' gotten your PhD so this is\nwhile I was still a student so this was\nabout my last year um I knew I was going\nto graduate soon now of course my\ngraduation schedule is a little messed\nup because I graduated in 2020 which is\nuh right when yeah and you know that\nsort of delayed a lot of things and\nthere was a lot of un things I weren't\nquite sure there um but uh yeah\nessentially towards the last year of my\nPhD I you know you start thinking about\nwhat it is you want to do and usually\nthe next step for a PhD when you finish\nit if you're going to stay in Academia\nis to go on and do a postto in another\nlab somewhere and um you know I I really\nlike the area here in uh the triangle\narea of North Carolina didn't want to\nmove um I wanted to sort of jump more\ninto my career stage rather than jump\ninto the postto life and then sort of\nthen go on to try to start a lab so\nindustry just seemed like the kind of\nthing I wanted to do so yeah during my\nlast year of um my PhD work looked for\ninternships um and you know you get\npermission from your pi to go and do an\ninternship if you're going to do it in\nthe middle of your PhD work and so we\nworked that out and uh yeah I just\nhappen to find this um\ncompany collaborations Pharmaceuticals\nwhich is where I ended up for the\nmajority of this far and they just had a\nlot of inter really interesting postings\nof what you could work on if you wanted\nto be an intern with them so I\nessentially emailed um Sean who's the\nCEO and he he I think on a Thursday off\nto double check but then he he got back\nto me within an hour and was like oh can\nyou start\ntomorrow wow very different experience\nthan most um yeah wow it's yeah I will\nsay there's a lot of yeah one of the\nnice things about the sort of company I\nwork at I kind of try to push this\nperspective a little bit is as I think\nwe have this sort of concept of what a\nbiotech or drug Discovery company looks\nlike and it's usually from a very large\npharmaceutical company point of view and\nyou know the pluses and minuses that\ncome with that sort of Industry\nviewpoint but um when I joined\ncollaborations as an intern I think we\nhad less than 10 people including the\nCEO\num so it's very very small um I almost\nwant to call it a mom and pop operation\nbecause it's literally um sha and his\nwife are the ones who run it and so\nreally yeah so it's very very different\nand I I think that's what really drew me\nto it is just how different it was from\nthe general sort of you know drug\nDiscovery monoliths that I think of in\nmy head um the the other thing that\nreally drew me to actually applying as\nan intern for them is that they focus\nspecifically on rare neglected\ndiseases oh really wow mhm so they they\nactually foro a um the typical funding\nmodel of of Biotech where you go to get\nVC funding and raise capital and you\nhave investors their actual original\ninvestment strategy was through grants\nfrom the government in order to create\nTechnologies and in order to find\nTherapeutics for rare diseases which are\ngenerally considered not profitable and\nso that was a real big sort of draw for\nme to them so okay you know I I really\nlike that they had such a very different\nmodel and that their focus was a lot\nmore on the actual therapeutic side and\nwasn't completely driven by sort of like\nan investor type um model which mostly\npharmaceutical companies are um so yeah\nso I I started the internship and um\nessentially just had a lot of random\ninteresting problems thrown at me and uh\nyou know tackled them and it was it was\na lot of new learning I will say mhm um\nabout yeah because it's so different\nfrom a from the biology that I studied\nit was very drug Discovery focused it\nwas very machine learning focused and\nwhile I was fairly good with the\nstatistical background machine learning\nwas fairly novel for me so there was a\nlot of learning in that sort of first uh\nfew months um completed the internship\ndecided I really wanted to work there so\nI you know essentially applied\nafterwards and they were happy to have\nme and then I spent the next few months\nfinishing my PhD defending and then\njoined the company essentially as a um a\npostto okay and we could go through the\nwhole process but essentially over the\nnext few years I sort of rose to the\ncompany until I eventually became an\nassociate director uh mostly overseeing\na lot of the machine learning a lot of\nthe software development and um sort of\nthe\ncreative experiment mental computational\nside of the company can you tell us a\nlittle bit about uh what the machine\nlearning what you guys are doing machine\nlearning what the uh what are you trying\nto accomplish with machine learning yeah\nso um we focus on what's generally\ncalled early stage drug Discovery and so\nwhat that means is uh we may have a\ndisease of some sort so one of the ones\njust to pick is a malaria for example\nand maybe we want to try to find new\nantimalarials\nand so what we do with the machine\nlearning is or I guess I'll start with\nthe traditional drug Discovery approach\nis usually you would take a current\nantimalarial compound or drug and a\nmedicinal chemist might take it and try\nto alter the structure of the molecule\nin order to try to find some uh maybe a\nbetter antimalarial that's kind of\nsimilar um the sort of alternative is\nthis naive strategy of you just kind of\nbrute force it by getting very large\nnumber of diverse compounds and you\ncreate an assay that can tell you\nwhether a compound has antimalarial\nactivity or not and then you just Brute\nForce this whole entire set of compounds\nand try to find something randomly okay\nand that's actually been a strategy\nthat's been one of the more successful\nwhich kind of tells you how um how\nnon-specific and kind of random drug\nDiscovery is a little bit yeah so that's\nthat's got to hurt the ego of the um\nwhat did you say structural yeah the the\nstructural medicinal chemist yeah so\nfinding new structures is difficult\nproblem yeah um and so machine learning\nis sort of this way to to dive in and\nkind of bridge the two gaps there and\nthe way we use machine learning is we'll\ntake our known uh drugs or compounds for\nexample we may we have a list of\nantimalarials um okay as well as you\nknow compounds that do not do anything\nto malaria they're dcri\nnegatives and then we can give these\nstructures to machine learning model and\ntrain the model to essentially take in a\nNew Drug chemical structure and decide\nor predict whether that new molecule is\nlikely to have antimalarial activity or\nnot okay and so then we can take this\nnew machine learning model and we can we\ncall it virtually screening all we have\nto do is take the structures virtually\nand essentially put them through this\nmachine learning model and predict which\nof these you know thousand 10,000\n100,000 virtual compounds might be\nantimalarials and then we can follow up\nthe predictions by actually testing\nthose compounds and this way instead of\ntesting thousands and thousands of\ncompounds kind of randomly we can narrow\nit down to 10 or 100 compounds and test\nthose and that sort of accelerates our\nability to find new drugs or new\nantimalarials in that\ncase off the top of your head like of\nthose 10 that you actually test do all\n10 show some efficacy or or how how\naccurate is this method yeah it's a\ngreat question so um we've had success\nrates where we've picked three compounds\nand those three were all active and\nthose are actually three that we've\ntaken forward so that was a really I\nthink that was I forget exactly which\nwhich um viral we originally tested in\nbut I think it went on to become we\nfound like three anti-ebola anti- some\nsort of antivirus but yeah we've had\nsuccess rates upwards of three out of\nthree to totally novel structures that\noh wow ended up being efficacious we've\nhad uh something I think I think like\nanother project we had something like\nseven out of 10 were very efficacious\nokay and then we have had projects where\nwe tested 10 and zero out of the 10 were\nactually efficacious so we tend to have\na very nice enrichment rate meaning 10\nto 100 or a thousandfold\na better chance of finding new compounds\nnew drugs but it is still a um challenge\nonce you build a machine learning model\nto decide is this going to be applicable\nare we going to actually find what we\nare looking for but yeah we have had\nsome pretty good success rates using\nthese machine learning models have have\nyou learned anything you know when you\nwhen you get zero out of 10 hits has\nthere been like oh\num you know this disease has this\ncharacteristic that wasn't part the\noriginal training data or is there some\nindication as to why uh it didn't\nwork yeah um one of the things and this\nis something we deal with is um we call\nit a coverage or sort of an\napplicability domain which is this idea\nof our you know training sets for these\nmodels are compounds and drugs\nthemselves but sometimes those drugs and\ncompounds in the training set cover a\nvery very narrow chemical range and all\nI mean by that is they all look very\nsimilar to each other structurally yeah\nand so when we then build this machine\nlearning model it might look really\nreally nice on paper because what it's\nactually doing is just learning a very\nnarrow chemical space and then when we\ngo on to predict on maybe something way\noutside that chemical space the model is\nmuch more confident than it actually\nshould be because it's only learned on\nsuch a narrow chemical space so that's\nthat's one challenge we generally face\nis how diverse is our training set uhhuh\nand um that tends to be the main problem\nin in our field unlike you know text\ngeneration or image generation is it's\nreally really expensive to generate data\nsets usually have to do full experiments\num on a single compound in order to get\neven one data point so usually it's the\nspareness of the data that ends up being\nthe the issue for the most\npart which uh kind of leads to another\nquestion which is usually when you do\nwhen I think of machine learning I think\nof Big Data huge data sets um how large\nare the data sets that you're working\nwith I I just assume that they must be\nmuch smaller because obtaining them is\nrelatively expensive it's you don't just\nsuck down the entire\nWikipedia um and then work from that or\nor or yeah so what can you tell us about\nthat yeah I mean you're 100% right the\nour data sets are considered tiny by\ncomparison to what do you think of as\nlike you know maybe what open AI is\ndoing with chat GPT they're like you\nsaid consuming terabytes of data we work\non the order of hundreds of compounds or\nhundreds of data points up to maybe I\nthink the biggest data sets we get are\ngenerally about\n100,000 data points and so we are in\nkind of a small range and that that does\ntend to present challenges that we've\nbeen investigating a lot of uh potential\nanswers for especially on the the\nextreme end we actually recently\ncompleted a project which we're\nhopefully going to publish on soon where\nour data set was actually 15 compounds\nin size oh wow holy smok we managed to\nengineer a model that could actually\ngive some predictive power and actually\nfound some new compounds on on a\ntraining set that's only 15 data points\nnow that is fascinating that's almost\nlike a nano data set especially for\nmachine learning I think for traditional\nstatistics 15 is large but for machine\nlearning that is the smallest data set\nthat I other than like the really simple\nexamples that I use in my little videos\nuh which I intentionally make as small\nas possible just so we can see the math\nbut I didn't actually think it was\npossible to have a data set that small\nI'm going to be honest that that just\nsounds like you're you're kind of\nblowing my mind yeah it's a it's a field\nthat doesn't get as much attention\nbecause it's you know you don't get\nthese massive impressive models but a\nfew shot learning models or even zero\nshot learning models is is kind of what\nthese are generally called and the idea\nis if you approach from the extreme end\nwhat how can you extract sort of a\nmaximal information from these tiny data\nsets and they're not going to be super\nimpressive but where you can focus their\napplication and their predictions they\ntend to be pretty accurate so yeah we\nwere we were very excited to kind of get\nthose results um can you share any of\nthe tricks that you used uh can you tell\nus about the model one the type of model\nthat you trained and share some tricks\nthat you Ed to make it work with such a\nmicroscopically small data set yeah um\nso the the model type that we use um\nit's called a prototypical network I'm\nsure most people haven't heard of it\nunless they've actually worked in this\narea so I've never heard of it yeah so\nit's a it's a very you know it's\nactually a fairly simplistic model\nunderneath the underneath the hood and\nit's designed that way on purpose\nbecause the more the simplistic models\ntend to have a lot of um bias within\nthem which we generally think of as as\nhurting modeling but actually um bias\ncan actually be a good thing if the\nmodel's bias is sort of correctly\naligned with what you think the data set\nstructure is so so the simplest um I\nguess uh example of this is for example\na linear model um anyone done yal MX\nplus b you know that's the line and\nslope that's a linear model if your dat\nset is actually linearly correlated that\nmodel is most likely going to outperform\njust about any other model you could\nthrow at a small data set because it's\ngoing to draw a straight line which is\nyour actual Association so that's for\nexample a bias of a linearity within\nyour data um yeah with the prototypical\nnetworks the general bias is you know we\nwant to give it two classes of compounds\none that we know are uh I'm going to\nstick to my malaria example\nantimalarials one we know that are not\nand so the thought is it's going to try\nto figure out what is the maximum\ndifference between these two classes\nwhat is so extremely different between\nthese very small data points and so we\nyou know we tried a couple of strategies\num one of them\nbeing can we choose negative compounds\nthat look very very dissimilar\nstructurally from the positive compounds\nbecause when we have a data set of uh\nyou know 15 compounds it we could\nessentially give the entire data\nset and sometimes if the structures are\nkind of on a spectrum where they all\nkind of look similar with ones and zeros\nuh or sorry with um the uh antimalarials\nand non- antimalarials looking too\nstructurally similar that could be a\ndifficult problem for the machine\nlearning model to do so we actually even\npruned away some of the data a little\nbit to give it only the things in the\nnegative that look as dissimilar as\npossible from the things in\nthe positive which was the antimalarials\nand it sounds almost a little bit like\ncheating with your data but when you\napply it to outside data test set you\nknow finding new compounds it tended to\nwork really really nicely in that in\nthat\nfashion so the so the it's sort of like\nthe ends justify the means like maybe it\nlooks like you were cheating but as a\nresult you had something that\ngeneralized a little better uh than it\nwould have otherwise\num absolutely I mean that sounds good to\nme yeah um and that's fantastic and can\nyou tell us a a little\nbit about the T I've already forgotten\nthe type of model you're using uh can\nyou tell us a little bit more about that\ntype of model I've never heard about it\nso I'd like to kind of understand sort\nof what how what what how does it work\nyeah so the Mel type again is a\nprototypical network which sounds kind\nof funky um yeah and and there's\nactually a second part of the model\nwhich I won't get into too much detail\ncuz it gets a little too into the weeds\nokay um so the way our model is set up\nis it's composed of of two pieces one is\nwhat we call an embedding model and the\nsecond is the actual prototypical\nNetwork so what the embedding model is\ntrained to do is to take in a compound\nstructure and generate a vector of\nnumbers uhhuh that represent that\nstructure yeah and the important thing\nthat it's required to do when we train\nit is that structures that are similar\nto each other so if you have two\nmolecules that look really close to each\nother the vector of numbers they produce\nshould look very very similar so the\nnumers should align really really well\num whereas structures that are very\ndissimilar should look very different so\nif you okay lined up their numbered\nvectors they should look different they\nshould have very different numbers so\nthis allows us to we call it embedding\nour molecules into Vector space which is\nto say we just numerically put these\nmolecules in a in a high dimensional\nnumber space where structurally similar\nare numerically similar and you know the\nopposite so once that model is trained\nin order to do that we take all of our\nmolecules we embed them into this Vector\nspace and the prototypical network the\nway it works is it takes these number\nvectors and it sounds so simple because\nit is it takes the two classes our\nantivirals and our non- antivirals and\nit essentially finds the average of\nthese Vector numbers in this space\ncalled a prototype what you're doing is\ngetting the average Vector of the\nantivirals and an average Vector then\nit's so simple but it's so powerful when\nyou apply it to small data once you have\nthese uh sort of they call prototypes\nthese mean vectors that represent the\naverage antimalarial non- antimalarial\nthen in order to predict a new compound\nas you know maybe having antiviral\nactivity is you would put it through\nthis embedding vector and you would\ndefin which prototype is your vector\nclosest to so is your a new molecule\ncloser to the zero class where there's\nno antiviral Properties or to the\nantiviral vector or prototype yeah and\nit's very simple but very very powerful\nso it's almost like a nearest neighbor\nalgorithm yeah it I believe it's based\noff of a k nearest neighbor actually is\na very is essentially the the origin of\nthat strategy so from my perspective um\nso I've got a I hate to I'm like tooting\nmy own horn here but I have a video on\nsomething called word embedding uh which\nis essentially uh what you described but\napplied to molecules instead of words\nand it does the same thing and then you\ntake those those vectors of numbers that\ncome out of the embedding Network and\nyou obey basically apply K nearest\nneighbors to it uh to find and this to\nfind sort of to to then classify\nwhatever your new molecule is this\nsounds fantastic I absolutely love it I\nlove the Simplicity of it um I know I\nknow right now when people think of AI\nand they think of ml they think of chat\ngbt and they think of these huge\nmonolithic\nmonsters um and these from what I know\nare incredibly expensive to run they\nlike just you know they've got so many\ngpus just chewing up electricity\ngenerating heat and need to be cooled\ndown everything about them is expensive\nexpensive expensive and it's and and\nthat's awesome and whatever fine but it\nsounds like what you've got is\nincredibly simple uh and I'm I'm going\nto guess that the hardware you run this\nmodel on is relatively modest compared\nprepared to say what chat GPT runs on\nyeah that's understating it I I've run\nthese on like a 2015 Mac like there's\nnot not even gpus involved in half of\nthis um it it actually is yeah you you\nhit a point there which is um something\nreally nice in I'm going to say our\nfield my field of drug Discovery and\nmachine learning is the one plus side of\nthese data sets being so small is that\nwe can run them on pretty much any\nmachine so while we do have um you know\nfairly large GPU clusters for running\ndeep learning models where we try to\naggregate much larger pieces of\ninformation um when it comes to some of\nthe smaller models uh yeah we tend to\njust be able to run these very uh simply\nlocally on our own machines most people\nwould be able to run these machine\nlearning models themselves on their own\nyou know Hardware desktop you just need\na your your modern day CPU and it'll\nwork just\nfine I love it I I just love it it's\nthis to me it's one of the unexpected\nthings like like for me I just assume\nthat everyone is and maybe other people\nare like this too but I just assume that\neveryone is doing something more fancy\nand more complicated than I could ever\nwrap my brain around and this sounds in\na way it's like oh I could do this it's\nlike it's and and I know that I I don't\nwant to make it I don't want to belittle\nwhat you're doing I'm in a way it's like\nI feel like I've been empowered by\nlistening to what you've said I I felt\nlike like the crazy world out there\nisn't as crazy as I assumed it did you\nyou kind of like brought it down and\nsaid\nhey get rid of all those silly extreme\nideas that you have let me let me tell\nyou what it's really like and it's and\nit's not so scary yeah I I completely\nagree I I don't think it's belittling at\nall I think I think you're right it's\nit's something that I wish more people\nwould realize is you don't have to go\nand try to chain train your own\ntransformer model with 96 layers in\norder to do machine learning or in order\nto you know even dabble and get\nsomething useful you know uh a lot of\nit's no secret that a lot of fields\nstill have a very small amount of data\nyou know the reason these uh image\ngeneration models and um text generation\nmodels are so big is because they're\nessentially the only Fields where you\nhave a huge amount of data and in any\nother application you know it's a I'll\nsay it's actually a common thing for\nwhen uh someone kind of new wish to\nmachine learning comes into our field\nthat the first thing I want to do is the\nlatest and greatest published model\nand what's really funny is a lot of\ntimes if you just train some of the\nsimple old models support Vector\nmachines you may have heard of from\ndecades ago random forests even um they\noutperform some of the newer model types\nM very easily actually in our field and\na lot of that has to do with you know\nhow much data there is and how um each\nof the algorithms sort of treats the\ndata so so it is a bit humbling because\nI did the same thing I came into the\nfield I said oh this is how what we're\nusing these are the featur we're using\nrandom forest for Vector machines oh I\ncould I could out do this come on you\nknow and I start all these models and\nnope I completely failed in my ability\nto outperform what was there so that it\nwas very humbling to come in and do that\nand to really have to sit down and\nreally engineer the problem a bit well\nhooray for the old models I'm I'm a\nsecret I'm a secret fan of all those old\nmodels I love it um um well to be honest\nthat's a cool that we could end on that\nif if if if that's all we got but if you\ndo have any you know final words of\nwisdom uh anything\nfor just you know a a student or someone\nwho's maybe changing their career if you\nhave any advice for them um we can you\nknow we can just let me know what you\ngot okay um you might have to cut me off\nbecause I can talk forever but so yeah\nyou're doing great you're doing great I\nguess you know the the biggest thing I\nthink I see people struggle with is\nthinking that they either don't know\nenough or they're really uncomfortable\nwith you know encountering a problem\nthey can't understand right away and you\nknow I I've trained a number of um\nstudents under me and I do see I think\nthe biggest struggle is yeah you you\nencounter something you don't know what\nit is and you shy away from it because\nit makes you feel bad because you don't\nyou don't like that uncomfortable\nfeeling of not knowing of not\nunderstanding it makes you feel like\nmaybe you you're you're not smart enough\nto understand\nsomething and that's such a small thing\nit sounds like but I think that's\nactually one of the biggest sort of\nbarriers to people kind of going on and\nbeing a lot more successful I guess or a\nlot more um a lot more willing to plunge\nahead into the\nunknown um and I say this only because I\nthink I have something a little broken\nin my brain that I've never had that\nissue where if I don't know something if\nI don't know something I I just don't\nknow it and if I try to understand it\nand I don't understand it I don't\nunderstand it but I'm going to keep\ntrying until I eventually do and so that\nyou know I I don't think I guess what\nI'm trying to say is it's the work you\nput into finally understanding it not\nhow long it takes you understand it\nthat's really important and most people\nwho are in fields who understand and\nknow a lot I think most of them they're\nnot necessarily smarter than anyone else\nI think they just are okay with\nrealizing they don't understand it they\ntry to learn it they still don't\nunderstand it they try to learn it again\nthey still don't understand it and they\njust keep going until they finally get\nit or know it enough that they can give\nit to somebody else to take\noveration SS but yeah um so you know I\nthink getting comfortable with feeling\nreally uncomfortable can be a really big\nBoon to people especially when they're\nthinking about changing Fields one of\nthe difficulties I know especially if\nyou're coming from a very very different\nbackground than what you're transferring\ninto is you're essentially starting over\nfrom scratch you become an expert in\nyour field over the course of say your\nPhD or even your undergrad you feel like\nyou become kind of an expert in the\ngeneral area and so you feel like you\nknow a lot and you've done this whole\nprogress for four plus years six and a\nhalf if you're doing a PhD sometimes and\nyou feel like you know it a lot of stuff\nand then as soon as you switch Fields\nyou all of a sudden are almost starting\nfrom scratch you don't know the lingo\nyou don't know the acronyms you don't\nknow what people generally do in the\nfield and that can be pretty\njarring and so that is also one of those\npain points where as long as you're find\nsort of going back and struggling\nthrough things again like you did at\nfirst you know it won't take long before\nyou eventually pick it up and are able\nto sort of run with it and because it's\nnot your first rodeo going down the\nresearch path it tends to be a lot\neasier it'll still take you a while but\nit's tends to to not be too difficult\nand so that's sort of what I found when\nI switched over from cell biology to\nsort of a drug Discovery machine\nlearning well yes I had some experience\nin the past felt like I was sort of\nreading the original papers in the field\nagain it was consuming so many reviews\nand so many papers right I didn't quite\nunderstand it and then having to find\nthe you know go back into the uh\nliterature and down more references\nrabbit holes all that um but you know\nit's one of those things where I think\nif you can get past those feelings of\nalmost an adequacy and can kind of push\nthrough it and recognize that you know\nyou can learn it you just have to put in\nthe time and you just have to be\ncomfortable with the uncomfortable then\nyou can switch Fields pretty much at any\npoint in your career I think it's scary\nbut if you have the capability of\nlearning once a very specific set of you\nknow topics you can do the same same\nthing over and over again there's\nnothing different about your brain from\nwhen you first started except maybe you\nknow maybe you're a little bit older and\ntakes a little bit longer but you know\nthere's nothing different about it from\nfrom the first time you did it so it you\nknow jumping Fields can be scary but\nit's 100% worth it if it's something you\nreally really think you want to\ndo I love it well on that note thank you\nvery much for being with us today Fabio\nuh it was great talking to you and\nhearing what you're doing I learned a\nton uh which is why we're doing this\npodcast to begin with so thanks again\nthank you",
  "transcript_chars": 33058,
  "ingested_at": "2026-05-15T10:54:30.586625+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 5288,
    "like_count": 86,
    "channel_id": "UCtYLUTtgS3k1Fg4y5tAhLbw",
    "categories": [
      "Education"
    ],
    "tags": [
      "Josh Starmer",
      "StatQuest",
      "Machine Learning",
      "Statistics",
      "Data Science"
    ]
  }
}