{
  "video_id": "L3nlGfSyHV0",
  "channel_slug": "statquest",
  "channel_handle": "statquest",
  "title": "False Discovery Rates, FDR, clearly explained",
  "duration_seconds": 1107.0,
  "url": "https://www.youtube.com/watch?v=L3nlGfSyHV0",
  "upload_date": "",
  "transcript": "Holy freaking smokes.\n>> What?\n>> It's time for Stat Quest.\nHello and welcome to Stat Quest. Stat\nQuest is brought to you by the friendly\nfolks in the genetics department at the\nUniversity of North Carolina at Chapel\nHill.\nToday we're going to be talking about\nfalse discovery rates or FDR.\nIf you've ever seen or done anything\nwith high-throughput sequencing, chances\nare you've heard of false discovery\nrates, FDR, before.\nYou may have even used them.\nBut where do false discovery rates come\nfrom?\nAnd how do they work?\nBefore we get down to the nitty-gritty,\nlet me blurt out the main idea of this\nwhole Stat Quest.\nFalse discovery rates are a tool to weed\nout bad data that looks good.\nNow let's get down to the nitty-gritty.\nLet's start with an example of measuring\ngene expression using RNA sequencing.\nHere, we're going to plot the\nmeasurements or the read counts for a\ngene called gene X, which is an\nimaginary gene,\non a graph with the Y axis being gene\ncounts and the X axis being the samples.\nFor this example, imagine that we are\nlooking at normal wild-type mice.\nLater on, we'll be comparing them to\nmice that have been treated with a drug.\nIsn't it funny [clears throat] that\nnormal mice are called wild-type?\nIf someone said I was a wild-type, I\ndon't think they would also think I was\nnormal.\nAnyway,\nhere's our measurement for the first\nmouse that we do this RNA sequencing on.\nAnd here's the measurement for the\nsecond mouse that we do the RNA\nsequencing on.\nRNA sequencing isn't perfect and\ndifferent samples are always a little\ndifferent. So each time we measure\nexpression, we'll get slightly different\nvalues.\nHere's the measurement for the third\nmouse.\nAnd here are the measurements for a\nbunch of mice.\nIf we measured all normal mice, we'd be\nable to calculate the average for all\nnormal mice.\nMost of the values are going to be close\nto the mean.\nRarely, we'll get a value that is much\nlarger than the mean.\nAnd rarely, we'll get a value that is\nmuch smaller than the mean.\nWe can summarize the distribution of the\nmeasurements using this bell-shaped\ncurve.\nMost of the measurements, which are\nclose to the mean, will come from the\nmiddle of this curve.\nThe rare measurement that is\nsignificantly larger than the average\nwould come from the right side of the\nbell-shaped curve.\nAnd the rare measurement that's\nsignificantly less than the average\nwould come from the left side of the\nbell-shaped curve.\nNow, imagine that we do RNA-Seq on three\nmice.\nCollectively, we'll call these three\nmeasurements sample number one.\nBecause these measurements are close to\nthe mean, they come from the middle of\nthe distribution.\nNow, imagine we compare sample number\none to another three measurements taken\nfrom normal wild-type mice.\nWe'll call these new measurements sample\nnumber two.\nAgain, these three measurements come\nfrom the middle of the distribution.\nIf we did a statistical test to compare\nsample number one to sample number two,\nthe P value would be large, greater than\n.05, because the two samples overlap.\nVery rarely, we'll get two samples that\ndo not overlap.\nWhen this happens, the P value will be\nless than .05.\nThis is called a false positive\nbecause the small P value suggests that\nthe samples are from two types of mice\nor two separate distributions, and this\nis false.\nNormally, false positives are rare\nunless you're a P hacker, but that's\nanother StatQuest already on YouTube.\nAnyways, normally, false positives are\nrare. 95% of the time, the two samples\nwill overlap. This will result in a P\nvalue greater than .05.\n5% of the time, they don't.\nThis will result in a false positive\nwith a P value less than .05.\nBut human and mouse cells have at least\n10,000 transcribed genes.\nIf we took two samples from the same\ntype of mice and compared all 10,000\ngenes,\nwell, 5% of 10,000 equals 500 false\npositives.\nThat means there will be 500 genes that\nappear to be interesting even when they\nare not.\n500 false positives is a lot.\nCan we do something about them?\nThe false discovery rate can control the\nnumber of false positives.\nTechnically, the false discovery rate is\nnot a method to limit false positives,\nbut the term is used interchangeably\nwith the methods. In particular,\nit is used for the Benjamini-Hochberg\nmethod.\nNow, there's a high probability that I\njust mispronounced Benjamini or\nHochberg, and if I did, I apologize.\nBefore we talk about the details of the\nBenjamini-Hochberg method, let's review\nthe concepts that it's based on.\nWe'll start by generating 10,000 P\nvalues from samples taken from the same\ndistribution.\nThat is to say, we'll start with test\nnumber one\nand we'll use wild-type mice\nand we'll take two samples from them.\nWe'll then compare the two samples with\na statistical test and calculate the P\nvalue. In this case, the P value is\nlarge. It's 0.83.\nThis is exactly what we expect because\nboth samples are taken from the same\ntype of mice.\nAnd then we repeat the procedure for\ntest number two and calculate another P\nvalue. This time it's 0.98. Again, this\nis as expected.\nTo make a long story short, we just\nrepeat this procedure 10,000 times.\nNo big deal.\nHere, I've drawn a histogram of the\n10,000 P values generated by testing\nsamples taken from the same\ndistribution.\nOn the X axis, we have possible values\nfor P values.\nOn the Y axis, we have the number of P\nvalues in each bin.\n510 P values or 5.1%\nare less than 0.05.\nClose to 5% of the P values are between\n0.5 and 0.1.\nActually, each bin contains about 5% of\nthe P values.\nAbout 500 P values per bin.\nSince the P values are uniformly\ndistributed, there's an equal\nprobability that a test P value falls\ninto any one of these bins.\nNow, let's look at how P values are\ndistributed when they come from two\ndifferent distributions.\nAnd by two different distributions, I\nmean two different types of mice, where\nwe have wild type versus knockout, or\ncontrol versus drugged. We're just\ncomparing two different situations.\nLike before, we start off with test\nnumber one.\nBut now we have two different\ndistributions.\nThe black distribution is for our\ncontrol mice.\nThe red distribution is for mice that\nhave been treated with a drug.\nIn this example, the drug increases this\ngene's transcription.\nLike before, we take two samples.\nSince the samples are now coming from\ntwo separate distributions,\nthere's a higher likelihood that the two\nsamples will be separated and not\noverlap.\nWhen we do the statistical test, in this\ncase, we get a P value that equals 0.03.\nAnd then we do the exact same thing for\ntest number two.\nNotice that both of the P values are\nless than 0.05,\nso they're statistically significant.\nSince the samples were taken from two\nseparate distributions,\nthis is what we'd expect.\nLike before, we repeat this process\n10,000 times.\nHere, I've drawn a histogram of the\n10,000 P values generated by testing\nsamples taken from two different\ndistributions.\nMost of the P values are less than 0.05.\nThis is what we'd expect.\nThe P values greater than 0.05 are false\nnegatives from where the samples\noverlapped.\nYou can reduce the number of false\nnegatives by increasing the sample size.\nTo summarize what we know so far,\nwhen the samples come from the same\ndistribution,\nthe P values are uniformly distributed.\nBut when the samples come from different\ndistributions, the P values are heavily\nskewed and closer to zero.\nNow, imagine we're doing an experiment\nwhere we are testing all of the active\ngenes in neuronal cells.\nOne set of neuronal cells is treated\nwith a drug, the other is not.\nThe drug might affect 1,000 genes.\nThe measurements for these genes will\ncome from two different distributions.\nThe black sample is from the control\ncells,\nand the red sample is from the cells\ntreated with the drug.\nSince the samples come from different\ndistributions, the P values are skewed.\nThe remaining 9,000 active genes might\nnot be affected by the drug.\nThis means the measurements for most of\nthe genes will come from the same\ndistribution.\nThe P values for these genes should be\nuniformly distributed.\nThe histogram of P values we obtain from\nall 10,000 genes is the sum of the two\nseparate histograms.\nThe uniformly distributed P values come\nfrom the genes unaffected by the drug.\nThe P values on the left side are a\nmixture from genes affected by the drug\nand genes unaffected by the drug.\nBy eye, we can see where the P values\nare uniformly distributed and determine\nhow many tests are in each bin.\nHere, I've drawn a line indicating that\nabout 450 P values are in each bin in\nthe uniformly distributed part of the\nhistogram.\nWe can extend this line and use it as a\ncutoff to identify the true positives.\nSince we usually use a cutoff of 0.05,\nwe're going to focus on these P values.\nRoughly 450 P values less than 0.05 are\nabove the dotted line.\nAnd roughly 450 P values less than 0.05\nare below the dotted line.\nOne way to isolate the true positives,\ngenes affected by the drug, from the\nfalse positives, would be to only\nconsider the smallest 450 P values.\nThis procedure works fairly well because\nthe P values within the bins are skewed\nfor the genes affected by the drug.\nNote, this histogram is for P values\nbetween 0 and 0.05\nand spread evenly for the genes not\naffected by the drug.\nBam!\nIf you can understand these concepts,\nthen you understand more about false\ndiscovery rates and the\nBenjamini-Hochberg method than most\npeople who use it.\nAll Benjamini and Hochberg did is\nconvert this procedure that we just did\nby eye into a mathematical formula.\nSo, now let's talk about the details of\nthe Benjamini-Hochberg method.\nLike I just said, it's based on the\neyeball method we just saw.\nThe Benjamini-Hochberg\nmethod adjusts P values in a way that\nlimits the number of false positives\nthat are reported as significant.\nAdjust P values means that it makes them\nlarger.\nFor example,\nbefore the false discovery rate\ncorrection,\nyour P value might be 0.04,\ni.e., significant.\nAfter the FDR correction, your P value\nmight be 0.06.\nNo longer significant.\nIf your cutoff for significance is FDR\nvalues less than 0.05,\nthen less than 5% of the significant\nresults will be false positives.\nIn other words,\nthese are the genes with P values less\nthan 0.05.\nThe black box shows the genes with FDR\nmodified P values less than 0.05.\nNotice that not all of the true positive\ngenes are inside the box.\nHowever, only 5% of the modified P\nvalues in the box are false positives.\nThe remaining 95% are true positives.\nWhy don't all of the true positive genes\nhave adjusted FDR P values less than\n0.05?\nBecause not all true positive genes will\nhave super small P values.\nHere's the histogram of true positive P\nvalues less than 0.05.\nThese genes on the right side of the\nhistogram probably won't remain\nsignificant after adjustment.\nSurprisingly, the math behind the\nBenjamini-Hochberg method is simple.\nLet's take a look.\nLet's start with another simple example.\nWe'll take 10 pairs of samples taken\nfrom the same distribution.\nI.E. 10 genes that were not affected by\nthe drug.\nAnd here the P values from those 10\ntests.\nThe first thing we do is we order the P\nvalues from smallest to largest.\nNotice that one of the P values is a\nfalse positive. That is to say, it's\nless than 0.05.\nLet's see what the Benjamini-Hochberg\nmethod does to it.\nThe second step is to rank the P values.\nAnd let's make spaces for the FDR\nadjusted P values that we're going to\ncreate.\nThe largest FDR adjusted P value\nand the largest P value are the same.\nThe next largest adjusted P value\nis the smaller of two options.\nEither the previous adjusted P value,\nwhich in this case equals 0.91,\nor the current P value times the total\nnumber of P values divided by the P\nvalue rank.\nIn this case, the current P value is\n0.81.\nThe total number of P values is 10.\nAnd the P value rank is nine.\nPluging these numbers in, we get 0.90.\nSince we select the smaller option,\nwe're going to go with 0.90.\nFor the next largest adjusted P value,\nwe just repeat step four.\nThat is to say, we select the smaller of\nthese two options.\nPlugging in the numbers gives us the\nchoice of 0.90\nor 0.89.\nSince we use the smaller value, we go\nwith 0.89.\nAnd for the next largest adjusted P\nvalue,\nOkay, you get the idea. We just repeat\nuntil we've adjusted the remaining P\nvalues.\nAnd here I've plugged in the remaining\nadjusted P values.\nThat false positive P value\nis no longer significant.\nHooray!\nNow let's look at a huge example.\nThe blue boxes represent the P values\nfrom when the samples came from two\nseparate distributions.\nThat is to say, these P values are for\ngenes that were affected by the drug.\nI've made these P values relatively\nsmall to reflect the normal skew.\nThe P values in red boxes came from\nsamples taken from the same\ndistribution.\nNote, we've got some false positives.\nThe eyeball method suggests we draw a\nline at the top of the uniformly\ndistributed P values.\nAnd extend it to separate the false\npositives from the true positives.\nThese are the P values that the eyeball\nmethod suggests are true positives.\nNow, let's see what the\nBenjamini-Hochberg method does.\nHere, I've shown you the adjusted P\nvalues.\nThe false positives are now all greater\nthan .05.\nBut these true positives remain less\nthan .05.\nDouble bam.\nHooray, we made it to the end. Tune in\nnext time for another exciting\nStatQuest.",
  "transcript_chars": 13294,
  "ingested_at": "2026-06-18T16:32:44.490507+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": null,
    "like_count": null,
    "channel_id": null,
    "categories": null,
    "tags": null
  }
}