{
  "video_id": "R11ESdfVX64",
  "channel_slug": "machinelearningstreettalk",
  "channel_handle": "machinelearningstreettalk",
  "title": "Why Humans Are Still Powering AI [Sponsored] - Phelim Bradley",
  "duration_seconds": 1460.0,
  "url": "https://www.youtube.com/watch?v=R11ESdfVX64",
  "upload_date": "",
  "transcript": "There's a dirty secret, isn't there, in\nin Silicon Valley and in the tech world\nthat there is, and I don't think people\nrealize the extent to this, there is an\nabsolutely huge importance on human data\nand human expertise and human\nunderstanding and that is completely\nglossed over.\nYeah. I mean, fundamentally, uh,\nartificial intelligence is founded in\nhuman intelligence. uh and I think in\nthe the stack of data algorithms uh and\ncompute I think the the human data\nelement is often the the least spoken\nabout maybe the least least glamorous.\nPeople want to imagine that there's a a\nsimple kind of input output equation,\nbut ultimately there's a a messy\nlayer in the in the stack of human\nbeings who are providing their data to\nyou know either uh label data provide or\nHF post training data and then\nultimately the evaluation and the\nassessment of the the model model\nperformance all ultimately is kind of\nhas an element of of human human data in\nit. I I can't tell the difference\nbetween a doctor and someone who\npretends to be a doctor. The only way\nthat you can actually tell the\ndifference is if you [music] have this\ndeep abstract understanding and you know\nthat they're breaking the rules.\nIt's increasingly clear that Frontier\nmodels as a platform are going to be\ncentralized and controlled by a\nrelatively small number of of players\nwhich at the moment is almost\nexclusively these US tech companies. So\nI think there is a bit of a a wakeup\ncall. What the future will hold is\nbasically a marketplace of intelligence.\nSo in the past we had a marketplace of\num oil for example or electricity.\nIntelligence is going to be the new\ntraded thing.\nI'm the co-founder and CEO of prolific\nand uh prolific is a human data uh\ninfrastructure company. So we make it\neasy for people developing frontier AI\nmodels and running research to get\naccess to trustworthy high-quality\nparticipants for high quality online\ndata collection. Prior to the chat GBT\nmoment, the primary modes of data\ncollection, the people were fairly\nfunible, right? So, you're optimizing\nfor cost and scale, maybe offshore,\nlower cost labor, which I think created\nthis dynamic of human data for AI being\na bit of a a dirty secret.\nLet's talk about your core technology.\nNow, you you've solved an interesting\nproblem that I've tried to solve in the\npast. So, I I I started a company called\nMerge and it was a code review platform\nand it was exactly the same thing. I\nrealize that code review has to be done\nby humans and people talk about\nautomation and the software engineering\nlife cycle. It's mostly It's\nit's actually orchestration. You need\nhumans involved in every single step.\nMost importantly, code review. We had a\nskill matrix and we could learn their\nskill and and we could, you know, a pull\nrequest would come in from one of our\ncustomers and we would dynamically\nassign it to an expert in that field and\nthey would, you know, because you don't\nwant a superficial rubber stamping pull\nrequest. You need to have someone\nactually who understands the code and\ngoes into it. And this this I think is\none of the biggest problems in business\nin general which is that you know we\nhave all of this expertise out there. We\nhave these problems over here. How do we\nmatch them together? How have you done\nthat?\nFirstly I appreciate that you understand\nthe uh the challenges and the complexity\nof of dealing with uh with human data\nand and uh the orchestration and the\nrouting of of tasks to the uh the right\nhuman. I think often people want or\nexpect this data to be uh as simple as\nas calling uh an API and getting a kind\nof a a CI/CD\nuh style response where it's fully fully\nautomated uh easy clean uh and simple uh\nbut the reality is that uh humans are\nare messy uh and uh dealing with with\nhumans especially at the scale uh that\nwe're dealing with them where we have\nhundreds of thousands of of active uh\nparticipants or raiders uh on the the\nplatform is challenging.\nAnd I guess fundamentally the the value\nthat we that we add is the deep kind of\nuh verification uh vetting of of these\nparticipants. So building up a a deep\nprofile um of the the nuances of of uh\neach of our participants and the\nbehavior that they uh provide in in the\ndata collection tasks uh so that we are\nable to route the most appropriate task\nor the right uh uh project to the right\nuh humans. Um I think increasingly uh\nunderstanding how to incentivize these\nparticipants and make sure that the um\nuh we're we're incentivizing the the\nright right behavior and it's kind of a\nwin-win-win dynamic across the uh the\nthree different relationships on the\nplatform. sort of data collector uh us\nuh and the participants and not treating\nthis as a strict supply chain where\nwe're trying to uh commoditize uh or\naggregate the participants and their\ntheir data and get it for the the lowest\nlowest cost. Ultimately we we believe\nthat the highest data quality is\nproduced uh by people who are properly\nincentivized like understand the impact\nof their their work and are are going\nkind of beyond just the financial uh\nfinancial incentives. I mean, obviously,\nwe could use the the Uber analogy, but\nit doesn't quite work because we're\ntalking about very deep expertise here.\nWe're talking about very, very specific\nthings. So, um, you need to find people\nand verify that they have that\nexpertise. You need to check that\nthey're not gaming the system. So, I'm\nnot sure you must have some kind of\noperational analytics where, you know,\nlike this is what a normal behavior\nprofile would look like. And you need to\nincentivize them. And certainly when you\nemploy people in the real world, you\nincentivize them in terms of things like\nautonomy and cultural fit and lots of\nlike humanlike factors. How do you how\ndo you do all of that?\nSo the the onboarding uh of of of the\nparticipants, everyone is kind of ID\nverified, make sure they are who they\nsay they are, they are where they are in\nin in the world. Secondly, there's a\nfeedback loop from the from the\nresearcher. So this is analyzing and\nassessing the kind of QA of of the data\nuh feeding that back into the the model\nand ultimately using that information to\nuh rank participants. So you're not just\nlike selecting for uh an audience but\nwe're able to um uh preferentially\nprovide the participants who are going\nto provide the highest data quality for\nthat uh task uh context. And then I\nwould say the third is the network\nanalysis. So uh looking at this\nparticipant as the our participants um\nas a network uh understanding the\ninterconnections between those and\nfinding you know pockets of behaving\nparticipants or or people who are maybe\ntrying to game the game the system and\nusing that information to kind of like\nfilter filter the pool or or um uh dank\nthose participants uh so that they are\nuh providing uh less data over time or\nultimately removed from the from the\nplatform if uh if if required. And then\nthe other other thing is I think is the\nincentive piece which I think is like\nvery very interesting kind of game\ntheory around where uh from kind of\nbehavioral research you mentioned Dan\nDanny Canaman and Cole we know that when\nthe opportunity arises people tend to\ncheat uh a bit um and especially when\nyou have single shot relationships. So\nthere's the the classic game theory\nshare or or steal. I don't know if you\nyou're familiar with that that\nexperiment.\nNo it's a bit like the prisoners\ndilemma. Yeah, prison. That's the that's\nwhat I'm thinking about. So there\nthere's the uh the game theory idea of\nprisoners where ultimately if you if you\num see that as a single uh shot\nrelationship, a single point in time,\nthe incentive is for both participants\nto uh steal. The things that change is\nchange that into a dynamic where\nboth participants are incentivized to uh\nshare or in our context kind of um not\ncheat or not gain the system is if you\ntreat it as a relationship and you have\nmultiple touch points over a long period\nof time. uh you have high communication\nbetween the two sides uh of the uh of\nthe platform and ultimately you drive\ntowards this kind of like win-win-win uh\n[snorts] dynamic uh rather than a uh\nsingleshot um uh experiment or or or\ndata collection.\nYou might be, you know, measuring how\nlong they spend to take to do tasks. So\nhow how do you sort of bring the the\nhuman component into it?\nThat's a great question. So you're\nprobably familiar with the analogy of\nthe the mechanical uh mechanical Turk.\nOh, yes. Tell the audience about that.\nWe had a a chap who was I think it was a\na robot playing uh chess was a was a\nclassic uh example and it was perceived\nfrom the audience as this was a fully\nautonomous uh process and then behind\nthe scenes ultimately uh it was a human\ncontrolling the the robot. I I think the\nthe analogy of the mechanic truck is is\nsuper uh interesting because people want\nthis process to be to be simple and for\nyou to be able to like fully abstract\nthe humans behind uh an API and for you\nto do be able to call human intelligence\non demand vine API and that is the value\nthat we want to provide uh ultimately\nand we try to abstract away as much\nmessiness and as much complexity uh as\nwe can and I think how we bring the the\nhuman human element back in is by trying\nto get out of the way as a middleman.\nAnd uh our philosophy is to try to build\na direct connection between the people\ncollecting the data um and the\nparticipants providing the data. Uh so\nthey're able to communicate, provide\nfeedback, uh ultimately kind of\nunderstand the impact that their work is\nhaving uh on the um on the data\ncollections whether that's research or\nor model development uh etc. Um so\ngetting this feedback uh mutual mutual\nfeedback across the platform this\npeer-to-p peer uh messaging we think is\nlike crucial to uh the the trust and\nmaking sure that there's a sufficient\namount of uh empathy for the people who\nare providing this um uh this super\nvaluable uh data. Now, I've got a lot of\nexperience with Upwork and I um it's\ngood and bad in a way. I I don't like it\nbecause as we were just saying that\nthere's this huge epistemic history.\nEven if someone is a creative\nprofessional, they've been doing it for\nyears, I still find that quite often\nthey're just unccalibrated and the onus\nis on me to specify what I want to an\ninsane level of detail. And the onus of\ndoing that is often greater than just\nthe the cost. I might as well just do it\nmyself. So, how how do you guys overcome\nthe specification problem? And do the\ntasks that you do on Prolific, do they\ntend to be quite close-ended? What I\nmean by that is is you have specific\noutputs or sometimes are they quite\nambiguous and open-ended in terms of the\noutput?\nYeah, that's a that's a great question.\nI think it ties back to um our obsession\nwith with data quality and I think\nthere's two uh two aspects to to data\nquality is obviously the uh profiling\nthe quality of the uh audience. So how\num what what expertise and and\nspecialism training that the\nparticipants have h and then also the\nquality of the the task uh task design\nor even the specification of the people\nyou're looking for right we'll often get\nrequests for uh PhDs in in biology like\nokay there's there's quite a lot of\nnuance within uh within biology you're\nlooking for uh genetics uh expertise uh\nuh bionformatics uh healthcare etc etc\nuh so So supporting the um the\nresearchers in in specifying uh the\naudience with with sufficient detail. We\nhave a mix of of tasks. So the the\nplatform is fairly use case agnostic. So\nsome of it is is fairly fairly\nself-contained game theory dynamic of of\nmultiple touch points with the same uh\nparticipants tends to lead to kind of a\nrelationship between between the\nresearcher and the uh and and and the\nprincipal which leads to better better\ndata quality. Uh we have many many\nprojects that are long running where the\nfirst uh part of the project is is uh\ntraining or providing that context to\nthe uh the audience uh so that they're\nable to uh learn over time what what um\nuh what what good data quality means for\nuh for this for this project and that\ncan uh um be very very long running. So\nsometimes uh weeks or months of of\nlongitudinal or multi- uh multi-step\nstep data collection.\nVery cool. So essentially you built this\nplatform which gives you human expertise\non demand. What does this mean for the\nfuture of work?\nI think yeah so we've optimized the\nplatform for breadth of audience choice\nalthough you can earn a uh a great um\nkind of side hustle uh on on prolific\ndon't necessarily want to optimize for\nuh prof like very very prof\nprofessionalized uh raiders or uh\nresearch participants right we want to\ntap into uh real world users uh so let's\nsay for example you're looking at\nhealthare workers to evaluate your uh\nmedical chatbot. We want to we want to\ntap into folks who are actively working\nin in the field uh and not people who've\nleft the field and are now kind of\nprofessional uh professional annotators.\nWe're trying to reflect the real world\nand real world users uh as much as uh as\nmuch as possible. Uh so we we absolutely\nsee our platform as as an augmentation\nuh to work necessarily than a uh than a\nreplacement though increasingly this\nwork of human data for for AI uh is\nbeing professionalized. Um\nwhat I like about it is it it increases\nthe market efficiency you know like the\nmarket is all about we have these\neconomic tasks that that are valuable\nand we have these folks over here with\nskills. Uh yeah, absolutely that that is\nsomething we're explicitly thinking\nabout is is uh how do we train folks uh\nwith the skills that are uh useful for\nthese frontier model providers. So for\nexample like how do we take uh a general\num general participant and turn them\ninto a high taste uh evaluator. So it's\nnot necessarily just uh domain expertise\nor domain knowledge which is which is\nvaluable uh but also kind of general\naudience uh who are uh skilled uh in in\num uh in overcoming the kind of typical\nbiases of of preference valuation.\nYou you know like um Uber for example\nthere is a critical mass maybe even\ndating websites is another example like\nyou have to bootstrap it. So Uber\nwouldn't work if there if there was just\nno density of cars in my area. So I I\ndon't know where is it the case that you\nyou you had to bootstrap it and it got\neasier or was it easier at the beginning\nbecause even when you have loads of\nparticipants you might have problems\nwhere there's just not enough work to go\naround because it's very stratified you\nknow for example if you're doing some\ndemographic research and I think you\nsaid you know 6% of people of a major so\nyou'd want to like get 6% of people you\nknow so you might have this sparity\nproblem but by the same token when you\nactually have this scale like do you\nfind that it kind of works better? Yeah,\nwe have the chicken and egg problem of\nof all marketplaces and I think um you\ncan think of this again analogist to the\nlike Uber cities uh problem, right? So\nin in early the I think the analogy of a\nof a city in the Uber context for us\nwould be a um a segment of the of the\naudience or people with a with a\nparticular particular skill set where we\nmight have bootstrapped up to um from an\natomic network to a kind of a scaled\nnetwork. say for example for a US\ngeneral audience uh population or UK\naudience uh population uh but then for\neach uh new um expert or or kind of\nsegment that's in in demand uh we need\nto go through that that kind of atomic\nnetwork um uh scaling scaling that\nnetwork up to a point where we're able\nto provide kind of incident uh uh data\nyou know we we optimize for for data\ncollection kind of in a space of hours\nrather than uh days or uh days or weeks.\nSo ultimately we we look at that kind of\nmarket uh liquidity as a user\nsegmentation uh problem where we might\nhave um uh scale networks for particular\naudiences and then it's like thinking\nabout what what's the next marginal uh\nuser who can kind of incrementally add\nto that uh network the the most value\nand we kind of drive the the growth of\nthe network in in uh in that way.\nThe matching algorithm itself could you\ncould you tell me about that?\nUh that does yeah secret secret source\nfor sure. Yeah, I'd be so fascinated to\nknow because, you know, I'm thinking,\nyou know, maybe there's an analogy to\nGoogle search\nand they have like a page rank algorithm\nand that is an example of this kind of\nsocial graph type metadata because, you\nknow, you you you put a hyperlink to\nanother page, you know, if if you\nactually like that page and then of\ncourse you can build up a ranking from\nthat, but then there's like intrinsic\ncontent type metadata which you know\nobviously you guys have come up with a\nway to like mix all this together.\nYeah, exactly. I think the other analogy\nis like to I think it's they're called\ntwo towers uh algorithms. So, uh, Tik\nTok, Instagram, reels where you have the\nthe context of the, uh, the user who's\nwho's filtering, uh, for the, uh, for\nthe content, uh, and then you have all\nof the choice of of content. I think the\nanalogy here is the, uh, you have the\nthe task context, uh, and then you have\nthe, uh, the human context, and you want\nto be able to, uh, rank the, um, uh, uh,\nrank the the humans so that you're\ngetting the the kind of optimal uh,\npeople floating to the uh, floating to\nthe top. And in a similar way that when\nyou open up uh YouTube or Instagram uh\nreels or Tik Tok uh you get the content\nthat's most relevant for uh for for for\nyou.\nWho controls these AI platforms controls\nquite a lot.\nThe internet as a fully decentralized\nplatform not particularly owned by\nanyone. uh it's increasingly clear that\nAI infrastructure and a frontier models\nas a platform are going to be\ncentralized and controlled by a\nrelatively small number of of players.\nUh and predominantly US players at the\nmoment uh maybe China's playing a role\nuh unfortunately not a very very\nsignificant role uh being played in the\nthe UK and and Europe as a whole right\nnow. I I think I think we're lucky that\nmany of the folks who who work for these\nuh uh global labs even if the capital is\ncoming from the from the US are\ninternational. They have a global\nperspective. Uh and we find when working\nwith our our customers in the frontier\nlabs uh have extremely uh positive\nintent. uh they want these models to be\nglobally useful and reflect the input\nfrom a wide uh variety of uh opinions\nand and subjectivity. Though I think\nthere is obviously a risk if these uh\nmodels do become super intelligent and\nthere is massive uh labor impact that\ncost is going to be felt uh\ninternationally and will be felt by uh\nus here in in the UK and and Europe and\nthe the value of that efficiency is\ngoing to flow to the owners of the\nplatform uh which at the moment is\nalmost exclusively these US tech\ncompanies. So I think there is a bit of\na wakeup call for the for the UK and\nEurope even though there's we're maybe\nmaybe late uh to play a more significant\nrole in the life cycle of of these\nmodels uh whether it's like uh owning\nmore of the the training having uh more\nlocally produced models data centers\nenergy abundance in order to power these\nuh these energy hungry models definitely\nthink uh there's there's room for more\nUK EU dynamism and accelerationism in in\nthis space we are still hiring software\nengineers aggressively including\nincluding junior uh software engineers.\nSome people have a philosophy of junior\nengineers and software engineers won't\nbe won't be hired uh anymore because the\nsenior engineer uh plus AI agents will\nbe 10 times uh more efficient. I think\nthat neglects the fact that these models\nare extremely powerful uh teachers and\nand coaches and junior engineers um much\nmore rapidly become as competent as\nsenior engineers with this co-pilot uh\ntraining to be better better software\ndevelopers. You know when I when I\nstarted Prolific one of the motivations\nfor for starting Prolific was to uh\nlearn about web development and building\na product and uh this was a relatively\nsort of learning process than it would\nbe would be now and I would have I would\nkill for a near super intelligent\nco-pilot in order to uh build more\nproduct uh faster and I think there's a\nvery elastic demand for many of these\nthings like software that uh models are\ngoing to um uh improve our efficiency to\nbuild. There is something special about\nlocal situated human expertise and what\nI see is that this could make the pie\nbigger because assuming that these\nmodels will be limited in how they\nunderstand different domains, we might\nin the future need to have an\noperational loop on top of language\nmodels just so that people can actually\nverify the in you know it might be\nconsequential health advice or something\nlike that and the user doesn't know\nwhether it's whether it's correct or\nnot. So we could have a prolific type\nplug-in system where a user could press\na button and say is this legit?\nThis is definitely where we see\ndirection uh of of travel also like you\nuh AI human interaction or like agent\nhuman uh interaction uh analogist maybe\nto like human computer uh interaction. I\nthink that's like the next uh next phase\nof of research. Yeah, you could you\ncould imagine for example a a deep\nresearch uh style agent going off and\nand doing a longunning task and one of\nthe steps along that workflow is to uh\nget a review from a human expert that's\nrouted to the the most appropriate\nperson for for that for that task. Uh so\nagain I think something that we're\nlooking at at at prolific is uh how do\nwe build these systems where agents and\nhumans can uh collaborate uh and where\nwe can go move beyond this kind of\nrelatively simplistic uh preferencebased\nevaluation in order to build out uh\nsystems which better simulate the\nultimate uh objective that we're trying\nto aim for. Right? We we know that\nmodels are when they have a goal because\nof good heart's law they're very very\ngood at optimizing for that for that\ngoal. So therefore, choosing the right\ngoals is uh increasingly important and\ndeveloping the the tools to effectively\nuh simulate those those goals and get as\nclose to the the true objective as\npossible. What whatever I think you mean\nby by that objective, but I think\nultimately it's real world performance\nfor real world users and like real world\nuh context is ultimately the thing that\nwe're trying to simulate with all of\nthese uh evaluations and and benchmarks.\nWe are still building a marketplace of\nintelligence by creating this link\nbetween all of these human situated\nexperts and the system which matches and\nlearns and and maybe there'd be a middle\nway. Maybe we can we can create some\nautomation against that. But\nfundamentally I believe that we're going\nto need human expertise more than ever.\nI I I think the pie has gotten bigger\nbecause so many people are just they're\ngetting their appetites wetted. They're\nthey're generating videos. They're\nwriting code. They're building\napplications. They're doing all these\nthings they couldn't do before and then\nthey they they hit a brick wall because\nthey realize they don't actually\nunderstand it deeply enough. So they\nneed to bring the experts in. All of\nthese experts are in in in more gainful\nemployment in my opinion than they ever\nwere before and this could create a\nvirtuous cycle. I mean I I think that\nwould that's a positive outlook of the\nfuture.\nWhat do you think?\nI uh yeah I tend to agree. I think this\nlike reminds me of the dynamic between\nuh synthetic data versus human data and\npeople often like kind of frame that as\na eitheror proposition whereas uh I say\nwe we're very bullish both on synthetic\ndata uh and human data and and\nultimately uh augmenting human data and\nhuman expertise is expensive and models\nare able to make that that cheaper uh\nand maybe more effective. But when you\nreduce the the cost of something you you\ntend to increase the uh demand to to\nkind of compensate for uh for that\neffect of often comes up with um these\nuh AI models as uh as as infrastructure.\nuh but I think you'll end up with a with\na kind of a similar effect even if we're\naccelerating uh human data with you know\nin the in the in the context of of post-\ntraining or model evaluation with LLM as\na judge or synthetic data uh because of\nthe explosive demand for these uh models\nuh in the first place even if the\nproportion of human data decreases over\ntime the actual scale and importance of\nthat that data uh is is likely to to\nincrease the Financial Times uh the New\nYork Times YouTube Reddit etc uh they\nget a recurring payment for for their\ndata and it's it's licensed and I think\nyou could also imagine that human\nexpertise uh being licensed in a in a\nsimilar similar way where you get an\nongoing passive incentive for for you uh\nproviding data to to improve the improve\nthe model maybe analogist to like a a\nSpotify right where you have the\nproceeds of all of your subscriptions\ngoing out to all of the people who are\nuh providing music. You could imagine a\nsimilar approach to incentivizing uh\nhuman data in the future as well where\nwhere we get this kind of more ongoing\nincentive to to continually improve\neither like a very personalized model as\nyou mentioned, right? Kind of a digital\ntwin of your expertise. I think that's\nmaybe the the more obvious example or\neven like to improve the infrastructure\ncentralized uh models.\nYeah. [music] It's almost a bit like you\nbuy a solar panel and then you can give\nyou can give energy back to the grid.\nYeah.\nYeah. Similar thing to that. Awesome.",
  "transcript_chars": 25439,
  "ingested_at": "2026-05-12T00:41:55.302252+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 3272,
    "like_count": 72,
    "channel_id": "UCMLtBahI5DMrt0NPvDSoIRQ",
    "categories": [
      "Science & Technology"
    ],
    "tags": []
  }
}