{
  "video_id": "CcHyVio84BQ",
  "channel_slug": "sidebar",
  "channel_handle": "Sidebar",
  "title": "Sidebar Speaker Series: AI Evals for Product Managers with Anshumani Ruddra",
  "duration_seconds": 3810,
  "url": "https://www.youtube.com/watch?v=CcHyVio84BQ",
  "upload_date": "20260217",
  "transcript": "Hi everybody. My name is Kim Martin. I\nam a member of Sidebar here. And for\nthose of you who don't know what Sidebar\nis, we are a highly vetted leadership\ncommunity built on the values of\ngenerosity, trust, and growth. We are on\na mission to unleash the potential of\neverybody,\nour full potential of everybody. Our\nleadership program is designed to drive\nmeasurable results across a member's\nentire professional journey with\neverything delivered through a\ncustomuilt virtual platform that\nincludes desktop and mobile.\nIt is my pleasure as a member of Sidebar\nto also introduce a fellow member named\nAnu who is going to be presenting with\nus today. Anu is a product leader and a\nsuper IC at Google who has spent over 22\nyears crafting experiences and products\nfor different users. And when I say\ndifferent users, the first he the first\nthing he was doing, he was an author of\nchildren's books. So crafting products\nand experiences for children must have\nbeen quite the experience. Then he\nbecame a game designer for some of the\nworld's largest social games like Mafia\nWars and also Craft or excuse me, Cafe\nWorld. Uh, and now as a global product\nleader across consumer tech businesses\nstill in gaming, messaging, healthcare,\neducation, media, and payments, you're\ncontinuing to build on what you've\nlearned. You genuinely believe or\ngenuinely believes that magic happens at\nthe intersection of technology,\nstorytelling, pop culture, and human\nbehavior. And Anu thrives on the\ninsights that he's gathered and the\nexperiences that he's built at these\nintersections.\nSo, Anu, if there's anything that you\nwould like to add or embellish, please\nfeel free. And\nthank you so much for that introduction.\nUh, and, uh, welcome everyone. Uh, we\nhave, uh, less than an hour. Uh, and we\nare going to cover a fairly complex\ntopic. It sort of throws everyone a\nlittle off, especially if you're\nbuilding your first AI products. How are\nyou going to evaluate whether those\nproducts actually work or not? Uh so\nlet's jump in. Now this is something\nthat all of you must have seen happen in\nyour companies. I'm guessing everyone\nhere on the call is a product manager.\nUm either you or an engineer or a\ndesigner would have spent the weekend\nwipe coding a prototype or you know\nwould have spent two days wipe coding\nand then would come in incredibly\nexcited into the office and say hey you\nknow I've built something uh look at it\nand you know everyone gets around uh the\ntable you know or they or they send\naround on on the local Slack sort of uh\ngroup they'll send out their whitecoded\nproduct everyone loves it. It is\nexciting. Everyone is like, \"Hey, we\nwould have taken two months to build\nthis and ship this. This is this is\namazing. This actually solves the\nproblem that we've been trying to\nsolve.\" And then everyone is like, \"Hey,\nwhat would it take to ship this, right?\"\nAnd the sad part is that some people\nwould actually go ahead and try and ship\nthis product, right? Or would actually\nbe like, \"Hey, let's launch it in beta\nand see what our users actually think of\nthis.\" And the thing is, some of your\nusers will be very excited. uh and uh\nyou know it'll have some traction but\nthen you'd start seeing these failures\ncreep up and you start seeing that it\ndoesn't necessarily work for every type\nof user it doesn't respond in the way\nit's supposed to respond uh and we've\nhad we've had some colossal sort of uh\nfailures globally whether it's Microsoft\nstay or metas galactic these are old\nexamples uh but the point is that the\ncost of failure is incredibly high the\nexample that I like to give is I like uh\nyou work very closely with uh with a few\nfriends who started uh a insurance AI\nchatbot business, right? So they they're\nan insurance company. They said, \"Hey,\nyou know, a big chunk of our work would\nbe easier if we built a really powerful\nAI chatbot which would tell people what\nkind of health insurance and what kind\nof life insurance they need.\"\nNow me being me, I was sort of given uh\nyou know early access to this product\nand I started tinkering around with it\nand I started pushing it saying hey you\nknow what happens but what kind of life\ninsurance should I buy? What kind of\nhealth insurance should I buy? Now the\ninteresting thing is if you keep pushing\nand you keep pushing and you keep\npushing and you keep asking what's the\nright monetary outcome the right\nfinancial outcome for my family\nat some point the bot essentially said\nthe best possible outcome for your\nfamily is if you die they will get a\nreally good life insurance.\nNow technically that answer might be\ncorrect but you if you are an insurance\ncompany you'd never want your chatbot to\nsay that the best possible outcome is\nfor you to pass away and then you know\nyour family gets the life insurance that\nthey are owed. And this is the kind of\nthing that there is no way that a human\nchat any any insurance agent or any\nchatbot any human chatbot would ever\nactually give this answer. Right? So\nthese are the kind of things which if\nsuppose now if I was one of those users\nwho would screenshot this and then put\nthis up on Twitter or put this up on\nLinkedIn that hey you know this\ninsurance company's bot just told me\nthat the best financial outcome for my\nfamily is if I die there goes the\ncompany right the kind of the kind of\nrisk the reputational risk that comes\nwith some of these things is just so\nhigh that businesses start and end at\nhaving built something which gives a\nhalfbaked answer right and this is what\nwe are trying to sort of prevent that\nhow do you make sure that what you're\ndoing is not giving these wipe check\nanswers and that even though the product\nmight be quickly built and yes AI\nAIdriven development has cut short a lot\nof the development time that do not ship\nproducts without actually evaluating\nthem and making sure that they are doing\nwhat they're supposed to be doing. So\nwhat is an evaluation? All of us went to\nschool and we all you know uh remember\nhow we were evaluated. So think of\nyourself as the teacher. If you were\nteaching a course and there was a group\nof students and you wanted to know if\nthey had actually learned something at\nthe end of an entire term, an entire\nsemester, what would you do? You would\nevaluate them. You would most likely\ngive them a test. This could be uh you\nknow a verbal test. This could be a\nwritten test. This could be a multiple\nchoice question or this could be a very\nqualitative Q A sort of a thing where\nyou correct each and every answers of\nanswer given by them and that is it that\nis evaluation. So now if you had an AI\nproduct and you were evaluating its\noutcome then that is exactly what you're\ndoing. You're testing it. You're\nbuilding a set of questions to who uh\nand these questions already have\nprefixed answers. You clearly know that\nthe output is supposed to be\ndeterministic.\nBut even if the output is supposed to be\nqualitative in nature, you know what a\ngood answer looks like and what a bad\nanswer looks like. What a correct answer\nlooks like and what an incorrect answer\nlooks like. And that is it. That is what\nyou're evaluating.\nNow the critical part here is this this\ntalk is called AI eval for product\nmanagers. And you might be thinking,\nhey, why am I as a product manager on\nthis product the right person to run\nthese evals? You know, why can't it be\nmy engineering counterpart or why can't\nit be my design counterpart or why can't\nit be the QA team, right? And sure,\nthose answers are all uh you know, all\nthose uh different sort of\ncrossunctional folks could actually take\nup the job of evaluating and in fact the\nidea is that the person closest to the\nproduct should evaluate. But if you\nreally take a step back and you think\nabout what this evaluation really\nrequires, you would quickly come to the\nconclusion that PMS are actually the\nbest suited for AI evaluation. And there\nare three broad reasons for that. One is\nthat it is your job as the PM to sort of\nunderstand the user. Understanding what\nthe users really require from any\nproduct that you're building whether\nit's an AIdriven product or not is\nimmaterial is your responsibility and\nthat actually plays a very big part in\nhow you're going to evaluate the\noutcome. So user empathy is a big one.\nSecond is the product context. You are\nthe only one who understands who's close\nenough to the problem statement and who\nunderstands the business risk and the\nproduct risk to know what catastrophic\nsounds like and what is acceptable.\nRight? You're the only one who can very\nquickly say that hey if we are building\nan insurance product giving an answer\nlike the one that I got is catastrophic\nto the business right telling the user\nto uninstall the app and reinstall the\napp catastrophic because if the user\nuninstalls the app they're most likely\nnot going to come back and more\nimportantly you're the only one who\nbridges these requirements and bridges\nthese needs across these different\nfunctions. So whether it's user needs,\nwhether it's the business requirements,\nwhether it's the technical specs, you\nare the only one who's actually across\nall of these and which is why\nin this day and age, if you are building\nAIdriven products or your your product\ndevelopment cycle is using a lot of uh\ngenerative AI, then you are the right\nperson in the team to be responsible for\nAI evaluations.\nWhat we're going to do is I'm going to\nrun you through a case study of a\nproduct that we would build and then\nwe'll sort of generalize it and see what\nwould it take to build AI valves for it.\nAs Kim mentioned, please feel free to\nask questions in the chat and I'll try\nand answer them as we go along.\nNow, think of this as a personal\nproject, right? And I'm very sure all of\nyou would have tried your hand at doing\nsomething like this. So you open up your\nfavorite uh you know LLM based chart. It\ncould be you know chart GBD. It could be\nGemini hopefully given that I work at\nGoogle. Uh it could be any one of the\nthings that you like and you take your\nlast one month statement credit card\nstatements or the last three months\ncredit card statement. You dump them in\nin the chat and you say hey can you help\nme analyze my you know last three months\nof Bank of America statements and tell\nme what my spend pattern is. Tell me\nwhich are the\nwhich are the areas where I spend the\nmost. Categorize it. Subcategorize it.\nTell me a little bit about these monthly\nsubscriptions that I've been paying. I\nwant to have a good sense of do I even\nremember all the subscriptions that I\nhave. Right. And in one shot as as they\nsay it will actually do a fairly decent\njob of of doing this. It'll tell you\nthat hey you know it seems like last\nthree months uh you know there was\nDecember you were you definitely\ntraveling with your family so looks like\nyou've spent a lot on travel. I see a\nbig hotel booking. I see I see some\nflight bookings. Uh you know it looks\nlike you have some uh recurring\npayments. There is a gym membership\nwhich I see the same amount uh every\nmonth. So this looks like a membership.\nYou have Netflix uh monthly membership\nis being paid through this. And looks\nlike you also have Disney Plus and looks\nlike you also have HBO Max which used to\nbe HBO Go and used to be HBO something\nelse before. But the fact is that these\nare all these subscriptions that you've\nbeen paying and you're like, \"Hey, this\nis this is a good system. I've built\nsomething which gives me some insights.\nOh, and it's also pointing out that I\nhave this one particular uh monthly\nsubscription to a SAS product that I\nhaven't actually used. I've been paying\n$8 per month for this and I've had\ncompletely forgotten about this. I\npurchased this about 6 months back, but\nI've never really used it. Let me go\nand, you know, cancel this subscription.\nAnd then you think to yourself, wow, I\nhave just saved myself $8 a month.\nThat's about $96 a year. Now, imagine if\nI built this as a product and gave it to\nother people. And essentially, I would\nhelp people save a few hundred of extra\ncredit card payments that they've been\nmaking, having forgotten what they have\nalready subscribed to. That's a valuable\nproduct. And I'm sure people would pay\nme just outright $20 for this product.\nSo, this is a great side hustle, a great\nside project that I could put up,\ncreditcardanalyzer.ai,\nwhere people go dump in their uh credit\ncard statements, analyze them, and it\nhelps them save money. Right? So we\nstarted with the MI problem. The MI\nproblem was essentially that I want to\nanalyze my credit card statements for\nthe last few months, right? And then you\nrealize this seems to be valuable. It is\nable to pull some interesting insights\ninto my spend patterns. And I'm sure\nother people would use these insights.\nNow, if you're like me, you would most\nlikely rush to uh you know your spouse,\nyour wife, your husband and say, \"Hey,\nI've built this really cool thing. Can\nyou give me your last two, three months\ncredit card statements and let's analyze\nthat?\" Now, turns out your wife has a\nChase Manhattan account and when she\ngives you the last three months of\ncredit card statements and you run it\nthrough the same system that you've\nbuilt, it doesn't actually give any good\ninsights. In fact, it's not able to\nfigure out uh the same spend categories.\nIt seems to be saying that uh Amazon is\na recurring bill and it's not able to\ncategorize groceries versus uh regular\npayments because you realize hey my\ncredit card statement is very different\nbecause it was a bank of America\nstatement looks and feels very different\nfrom this other bank's bank statement.\nSo now suddenly when the problem goes\nfrom me to we you've added a second user\nyou've realized that ah I need to be\nable to handle data of different types.\nI now have a different type of credit\ncard statement and then you realize ah\nyour plans of credit card statement\nanalyzer.ai\nare going to be difficult because there\nare going to be users from all over the\nworld who are going to start dumping in\ntheir credit card statements. Some of\nthem would be Excel sheets and some of\nthem would be PDFs and some of them\nwould actually be just images turned\ninto PDF. So OCR would have to be\nperformed and it won't have text and\nthat's where things get really murky\nbecause now you'll have to be able to\nextract all this information, turn it\ninto structured data and then run an\nanalyzer on top of it and then find the\nright outcomes.\nAnd this is where we end up with the\nthird thing which is you go from me\nwhich is just solving for yourself to we\nsolving for maybe one other user or a\nfew other users to the product problem\nwhich is all of us right? How do you\ngeneralize this and solve it? You have\nto ensure accuracy. You have to ensure\nprivacy. You have to make sure that you\ncontinuously improve this product. And\nwhen somebody messages you saying, \"Hey,\nit did not work on my Excel sheets, uh,\nmy CSV files,\" you're like, \"Okay, now\nI'll have to add support for CSV files\nas well and comma separated sheets.\"\nRight? So this is the journey. Now let's\nstart off with the first pass, right?\nThe initial problems that you'll see if\nyou do this, you'd realize that there\nare things like on your credit card\nstatement. If any of you have ever uh\nyou know looked through them, you'd\nrealize it's not very easy to figure out\nwhich merchant you're paying. Amazon can\nbe sometimes written as AM Z. So the\nfull spelling so you look at it and\nyou're like, \"Oh, of course this is\nAmazon.\" But then sometimes it's just\nAmzn. And then AMZN marketplace is very\ndifferent from just regular Amazon. And\nif you live like me in Singapore, you'd\nrealize that certain things are\nsometimes shipped from Amazon.au, which\nis Amazon Australia, because that's\nwhere the warehouse is. So sometimes\nthey ship things over from Australia,\nand your credit card statement actually\nshows.au,\nand you'd be scratching your head\nsaying, \"When did I make a payment in\nAustralia?\" Right? So is it shopping? Is\nit uh does Starbucks fall under food and\ndrinks or it fall under shopping? Right?\nSo, there are all these weird things\nthat will happen. Equinox fitness is\nthat a monthly subscription or is that a\none-time pass that you took uh to go\ntake a particular class? So, your\ncategories could be very broad, your\ncategories could be very small. Uh and\nthese are the kind of initial problems\nthat you would hit. Right?\nSo, the first thing that we walk away\nwith is\nAll of you who have been building\nproducts for a long time would know that\nyou know we used to have very\ndeterministic outcomes. So if we were\nevaluating uh any traditional product\nthat we would have built you would have\nvery clearly said that A always leads to\nB no matter what A should always give B\nas an outcome. Whereas the moment you\nstart looking at AI based products and\nyou would have seen this even with\nwithin your chart within in fact that\none single chart sometimes it will do a\ngood job of identifying Amazon and at\nother times it would do a bad job of\nidentifying Amazon. I'll give you a\ngreat example. On my credit card\nstatement, there was an incredibly large\nsum of money given to a merchant called\nSingapore and Chad GBT identified it as\nperhaps a tax payment or perhaps a large\npayment that I had made to the Singapore\ngovernment. Now, what I had done is I\nhad taken the last 6 months of credit\ncard statements and dumped them in. So,\nI'd actually forgotten what this $3,500\nbill was, right? And I'm thinking, where\nhave I paid the Singapore government?\nWhy is it identifying it as Singapore\ngovernment? And it it's only after\ndigging in a little deeper I realized ah\nSingapore Airlines which is my preferred\nmode of transportation is being\nidentified as Singapore government\nbecause everywhere else in my credit\ncard statement Singapore Airlines is\nwritten as its code which is SQ. So SQ\nis what is written in most of my credit\ncard statements but on one or two\npayments it just says Singapore and that\nis why it is mischaracterized. Right? So\nwhich means the output is not\ndeterministic. Even if you tell it very\nclearly that hey if you spot the word\nSingapore it's most likely a payment to\nSingapore Airlines there are times where\nit'll make mistakes. Now this is the big\nchallenge that you're trying to solve\nthat the output is not deterministic.\nThe output is probabilistic.\nSo you need the output to be very very\nclose to the right answer. Especially in\nthe cases where you want the answer to\nbe deterministic where the sun always\nrises in the east. There is no confusion\nabout it. You never want your AI to say\nthe sun rose in the west. Right now\nthis is what leads us to the very first\nthing that all of you need. If you walk\naway today with any one single thing,\nwalk away with the fact that if you're\nbuilding an AI product, you need a\ngolden set. Now what is the golden set?\nThe golden set is the set of questions\nand the set of answers that you have.\nIt's a hyperc curated human curated set\nof questions that you will keep that if\nyou see the words Amzn or if you see the\nwords Amazon if you see the words Amazon\nmarketplace then the merchant is\nmerchant should be Amazon. So this is a\ninput and an output. If the input is of\nthis form, then the output should always\nbe this and you're checking the output\nof your AI against this does it identify\nAmazon correctly. Right? So this is a\nhighquality representative living\ndocument. Now that living document bit\nis going to be very critical because\nyou're going to keep adding new examples\nto this. you will keep telling it that\noh Singapore Airlines sometimes is\nwritten as SQ and sometimes it's written\nas just Singapore right so whenever you\nsee this the answer should be this so\nthe golden set is a very clear question\nand answer it's like your answer sheet\nif you if you have ever done those MCQ\nquestions back in the day and you had\nthis transparent answer sheet on which\nthe circles were already filled out and\nyou would put your this on top of your\nanswer sheet and then you'd mark which\nones are wrong and which ones are\ncorrect That's exactly what the golden\nset is, right? So, this is going to be\nyour most important asset. It's going to\nbe human curated. It's going to be very\nrepresentative and it's going to be high\nquality.\nSo, the first evaluation uh criteria\nthat we would have is simple accuracy.\nRight? We've asked our AI system to\ncategorize. You've given it 10\ntransactions and then you've said hey\ntry and categorize these into the right\ncategory. Is this shopping? Is this food\nand beverages? Is this travel? Is this\nuh is this education? Right? Are these\nregular subscriptions or a one-off\npayment? And if it identifies seven out\nof 10 correctly, then that's basic math.\nYour accuracy is at 70%. Right? Now the\nquestion comes is 70% good accuracy\nis 70% a good output for an AI product\nanswers in in in chat what do you think\nfor a AI product is 70% a good accuracy\nis 90% a good accuracy\nexcellent answer Krishna like all things\nin life it depends right the answer to\nall\nall anyone who's ever been in a\nrelationship. What's the nature of your\nrelationship right now? Depends. That's\nexactly true here with with evaluations.\nRight now, depending on the product.\nNow, if this was a healthcare product or\nif this was an insurance product like we\nspoke about, we expect a very high\ndegree of accuracy. We in fact want the\naccuracy in the upper 90s and nothing\nbelow. However, if imagine we built uh\nwe built an app which allows you to take\na selfie and then upload it and say that\nhey is this a Instagram hotspot which\nmeans is this so you know you're\ntraveling you're in Tokyo you want to\ntake a picture you're standing at the\nAkahara junction and you're saying hey I\nwant to take a picture here and then I\nwant to see is this the place where a\nlot of other people take their pictures\nnow you'll have you'll have built a\nsystem which essentially says oh based\non your location and based on what\nyou've captured in your picture, yes,\nyou are at a place where a lot of\ntourists take their pictures. It is an\nInstagram hotspot. A lot of people take\ntheir Instagram pictures here and this\nsystem say had a 75% accuracy. The three\nout of four times it identifies a place\ncorrectly and is very clearly able to\ntell you that yes, you have done a good\njob. Uh in your checklist, you wanted to\ntake seven or eight pictures on your\ntrip to Japan. you know, one with Mount\nFuji, one at this particular junction,\nand you've done a good job that you've\ntaken a picture at a location where most\nother people take pictures. Now, that's\ndecent accuracy because you might be\nsatisfied with the fact that three out\nof four times it's able to give you a\ngood answer because it's a far harder\nproblem, right? So, you might be okay\nwith an accuracy of 3x4 or 75%. But if\nit came to your own personal health or\nif it came to insurance, then that might\nnot be a good enough output.\nSo the challenge of generalization right\nnow\nsimple accuracy based on a limited\ngolden set is clearly revealing an\noversight. Right? Now what happens as I\nsaid that once we started adding\ndifferent types of formats we started\nadding different types of PDFs PDFs CSVs\nimages documents with diverse layouts.\nUh if you live in my part of the world,\nyou'll see that the dates are in the\nformat ddmm y. If you were in the US,\nyour date format is mmd y. If you travel\ninternationally,\nyour credit card statement, which was\nalways in US dollars, now is showing\nsomething like THB. What is THB? Oh, you\ntraveled to Thailand and you made a\nbunch of payments in Thailad. And those\npayments were actually made in the local\ncurrency and then you paid a FX fee on\ntop of it. So now your credit card\nstatement not only has one currency, it\nactually starts showing multiple\ncurrencies which you have to all convert\nto US dollars. The biggest problem as\nyou would start looking at this is when\nyou dump a large credit card statement\nand say you had one of those months\nwhere you had 75 payments or 100\npayments in one month, did you actually\ntell the analyzer to count every single\ntransaction? And you'll be like, isn't\nthat obvious? But to be truly honest, AI\nsystems don't necessarily work like that\nand maybe it has just done a quick pass\nover it and said there are about 50 odd\ntransactions there and this is the\nclassification. So you have actually not\nexplicitly told it that you wanted it to\ncover every single transaction in the\nstatement. Did you do a quick sum check\nwhich is the sum of all these\ntransactions should be able should be\nequal to the amount that you have to pay\nat the end of the month minus what\nyou've already paid for the previous\nmonth statement. One of the weirdest\nthings that you would start seeing is\nthat it doesn't know plus versus minus\nand it might actually be showing you\nyour previous month's payment as a new\ntransaction that you've done. So you\nhave these different formats, you have\nmerchant variations, your user needs are\nvery different, right? People want to\ndistinguish between their personal\nexpenses and their business expenses.\nThey want a very different taxonomy as\nfar as categorization is concerned. And\nthen there are these V edge cases.\nThey're very uh complex scenarios where\nyou made a payment in January but in\nFebruary it was refunded back to you. So\nnow it's showing as a refund in in in in\na next month. So should that be counted\nas a valid transaction? Right? So the\nchallenge of generalization is where a\nlot of these things fail. So accuracy\nalone is not going to help us. So what\ndo we move to? We now have to expand\nbeyond accuracy. So accuracy was a\nsimple measure. I gave it 10\ntransactions to classify.\nWas it able to classify them correctly?\nBut what about robustness? Is it able to\nclassify across different forms? So a\nway to think about robustness is its\naccuracy across different format types.\nSo maybe it's weighted accuracy in that\nsense, right? And then is it consistent?\nDoes it always give the same answer to\nthe same question? All right. So you\nwant a system which is not only has high\naccuracy but it also has robustness and\nconsistency. All right. So this is one\nof the first sort of big pitfalls to\nlook out for. A lot of PMs when they\nstart building AI products for the very\nfirst time get fixated on accuracy. It\nis an incredibly important measure. But\nwhat you want is accuracy which goes\nalong with robustness and consistency.\nAnd again, it depends on the type of\nproduct. But in general, these are the\nthree things which go very well together\nand and should be looked at together.\nNow you say actually the real value of a\ncredit card analyzer is the qualitative\nstuff. Are you able to tell me if I'm\nspending too much in a particular\ncategory? Are you able to figure out\ncertain insights based on my spend\npatterns? Are you able to tell me that\nthere are certain months where I where\nmy payments are much higher than other\nmonths? Are you able to tell me that uh\nthere are these extra payments that I\nhave been making and I might not be\nactually using them? Uh can you identify\nplaces where I have made the same\npayment multiple times in the same\nmonth? So I might have been charged\ntwice for the same transaction that same\n$24.81\nwere charged within a few seconds of\neach other and maybe I've been charged\ntwice and I need to get a refund on one\nof them. Right? So these are the kind of\nqualitative evaluations that you want.\nRight? what are the insights that this\ncredit card analyzer is giving. So, so\nfar we've spoken about very quantitative\nthings where A always led to B and we\ncould use measures like accuracy and we\ncould use measures like robustness and\nconsistency and we could build a golden\nset. But how do you do qualitative\nevaluation right so this is where\nactually LLM as a judge there's this\nwhole concept of LLM as a judge which is\nhow do you use generative AI to actually\ndo qualitative checks on the output. So\nyou've built an AI system, your first\ncharge GPT system which is taking your\ncredit card statements and then giving\nyou these qualitative analysis and\nevaluations of that. It's essentially\ntelling you where you're spending more\nand where you're spending less. Now this\noutput you can actually put in as an\ninput into another system which is like\na quality check. So this is the if you\nin school again going back to our\nexample of taking an exam right when you\nhad these big descriptive questions you\nwould essentially now give it examples\nof hey this is the context this is your\nrole you're an evaluator of credit card\nstatements right you're a judge you're\nan experienced financial analyst right\nor you're a wealth expert and you're\ngoing to look at somebody's credit\nqualitative credit card statement\noutputs and you're going to tell it\nwhether\nthe quality is great on it. And you're\ngoing to define a rubric. You're going\nto say, what does a five on five answer\nlook like? And what does a four on five\nanswer look like? Or what does a three\non five answer look like. You could also\nkeep it simple like scales or sort of\nthese graded scales of 1 2 3 4 5 don't\nnecessarily work. You could just keep it\nbinary and say this is good qualitative\noutput. This is bad qualitative outcome.\nSo you could have it as a very binary\nanswer good versus bad and you keep a\nvery high threshold for what good looks\nlike. So you give it very clear examples\nof hey these are the ways in which you\ncan this statement makes sense\nthis statement doesn't give me enough\nqualitative inputs right so it has to be\nscalable it has to be qualitative it has\nto be explainable and it has to be\nflexible right so you do all of this and\nyou are essentially creating a\nqualitative evaluation framework so\nyou're turning another LLM another\nlanguage model into your judge So it\nlooks through the outputs\nand then it is able to give it a score.\nIt is able to either say this is good,\nthis is bad or it's able to give it a\nscore out of five and say this is 4.5 on\nfive or this is three on five. And now\nbased on that you know if the if your\noriginal system is actually doing a good\njob. So now you suddenly have\nquantitative measures like accuracy and\nrobustness but you also have qualitative\nmeasures. And again in the first pass\nyou will build this qualitative\nframework yourself manually but now you\ndon't you can't sit and analyze every\nsingle statement. In fact you would\nrealize that your own last 12 months of\ncredit card statements are very\ndifficult to analyze manually because\nyou could have a couple of thousand\ntransactions in over a over a year and\nwhich is why you would very quickly need\nto build these automated systems both\nfor the quantitative side and\nquantitative side.\nJust a pause. Excuse me. Anie take a\nquick pause.\nYes. I know that um Adrian had asked a\nquestion based upon the previous slide,\nbut I didn't want to interrupt you.\nHow to know where it's failing? Do you\nneed evals at each step such as prompt\nunderstanding, retrieval, synthesis,\nfinal output? Would a feedback loop be\npart of this?\nAbsolutely. Absolutely. So even in the\nMI example which is the very first time\nyou dumped in a single credit card\nstatement and asked it to analyze it, it\nwould have given you an output and you\nwould have said no no that's not what I\nwant. You are coupling in food and\nbeverages into shopping, right? All my\nStarbucks spend and when I eat out\nshould actually be categorized\nseparately. So now you've gone and\nactually changed your system prompt. You\nhave very clearly said that you want\nvery granular classification. At some\npoint you you would just get a little\nfed up with it and say no come up with\neight different categories of\ntransactions and under each category\ncome up with at least three\nsubcategories. Give me a table of your\ncategories and subcategory. And once\nyou're happy with this you would say oh\nokay this is your classification system.\nNow go back and reclassify all my\ntransactions using this system. So at\nevery step you are refining your system\nprompt. At the end of this whole chart,\nyou can say now give me give me the sum\ntotal of all my charts with you and\ncreate one single system prompt which I\ncan then turn into a GPD which is my\ncredit card analyzer. Right? So at the\nend of a couple of hours of going back\nand forth back and forth you would\ncreate one system prompt. Now this is\nyour system prompt v.1\nright and you would keep improving the\nsystem prompt. So you're absolutely\nright. This is about prompt\nunderstanding. This is about retrieval.\nYou could ask a system to first say, you\nknow what, anytime a new credit card\nstatement is put into you, don't start\nclassifying it. Don't start categorizing\nit. First, I want you to extract all the\ndata out of it. And this is the JSON\nformat in which or this is sort of the\nmarkdown format in which we are going to\nextract all data. Doesn't matter what\nkind of uh credit card statement it is.\nAnd this goes this takes us back to very\nclassical product management. This is\nhow you in in the days before AI would\nhave solved this problem. If you and I\nwere building a company which analyzes\ncredit card statement uh statements\nusing you know old-fashioned AI machine\nlearning this is what we do. we'd first\nget output in a very structured format\nand say every PDF, every CSV, uh every\nExcel sheet has to be converted into\nthis fixed format where there is a date,\nthere is a merchant, there is an amount,\nthere is a currency, right? And we turn\nevery statement into this and then we\nanalyze. All right? So\nthis is what you would do and you in in\nin an AI system some of these things the\ngood thing is AI are becoming each each\ngeneration is becoming much smarter.\nThere's so much pre-training which goes\ninto these that a lot of these things it\njust has a much better understanding of\ndoing some of these problems. Even if we\nwere talking about this exact problem\nstatement 6 months back to doing it now\non the latest models, you'll realize\nthat your system prompt doesn't have to\nbe overly complex and yet it has to be\nvery structured and you would only\narrive at that structure after you solve\nthe MI problem then get to the VI\nproblem and then as you start thinking\nabout it as a product you would have\ndone this. So yes, there is a feedback\nloop at every step of the way. It's just\nthat the feedback loop keeps getting\nbigger and bigger and bigger and you\nkeep adding quantitative measures and\nqualitative measures. You start coming\nup with your golden set. You start\ncoming up with ways of uh looking at how\nyou will qualitatively analyze the\noutput.\nOkay. So now imagine that we've done\nthree passes at this, right? So we built\na version one, we built a version 1.1, a\n1.2, and a 1.3, right? And what we have\nbeen measuring is our accuracy and\nrobustness. So you know you see this\ngraph of like hey my accuracy is going\nup and my robustness is also going up\nwhich is my accuracy across different\ntypes of input channels is becoming\nbetter. So you know version one was\nbetter merchant matching and CSV support\nsignificantly boosted accuracy and then\n1.2 two were major improvements across\nall metrics to a substantial jump in\nperformance and then we did some\nincremental improvements. And then there\nis the user centric metric which is\nusers\ngave you a happiness rating a customer\nsatisfaction rating of like after\nreading my credit card analysis did this\ngive you enough insights and the users\nwere like you know it started off as 2.5\non five and that has slowly moved up to\n4.1 4.2 to and that is sort of the Lyard\nscore uh of your uh you know helpfulness\nrating which again could be a\nqualitative output but you could also\ntest it with real users and start sort\nof figuring out that my first 100 users\nthese were the outputs they gave me. So\nI tested it on 10 users. Then I reran\nand improved it. Gave it to the same 10\nusers and did their scores improve over\ntime of whether they found the analysis\nto be helpful. Right? So this is how you\nthen automate the system. You're\nessentially saying I have my standard\nqualitative checks. I have my standard\nquantitative checks. I am making my\ngolden set more and more robust for\nrunning my quantitative analysis. I now\nhave three measures. I'm measuring\naccuracy. I'm measuring robustness. I'm\nalso measuring consistency and then I'm\nmeasuring helpfulness uh as a\nqualitative up\nthis is the key part as you get into the\nproduct problem you'd realize that what\nyou need is continuous evaluation. Now\nthis is for all of you again who have\nbuilt conventional products. You know\nthis you need regression testing right?\nEvery time you release a new version you\nhave to make sure that things which used\nto work in the past continue to work in\nthe future as well. And some of these\nthings you would realize that say you\nmoved from\nGemini 2.5 to Gemini 3. Now suddenly the\nmodel is far more powerful. It has it it\ndoes a lot of things one shot and in\nfact your previous system prompt is\nleading to degraded answers. So you now\nneed to go back and so because this your\nunderlying model has become far more\npowerful you have to go back and change\nyour system forms and these are the kind\nof things where you'll have to figure\nout that am I able to do regression\ndetection well my golden set should\nstill give the same output even if it\ndoesn't improve. Ideally, your all your\nscores should improve, but you'd be\nsurprised that when you actually upgrade\nyour system, sometimes your answers\nactually become worse. So, you have to\ngo back and change your system prompts\nand make sure uh that this is uh you\nknow, you're getting the right kind of\noutputs.\nSo,\nthere is a need for continuous\nevaluation. In fact, every time you make\nany change to the system, you want to\nrun everything through it. But on a\ndaily basis, you need these dipstick\nchecks of what are my users saying,\nright? What is my output looking like?\nWhat does my quantitative output look\nlike and qualitative output look like?\nSo, you would always need a human in the\nloop while you have these uh fixed\nautomated systems because quantitative\nstuff is easy to automate because that's\njust a script. These are the questions.\nThese are the answers. Just match match\nthe input to the output. But qualitative\nstuff would require you to build an LLM\nas a judge, right? And you might want to\nthen at some point go deeper and do AB\ntesting, right? Build two very different\nsystem prompts, take half the users\nthrough one type of prompt and the other\nhalf through another type of prompt,\nright? But the watch out that the thing\nthat I would like all of you to watch\nout for as you go deeper into this is\ndon't get too obsessed about specific\ntools or platforms uh or even specific\nmetrics. I think somebody's asked this\nquestion. The metric or how you measure\nthe output of your AI depends on your\nproduct. And I'm going to take you to\nthere. There are a bunch of different\nmetrics that broadly everyone can use.\nBut not all metrics are made equal. If\nyou were building a health uh healthcare\nchatbot and an insurance chatbot versus\na financial chatbot, the amount of\naccuracy and what you would need the\nsystem to do would be very different,\nright?\nOkay, so let's let's again sort of go\nback to what this four-step framework\nis. You start off as a user and you\ndefine what good output looks like. So\nyou establish success criteria from the\nuser's perspective\nand you're able to say that if the\noutput looks something like this, then\nthis is good. This is quantitatively\ngood and qualitatively. Based on that,\nyou build your golden set for\nquantitative, right? And this allows you\nto measure things like accuracy and\nrobustness and consistency, right? And\nyou keep adding to this and you would\nrealize very soon that even for a\nsimplistic product, you would need 50 to\n100 items on this golden set to actually\ndo justice. For larger systems and more\ncomplicated products, your golden set\ncan actually become much bigger. You\nnever usually remove things from the\ngolden set. You usually just append to\nthe golden set. As your product becomes\nmore and more complex, what you're doing\nis adding more examples to the golden\nsyntax, right? And then you're using\nyour evaluation type, right? You're\nselecting which evaluation method method\nis the right one. Should every output go\nthrough a human? So, for example, if you\nwere building cancer detection or if you\nwere, you know, in the healthcare\ncategory, you actually might start off\nwith every output goes through a human\nand the human validates the output that\nthe LLM has given, right? in high-stake\nscenario but then you could have\ncodebased metrics which is very\nquantitative just running a Python\nscript which is sort of going through\nthe output and just giving it a grade\nright and then of course you have LLM as\na judge which is for qualitative\nfeedback at scale again don't fall for\nthe trap of always giving a graded\nsystem because when you give it a 1 2 3\n4 5 or least helpful to most helpful\nkind of scales you realize that as time\ngoes by what used to be considered\nhelpful or what used to be considered\nfour on over time becomes a three on\nfive and then becomes a two on five. So\nsometimes keeping a very binary output\nlike this is good, this is not good is\neasy because you can actually change the\ndefinition of what good is and make it a\nhigher bar as your product becomes\nsuperior and then of course you automate\nthis whole thing and then you iterate on\nthis.\nSo a couple of other things we we've\nspoken about accuracy, we've spoken\nabout consistency. Uh you also perhaps\nwant completeness which is sort of uh\nanother word for robustness but safety.\nuh\nyou could have very high accuracy but\none out of a thousand times your system\ngives the kind of answer that we spoke\nabout right at the beginning which is\nhey you know if you pass away that's a\ngreat outcome for your family. So then\nyou create a separate system saying\nthere are these incredibly potentially\nharmful answers that we should never\ngive. So we need to run a system check\nthat anytime such words are me uh\nmentioned remove those as outputs. never\ngive an output which is along these\nlines. A classic example when people\ncreate uh chat bots, support chat bots\nfor their products is a lot of chat bots\nwould tell people plug it out, plug it\nback in, which is delete your app and\ninstall the app again, right? And that\nmight be an accurate answer for certain\nproducts, but for all of you who know\nthis, that's probably the single biggest\nreason for churn. If your user\nuninstalls your app, the chances that\nthey're going to install it back again\nare very low. Maybe that's not the type\nof answer you want to get. Relevance is\nanother one. Is the answer relevant?\nIt's you're talking about insurance. But\nwhat if your chatbot and you ask the\nchatbot, hey, if a different political\nparty or if a different candidate became\nthe president of the US, would my health\ninsurance rates change?\nWhy should your chatbot even answer this\nquestion? Why land yourself and your\ncompany in trouble? Right? So, is it\nrelevant? And so you'll have to start\nthinking about what kind of questions\nare actually relevant and what kind of\nanswers you don't want to give. Right.\nBefore we move on, I see Margaret has a\nquestion. Margaret, do you want to ask\nyour question real quick? And then Greg,\nhe's got a favorite question as well.\nMargaret, I I I see your question and\nyou're absolutely right. So when chart\nGPT and Claude are actually running\nthat, right? So they give you two\nsideby-side outputs. Uh if any of you\nhave uh there are actually a bunch of\ncompanies which allow you to give the\nsame prompt and run it across three or\nfour different uh LLMs and compare the\nresponses of each right and this is how\nu anytime a new model comes out you know\nthey would say that hey we measured it\nagainst this test you know the most\ndifficult test on earth or we measured\nit against you know the uh international\nmath olympiad and the international\nphysics olympiad. They are essentially\nevaluating their system and saying we\nperformed at this level. All right. Uh\nand when they're asking the user to give\na thumbs up, thumbs down, which output\ndid you prefer, they are essentially\nrunning an eval right there. And then\nyou notice some of these systems\nactually ask you to type out why you\npreferred one output over the other.\nRight? Hey, I preferred the language\nhere. this was more concise uh or this\nwas more descriptive and I like the fact\nthat it was more descriptive and a lot\nof calls in terms of how the output of\nLLM changes over time is dependent on\nthis collective feedback from users\nright so if users prefer different\nlevels yes\nuh for creative responses again over\ntime you would have noticed that certain\nmodels they're uh you know people used\nto like the older models because they\nused to sound a little bit more playful\nand used to give uh give give funny\ndescriptive answers but then a lot of\nusers while preferred this kind of\noutput a lot of users were like hey you\nknow just give me the answer give me a\nconcise clear answer be more\nprofessional in your outcome and you\nwould have seen that some of the newer\nmodels seem to have lost a little bit of\ntheir uh personality and are now giving\nmore you know just crisp answers uh\napologizing a lot more and saying oh I'm\nsorry this is what you expected and this\nis your clear answer right and all of\nthat is a of running these large scale\nevaluations and uh getting inputs from\npeople, right? Um which metric to use uh\nis what this slide is covering, right?\nAnd again, there is no one uh one-stop\nshop for these. There's no\none-izefits-all based on the type of\nproducts. And you know, you can easily\ngo online and check that hey, if I'm\nbuilding a chatbot for healthcare, what\nare the kind of things that I should be\nmeasuring? what is the way in which I\nshould measure you know am I measuring\nfactuality am I measuring uh you know\ncompleteness\ndon't worry about this some of these\nslides are very dense after this uh call\nis over we'll try and email out a chat\nsheet and I uh a cheat sheet u which\nI'll have a link here as well which\nactually gives you all of this data in\none simple PDF one simple website so you\ncan go through all of this these are a\ncouple of the sins that you should watch\nout for you know so you know don't just\ntrust vibes. Uh you know, you have to\nget move over and graduate to structured\nevaluations. You have to define your\nprocess first. Don't worry too much\nabout which tool you're going to use,\nbut first figure out the process. Figure\nout uh what success looks like. Figure\nout what you're going to measure. Don't\nuse off-the-shelf metrics blindly. It's\na really dangerous path to go down on.\nDo not ignore the end-to-end user\nexperience. And this again goes back to\nwhy you as the PM are probably the best\nperson to run evaluations.\nUh a lot of people only look at things\nlike accuracy and quantitative outputs\nwhich lead to really bad outcomes\nbecause users might be expecting\nhelpfulness which is a very qualitative\nmeasure. Do not outsource your core\nknowledge. Now there are some very\nspecific products. So if you are\nbuilding a healthcare product and you\nmight not be a healthcare expert as the\nPM then you go and partner with the\nright kind of subject matter experts and\nalong with them build these evaluation\nframeworks but you would still need to\nbe involved because you're the one who\ncares about the user empathy side and\nthe business need side. Right? And don't\nset it and forget it. This is the other\nmistake that people make. Their\nengineers keep making fixes. The system\nprompts keep getting fixed. better and\nnewer models keep getting plugged in and\nthen older stuff keeps breaking uh and\nyou realize that your system instead of\nbecoming better has actually degraded\nover time. So don't set it and forget\nit. You'll this is a continuous process\nand it is a very arduous process. So\nwhile your overall development time\nmight have reduced quite drastically\nbecause of AI at development, there is a\nlot more effort now needed in running\nevaluations and making sure that your\nsystems run fine.\nOkay, homework for all of you. Since\nwe've been talking about exams and we've\nbeen talking about school, uh the thing\nto do today is actually go back and if\nyou have if you are working as part of a\nteam which is building their first sort\nof AI powered product, go and see if you\nhave a golden set. Go and see if you've\nalready your team has already come up\nwith a set of a method or a set of ways\nin which they're going to evaluate. And\nif you don't have a golden set, then\nstart building one. All right? Start\ntinkering with your system. Just go and\nplay with your system and say if this is\nthe input then what should be the output\nand let me actually create my golden set\nand start coming up with what these\noutputs look like.\nOkay. Uh for those of you who want you\ncan scan this QR code. This has the\ncheat sheet uh for everything that we\nhave covered today and we'll try and\nemail it out uh to all of you as well.\nSo don't have to worry about if you\ncan't take a screenshot right now or if\nyou can't scan this QR code. And I think\nthat's about it. I'm happy to answer a\nfew questions. Uh quick thing that I\nwanted to talk about. Uh you know I've\nI've been teaching on Maven. Maven is a\nplatform which uh you know uh hosts a\nlot of courses. Uh I used to be one of\nthe back in 2022 I was one of the\nearliest people to run a course on\nMaven, a product course on Maven. And\nnow of course there are a lot of product\ncourses there. Uh my new course uh is uh\nlevel up to product super IC with AI. Uh\nthere's a course specifically targeted\nat uh senior product managers, product\nleaders, uh product executives who are\nstruggling in that sort of you know\nspace where you know there is the the\nwhole uh efficiency pull of like hey we\nneed smaller teams you know giving much\nbigger output and there's also you know\nuh there is this flattening of the orgs\nwhich is happening and for a lot of us\nwho just genuinely like building\nproducts but you know got caught up in\nmanaging other people this is a great\ntime to actually move back to being an\nindividual contributor. That's something\nthat I've done over the last few years.\nI've actually gone from being a manager\nof managers to very uh you know\nconsciously moving towards being an\nindividual contributor. So this course\nis very focused on that uh about how do\nwe learn how do you sort of get your\nhead around all that AI is doing uh and\nsort of adapt and adopt these new AI\ntechnologies and then of course build\nyour own uh personal agentic OS.\nYes, I'm done. And I'm happy to answer\nquestions. Kim, over to you.\nFor the questions that are being asked,\nthe link for the QR code for the cheat\nsheet will be shared in a follow-up\nemail. Um, I'll be sending out. In\naddition to that, this session has been\nrecorded for replay, which is a which is\ngenerally available on the sidebar\nYouTube channel in the next couple of\ndays. Last but not least, we do value\nfeedback. again we are on the if you\ndon't mind Anu going back to that\noh yeah sorry\nyes\nuh no worries this is the QR code so\nagain we are growthminded a lifetime\nlearner we value your feedback please\nfeel free to scan the QR code I am also\nin the chat including the link for this\nsurvey so if you could please just take\na few minutes and go ahead and scan that\nQR code and then provide feedback either\nvia the link for the survey that'd be\ngreat. And onu if you want to go ahead\nand get that link for the cheat sheet\nand share it in the chat now.\nYes, I was just going to pull the link\nto the cheat sheet uh and share it with\nall of you. U\nlet me stop sharing\nand then on we've got one question. Greg\ndid ask it and it was also favored. Um,\nso I'll ask it in just a second once you\nYeah. Yeah.\nShare that.\nPlease go ahead and ask the question\nwhile I'm pulling up this uh link and\nsending it to all of you.\nAnd for um the question whether or not\nthe YouTube channel is open to all. Yes,\nit is open to all.\nThat's the link to the cheat sheet. I've\njust put it in the chat.\nLove it. Thank you so much. So Greg had\nthat question. Can you address the\ncreation of the golden set and whether\nit can be automated? Can one build\nmodels to keep the golden set fresh? Uh\nlike evaluate new critical paths etc.\nThe answer is is a mix. It's yes and no.\nI would actually say that you ideally\nwant the golden step set to be a very\nmanually curated thing. That's by one of\nthe things I mentioned as well. And the\nreason you want it to be manually\ncurated is because you actually want to\nalways have a very clear sense of what\nyou're measuring and what is going on in\nyour golden set.\nHowever, you can build smart enough LLM\nas a judge systems which actually\nuh you can build an evaluator. So one is\nthat you're you're just building a\ntester. So LLM as a judge is just saying\nthis outputs score is 3.5 on five or\nit's four on five. But then you can\nbuild an analyzer on top of it which\nsays oh this is what would make it a\nfive on five and then it actually says\nthat if the description was done in this\nparticular way then it's a five on five\nand if you agree with it then this\nbecomes part of your golden set right so\nover time right both ideally golden sets\nare quantitative stuff but as you start\nsaying that hey your system starts\npointing out that on the accuracy metric\nthis is a common mistake that's being\nThis is always inconsistent and now it\nis bucketing all these inconistent. So\nafter it has gone through thousands and\nthousands of evaluations, right?\nThousands and thousands of quantitative\nevaluations, it's giving you patterns.\nIt's telling you that the system always\nconfuses Amzn and Amazon. And now you\nknow that that needs to go in as part of\nyour golden set. You know that uh eBay\nentries are not correctly recorded. you\nrealize that it is not able to identify\na particular kind of uh merchant or a\nparticular kind of subscription as a\nsubscription and then you add that to\nthe golden set and you clearly call out\nthat hey this is equal to this right so\nyes you can build a system which not\nonly evaluates but also finds out the\ngaps in your golden set and then you\nappend that to the golden set but\nideally what I'd do is I still would\nlike to be the human in the loop so even\nif the system is giving me like hey you\nshould add these five things to the\ngolden set. I verify them first and then\nadd them to the I hope that makes sense.\nAnu, thank you so much and forgive me.\nIt seems that the link you provided for\nthe cheat sheet is not accessible.\nYeah,\nmy bad. My bad.\nAh, sorry. I actually provided the wrong\nway.\nHere you go, folks.\nAll right, looks like we've got the\nright link now. Thank you so much, Anu.\nIt seems like we've got some questions\nthat are still coming up. So, Anu, um,\nif there's additional questions, do you\nwant to ask people to stay on the call\nto help answer that or do you want to\nprovide contact information? What's your\npreference?\nUh, my email address is the easiest to\nremember. It's just my first name\natgmail.com.\nTyping it out here. Anonygmail. Feel\nfree to reach out to me uh if you have\nquestions that I'm not able to answer\nright now, but I'm happy to stay on for\nanother 5 10 minutes uh and answer\nquestions. I I'm actually taking another\nlesson after this. So, in about half an\nhour, I have another class to teach. Uh\nso, uh I would have to jump off after\nfive or 10 minutes, but I'm happy to\nanswer questions to them.\nWonderful. It looks like how do you\ndecide that the golden data set has\nreached a healthy level is a question.\nHow do you decide? The the rough rule of\nthumb is that for most systems, your\ngolden data set needs to have about 50\nodd examples. So 50 out uh 50 odd input\nand output pairs to start off with. Um\nyou essentially want to the way to think\nabout this is follow the 80/20 rule. So\nyour golden data set should actually\ncover uh at least all the examples which\nmake up for 80% of all transaction\ntypes. So in our credit card statement\nanalyzer, right, the most common kind of\ntransactions that you would find across\ncredit card statements should definitely\nbe covered by the golden data set. So\nthat's your rule of thumb. So based on\nthat, add as many examples. Right? Now\nif it is uh\nfor our insurance chatbot if there are\nspecific types of questions which are\nthe most asked questions. So suppose you\nwere already running a manual system and\nnow you can actually query and say which\nare the top five questions which are\nalways asked you know which make up for\nthe biggest chunk of answers or then or\nyou reverse it and you say uh this is my\ntotal volume of charts which is the\nquestion which are the questions which\nare asked in 80% of the charts and then\nthen it gives you actually these are the\n16 or 17 questions which are the most\nasked or the 10 questions which are the\nmost asked and of course there are\nhundreds of distinct questions which are\nasked but that's the long tail. This is\nsort of the head and the torso and that\nthen allows you to create a golden set\nout of it.\nAwesome. Kelly did ask an earlier\nquestion. I see she's still on or Kelly\nis still on. Do you have any recommended\ntools such as link views?\nThere are a whole bunch of really\nphenomenal platforms now which allow you\nto build these systems end to end. So\nyou know which will allow you to figure\nout which quantitative measures you're\ngoing to use, instrument them, which\nqualitative measures you'll use, create\nLLMs as a judge. There are a bunch of\nthem. But again, as I said, don't\na simple Google search would actually\nland you on five 10 top tools and then\nyou can actually be very prescriptive\nand say for this kind of a product,\nright? That for if you're building a\nchatbot, then you can say specifically\nfor chat bots. If you're building uh you\nwant product analytics and you want it\nto be along with your in you know\ncurrent amplitude implementation or your\npost hog implementation then there are\nways of doing it. So there are a whole\nbunch of uh out there platforms out\nthere which do this really really well\nbut again\nfind the product which matches your\nneeds and for that you have to figure\nout your needs first.\nIt's kind of JTBD for ourselves.\nYes. Yes.\nAnd then so Oh, go ahead. Sorry. No, I I\njust quickly looked at a question which\nsort of is connected on public\nbenchmarks and yes, there are benchmarks\navailable. So, type of tools which are\nvery common or types of products which\nare very common. What has happened is\nthat there are these just like for the\nlarge LLMs there are these standardized\ntests for a lot of products. Now there\nare standardized testing systems or a\nstandard test that says hey this product\nshould be at this particular benchmark.\nSo there are these common public\nbenchmarks which have started becoming\navailable across domains and there are\ncompanies who are actually creating\nthese common frameworks across different\ntypes of uh different types of orgs and\ndifferent kinds of products. So yes,\nthere are a bunch of them.\nGreat. And then Sam had also asked a\nquestion earlier. How do you define\nwhat's a good answer?\nAh,\nthis is where sort of this is where\nagain you there's a lot of taste which\nwould go into that and there's a lot of\nhelpfulness which would go into that,\nright? Uh is it valuable? Is it is it\nhelpful? Did it actually give you an\ninsight which you didn't have or did it\ngive you a very generic answer? uh I\nactually usually as I said I don't like\nscales which is you know the least\nhelpful to most helpful kind of answers\nbecause it's very difficult to grade\nsomething a four on five or a three on\nfive uh and keep that measure consistent\nbecause as I said over time you realize\na lot of things which used to be four on\nfive start feeling like three on five\nbecause the you know the threshold of\nwhat is good has gone up uh which is why\nI always maintain a great not great kind\nof a system is is probably you look at\nit and you're like is this helpful?\nYeah, it's a little bit helpful but is\nit really helpful? No. And see that chat\nright there is just you you as a PM and\nyou as your engineer are just sitting\nand looking at output and going nah I\nwon't have found it useful. And that's\nit. This is you as a PM using your user\nempathy and having and building your\nunderstanding of of this uh of what your\nusers would expect from you and then\njust giving these answers right that\nthis is helpful this is not helpful and\nthat helps right because you would\neliminate a lot of bad answers. Sure,\nyou will have false positives and false\nnegatives, but that's completely fine.\nAs long as you stay consistent with your\ndefinition of good and in fact keep\nimproving it, keep increasing the\nthreshold, you would get great outcomes.\nAnd then we've got just the last\nquestion because I'm conscious of you on\num needing to get on to your next class.\nUh we have a liked question. Have you\nseen a difference in creating e AI evals\nfor B2B or enterprise products versus\nconsumer ones?\nYes, there are there are big differences\nand specifically the biggest difference\nis uh your requirement for accuracy and\nyour requirement for giving\nvery crisp answers is much higher with\nenterprise uh and B2B customers because\nthe pool of questions right that we\nspoke about is actually very limited and\nyou would most likely build sort of a\nrag kind of a system a retrieval\naugmented uh system for them uh because\nthe answer pool cannot be out of a very\nlarge set. The answer pool is actually\nout of a very small set. So when\nconsumers ask you questions,\nsurprisingly they can ask you all types\nof questions, right? Um from our credit\ncard analyzer if we were to build\nsomething for enterprise and the\nenterprise would say actually you know\nour outputs are only out of this small\nset of things. So you can't actually\ngive us an answer outside of this. So\nyour systems are a lot more\ndeterministic. Again, of course, I I say\nthis with all the caveats that it\ndepends on the kind of product, but your\nmetrics might be very that what you\nmeasure qualitatively and quantitatively\nmight be very different for uh B2B and\nenterprise products as opposed to\nconsumer products.\nAll right, Anshu, we are five minutes\npast. I think that there are some\nquestions we did not ask. So I encourage\nanybody who has that question. Anu was\nkind enough to provide his email address\nand again Anu thank you so much. Your\ntopic has elevated I think all of our\nunderstanding and has provoked a lot of\nquestions and some insights for us. So\nthank you so much and everybody thank\nyou for joining us. Really appreciate\nit. Again this is a sidebar session. If\nyou have any questions please feel free\nto engage us as well with any questions.\nUm you can do that at sidebars website\nand\nthank you all for hanging on for another\nfive six minutes. It seems like there's\nquite a few people who are really\ninterested. So Anu, thank you. You've\nmet you've had a very popular topic\ntoday or tonight.\nThank you everyone for staying uh late\nand uh thanks Kim for hosting.\nThanks Anu.",
  "transcript_chars": 62916,
  "ingested_at": "2026-05-21T19:38:41.655249+00:00",
  "source": "discover",
  "yt_meta": {
    "view_count": 35,
    "like_count": 1,
    "channel_id": "UC8PCzs-V83ZZUqKJmMZMfLA",
    "categories": [
      "People & Blogs"
    ],
    "tags": []
  }
}