{
  "video_id": "t99qbVBVGmg",
  "channel_slug": "arizeai",
  "channel_handle": "arizeai",
  "title": "Inside Typeform's AI Agent Stack",
  "duration_seconds": 501.0,
  "url": "https://www.youtube.com/watch?v=t99qbVBVGmg",
  "upload_date": "",
  "transcript": "I'm Marta Lawrence and I'm a senior data\nscientist at Typeform.\nSo, uh at Typeform uh we have plenty of\ngenerative AI use cases and they span\nacross like three pillars. So, we have\ncreator AI uh that one helps users uh\ncreate their forms and flows. We have uh\ninteraction AI which improves the\nrespondent experience with these more\nnatural interactions and we have insight\nAI which helps users u get insights from\nresponses and extract some patterns. So\num I mostly focus on type form AI for\ncreators uh which is where um we let\nusers uh describe their goal in natural\nlanguage and we generate a structured\nform in seconds for them with the the\nbest question types the copy and uh like\nlogic different different question flows\nand the goal is that we create high\nperforming forms in minutes for our\nusers. they don't have to struggle with\nthat and they can spend time and focus\non on the insights on the decisions that\ntruly matter to them.\nSo one of the lessons since launching\ntech from AI is that evaluation is an\nongoing task and it's not a one-off and\nwe learned that evaluation can quickly\nbecome obsolete in uh in these you know\nfast evolving applications. So our\ntakeaway is that evaluation is part of\nthe product experience. So you have to\nrevisit the evaluation continuously\num as your future uh as your feature\nevolves and well because otherwise you\nwould have you would end up having\nmetrics that look good on paper that can\nquickly start reflecting real value in\nproduction.\nSo apart from there were several\ncapabilities that we had to build\nourselves to make AI generally useful.\nSo first of all you know turning a user\nrequest into this structured form that\nrequires its own AI system that has to\nbe integrated with our internal\nservices. Uh we also built uh layers of\nsafeguards uh so customers can use AI\nconfidently and we respect their privacy\nand compliance. Of course uh we build\nevaluation pipelines um because we want\nto uh continuously improve behavior and\num lastly you know traditional product\nanalytics is not enough for AI\napplications. So we needed to track all\nthe detailed traces of AI calls\ngeneration uh latency failures and\nwhether AI actually helped our users\ncomplete their tasks. So these are the\ncapabilities where Arise definitely was\nwas very useful.\nI think evaluation is critical because\nwithout it uh we make decisions are\ndriven by intuition rather than\nevidence. We don't want subjective\nimpressions because because they cannot\nreally tell you if your feature is\nworking right. So uh without evaluation\nyou risk hallucination uh irrelevant\noutputs mismatches negative impact on\nuser trust. So our approach to\nevaluation is about measuring behavior\nacross many cases and turning it into\nsubjective uh feedback turning the\nsubjective feedback into object\nobjective signals. Um\nuh we track outputs, user interactions\num\nall the outcomes to verify AI generated\nuh forms meet our expectations\nand I think that good evaluation uh\nimpacts everyone because customers get\nhigher quality results. They get more\nrelevant AI responses. Teams building AI\nalso gain visibility into quality and\nreliability. And Tyiform also benefits\nfrom stronger user trust, engagement,\nscalable and responsible AI features.\nOur\nprimary cloud provider is AWS and they\nplay an important role in our strategy\nbecause they provide scalable\ninfrastructure, manage managed services\nand uh well reliability for running our\nmodels and their platform uh allows us\nto focus on building the AI experiences\nand orchestration layers and and we\ndon't have to manage low-level\ninfrastructure.\nWe are definitely seeing seeing that\ntrend and uh at type form we treat\nevaluations as part of the product\nitself. So it's just not so it's not\njust this backend task. We have PMs,\nengineers, designers, everyone uh is\ncollaborating on what good looks like.\nEveryone contributes with examples. uh\neveryone is helping shape uh those test\nsets. Um and I think um this\ncollaboration improves the evaluation a\nlot because it brings this diverse uh\nperspectives. Um it can help uncover\nblind spots, you know, more people, more\nideas. Um and we can thanks to that we\ncan ensure that our AI outputs align\nwith user needs, our protocols, brand\nstandards and uh we we keep the entire\npro process uh measurable.\nWell, so at Tatform an agent engineer is\ninvolved end to end um in uh AI future\ndevelopment. So we have from ideation uh\nphases to to execution. Agent engineer\ncollaborates closely with PMs,\ndesigners, engineers, we make sure that\nAIdriven product is technically physible\nand also delivers great UX.\nUm so we we build and orchestrate the AI\nlogic. We define interfaces, guard\nrails, integrations. We have decision\nrights around technical tradeoffs from\ndesign, system behavior and evaluation\nof course so we can ensure that the\nfeature is reliable and aligns with the\nproduct vision.\nOne of my favorite AI practitioners uh\npoints out that the fastest teams are\nthose that maintain a disciplined\napproach to evaluation and error\nanalysis. And I really find this to be\ntrue. Um if I was to start from scratch\ntoday, I would definitely focus on small\ntargeted evaluation first in any\napplication because with this uh you\nknow few test cases, you can already see\nsome patterns and uh uncover some like\nkey issues um and you can iterate\nquickly. At the same time, I would\ndefinitely avoid overbuilding and trying\nto handle like every edge case at the\nbeginning because that is a lot of\ncomplexity and it will definitely delay\nany learning before you know you\nactually know what what creates value.\nSo at Typeform, we run this extensive uh\nresearch and like pro proof of concept\nphase where we evaluated multiple LLM\nevaluation and observability frameworks\nand we collectively decided that Arise\nwas the best fit for our needs and uh so\none of the key reasons was that it was\nvery easy to set up and customize which\nwas crucial for us because we have\nspecific evaluation requirements and and\nneeds.\nand it also meets our enterprise needs.\nSo everything related to security,\ncompliance, ongoing support and um on\ntop of that uh the AR team that was was\nsupporting our on boarding has been\nfantastic. They helped us uh set up\neverything very quickly, guided us on\nbest practices. Uh they were always\neager to answer our questions and the\ncoll collaboration was an absolute\npleasure.",
  "transcript_chars": 6369,
  "ingested_at": "2026-05-12T00:39:30.647715+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 77,
    "like_count": null,
    "channel_id": "UCrVHzD-psX5IMCGoEWHXmGw",
    "categories": [
      "Entertainment"
    ],
    "tags": []
  }
}