{
  "video_id": "errTnC59gVM",
  "channel_slug": "arizeai",
  "channel_handle": "arizeai",
  "title": "Keynote | The Future of AI Agents | Arize Observe 2026",
  "duration_seconds": 1617.0,
  "url": "https://www.youtube.com/watch?v=errTnC59gVM",
  "upload_date": "",
  "transcript": "Welcome. I I'm so excited to to have\neveryone here today. I'm I'm Jason\nLateki, founder CEO of Arise. Uh\n>> hi everyone. Parta, co-founder here.\nSuper excited to welcome Observe.\n>> So welcome, welcome to Rise Observe. Um\nlike like that said, we we created this\nconference uh for you know I think this\nis the fourth annual one. There's one on\non Zoom which we don't count um that\nthat to to really speak to to us frankly\num to to be the the the place we would\nwant to to attend to learn. Uh hopefully\nit's deeply technical. Hopefully it's\njust not marketing speak. Hopefully you\nlearn something. Uh hopefully there's a\nlot of builders here building an\nincredible amount. Um that's what we do\nevery day. Uh so super excited to be\nhere today and um and and I think we've\nall felt the kind of magic of AI, the\nmagic of agents over the last year in a\nway and and every year I've been just\namazed by how fast this industry is\nworking, how fast it's moving and and I\nthink this year was the the fastest\nwe've you we've moved to date. Fastest\nwe've rolled out product. Um the the it\nfeels like things are accelerating. So\nexcited to to tell you a little bit of\nthe vision. We see um we see a lot.\nWe've, you know, hundreds and hundreds\nof teams of some of the the largest tech\nfirst enterprises, largest companies in\nthe world use us. We see everything. We\nsee an incredible amount. And what we'll\ntry to do is tell you what's really\nworking, what's not. Um also where\nthings are going as much as we can see.\nAlso want to thank our our sponsors from\nfrom Microsoft AWS uh Crew AI and and\njust one of my favorite awesome\nframeworks uh for agent frameworks out\nthere. Swift Ventures bet on us in the\nearly days. Thank you.\n2025\nit this is the we've been talking about\nagents for for years and and really it\nwas the first year that that this is the\nfirst year that they're actually doing\nsomething. Let's let me be honest. It\nreally is the BL first year that these\nthese agents are actually actually\nuseful. Um harness, how many how many of\nyou have heard the word harness?\nOkay, that it didn't exist in December\nof last year. You know, like that that\nwas not a word that people were using.\nBy January, it started to get used. Um\nand and and now they're probably some of\nthe most important products that we use\nevery day. You've got Hermes, one of the\nbest open source harnesses out there,\nstarted from, you know, coding harnesses\nand the these are now general purpose\nagents. Uh, Open Claw taking over the\nworld. Uh, Claude Code, I saw Boris's\nspeech last year at the AI engineer\nconference. It was kind of the la It's\ncrazy. It's been maybe a year these\nthings have been out there. And myself\num you know I I I view these as I'm a\nproduct person. I'm someone who who\nloves products and these are some of the\nmost highly I like best products I've\never used. The utility is incredibly\nhigh really these are the first agents\num these are the first agents that are\nactually doing something for you every\nday. Um, and the the thing is though,\nthese agents still are just are are\nrunning locally on on our laptops. Most\nof the time they're running with a human\nin the loop. There's a person maybe\nstepping this along in some way. Okay.\nAnd this is this is what you know from\nour case we might a customer might pull\ndown traces might step through a fix of\nan agent that they're deployed um a and\nthey're doing it manually locally on\ntheir laptop and and that that's kind of\nthe way the last year has felt or looked\num is is these these agents are running\nlocally uh and and and then there\nthere's kind of a human in that in that\nloop and I think our vision is is kind\nof going kind of upstream that that\nthere's you know a human uh human steps\nare becoming more automated and so in\nthe last year a loop might have looked\nlike you you send your traces to a\nplatform you have some manual evals that\nyou're writing uh you have all these\nsteps that that you're taking as a\nperson and then as the year evolved we\nwe built more automation we built more\nautomation into our products uh we had\nAI we have our our assistant Alex\nactually helped write evals for you. Now\nwe have automation our product that\nhelps you find the issues. Um we have\nskills now that help you make the fixes.\nSo the humans kind of you're building\nautomation where the humans kind of\ndirecting the things but still um these\nloops are being done a little bit more\nfor you. Um, so this is kind of what\nit's felt like through the year. Um, as\nas we've been kind of rolling out\nproducts and a lot of us ask ourselves\nevery day, me and our partner, literally\nevery day, um, where does it go? Like\nwhere where what are the where does this\ngo? Well, you know, the first step we're\nfeeling is that this local thing is\ngoing to be um, agents kind of run, you\nknow, these agents that are running\nlocally are going to be running as\nlong-term processes. You kick off an\nagent for every issue. Some of you are\nalready doing that today. Um, an agent\nfor uh for for looking at you reviewing\nthe issues fixed. An agent for um for\nevaling your system constantly every\nevery five minutes. So longunning agents\nand agents running um you know agents\nthat were running locally now running in\nthe cloud and doing things. Absolutely.\nWe're seeing it already. We see that\nrunning forward. So um think of it as\nagents everywhere and and these aren't\njust agents to fix your product. These\nmight be agents to do back-end processes\nin your business. These might be agents\nin your enterprise. These a and but but\nit is there's no doubt in our mind that\nthese there there is going to be\nthousands and thousands and thousands of\nthese agents deployed in every business.\nAnd I think that's that's much more\nclear than it was last year. and and\nreally the the humans playing this role\nof of of managing these making sure\nthey're doing the right things. Um but\nreally trying to to to control this\nsystem so that so that the the system\ncontinually gets better.\nOur vision when we started this company,\nit's literally on our on our seed deck\nwas we want to make the world's AI work.\nUm it literally was on our seed deck. Uh\nand originally that was that was\nobservability. It was traces. was\ntracing the these systems um evolved\ninto eval which are are really AI\nhelping you understand what AI is doing.\nAI is sitting on top of these traces. Um\nnow what I feel in the last year is even\na higher level of just can you automate\nthe loop? Can you can you can your error\nissue turn into a fix automatically? Um\nreally that's what we're doing is\nbuilding systems just like humans learn\nand get better that will automatically\nget better. You're you systems that\nautomate the shipping of of agents. So\nAaron is going to tell you a little bit\nabout this.\n>> Awesome. Thanks Jason. So where does\nthis all start? I mean this is this is\nsomething we believed really from the\nbeginning which is that the traces are\nthe foundation for what continual\nimprovement for agents starts with. And\nwe believed in this literally since I\nthink 2023 when we rolled out open\ninference which has now become the\nstandard in how people trace agents.\nIt's the most widely used set of\nsemantic conventions for geni. And what\nis a trace? I mean the trace is the\nblueprint for what actually happens.\nIt's the LLM call. It's the tool call.\nIt's the reasoning behind the model.\nIt's every single step that the agent\nactually had to take in order to to kind\nof complete its action. And we already\nsee teams do this today. I have so many\nof you that have started using our\nskills to go pull down the traces that\nyou're collecting. Use it to actually,\nyou know, pass it to an agent to go\ndebug and fix an issue. And we started\nto think about kind of how does this\nevolve today. A lot of us were kind of\ndoing this in a very local environment.\nYou pull down the traces, you give it\nyour coding agent, the coding agent\n>> what the issue is. Um, but there's\nalready things changing with that. The\nfirst thing is well now agents can\nactually read the traces. I don't have,\nyou know, there's just the volume and\nthe scale of traces has gone from\nthousands to millions in in literally\nthe last year. And so you have agents\nreading the traces with skills and MCPs.\nAnd then second, this is probably the\nmonumental, but we no longer, you know,\nthe the launch of kind of clawed managed\nagents and cursor cloud agents means\nthat you don't have to run these things\nin your local environment. You can\nactually run these things in the cloud.\nAnd so where this is shifting is now you\ncan orchestrate a team of agents.\nSomeone to actually read the traces,\nsomeone to actually go put up the fix,\nsomeone to go review the fix that\nanother agent put up. And so now you\nhave a team of agents that are actually\ndoing the workflow that you used to do.\nAnd where this is going is really a\nfleet of agents kind of um reading your\ntraces, putting up fixes, all running in\nthe cloud. And there's there's just a\nwhole organizational shift that's about\nto happen here. You have agents that\nhave to review other agents work. You\nknow, you have a I don't know, senior\nagent reviewing a junior agent's work.\nI'm just kidding. But, you know, it's,\nyou know, you have to think about how do\nagents share information with each\nother? How do they review each other's\nwork? How do you share context of what\none agents learned with another agent?\nAnd so all of these things that we've\nkind of figured out with humans, we're\ntrying to figure out how do we build um\nevery single one of you in this room is\ngoing to become like forget AI manager,\nit's like director of you're owning orgs\nof agents that are doing work. And so\nthis is this has changed the way that we\nbuild and improve agents. It's no longer\nyou in your local environment reading a\nhandful of traces to figure out what to\ngo improve. You're now having an agent\nread your traces. an agent go put up\nfixes and you are now managing a fleet\nof agents to go do this at scale in\nparallel and there's so many challenges\nas we start to scale up to millions of\nagents I mean you have to know what the\nhell are these agents doing are they\nkicking off on the right problems are\nthey solving the right thing um yeah\nsecurity is a massive concern there's\nprobably like an AI security issue every\nday it seems like right now it's just\ntoo easy to make a mistake it's too easy\nto make an issue. Um, and then and then\nthere's cost. We all were just looking\nat like token maxing and being the first\non a token leaderboard internally to now\nthere's token budgets per employees. Um,\nand so you know the these are all\nchallenges that we're all going to face\num kind of very quickly.\nToday, there's a whole series of\nreleases we're actually going to make\naround the agent improvement loop. And\nI'm actually really excited because\nevery single problem I just shared about\nhow do you go observe your fleet of\nagents, how do you go continually\nevaluate it, how do you go orchestrate\nagents to go make fixes, we're all kind\nof releasing a whole suite of product\nlaunches today to help with that\nprocess. So, let me start off with the\nfirst one. So many of you have come to\nme and asked about a partner the the\nvolume of kind of traces is just too\ndamn high. Uh it went from thousands to\nmillions to now so many of you just the\nthe scale of agent telemetry data that's\nbeing collected is too much to humanely\nkind of go through and read. And so\ntoday we're launching Signal. Signal is\na longunning agent that actually reads\ntrace data. signal is actually going to\nrun. It's a long writing agent that\nactually runs on top of your telemetry\ndata and surfaces up categories of\nissues. So things like, you know, wrong\ntool calls, if there's gaps in\nfunctionality, if the agents passing the\nwrong context to actually go solve a\nproblem, every single issues are\nactually detected can be kicked off to\ngo and put up a fix. But signal is an\nagent now reading your traces to go find\nand discover patterns of problems that\nyou can go fix. Well, that's first half\nof the problem which is the discovery.\nThe second half is well, how do I\nactually go make the fix? And so today\nwe're actually releasing a way in Arise\nto go orchestrate\nagents, managed agents. So, and and the\nthing that I'm excited about this is\nfrom day zero, we've been thinking about\nhow do we make sure it's super\ncustomizable. Bring your own harness.\nBring Claude, bring hers. Codeex is\ncoming soon as they launch manage\nagents. Um, you can bring your own\nskills. So, the exact experience that\nyou have debugging locally can also be\nused to actually debug in the cloud. You\ncan bring, you know, telemetry data. You\ncan bring up um you know any any sandbox\nthat you're actually using. So the\nDaytonas of Versel's bring all of these\nto actually go put up a PR and go put up\na fix for you so that you can go from\nissue detection to fix all within the\nplatform and manage kind of thousands of\nagents to go fix and debug issues. Um\nthis is going to scale to a lot of\nagents very quickly. So I'm actually\ngoing to invite Sally our head our\nproduct to come up and share how you're\ngoing to manage and scale these fleets\nof agents.\n>> Thank you so much Aerna. Super excited\nto walk you through our third product\nrelease here. So our third release here\nis fleet observability. So Aparta has\nbeen talking about how we're going to\nscale from a single agent to a fleet of\nagents orchestrating work across your\nentire stack. And so as we do that we\nneed our tools that we're using to\nmanage them to evolve as well. And so\nwith our fleet observability,\nwe have the ability to understand\nexactly what our agents are doing,\nwhether or not they're being\norchestrated by Arise. So it's full\ncontrol, full visibility into all the\nwork that's being done. So if you have a\ncontainer that's running that's maybe\nstopped that you need to restart, you\ncan do that. And maybe your agent is\nstuck in a loop and needs to or perhaps\nyou just want to be able to terminal in\na certain direction. And that's exactly\nwhat you're able to video with our\nfleet. Um this is full of control but um\nyou know as we're launching millions of\nthese agents cost is surely going to uh\npick up very quickly and so we need our\nagent console to do the cost tracking\nbit of this. We have uh harness support\nfor tracing across all of your popular\nharnesses. So your cloud codes your\ncursor your codecs etc. All of these can\nbe traced into a rise and so now we have\na single pane for how our pulse track is\ndoing. So, similar to how a partner had\ntalked about signal helping us find\nquality issues, our aging console can\nhelp us find cost issues. So, maybe you\nhave a user who is kind of mishandling\ntheir tokens a little bit, maybe they\nhave a very expensive query that's\nrunning in a loop or maybe they're\ntrying to build an entirely new\ndatabase. And so, we really want to use\nour agent console to pinpoint those\nexact sessions that are increasing our\ncosts so that we can do something about\nit. Now, for our cost console here, it's\nreally going to be your your best friend\nwhen it comes uh to token maxing, but uh\nwe really also want to think about all\nof these issues that we've been talking\nabout. Uh with these complex agents,\nthere's tons and tons of issues that\ncome with that. There's quality issues,\nsecurity issues, cost issues. It really\ngoes on and on. There's tons of unknown\nunknowns. In 2023, we released LLM as a\njudge, which is super powerful for when\nyou have a known issues. Uh but what\nhappens when the surface area is ever\ngrowing? That's where we're excited to\nlaunch our fifth product release harness\nas a judge. So with traditional um Alen\nas a judge, you have this static prompt,\nbut with our complex agents, it's just\nreally not going to cut it any longer.\nWe need something that can evolve with\nour data and that's where this feature\nreally comes into play. So with harness\nas a judge, um instead of needing to\ndefine your exact prompt, all you really\nneed to do is define what you want your\neval to do and the harness manages the\nrest. So I'm sure plenty of you have\nasked your product person or somebody on\nyour team to define an eval and then it\nquickly becomes irrelevant. Well, with\nharness as a judge, we're using\nharnesses to evaluate our harnesses and\nthen it can manage that complexity and\nevolve over time with your use case. So\nreally powerful and I think the the\nfuture of evaluation.\nNow, evaluation is really important for\nproduction, but you're also going to\nwant to use that with your\nexperimentation, which brings me to our\nsixth product release, agent\nexperimentation. So, uh, agents are no\nlonger single prompts. They're not even\na single tool call. Uh, they're really\nthese harnesses that are running.\nThey're reasoning. They have all these\nreally complex trajectories, and we need\nan experimentation that matches that.\nAnd that's exactly what we built with\nagent experimentation. So, um, as you're\nbuilding these various variations, I\nthink we probably all experienced, you\nknow, making one small change that\ncauses your whole agent to behave\ncompletely differently. Um, and we're\nall iterating really fast and we need a\nway to test that and that's what agent\nexperimentation aims to do. So, you're\ntesting fast with your real agent and\nyou're being able to do this all from\nthe Arise UI.\nNow, we built all of these tools\ntogether um to build our agent\nimprovement loop. So, we've really\nthought through all of these workflows.\nYou know, we talked about signal\nidentifying issues, workers kicking off\njobs for you to fix PRs um or, you know,\ntriage things on your behalf. We have\nthe fleet observability to help you\nmanage all of this that's happening. And\nthen we're evaluating that state with\nour harness as a judge and then finally\ntesting it end to end with agent\nexperimentation. Now, all of these are\ndesigned to be able to work together,\nbut they're also standalone products.\nYou don't need to use them. You can\nsimply choose the ones that make the\nmost sense for you and your team and\nbuild out your agent improvement loop.\nNow, I do have one more release for you\nall here, uh, which is our voice agents.\nUm, a lot of us are still leveraging our\napplications with text, but I think\nthere's a full frontier of applications\nthat are going to be untapped with voice\nand voice comes with its own set of\nissues. um there's interruptions,\nthere's containment issues, uh\nescalation quality, etc. And so with\nArise, our full endto-end voice support\nallows you to understand what your\napplication's doing, evaluate the audio\nquality, and automatically detect\ninterruptions all out of the box. So\nsuper powerful for those voice\napplications um and the issues that go\nalong with them. And all of the issues\nthat we talked about today um or all the\nworkflows that we talked about today\nalso extend to voice\nNow, as a as a company, we're really\ninvested in open source. Over the years,\nwe've invested in our packages like open\ninference, um,\n>> as well as our open source tool,\nPhoenix, with millions of users. Um, so\nI'm super excited to introduce you to\nour head of OSS, U Mio, and he's going\nto tell you a little bit more about what\nwe've got working for the community.\n>> Thanks, S. This working?\n>> Okay. All right. So, as Sally alluded\nto, um I work on a fully open source, I\nthink it's a unique category of\nobservability platform and framework in\nPhoenix. Like everybody knows, you know,\nthese agents are becoming more and more\ncapable. They're becoming more reliable,\nbut also they're kind of brashly with\nconfidence failing. And so the\nobservability layer that we have is sort\nof you know a happy accident I guess is\nbecoming you know the critical piece of\ninfrastructure that's necessary not only\nto observe these systems to be able to\nevaluate them but also to put them into\nhow we improve these agents as Sally uh\nwalked you through. And so what's really\ncritical is that you are able to trust,\nyou know, this system, right? Because\nthis is critical pieces of\ninfrastructure. You need to be able to\nrun it. You need to be able to fork it.\nYou need to be able to customize it and\nbe able to own it. And it needs to be\nable to run in any environment, right?\nIt needs to be able to run locally. It\nneeds to be able to run in airgapped\nenvironments. It needs to be able to run\nin any virtual private cloud. And we\nreally believe this deeply and this is\none of our core values and it runs in\neverything that we do. And so first I\nwant to talk a little bit about you know\none of the hardest problems is figuring\nout how your agents are failing. A lot\nof evaluation platforms and you know\nwe're we were one of them you know is\nkind of gave you this rolodex or this\nmenu of metrics right like you run it\nyou get a score but really you know that\nonly lets you evaluate the thing that's\non that menu. It doesn't really evaluate\nkind of the complexity that comes with\nthese harnesses and these agents. And so\nour solution to this is sandboxes. And\nsandboxes really lets you bring your own\nevaluator. Right? you really don't\nunderstand how these things are failing,\nyou need to be able to adjust them. You\nneed to be able to use your own\npackages, be able to, you know, consult\num whether or not it's grounded in web\nfreeze, whether or not you want to\nconsult multiple LLMs or even build uh\nas we alluded to agent as a judge or a\nharness as a judge as the next platform.\nAnd we really think that this is\nimportant, you know, that there's this\ncontinuous kind of evaluation sandbox\nevaluating um an agent and bringing back\nstructured annotations so you can help\nunderstand your agent. And you know, all\nof these run locally entirely on your\nPhoenix instance in most cases, but as\nyou want to scale, you know, we have\nbuilt-in support for Daytona, Modals,\nE2B, uh you name it. And so we really,\nyou know, believe that you shouldn't be\nevaluating your agent based off of, you\nknow, just some sort of menu. You really\nneed to evaluate the critical parts um\nthat you fully want the, you know,\nevaluate the behavior of your agents.\nSo you know, evaluation is one thing,\nbut still I think, you know, all of us\nhave come through, you know, thousands\nof thousands of traces. it's really hard\nto find the fossil fuel that's really\ngoing to help us improve this agent. And\nthat's where almost all the time you\nspend uh is going to be is in digging\nthrough that data. And you know, so\nthat's why I'm really excited to\nannounce our first open-source\num you know AI engineering agent that's\nbuilt directly into Phoenix. And you\nknow, AI is amazing obviously at combing\nthrough large amounts of data, but\nreasoning through it, right?\ntroubleshooting, you know, writing to\nfiles, snapshotting. That's really where\nthe power comes in with a lot of these\nthings. And so that's why we've, you\nknow, empowered Pixie with its own\nharness, uh, direct file system, its own\nbash emulator, its own set of skills,\nits own set of tools to really help you\nanalyze and figure out the information\nthat is necessary to understand your\nagent and troubleshoot it. And you know,\nour goal with this is really not to kind\nof, you know, build uh, you know,\nautonomy straight off the bat. It's\nreally to honestly give you, you know,\ngive you that time to be able to get it\nto investigate and surface up critical\ninformation that is necessary for you to\nbe able to troubleshoot and understand\nyour system so that you can improve it.\nYou know, it's an open- source uh, AI\nengineering agent. We really hope you\ntry it. We hope you try it out, use it,\nand then give us feedback. You know,\nwe're very receptive on GitHub, so we\nreally hope you try it out. This was\nPixie um automatically tuning a prompt.\nYou know, as we've been talking about,\nlike telemetry, it's the fossil fuel\nthat, you know, empowers your agents.\nBut but imagine, right, like these\namazing agents like Hermes, Plug Code,\nhopefully Pixie, right? You know,\nthey're only as good as the data. you\nknow, if the data is lost, you know,\nthere's no way to kind of fix that data\nproblem upstream. And our answer to that\nhas been really, you know, the bedrock\nof agent observability, which is built\non top of open telemetry and open\ninference. And, you know, with tens of\nmillions of downloads every month, it\nreally is the gold standard uh for agent\nobservability. But you know you know in\nthe true spirit of open source and also\nbecause you know self-coding agents are\nbecoming a thing with pi and things like\nthat it's really important for this data\nto continue to be flow and to be a\nshared thing and so we've donated you\nknow open inferences source code to the\nLinux foundation and you know uh cloud\nnative to make this you know a\ndemocrative thing and we don't really\nbelieve that one company should own the\ninfrastructure for telemetry that is a\nshared thing and that really does\ninclude us.\nAll right. So, you know, Phoenix is\nright up on the precipice. I hopefully\ntoday it crosses this 10,000 star mark.\nAnd you know, this number to us doesn't\nmatter that much other than it tells me\none thing, which is that, you know,\nagent observability and AI needs to be\nsomething that you just don't consume\nfrom some vendor. It's something that,\nyou know, you own and introspect and you\ncan understand. And you know this is a\ntestament to that number to us. You know\nwe know that every single one of these\nstars is a person. You know we know that\nfrom kind of the the steady growth that\nwe've seen rather than rapid spikes. And\nyou know to every engineer that filed a\nGitHub issue to every uh maintainer\nevery single person that's contributed\ntelemetry. We really thank you so much\nfor that. And um you know we believe\nthat AI engineering is built in open and\nthe data layer that you know streams and\nhelps us improve these agents should be\nowned by you. So with that I want to\nthank everyone and we're really excited\nfor this next chapter.",
  "transcript_chars": 25891,
  "ingested_at": "2026-06-18T10:30:27.167642+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": null,
    "like_count": null,
    "channel_id": null,
    "categories": null,
    "tags": null
  }
}