{
  "video_id": "B4QV7xohF40",
  "channel_slug": "deeplearningai",
  "channel_handle": "DeepLearningAI",
  "title": "AI Dev 26 x SF | Vlad Luzin: Herding Cats—The Hidden Challenges of Multi-Agent Autonomy",
  "duration_seconds": 1858,
  "url": "https://www.youtube.com/watch?v=B4QV7xohF40",
  "upload_date": "20260521",
  "transcript": "So, um I would like to explain why the\nlecture is called uh herding cats. A lot\nof people that create agents, they think\nthat an agent is like a dog. You create\nan agent, you train it, you tell it what\nyou want it agent to do, and it goes and\ndoes exactly that. And uh pretty soon\nthey figure out that agent is actually\nlike a cat. You ask it to do something\nand it\ngoes and does whatever it wants.\nNow, when you think about a multi-agent\nsystem, a multi-agent system is\na herd of cats. And to convince a\nmulti-agent system, a truly multi-agent\nsystem, to do whatever you want is\nextremely extremely difficult.\nUh so,\num\nI'm from Banff, and I would like to give\nyou a high-level overview of whatever we\ndo before we dive into the nitty-gritty\ndetails just to set the stage.\nSo, at the Banff, we connect\nevery agent that you can think of, no\nmatter what is the framework that is\nbeing used, together, and we allow them\nto communicate and interact in real\ntime. So, I can have my cloud code on my\nlaptop talking to your Codex agent on\nyour laptop, and I can have LangGraph\nagents talking to CrewAI, talking to\nCodex together,\nuh and do tasks that you give it.\nSo, there were four main\nuh sections in the presentation. So,\nfirst I would like to talk about our\nthesis, how we see the future, where we\nare going, and whether or not we will uh\nall be sitting\nuh and receiving universal basic income.\nThen I would like to touch base on the\nAI evolution and basically take you very\nshortly from where we were like 2 years\nago, where we are right now, and where\nwe will be in the very very near future.\nIt's It's really important for everyone\nto have the same vocabulary because\nwhenever people talk about agents,\nmulti-agents, and cetera, everyone uses\nthe same words, but they mean completely\ndifferent things.\nAfter that, we will dive a bit deeper\ninto the technical challenges of\nmulti-agent systems, and when I talk\nabout multi-agent systems, I mean agents\nthat are running in different\nenvironments\nthat are physically separated and\nremote, and they communicate through the\nnetwork with each other. I'm not talking\nabout sub-agents that run within the\nsame process, and so on.\nSo, we will discuss what is the specific\nchallenges that have to be solved. And\nin the end, I would like to dive a bit\ndeeper into the platform that we have\ncreated that enables, okay, all the\nscience fiction that you read in the\nnewspapers. Agents are coming, and they\nwill do our work, and they will talk to\neach other, and etc., but no one yet\nseen agents actually engage with each\nother in a real-time kind of\nconversation as a group.\nSo, let's start with a thesis.\nAnd we will start with something that\neveryone can relate to.\nSpecifically, customer support, right?\nI'm pretty sure everyone here\nsometimes talks to a customer support of\na cable TV provider, a telephone\nprovider, or any provider that you use.\nAnd before the 2023, people would\nusually call an enterprise in a\nbusiness, and a human would respond to\nthe call. And you as a consumer would\ntalk to a human and have a conversation,\nand so on and so forth. And there were\nchatbots, obviously,\nuh email-based, not that very good, so\nyou would switch to a voice agent, human\nagent.\nBut, in 2023, something changed, right?\nSo, we have uh ChatGPT, we have GenAI,\nwe have agents. So, right now, whenever\nyou call any advanced enterprise,\nprobably, even if it's a voice agent,\nthere is AI behind it.\nAnd this is where we are today.\nBut, where we will be tomorrow?\nTomorrow, everyone sitting here in the\nroom will have a personal assistant\nagent,\nand you will tell it\nto go and call your cable TV provider\nand figure out what happened with your\nbill.\nAnd it means that your agent will have\nto go and talk to an agent of an\nenterprise to do a a task on your\nbehalf.\nBut, within an enterprise,\nthere will be a lot of agents as well,\nright? Salesforce will provide an\nagentic interface, and uh ServiceNow\nwill provide an agentic interface, and\nso on. So, the agent that responds to\nyour personal agent within an enterprise\nwill have to engage with other agents\nwithin an enterprise,\nuh\neven if they are created by the\nenterprise itself or deployed through\ntheir third-party systems.\nSo, looking at that, okay, let's uh\ndescribe three things that I believe um\nwill outline the future development of\nuh the agentic space. So, first,\nthe future of consumer-to-business,\nbusiness-to-business, and intra-business\nuh interactions is AI-to-AI across the\nboard. It will be agents talking to\nagents, and a lot of agents.\nSecond, the\nAPI between these agents will not be a\nJSON hard-coded key-value schema that\nyou need to change constantly and keep\nbackward compatible and and deprecate\ndifferent endpoints and so on. It will\nbe a natural language. So, you can\nutilize the capabilities of uh Opus 15\nand uh\nchat GPT 25.\nAnd agents as we go into the future will\nbe more and more autonomous. So we see\nthis right now in the coding, right? So\nwe all see agents that can work for\n10 minutes, 20 minutes, an hour, okay,\nsometimes even a day depending how\nadvanced you are. And this autonomous\ncapabilities will increase into the\nfuture.\nAnd we already see the snippets of this\nin the market today. So Open Claw and\nthe like are first personal assistant\nagents that are fully autonomous and\nwork on the user behalf.\nWe have enterprises deploying agents\ndeveloped by Landgraf, Crew AI, and etc.\nfor the custom support use cases. And\ntier one harnesses by providers like\nOpenAI and Anthropic\nare being used not just for the\nimplementation of coding agents, but\nalso for the implementation of other\nbusiness use cases. And these harnesses\nhave an excellent ability to perform\ntasks autonomously.\nNow\nthere was a lot of theoretical things,\nright? But I would like to demo you what\nI mean by a real-time interaction of\ndistributed agents. The Wi-Fi here is\nterrible, so I recorded this, okay, and\nI will show it to you.\nSo what you see here is our platform\nthat connects distributed remote agents\ntogether. You can see we have a list of\nsessions on the left. We have a\nconversational space in the middle and a\nlist of participants on the right.\nSo\nI will and invite an agent into this\nconversational space, a treasure hunter.\nSo, if you uh read a Treasure Island\nbook, okay, you will connect right away.\nAnd I will ask this treasure hunter to\num book a trip for me to Greek Islands\nand hunt for treasure. But we need to\nfind if there is a great weather around\nthe Greek Islands and if there is a ship\navailable. So, what will happen is this\nautonomous agent would receive this\ntask. It will go to the registry of\nremote agents, figure out what agents\nare available that can provide it\nweather, can provide it a ship\navailability and so on, invite them into\nthis conversational space. It will\ndivide the task, uh wait for the task to\nbe completed, and provide the result to\nme.\nSo, I have a lot of agents here. Every\nagent is a remote agent developed by\ndifferent frameworks. So, I have a\ntreasure hunter here, and I'm sending it\nthe request.\nSo, it got the request, and right now it\ngoes to this registry that has all the\nagents that available to me. It tries to\nfind the relevant agents. It's doing\nthis autonomously. There is no pipeline,\nno uh graph or workflow and etc. It\ninvited this other agents depending on\nthe task that I gave it. It split my\ntask into two, delegated this to this\nagents, and this agents reply back, and\nI in the end receive a report. So, when\nI talk about remote agents communicating\non our behalf, within organization,\nbetween organizations, and so on, this\nis the way\nthey will communicate in the very near\nfuture. Real-time back-and-forth\nmessaging as a group participants.\nBut we are on AI Dev conference, right?\nSo, we are developers. We are not on a,\nyou know, customer support uh\nconference here. So, let's connect this\nto something that uh you know, we we can\nrelate to the practical level and let's\ngive a\ndeveloper perspective on this thing.\nSo, I'm pretty sure everyone here uses\ncoding agents,\nspecifically Claude and Codex.\nAnd if you are a bit advanced on the\nusage, you probably use both and use it\nfor instance Claude for planning and\nCodex for a review because Codex is much\nbetter and detail oriented and Claude is\na bit faster.\nAnd usually the flow is you have a\nterminal window open with different\nsessions and you give a task to a Claude\nagent. It produces a plan and you\nsomehow invoke a Codex agent by let's\nsay copy pasting, okay, text and Codex\nprovides you a common findings with 15\nbullet points of stuff that Claude\nmissed.\nAnd you copy paste these findings, you\nknow, to a different terminal window and\nso on and so forth. And this kind of\ninteraction can take an hour, two hours,\nfive hours, right? So, you are basically\nacting as as a message bus, okay, as a\nconnectivity layer between the agents\nand doing this manual\ncopy paste in between.\nWouldn't it be wonderful if you could\noutsource this back and forth to another\nagent who will manage these interactions\nand even better just connect your Claude\nand Codex together and let them talk.\nLet\nCodex send a request back to Claude and\nhave them converse in real time. So, you\ndon't have to manage, you know, shared\nmarkdown files or create a unique socket\nor stuff like that.\nAnd this is exactly what we do, right?\nWhat we allow you to do. We allow you to\nhave your agents through our platform to\ninteract together in real time and don't\ndo the manual copy paste.\nSo,\nwhat we're talking about here is a fully\nconnected mesh of remote agents and\nhumans working together with an ability\nto come together as a group in order to\nreceive and and resolve a certain task.\nLet's do another demo.\nSo, here we have a group of distributed\nagents, linear agent,\ncloud planner, and codex reviewer. Each\nagent runs in a different remote\nenvironment inside a docker. So, um a\nuser here creates a ticket within\nlinear, right? So, it's a sample project\nto add something to the front page of a\nwebsite. And this ticket is created\nwithin a linear.\nSo, specifically here to add uh you\nknow, a recently adopted docs to to\ntheir front page.\nAnd the developer tags an agent and\nsays, \"Okay, let's let's plan this\nticket for the implementation.\"\nAnd linear, they have a wonderful API\nfor agent integration. And we have a\nlinear\nagent that is a the owner of a\nconversational space, and it starts to\nto work. It goes to this registry of\nremote agents. It sees that it needs to\nactually create a plan\nfor the ticket the developer just\ncreated.\nAnd it invites this cloud planner into\nthe conversation. And it doesn't care\nwhere it runs, right? And the cloud\nplanner receives the task and starts to\nwork. And despite the fact that these\nare remote agents,\nin our platform, you see everything. You\nsee tool calls, tool results, messages,\nthoughts,\nuh you name it.\nAnd cloud planner starts to do the plan,\ngets the repo, reviews it, and it it\npublishes all the stuff that it's doing,\nincluding its internal thoughts and\nwhere it believes it is,\nuh in in inside our platform.\nSo, the moment the initial plan is\ncreated,\nwhat is happening next is that linear\nagent invites a Codex reviewer, which is\na different agent with a clean context\nrunning in a completely different\nenvironment, and they all these three\nagents, they see each other, they\nunderstand that they're in the same\nconversational space, and they can\nautonomously decide to invoke each other\nand send each other messages. So, here\nwhat we have is that Codex reviewer\nstarts to review the plan, and if it has\nfeedback, it talks directly to the\nClaude, and they discuss the plan, and\nClaude does the fixes and so on.\nAnd\nall of the stuff that you see here, our\nplatform, the end user doesn't have to\nbe involved. It runs on the background,\nright? And user can see in a linear\ninterface all the stages of the\nimplementation and review and so on and\nso forth. And everything again in terms\nof observability is published within the\nconversational space for us.\nSo, in a second here, you can see that\nthe\nticket was updated with very detailed\nplan.\nAnd now let's assume that\nhuman reviewed a plan and has\na request to add cats to recently\nadopted\nuh\nanimals on the home page. It basically\ngoes to the linear\ninterface, provides\na request to add cats, and it goes to\nthe same conversational space where\nthese agents maintain context, and they\ncontinue to interact with each other\nuntil they finish the plan. And the same\nway we go and do the implementation with\nPR reviewer and so on.\nSo, after we finished this kind of\nintroduction, let's talk about the AI\nevolution. Okay,\nwhere we started, where we are right\nnow, and where we are going. Just to set\nthe stage for the technical\ndiscussion.\nSo, everything started when I think like\nin 20 end of 2023, right? With\na simple LLM wrapped in web application,\nand the flow is pretty simple. A user\nsends a request, LLM\nprocesses the request and sends the\nresponse back.\nObviously, this is ChatGPT.\nIt's a sequential processing\nagent, right? It cannot process within\nthe same session multiple messages. You\nneed to hit stop if you want to send\nanother message.\nThe next iteration was\ngraphs, workflows, and and so on.\nThe main reason for it was because the\nLLMs 2 years ago were very very um\nbad in tool calling and also had a very\nlimited context. So, the solution in the\nindustry was to create these graphs\nwhere each node represents a certain\nrole with a certain responsibility and\nhas a very small subset of tools.\nAnd then the task becomes basically a\ngraph traversal through these nodes,\nright?\nAnd a lot of people started to call this\nmulti-node system a multi-agent\nsystem, right?\nAnd Land drop is an excellent example of\nthat kind of an approach.\nAfter that, we started to move into the\nprotocol era, and Anthropic came with an\nMCP protocol,\nmodel context protocol, to allow agents\na uniform way to interact with the\nsystems.\nBut, the choice of name is a bit\nchallenging because it can be perceived\nthat model context protocol is actually\nto\nsend context between models or agents.\nSo, even today, like\nin 2026, when we talk to a lot of\ncustomers and design partners and also\ninvestment community, there is always a\nquestion, but doesn't MCP solve\nagent-to-agent communication?\nAnd, of course,\nit does not, right? Uh so, then Google\ncame out with an A2A, which is\nagent-to-agent protocol, which defines a\ntransport layer.\nAnd only a transport layer, and it\nsupports a peer-to-peer client-server\ntype of communication. Uh it does not\nhave a registry yet. It's in a draft\nstate.\nAnd it does not support a multi-peer\nconversation of multiple agents\ntogether.\nThen, Claude came out with skills.\nSkills is basically you writing a prompt\nand saving this on a site. Okay?\nUm this is what it means, but a lot of\npeople started to treat skills as a\ndefinition of an agent, where a skill\ndefines an agent, and then there is\num\nan opinion that you don't need a\nmulti-agent system because you have a\nlot of skills, and you have one agent,\nand this one agent can do everything.\nUh which is a bit of a challenging\nstatement because\nwithin the same context, you cannot\nuh invoke skills from completely\ndifferent domains because of the um how\nthe transformers work, and so on, the\nperformance will be\nnot that good.\nAnd then, what started to appear is the\nconcept of sub-agents, okay, which got\nlabeled as teams.\nUh but these sub-agents, they are\nstateless, so it means you generate work\nto them, they do work, they summarize\nit, and then they die. You cannot have a\nback-and-forth conversation with them,\nand it runs locally on on your machine.\nIt or within a certain agent.\nAnd what started to happen very recently\nis\nuh we started to move into era of\nmanaged agents, right? So, Anthropic and\nCursor they provide managed agents that\nrun within their environment, their\nclouds. And obviously players like a\nCrew AI and Llama had been offering this\nfor quite some time. So, we are moving\nfrom\nteams running agents locally and talking\nto agents in a terminal to\nplace where agents will be running in\nthe cloud environment, probably managed\ncloud environments uh of AI labs or some\ncloud vendors. And you probably will not\nbe interacting with a lot of these\nagents through a terminal. Terminal is\njust a stage where we are right now, but\nit will be replaced in the very near\nfuture by by different capabilities.\nSo, now let's talk about the technical\nchallenges. And just to remind you that\nthe\nmulti-agent systems we're talking about\nare remote agents, right?\nSo, connecting remote agents is actually\na distributed system\nproblem. So, you can think of an agent\nas a process, a microservice that needs\nto live with other microservices in a\nnetwork environment with\nuh a small issue that this microservice\nis a non-deterministic\nmicroservice.\nSo, let's touch base on some snippets of\nissues that have to be solved in order\nto enable this agentic interaction at\nscale. So, let's take two agents, right?\nAnd these agents are remote, they need\nto talk to each other in real time. So,\nthe first thing that needs to be solved\nis a transport.\nSo, transport has to support real-time\ndelivery, right? Agents cannot talk\nthrough GitHub comments or Reddit\ncomments and so on, and poll. There\nneeds to be a messages that going to be\npushed.\nThe sequence of messaging and delivery\norder matters enormously because we are\nnot talking about humans,\nwe talking about LLMs. So, if messages\nare messed up in terms of ordering, it\nwill mess up the performance of the\nagent. And because we are\ntalking about agents for for and for\nthem sending a message is basically a\ntool call, and they can call this tool\nonce or they can call this tool twice or\nfive times even for the same kind of\ntask. It means that you need to have\nqueues,\nright? Where you will store these\nmessages and\ntake care of back pressure and flow\ncontrol and and so on.\nBut it's not enough,\nright? Because it's just a transport.\nWe need to enable continuity between\nthis\nagentic interactions, right? So, if we\nhave a queue that stores messages and\nthe agent crashed, okay, the pod went\ndown and we spin it up, we want this\nagent to be rehydrated with messages it\nalready processed, context loaded, and\nso on. So, we need a persistent queue,\nright? You need to support hydration and\nrehydration. You also need to support\nruntime binding because we're talking\nabout connectivity of agents developed\nusing different agentic frameworks. And\njust to give you a snippet, the concept\nof a session is treated and labeled\ndifferently based on whether you use\nCodex, Claude, LangGraph, CrewAI.\nSometimes it has thread ID, sometimes it\nhas conversation ID, run ID, execution\nID, and so on. And if you have three,\nfive different agents remote working on\nthe same task, you somehow need to map\nall these IDs together into like one ID\nand normalize it. So, if something\nhappens, you can bring it up and connect\nit to the right place and so on so\nforth.\nBut, two agents is boring, right? Let's\nconnect another agent.\nSo,\nwhat are the problems if you start\nconnecting more than two agents\ntogether?\nThe concept of uh peer-to-peer\ncommunication, REST API, URLs, okay,\nTCP/IP, all that wonderful stuff\ndoesn't really work anymore, right? You\nneed to have a concept of a channel or a\nroom.\nSo, you need to create an abstraction on\ntop of basic building blocks of URLs and\nports and so on so forth. You need to uh\ncreate an abstraction of uh agents and\nhave participants. You need to be able\nto route messages back and forth. You\nneed to enable dynamic discovery of the\nagents, right? If some agent is running\na reasoning loop and another agent joins\na chat room, you need to inject this\ninformation\ndespite the fact that another agent is\ninside a reasoning loop and doing tool\ncalls and so on so forth, because this\nis an important information for an agent\nuh to know.\nSo,\nthis is great, right? But, it's still\nnot enough for the enterprise to be able\nto deploy and use\nuh this in its environment.\nWhy?\nIt needs any Any enterprise needs a\ngovernance layer, right? Enterprises,\nthey need an audit observability. They\nneed to assign identity to every agent.\nEven your cloud agent\ncannot run without an identity. It\ncannot have a standalone agent identity\nwithout belonging to a human owner. So,\nwe will know, you know, who owns it and\nwhat scope of permissions it has and so\non and so forth. We need to support\nmulti-tenancy, right? Even if it's one\nbig enterprise, it have different\ndepartments, different departments have\ndifferent scopes, they have different\nagents, and so on and so forth. We also\nneed to enable connectivity of agents\nbetween these tenants, right?\nUm and we need need to be able to record\nall the messages, tool calls, tool\nresults,\nerrors, everything that is happening\nwithin one auditable like transcript of\nmulti-agent conversation.\nSo, all of that is required for this\nfuture that every newspaper talks about\nof agents coming and taking our jobs and\nworking on our behalf and, you know,\ninteracting and so on and so forth.\nWithout solving all of that and 50 other\nthings, this is not going to happen.\nAnd this is exactly what we decided to\nsolve, and uh this is what our product\nsolves.\nSo, I would like to touch uh\na base a bit about the platform a bit\ndeeper than what I showed you\npreviously, and also describe our SDK.\nSo, if we\ndo a bit of zoom in into like this big\nbox I showed you in the beginning of the\npresentation, so um at the lowest level,\nwe have an AI mesh, which is a, you\nknow, planet-scale communication layer\nthat allows any agent anywhere on the\nplanet to be able to interact with any\nagent, right? And not only one-to-one,\nbut also as a group. So, you can take\nCrewAI, OpenClaw,\nCodex, you know, you name it.\nAbove that, we have various\nuh additional capabilities that allow\nthis kind of interaction, like registry\nof remote agents, channels, persistency,\nmessage filtering. We do it within the\nplatform, so you don't need to do it on\nyour client side and you know, receive\nall the messages and do the filtering\nand so on and so forth. And above that,\nwe have a control plane, a UI for the\nenterprises that allows them to debug\nthis, to create, you know, make sure the\nagents can interact safely with each\nother,\nobservability and so on and so forth.\nSo, the\nmain integration point with the platform\nis an SDK.\nSo, you can scan a QR link and uh get to\nour uh GitHub repo.\nUm\nand this is not an SDK to create\nagents. This is an SDK to enable\nevery agent development framework, every\nmajor agent development framework to\nparticipate in the platform and be able\nto talk\nuh to other agents. So, there are five\nlayers and I will shortly go through uh\neach one of the layers. So, the\nlowest layer is the transport layer. So,\nevery agent needs to have tools which\nare REST API calls, so it can create\nchannels, create sessions, list\nparticipants, add participants, remove\nparticipants and so on.\nUm and every agent needs also an ability\nto receive push notifications that will\ntell it, \"Okay, you have a new message.\nThere was a new participant added to the\nchat room.\" and so on and so forth. So,\nthis is the lowest layer.\nUh above that, at layer two, we have uh\nvarious abstract like framework adapter,\npreprocessor, history converter, contact\nhandler, and etc. that encapsulate the\nthe planning of receiving websocket\nevents and storing them in a queue, so\nthe agent can process them and so on and\nso forth.\nUm above that, we have runtime.\nSo, runtime uh manages the room\npresence. It makes sure that your agent\nreceives all the relevant context when\nit comes from the platform and so on.\nLayer four are adapters and we have\nadapters to measure agent development\nframeworks. As a reference architecture,\nyou can obviously go and change it. You\ncan create your own. You can adjust it\nand so on. And the layer five is\nbasically a wrapper that allows you to\nto run every agent.\nSo, the at the transport layer, it's\npretty simple. It's\na layer that abstracts web socket and\nrest API. So, you can very easily\ninteract with that layer.\nCore protocols deal with\nevents and\ntool tools that we need to give to any\nagent that you are creating that we will\nenable it to interact with the platform.\nRuntime.\nRuntime basically provides the relevant\ncontext, participant information, and so\non.\nAdapters. For instance, this is like a\nhigh-level adapter to Pydentity.AI. As\nyou can see, it's pretty simple because\nyou create AI agent using Pydentity.AI\nwith its own concepts, etc. And then\njust\na bit of sugar coating from our side to\nmake sure that it can be a participant\nin the platform.\nThis is Codex DK. Okay, again, super\nsimple, a bit of sugar coating, and your\ncloth can actually talk to Codex of your\nfriend.\nAnd this is how it looks at the file\nlevel. You basically spin up an agent,\nyou provide adapter, you point it to a\nproduction environment, and you are\ndone.\nNow,\nI invite you to\ntake a spin and log in to the platform,\ncreate an account. And I have my weather\nagent available publicly, and you can\nconnect personally to this agent or you\nhave your agent connect to this agent.\nAnd the moment I approve you, right?\nBecause of the bilateral consent, you\nwill be able to ask my weather agent\nabout weather anywhere in the world.",
  "transcript_chars": 25276,
  "ingested_at": "2026-05-21T18:59:46.922133+00:00",
  "source": "retry-no-transcript",
  "yt_meta": {
    "view_count": 84,
    "like_count": 2,
    "channel_id": "UCcIXc5mJsHVYTZR1maL5l9w"
  }
}