{
  "video_id": "_34LbyYNoqU",
  "channel_slug": "deeplearningai",
  "channel_handle": "DeepLearningAI",
  "title": "AI Dev 26 x SF: Emma McGrattan: Engineering the Context Layer",
  "duration_seconds": 926,
  "url": "https://www.youtube.com/watch?v=_34LbyYNoqU",
  "upload_date": "20260519",
  "transcript": "I'm Emma McGrattan. I'm from Actian. I'm\nChief Technology Officer at Actian. I've\nbeen in the data space for 30 years. You\ncan probably tell from the accent that I\ncame from Ireland. And I actually had my\nvery first job in America back in 1987\non a student visa working at the lab for\nartificial intelligence at MIT. So\nthat's how long AI has been around,\nright? But it got really, really good,\nreally, really fast and really\naccessible. So it's obviously the hot\ntopic and role here to learn more about\ndeveloping AI applications that we can\ndeploy across the enterprise.\nWhat I'm going to talk about is the\ncontext layer, right? Because as we\ndeploy AI into the enterprise, by\ndefault, LLMs know nothing about your\nbusiness, right? So you need to give\nthem the context of your business so\nthey can answer questions that are\ngrounded in your business reality,\nright? Not some default from the\ninternet. So what we're going to talk\nabout is the most important challenge\nthat we have to address today, which is\nhow do we engineer that data layer so\nthat AI can deliver answers and business\nvalue reliably and at scale.\nSo since I guess the end of '22 when\nChatGPT really burst onto\nour\nradars, right? We've all been playing\nwith it, right? We've been playing with\nAI applications. We've built some\namazing pilots. We've rolled them out\ninto production. Many enterprises now\nare seeing success both with the genetic\nAI and with generative AI. But for some\nof us, there's a real problem in\nbuilding out an architecture to support\nAI within the enterprise. And there's\nthree main pressures that are being put\non us that really spoil these beautiful\narchitectural diagrams that we had in\nour heads as to what the AI and data\narchitecture was going to look like. The\nfirst is regulatory pressure, right? So,\nif we think about the current\ngeopolitical climate, there's a lot of\ncountries and regions around the world,\nand even companies that are saying, \"We\nneed to unhook from our dependency in\nAmerican tech.\" Right? The US Patriot\nAct gives the US government the right to\naccess any data that's on US soil or in\na US based platform that they think that\nmight\nbenefit um the security of our nation,\nright? And that is very scary to other\nnations, right? So, they're looking at,\n\"How do we unhook from American tech?\"\nAnd sometimes that means moving to\nsovereign clouds. Amazon, for instance,\nhas a sovereign cloud in Europe, where\nEuropean governments can set up their\ndata and know that um it's never going\nto touch the confines of Europe. So, not\njust the data residency, where does the\ndata live, but also the people that\ntouch that, right? Any of the operations\nthat happen on that data will never you\nleave the confines of that sovereign uh\ncloud space.\nThe other team of people that love to\nmess up clean data architectures are the\nregulators, right? So, if we look at\nindustries like financial services and\nhealth care,\nthey have mandates that certain data\ncan't leave the confines of those um\ndata centers, right? So, in financial\nservices, a lot of the customer data,\nthey will mandate through regulation\nthat has massive fines attached to it,\nthat that data can't leave the data\ncenter. So, it means that we have to\nlook at on-premises deployment of some\nof our AI stack, right?\nAlso, when it comes to autonomous\nvehicles, fraud detection, the um\nlatency requirements that we have there\nmean that we can't go all the way to the\ncloud and back again, right? That could\ntake 50 milliseconds, that could take up\nto 200 milliseconds, and we need to be\nable to react in single-digit\nmilliseconds and in sometimes sub\nmillisecond time frames, right? Cloud\nprecludes that from happening. So, when\nyou need to make truly instantaneous\ndecisions in real real time, right?\nCloud is not an option here. So, we've\ngot to think about how do we deploy AI\nat the edge, right? How do we deploy AI\nso that it's sitting in the device where\nthe decision has been made. So, we're\ntalking about millisecond latency in\nmaking those decisions.\nAnd then the third is data gravity.\nSo, every enterprise has a very messy\ndata architecture, right? I heard\nreference to mainframes earlier, right?\nAnd when you go into financial services,\nwhat you'll find is that they've got\neverything from the oldest of mainframes\nto the most modern devices and they have\ndata across all of them, right?\nAccording to Gartner, the average\nenterprise has 400 different sources of\ndata. 400. Now, imagine where that's all\nrunning. Imagine trying to wrap your\narms around it and bring the AI to that\ndata. It's quite a challenge, right? So,\ndata gravity also plays a role.\nTypically, you're going to have data\nthat lives on premises. You're going to\nhave data that's in clouds. Maybe you're\nusing, you know, Azure, maybe using GWS\nor or AWS, right? You're probably using\nsome SaaS applications like Salesforce,\nright? You probably have some APIs that\nhave data behind them and you probably\nhave data pipelines, right?\nThe AI needs that data and the AI needs\nthat data quickly. So, your AI all of a\nsudden has gravity, right? It that data\ngravity is going to suck the AI along\nwith it. So, you're going to have to\nhave a very messy AI architecture to\nsupport the data and and support\ndecision-making where the data is.\nNow, many of you, I'm sure, are very\nfamiliar with rag, right? Rag stands for\nretrieval augmented generation, right?\nAnd basically, this comes back to the\npoint I made earlier where we say that\nLLMs don't have uh stateless, right?\nThey also don't have any knowledge of\nyour business. And what you need to do\nis to provide them with that knowledge.\nTo do that, you go to a vector database,\nright? The vector database is able to\nprovide semantic meaning so that you can\ndeliver that to the LLM so that the\nanswers that it provides are grounded in\nyour customers or sorry, in your\ncompany's data. I'll give you a quick\nexample, right? So, let's say you've got\nan insurance policy. Let's say it's your\ncar insurance and it went up by 20% this\nyear. And you haven't had an accident\nand you haven't had a ticket, right? And\nyou want to ask your insurance, \"Why did\nthis happen, right?\"\nWhat the insurance company's going to do\nis take your question, right? It's then\ngoing to look up all of the data that\ninfluenced the decision on your policy\nincrease, right? So, it may be that you\nmoved to a neighborhood where your car\nis more likely to be stolen. Uh it may\nbe that um your commute times have\nchanged, right? So, you're going to be\nin the car more. It may be that the\nroads in your neighborhood um have had a\nlot of potholes cuz you had a brutal\nwinter, right? So, what that that vector\nsearch is where we're going to get all\nof that context, right? So, we're going\nto take your question, \"Why did my car\ninsurance go up?\" We're going to take\nall of that answers we have that\ninfluenced why your car insurance policy\nwent up and we're going to provide you\nwith an answer um from the LLM interface\nthat's going to make sense to you,\nright? So, we're going to cite the\ndocuments that caused your insurance\npolicy to go up and we're going to give\nyou an answer that's grounded in the\nreality of why that decision was made.\nSo, that's rag, right? Retrieval\naugmented generation.\nAnd the challenge is that\nwhere all of that data that we look up\nto augment the uh the LLM, right? lives,\nright? Where that resides and where we\nretrieve it from is going to impact the\nperformance that we're getting when we\nask these questions um of our AI, right?\nSo, where the search run and where the\ndata lives is going to impact the\nperformance. And typically you want to\nget your answers back fairly quickly or\nyou go off and do something else cuz\nwe're easily distracted. So getting this\ncontext layer right, deciding on what's\nan appropriate architecture to answer\nquestions quickly is incredibly\nimportant.\nSo we've got three choices when it comes\nto the topology of our AI context layer,\nright? Cloud, my god, it's amazing,\nright? It It's has elastic scale, it's\ngot global reach, it's it can flex as we\nneed it to, right? If we're loading up a\nlot more information, if we're doing a\nlot more work, the cloud can scale to\nmeet those demands, right? It doesn't\nhave any If we don't have any restraints\naround where the data needs to live,\ncloud is a great answer, right? Um the\nchallenges with cloud is first off it\nhas latency, right? So 20 to 200\nmilliseconds round trip time to get any\nanswer.\nEgress co- egress costs can go up,\nright? So that is as we're taking data\nout of the cloud, there's a charge\nassociated with that. So depending upon\nhow much data is is being egressed from\nthe cloud, those costs can be quite\nsignificant. And then obviously if it's\ncloud, we've got a dependency upon\nconnectivity,\nright? But cloud is a really great\ndefault.\nNext up we've got on premises, right? So\nif you have data sovereignty\nrequirements, if the data must live\nwithin a data center either because, you\nknow, of you're a sovereign nation and\nyou required it or if because you're in\nfinancial services, health care,\ndefense, and some of these industries\nwhere they've got regulations that\nrequire data stays on premises, then\nyou've got to look at an on premises\ndeployment of this semantic layer.\nUh we here have um investments that need\nto be made in just the infrastructure to\nsupport this. We've got delays in\nordering all of that infrastructure. And\nthen we've all of a sudden we've got to\nbe able to manage, maintain, and patch\nthese environments. So that is um not\nnearly as attractive as cloud at where\nall of the infrastructure and operations\nare handled for you.\nAnd then the third\nuse case that we have for deploying our\nsemantic layer is edge, right? So if you\nneed to make decisions in milliseconds,\nright? Edge is where you need to be\nthinking, right? Edge devices also can\nhave no connectivity. So if you're an\nenvironment, let's say you're down in\nFlorida, our software is used for badge\nscannings on fast pass, right? And every\nafternoon when you're in Orlando,\nthere's a thunderstorm that comes\nthrough and every afternoon all of these\ndevices lose connectivity and people can\ncheat on the fast pass, right? But when\nthe internet connectivity comes back up,\nit all gets reconciled. And but being\nable to work from your disconnected is\nan important use case when it comes to\nedge, right? And also this is a a great\nuse case if the information cannot leave\nthe device, right? So maybe it's\npersonal health information and you\ndon't want that, you know, ever going to\ncloud. You want that to stay within the\nthe device, then edge gives you the\nability to do that. The challenge with\nedge is that the compute and the memory\ncapacity that's available to you is is\nquite limited. And also the freshness of\nthe indexes that you're working at\ncan be quite spotty because you've got\nto refresh those.\nAnd there's also if you've got a large\nfleet of edge devices, managing all of\nthat together can be quite difficult.\nSo getting back to what I said, you\nknow, cloud is always a really good\ndefault choice.\nOn prem, if you have on prem, we think\nof it as legacy, right? But really we\ndon't have an option when it comes to\nenvironments where we've got to think\nabout sovereignty, where we've got to\nthink about security, and when we've got\nto think about stability at scale,\nright? And so financial services when\nthey're doing things like trade\nsurveillance, like looking for the facts\nthat um\nwell, I won't even talk about our\ncurrent geopolitical situation, but you\ncan imagine that there's some very dodgy\ntrades that are going on and that we\nneed to be able to detect in real time,\nright? Also fraud, right? You want to\ndetect that as soon as possible. And and\nas I said earlier, in financial\nservices, there are heavy penalties\nassociated with breaching any of the\nregulations around data, so they've got\nto be able to prove that the data stayed\nwithin the confines of the data center.\nHealthcare and life sciences, similarly\nwith uh regulations like HIPAA and PHI,\nright? That data needs to stay within a\nvery hardened um stack within the um\nthe the confines of the health network,\nright? So it might be within a a data\ncenter, it might be within a doctor's\noffice, uh or hospital. And then in\ndefense, these these are kind of very\ncool deployments where we need to have\nair-gapped systems, right? So that\nnobody can infiltrate it. So you um we\nactually have software that's built into\nthe warheads of nuclear submarines,\nright? Once that's deployed, no one can\ntouch it, right? So you can't have\nthings like a license server that's\nchecking to see if they're still\nlicensed, right? So uh really important\nthat we understand as we deploy this, um\nyou know, what makes the most sense.\nSo, edge definitely when milliseconds\ncount, edge is the the best option.\nWhat we're going to find is that there's\nno one great solution that's going to\ncover every use case that you have, and\nhybrid is going to have to be what you\ndesign for. So even if today you're all\nin on cloud, as you design, design for a\nfuture where you're going to have\nsomething on an edge and you're going to\nhave something on prem.\nAlso think about intelligent query\nrouting, right? So if you have\nclassified data, right? If it's\nregulated data, then you're going to\nroute that to your on prem tier.\nIf the latency needs to be less than 5\nmilliseconds, you're going to route that\nto your edge tier.\nIf the freshness of the data is less\nthan a minute, definitely go into the\ncloud, right? And if you need to query\nall of it, right? If if your question is\ngoing to span multiple domains, then\nyou're going to have to fan it out, and\nthen you merge the results in your\napplication.\nUh how do you choose the right model? I\nwould take a photograph of this slide\ncuz I have 1 minute left to speak and\nI'm not going to get through all of it.\nBut but uh refer to this. Look all the\ncameras going up.\nUh yeah, refer to this, right? Make the\ndecisions for um each of the workloads\nthat you're dealing with, right? Map all\nof them and then you'll very quickly\ncome to a conclusion that you've got no\nchoice but to go hybrid. So plan for it.\nAnd where are they heading? Where are\nvector databases heading? At multimodal\nretrieval, right? So today it's mostly\ntext and text out. We're very going to\nvery very quickly going to get to a\npoint and and I think it's within 12 to\n18 months where when we're doing a\nretrieval, we will expect to retrieve uh\naudio and image data as well as text and\ntime series data. So multimodal\nretrieval is definitely in our future,\nso plan for it.\nUm managing indexes is difficult, right?\nTakes a little bit of brainpower. Uh I\nthink that's all going to be driven by\nAI again within the 12 to 18-month time\nhorizon. At governance-aware retrieval,\nright? So today, what happens is you\nretrieve the information and then you\nbolt govern on governance on after the\nfact, right? So after the fact you\ndecide who should have seen what and and\nhow you pass this on. That all needs to\nbe built into the index. And then, you\nknow, having to understand SQL and other\nlanguages um like graph is uh it's going\nto be a thing of the past. Uh what do\nyou say back to your team? Take a\nphotograph of this one as well, right?\nTopology, very important you get the\narchitecture right. Distributed AI\nrequires distributed retrieval. Silent\nfailures are incredibly dangerous,\nright? And the context layer is\nload-bearing, so treat it like your\nbusiness depends upon it. Uh I'm from\nActian. We're launching a new vector AI\ndatabase today. It's uh designed for\non-premises and edge deployment. We're\njust down here sandwiched between AMD\nand Oracle. I'm also very proud to say\nmy second book was published this\nmorning and uh you can download that for\nfree. It's on vector databases. Yeah,\nthank you. Uh it's available from\nO'Reilly. And thank you for your time.\nCheers.",
  "transcript_chars": 15897,
  "ingested_at": "2026-05-21T19:18:39.064027+00:00",
  "source": "retry-no-transcript",
  "yt_meta": {
    "view_count": 597,
    "like_count": 12,
    "channel_id": "UCcIXc5mJsHVYTZR1maL5l9w"
  }
}