{
  "video_id": "73UEJxAWzb0",
  "channel_slug": "techwithtim",
  "channel_handle": "Tech With Tim",
  "title": "Local Models Got a HUGE Upgrade - Full Guide (Ollama/OpenClaw)",
  "duration_seconds": 1131,
  "url": "https://www.youtube.com/watch?v=73UEJxAWzb0",
  "upload_date": "20260424",
  "transcript": "Local models just got a massive upgrade.\nJust over the past two months,\nwe've seen some really good local models\ndrop that are very capable\nand relatively easy to run.\nEven on just mid-tier hardware.\nSo in this video, I'm going to show you\nhow you can run local models on your own computer\nand then connect them to tools like OpenClaw\nso you can genuinely save thousands\nof dollars per month and no longer need to rely on\nany kind of cloud provider.\nNow, I am going to explain this extremely in-depth\nbecause there are some trade offs.\nYou need to pick the correct local model,\nand not everyone should be running local models.\nIt really depends on your use case,\nwhat you're looking for, security,\nand then obviously the type of hardware you have.\nAnyways, let's dive in.\nSo the experience that you're going to have running\nlocal models is really going to depend\non the model that you choose\nand the type of hardware that you're running.\nNow I want to go over both.\nAnd while it's going to seem a little bit\ncomplicated, it's really important to understand this\nbecause making the wrong decision here\ncan have a massive consequence and force you\nto go back to using these cloud based models,\nwhich are going to be really expensive.\nNow, the first thing to understand\nis that the local models\nare open source,\nor at least most of the case they are.\nWhat that means is that these are models\nthat are openly available,\nthat are free to use, that\nanyone can download, modify, view, doesn't matter.\nOkay.\nSo you're not going to be able to run,\nyou know, opus 4.7 locally on your own computer\nbecause that's gated by anthropic.\nAnd they want you to pay them a massive amount\nof money to use it instead of tools like OpenClaw.\nSame with GPT 5.4 or whatever the newest model is.\nYou can't run that locally.\nYou have kind of a limited selection\nwhile it is still quite large actually right now.\nAnd when you are picking those models,\nthere's only certain models\nthat are going to be compatible with tool\ncalling or agent kind of orchestration platforms\nlike Open Claw,\nwhich I promise you, we're going to get into\nnow in order to determine\nwhich local models you're capable of running,\nyou do need to understand\nthe hardware that you're working with.\nThese models are very demanding on the hardware\nand specifically on the Ram and the graphics\nprocessing unit that you have in your computer.\nSo before we go much further, I'm\ngoing to ask you to make sure that you understand\nthe specs of the computer\nthat you're going to use to run this locally.\nIf you're working with a higher end\nMacBook, what you're going to want to do\nis just go to the little Apple icon here,\ngo to about this Mac, and you're going to be looking\nfor this number right here, which is the memory.\nNow assuming that you have a mac,\nthat's no more than maybe 5 or 6 years old.\nThis is the unified memory that you have available\non your entire machine from your CPU,\nyour GPU, and effectively\nwhat a local model is going to be able to utilize.\nIt's not going to be able\nto utilize all of the memory.\nIn this case, I have 32GB, and probably the highest\nend model I want to run would take up maybe 20GB.\nAgain, we'll look at that in a second.\nBut what you want to find is, okay if I'm on Mac,\nhow much Ram do I have in my machine?\nAgain,\nassuming that you're running on a newer MacBook,\nif you have a machine that's six, seven,\neight years old, it's going to be very difficult\nto run local models, and you're\nprobably going to have a better experience with cloud\njust because they're going to be extremely slow\non that lower end hardware.\nIf you're on windows,\nthe situation changes a little bit.\nIf you're on windows or even a Linux device,\nyou're going to be looking for your GPU\nor your graphics processing unit,\nand specifically how much Vram is available there.\nSo the Ram in your computer doesn't really matter.\nIt's more about just the Ram. Again,\nif you're on windows.\nSo if you're running a 4090, for example,\nyou might have 24GB of Ram.\nIf you're running maybe an older GPU,\nyou might have eight gigabytes of Ram.\nNow, typically speaking, these local models\ncan utilize almost 100% of that Vram in your GPU.\nSo whatever that number is,\nis kind of the upper bound of the type\nor performance of local model\nthat you're going to want to be able to run.\nNow we're going to get into all of the details here,\nbut you want to understand what operating system\non my on for 99% of you it's going to be Mac\nor it's going to be windows.\nIf you're on Mac,\nyou're looking at the total Ram of your device.\nIf you're on a newer Macs on M-series Mac,\nand if you're on windows,\nthen you're going to be looking at the total amount\nof Ram that you have available in your graphics\nprocessing unit, assuming that that's an Nvidia GPU.\nAnd again, you're ideally going to want\nthe highest end hardware you possibly can have.\nBut even if you're running on a\nrelatively budget device, as long as you have Vram\navailable in the GPU or again on a newer series\nMac, you still can run these models.\nOkay, so now that we have that out of the way,\nthe next thing is model selection and actually\nsetting this up on our machine.\nSo what we're going to do in order to run\nlocal models is we're going to install a tool\ncalled Ollama.\nOllama allows you to pull down different models\nand then just run them natively on your own machine.\nThe only limitation for running these models\nis that hardware that I talked about before,\nit's completely free.\nYou don't need to pay for anything\nand then you can connect this to tools\nlike Open Claw,\nwhich we're going to do obviously in just a minute.\nHowever, before we go any further, I want to make you\naware of a really cool opportunity,\nespecially if you're someone who likes open claw in.\nIt's called ClawComp.\nNow Link Ventures is running it,\nand the short version is they're going to send you\na free Mac mini to build with Open Claw,\nwhere you can run local models.\nYou can keep it no matter where you finish.\nNo matter what happens,\nthe hardware is yours for free.\nNow here's how it works.\nIt's a multi-month build program, so it's designed\nto fit around whatever you're already working on.\nYou apply with a team of 1 to 3 people.\nYou hang out in their discord, you talk in their\ntality group, and if you stand out, you're in.\nAnd then you get three months to actually build\nsomething real using open clock.\nSo not a wrapper, not a demo that falls apart\nwhen someone pokes\nin an actual automation with measurable outcomes.\nNow the whole thing ends\nwith something called claw week.\nThis is from June 15th to June 18th in Cambridge,\nMassachusetts.\nAt Link Studios.\nThey fly you out,\nthey cover the housing, and you spend four days\npushing your project to the finish line in person.\nNow, founders and investors from the link ecosystem\nare there.\nResearchers from Harvard\nand MIT are there, and portfolio\ncompanies are walking around meeting teams\nand evaluating the pitches.\nLock now the prize for $17,500\nand the application date closes on May 8th.\nSo if you're a student building something\nor you have an open claw project\nthat you've been meaning to take seriously,\nor if you just want an excuse to go deep\nfor three months with a real deadline\nand real people watching, this is super cool.\nIt's an awesome opportunity. Again,\nit's free to apply.\nLink is in the description.\nGo apply, get in the discord stand out\nand I hope you guys get the free Mac mini.\nOkay, so that said, let's get back to the video here.\nI want to go through the setup process.\nSo the first thing we're going to have to do here,\nregardless if we're on windows, Mac or Linux,\nand if you're running on a virtual private server,\nthen same thing.\nYou're going to go over to your terminal\nand you're just going to run this command\nthat you can get from Ollama website.\nI'm going to leave a link to in the description,\nand it's just Ollam.com.\nYou're going to copy the install command.\nAnd even if you already have Ollama installed,\nyou should run this again to update Ollama,\nbecause you will need the newest version in order\nto use the model I'm going to recommend here.\nOkay, so you open a terminal or command prompt.\nYou paste this command, you hit enter\nand it's going to install Ollam for you.\nNow like I said, Ollama is just a tool runs on\nyour computer that lets you run these local models.\nSo we're going to wait for this installation\nto finish.\nOnce it's done, I'll be right back and then we're\ngoing to pull a local model to our machine.\nAnd then once we have the local model running, we're\ngoing to be able to actually connect it to OpenClaw.\nAnd good timing.\nIt looks like it's already installed.\nWhile I was just doing that speech.\nOkay, so we have Ollam installed now\njust to make sure that it's working.\nWe're just going to type the Ollama command.\nIf for some reason this Ollama command\nisn't working for you, just close your terminal\nand reopen it or close your command prompt\nif you're on windows and reopen it,\nand you should just see something popping up\nlike if the command does some output, you're good.\nAnd then what you can do\nis hit escape to exit out of that.\nOkay, get out of this interactive window.\nNow if you're on a virtual private server,\nwhich is how I typically recommend running OpenClaw,\nyou're going to want to make sure\nthat the virtual private server has the same hardware\nthat I talked about before, right?\nSo in the case of a Linux machine, you're going\nto want an Nvidia GPU with a ton of Vram.\nIf you have that, then you're going to be able\nto run the local model on the virtual private server.\nSo the steps that I'm showing you here\nwork on any operating system,\nwhether it's locally on your own computer,\nwhether it's on a VPNs,\nbut it requires that you have adequate\nhardware in order to have a decent experience\nrunning these local models.\nOkay, so now that we have Ollama installed,\nwhat we need to do is select\nthe model that we want to download.\nSo what I'm going to recommend is that we use\none of the newest best local models, which is GEMA.\nFor now. There's a bunch of other models\nthat you can choose from.\nAt least when I'm filming this video.\nThis is the current, smallest and best model\nthat most of you should be able to run.\nBut what we're looking for\nwhen we select these models\nis that they have the ability to cull tools\nfor OpenCL.\nSpecifically, you need this.\nYou don't just want a chat based model,\nyou want one that has tools thinking, right?\nAll of these different modes like Gemma for has.\nAnd if you want to see all the different models\navailable, you can just go to open claw.com/search\nor just go to the models tab and you can look\nand see all of the different models that are here.\nFor example Gwen 3.6.\nAlso another great model,\njust a little bit larger than Gemma four.\nSo I'm going to recommend\nthat we go with Gemma for for now.\nSo once we've selected that Gemma four\nis what we want, what we're going to run.\nIs this Ollama run or Ollama pull?\nI'm going to show you the command and a second\nGemma for this is a command\nthat's going to download the model to our computer.\nBut before you do that, you need to select\nthe size of the model that you want to download.\nNow you'll notice that if we go to the models down\nhere, we have a bunch of different options.\nWe have Gemma for latest, Gemma for 2 billion, Gemma\nfor 4 billion, Gemma\nfor 26 billion Gemma for 31 billion.\nAnd then the cloud one\nwhich we're not going to look at right now.\nNow the billion value here\nis the number of parameters that this model has.\nThe larger the amount of parameters,\nthe better performance of the model,\nbut also the larger size of the model.\nSo really the thing that you want to look at here\nis the size okay.\nSo we see nine gigabytes, seven gigabytes,\nnine gigabytes 18GB 20GB.\nNow you want to make sure that this size is smaller\nthan the amount of Ram\nthat you have in your computer\nor that you have available from your graphics card.\nAgain, based on what I talked about before,\ndepending on your operating system, in my case,\nI have 32GB of Ram, so I'm fine to go all the way\nup to this 20 GB model, but I wouldn't\nwant to go much higher than that because the bigger\nthe model gets, the slower it's going to be.\nAnd again,\nyou need to make sure that it fits inside of the Ram.\nNow theoretically, you can run any model you want,\neven if it was one terabyte,\nas long as you had enough hardware space.\nBut it's going to be so incredibly slow\nthat you're not going to be able\nto really even get any use out of it.\nSo that's why I'm emphasizing\nthat it needs to be small enough\nthat it fits into the Ram\nand gives you a little bit of a buffer.\nSo that's what we're looking for.\nSpecifically, you want to pick the biggest model\nthat your computer is capable of running.\nOkay.\nSo I'm going to go with let's just go with the 31 B.\nRight. Because that's going to work for my machine.\nEven though it might be a little bit slow.\nAnd what I'm going to do\nis I'm going to go to my terminal\nand I'm going to type the command Alama pull.\nAnd then I'm going to paste this Gemma for 31 B.\nNow, for most of you, you're\nprobably going to want to go with the smaller ones.\nYou might go with 4 billion right.\nOr E 4 billion E 2 billion,\nwhatever they've put here for this. Right.\nSo you might put E to be right.\nWhatever the name is that you see here,\nyou can just directly copy\nwhat this is going to do\nis then start downloading the model on your machine.\nIt's going to pull all nine, 15, 20GB, whatever.\nAnd once that's done we're good to actually right now\nI already have Gemma four on my machine.\nSo I'm not going to run this command.\nBut if you don't again you need to download it first.\nIt might take a few minutes\ndepending on your internet speed.\nNow once it's downloaded,\nwhat you can do to see all of the models\nthat you have available is to type\nthe command Ollama list.\nWhen you do that, it's going to give you a list\nof all of the models you've downloaded.\nYou can see that\nI just downloaded the latest Gemma for one recently,\nand this is the one that I'll end up\nusing with open Clock.\nBut you can see all of these.\nYou can also just go Ollama help.\nAnd if you do that, it's going to give you a list\nof all of the different commands\nwhere you can create a new model, you can pull\nmodels, you can sign in, you can copy a model.\nThere's all kinds of advanced stuff.\nI'm not going to go into all of the details.\nThe point is, awesome is super\ncool, is a lot of stuff that you can do with it.\nOkay, so now that the model is installed, we can just\nquickly test it so we can type Ollama run.\nAnd then we're going to go with whatever the name is.\nSo my case I'm just going to go Gemma for\nbut you would put whatever it is that you actually\ninstalled and it should take a second here\nand then allow you to communicate with the model.\nSo I'm just going to go hello world.\nAnd then you can see immediately\nI get the response, hello, how can I help you?\nYou know I'm good.\nHow are you? Right. Whatever I'm doing. Well.\nAnd then it gives me the response and you can see\nthis is actually quite fast\nbecause I'm running one of the lower tier models,\njust the nine gigabyte model on my machine,\nwhich is capable of running this.\nIf you want to get out of this window,\nyou can hit slash exit.\nThis is just kind of a terminal based view\nwhere you can chat with the model directly\nif you want to do that.\nOkay, so now that we have the model installed\nwe want to start configuring it inside of OpenClaw.\nSo first we need to make sure we have OpenClaw\ninstalled on our machine.\nAnd again\nif you're running this on a virtual private server,\nthat virtual private server\nwould need to have the hardware requirements.\nAnd you would follow the same steps I just did\nto install Ollama on the virtual private server.\nSo what I'm about to show you,\nyou just do wherever you have OpenClaw installed.\nOkay, I'm going to assume that you have it installed\nsomewhere if you're following along with this video.\nNow what you can do right is we go to open\nI just ran this command.\nSo I have OpenCL installed on my machine.\nAnd literally all we need to do to get Ollama\nworking with this is we can type OpenCL or\nconfigure.\nOkay.\nSo assuming this is installed on our machine right\nI don't recommend running it on your local machine.\nBut for this tutorial I will show you doing it right\nhere.\nWe're going to type OpenCL configure.\nMaybe you guys have a mac mini or something.\nSo you're going to do that directly on there.\nAnd what we're going to do is go through\nand we're going to select model okay.\nSo where it says model we're going to use our arrow\nkeys.\nWe're going to press enter.\nAnd we're going to select down here.\nMine's a little bit laggy when it first pops up.\nBut we're going to go down\nall the way to where it says Ollama.\nYou should see it popping up as an option.\nNow once it pops up\nwe're going to go with just local only.\nWe don't want to use the cloud one.\nI'm not going to get into that in this video.\nOllama has a cloud offering,\nbut I don't recommend using it.\nWe're going to go local only.\nWe're just going to leave the base URL as it is.\nWe don't need to change this and continue.\nAnd we're going to select the model\nthat we want to use with us.\nNow I apologize for the cut here,\nbut I just wanted to mention that if at this point,\nfor some reason Ollama is not appearing or showing,\nI can't reach the server\nand there's some issue, then what\nyou're going to want to do here is quit out of this.\nYou can hit Ctrl C\nand you want to just run the command alarmist serve.\nNow when you run Ollamaserve, this is just going\nto start the Ollama service on your computer.\nIt's going to run directly in this terminal window.\nSo just don't close it.\nAnd then you should be good to go.\nNow you probably don't need to do that.\nYou also can just go to your spotlight search.\nIf you're on something like Mac, and you can just run\nOllama, you can just literally run the application\nand you should see you get like a little Ollama\nicon popping up in the top,\nmeaning that it's running in the background.\nOkay, so you can see we have a bunch of options here.\nI'm just going to check this “ollama/gemma4:latest”\nas well as Java four.\nYou can select all these different models.\nYou can see I have a bunch of them right.\nBecause these are all the ones\nthat were installed on my machine.\nAnd I'm just going to go ahead and press on confirm\nokay.\nSo you just select the model that we just installed\nor whatever you picked.\nFor most of you it's going to be that gem of four.\nOkay. So we're going to go ahead and press enter.\nAnd now both those models or whatever ones\nwe've selected are going to be enabled in OpenCL.\nSo I'm going to press on continue.\nAnd then what we're going to do\nis just restart the gateways\nthat this model will now be available to open class.\nWe're gonna go open core\nand then gateway restart okay.\nNow when we run that command it's\ngoing to restart the gateway.\nBoth these models should then be available.\nAnd then what we can do is just start using them\ndirectly inside of OpenClaw.\nNow, the same thing goes for windows.\nYou can run the Ollama serve command,\nor you can just go down to your windows\nlike application search bar or whatever\nyou want to call it.\nAnd then you can just run Ollama and it should run\nautomatically in the background for you.\nOkay. So you want to make sure that's running\nas a background process.\nNow once we do that\nwe can go back to our open cloud control.\nAgain I'm going to assume you know how to get here\nbecause you already probably set up open\ncloud before.\nAnd what we can do is just asking the question so\nI can say something like, hello, how are you doing?\nCan you tell me the meaning of life?\nRight.\nAnd you'll notice that by default,\nif you haven't selected any model, it should just be\nusing GEMA four.\nAnd if we go enter here\nit should just be using this model.\nAnd you can see we get the response very quickly.\nNow if for some reason you have different models\nconnected, you\nprobably want to set the default model\nto be GEMA for.\nIn order to set the default model,\nyou can typically do that from the config.\nI don't know exactly where the configuration\nis for the default model, but I believe in the open\ncloud dot json file, there is a way to set\nthe default model to just use this local model here.\nAnd then you're kind of good to go.\nNow, what I typically will do\nwhen I'm using this is I might have a local model\nthat I use most of the time.\nAnd then I also configure a cloud model\nif I want to do something that's a little bit\nmore challenging,\nthat requires some higher intelligence,\nbecause while these models are good,\nthey are going to be stupider than like an opus 4.6\nor opus 4.7,\nso you probably want to use them in combination,\nunless you purely just care about privacy\nand you just want everything running 100% locally.\nSo my suggestion is you configure like an open\nAI model and a traffic model,\nand then the local model,\nyou set the local model as your default,\nand then you switch to the other model\nwhen you want to use that.\nNow what you can do is you can actually type\nslash model here.\nRight. And then you can change the model\nthat you want to use.\nSo I just do slash models.\nIt will give me a list.\nYou can see that right now we just have two.\nAnd then I can say slash model.\nAnd then you know Gemma for right.\nAnd then switch over to actually use that model.\nIn my case\nI'm already using it. So there is no switch.\nBut you also can just directly tell it, hey,\nI want to switch over to use Gemma for it.\nCan you switch the model and then open?\nClose should be capable of just running the command\nand automatically switching the model four.\nYou can see it called the tool\nand then it's going to switch it out.\nNow in this case it's saying you can't do it again\nbecause we only have that one model.\nBy the way,\nif you are wondering what I'm using to dictate here,\nI'm using a really cool tool called Whisper Flow.\nI can quickly click into it\nand kind of show you what it looks like here\nif we go to the settings, but effectively\nit's the best kind of voice dictation.\nIt uses AI in the background extremely fast,\nactually have a long term partnership with them.\nThey're free to try out and I just use them anyways,\nwhich is why we have the partnership.\nYou could see my words per minute here,\nand I'll leave a link to in the description\nin case you guys want to try it,\nbut especially if you're working with a lot of these\nAI tools, just much faster to speak\nrather than to type, right?\nSo if I want to say something, hey,\nI want you to go do this.\nHere's a bullet pointed list of Abcde.\nRight?\nAnd then, you know, we go and it just gives me\nwhatever and automatically formats\nit, fix the spelling, punctuation,\nall of that kind of stuff.\nSo anyways, that's pretty much all that I have for\nyou guys in this video.\nSetting it up with OpenClaw is very easy.\nIt's a matter of having Ollama installed\nand then just configuring them all.\nIf you want to go further,\nyou can set up inside of like your sold on md file\nas well as your agent start MD file\nand all of the configuration and open floor,\nand you can tell it when to use the local model\nversus when to use the cloud model.\nYou can install multiple local models. Right.\nSo maybe you want Gemma for for something.\nMaybe you want Gwen for something else.\nMaybe there's a specific image or video\nmodel you want and you install that.\nYou can go crazy with the configuration.\nThe important thing is really understanding\nthat hardware constraint.\nAnd again, the performance\nand speed that you're going to get\nis going to be dictated by that hardware.\nIn combination with that model selection.\nProbably you just want Gemma,\nbut you can mess around with some other ones\nand see what experience you get.\nIn my case, Gemma is much, much, much faster\nthan a lot of the other kind of best local models\nthat are currently out there.\nSo that's it guys. I'm gonna wrap up the video here.\nIf you enjoyed, make sure they like subscribe\nand I will see you in the next one.",
  "transcript_chars": 24058,
  "ingested_at": "2026-05-21T19:02:52.919465+00:00",
  "source": "retry-no-transcript",
  "yt_meta": {
    "view_count": 53913,
    "like_count": 1221,
    "channel_id": "UC4JX40jDee_tINbkjycV4Sg"
  }
}