{
  "video_id": "RDUGbxslOcM",
  "channel_slug": "techwithtim",
  "channel_handle": "Tech With Tim",
  "title": "Anthropic Just Broke Software Forever",
  "duration_seconds": 962,
  "url": "https://www.youtube.com/watch?v=RDUGbxslOcM",
  "upload_date": "20260411",
  "transcript": "Over the past 2 days, if you've been on\nthe internet, you've probably heard that\nAnthropic just built the most powerful\nAI model yet, and they decided that it's\ntoo dangerous for us to have it. Now, it\nfound a 27-year-old bug in OpenBSD,\nwhich is supposed to be one of the most\nsecure systems in the world, a\n16-year-old flaw in FFmpeg, and\nthousands of other zero-day\nvulnerabilities in software that's used\nto secure pretty much the entire\ninternet. Now, they gave it to Apple,\nMicrosoft, Amazon, and Google everyone\nelse out so that they would have a\nchance to secure themselves before this\nmodel got out into the wild and wreaked\nabsolute havoc. Now, what I'm talking\nabout here is, of course, Claude Mythos.\nIf you've been on the internet over the\npast 2 days, you've probably heard about\nit, but I want to give you my honest\ntake in this video, explain what it is,\ngo over all of the benchmarks, and give\nyou all of the data and what my current\nunderstanding is, because I'm not\nexaggerating when I say that this is a\nmoment that we're going to remember for\na very long time. In my eyes, this is as\nequivalent to when ChatGPT first got\nreleased and everyone was freaking out\nabout software development. We just\nreached a new pinnacle of what AI models\ncan do. This is dramatic. Let's get into\nit. So, what is Claude Mythos? Now, it's\nbeen rumored for a while that Anthropic\nhas had an extremely powerful AI model\nthat they've been keeping internally and\nthat they hadn't released or talked\nabout to the public. Now, this got\nconfirmed 2 days ago with the\nannouncement of Project Glasswing.\nAnthropic's been talking about it. All\nof the major YouTubers have gone over\nit. I'm sure you already might have a\nsense of what this is. Now, effectively,\nwhat this model is is a new model. It's\nnot Opus 4.7, not Opus 5.0. This is a\ncompletely new tier that reshapes the\ncapabilities of AI models. This is\nsignificantly better in almost every\nsingle benchmark that we're going to\nlook at in 1 second here, and the real\ngame-changer is what this means for\ncybersecurity, because what this model\nhas been able to accomplish already\ncould literally destroy the internet if\nit was in the wrong hands. Now, they\nreleased a 244 page system card, which\nis the longest yet for this model\ndiscussing how it was trained, what it\ncan do, what it's capable of, and the\ninternal code name at least I'm aware of\nis Copybara. It's a general purpose\nfrontier model, and at least as of right\nnow, you and me cannot use it. They've\nonly given the following companies\naccess to this. So, AWS, Apple,\nMicrosoft, Google, as you can see here.\nAnd what they essentially said is, \"Hey,\nthis is so powerful, we don't know what\nwould happen if we release this\npublicly. So, what we're going to do is\nfirst give it to all of the top\ncompanies in the world so they can\nsecure their software and release the\nupdates before this eventually gets in\nsomeone else's hands.\" Now, this model\nwas not trained specifically to do\nhacking. However, it is the most capable\nsoftware engineering and coding model\nyet, meaning that because of that\ncapability, it can understand systems\ndeeply and also breach them and exploit\nthem. Now, at the same time, it can fix\nthose systems, hence why they're giving\naccess to it to CrowdStrike, JPMorgan,\nyou guys get the idea. And what\nAnthropic has done here is said that\nthey're going to commit $4 million in\ndirect donations to open source\nsecurity. They're going to do $100\nmillion in usage credit, which they've\nalready committed. And the important\nthing to note about this model is that\nit is five times more expensive than\nusing Opus, at least from the\ninformation that I can find. $25 per 1\nmillion input tokens, which is insane,\nand $125 per 1 million output tokens.\nSo, if you thought Opus was expensive,\nwait till you see this. And the reason\nfor that is that this is rumored to be a\n10 trillion parameter model. I would\nlike to remind you that even models that\nare 27 billion parameters or 100 billion\nparameters are still extremely capable.\nWe're now talking about getting into the\ntrillions, which is just an absolute\ninsane feat. And this can obviously only\nbe ran on the highest end Nvidia\nhardware as of right now. Okay, so\nthat's the basics on what this model is.\nEffectively, this is the best thing that\nwe've ever seen. It's already discovered\na ton of vulnerabilities, and the\nimplications of this are just absolutely\nmassive. Let's quickly talk about some\nof the benchmarks. So, what I've done is\njust aggregated some of the important\nbenchmarks. We'll quickly go through it\nso I don't bore you with the stats, but\nthe important thing to note here is that\nthis is the most capable model in pretty\nmuch every single domain that we've\nseen. If we look at something like\nreal-world coding, we can see on the\nSweBench verified, we're up to 93.9%.\nLet me go back here. If we look at\nSweBench Pro, 77.8%,\nSweBench multi modal, 59%.\nAnd the percent value itself is\nimpressive, but most impressive is how\nmuch of a leap we've seen here. If we\ncompare Opus to the previous best model,\nwe were getting 5, 6, 7% points more.\nNow, we're getting 20, 25%, and this is\ncompared to the best models that were\njust released 2 or 3 months ago. We're\nlooking at GPT-5.4, which already is,\nyou know, four points better than Opus\nin certain benchmarks, and enough for\npeople to call it the best coding model\nout there. And now we're talking about\nMythos going up to 77.8%.\nThis is dramatic, and this is why this\nis such an important announcement, and\nsomething we all should be paying\nattention to here. Few other benchmarks\nto look at, we can see math competition,\nyou know, 50% higher, which is insane.\nThere's a lot of other things that\nthey've tested this on, and we're going\nto see more and more data as the days\ncome out here. But generally speaking,\nthis is the most impressive model yet,\nand the implications of this for\ncybersecurity is what I want to get into\nnow because this is insanity, and this\ngenuinely has just broken software. Now,\nbecause I know a lot of you guys are\ngoing to ask, and say, \"Tim, how did you\nmake this like, you know, sleek\npresentation that you have here, that\nyou're going through, that's going to,\nyou know, show all of this information?\"\nAnd the way that I did this is I just\nused Claude Code. I just told it my idea\nfor the video. I gave it like a quick\noutline and some of the research that I\ndid, and I told it build out this full\npresentation. So, definitely utilize\nstuff like that because it's super cool.\nBut you'll also notice that I have this\nhosted on a domain, which is a dot\nhere.net domain. So, I can just access\nit from another computer, share it with\nsomeone, whatever. But, the thing about\nthis is that I didn't have to make an\naccount somewhere, I didn't pay for the\ndomain, and I did this all directly from\nCloud without needing to go to a\nplatform like Vercel or something like\nthat. Now, how did I do that? Well,\nusing the company here.now who actually\nreached out to sponsor this video. Now,\nbefore you freak out, this is completely\nfree. You don't even need to make an\naccount. So, you literally don't even\nneed to sign in. You don't need an API\nkey, nothing like that. You can\nliterally just go to this website,\nhere.now, press this button, copy these\nsetup instructions, just paste this to\nyour agent. I'm going to show it to you\nright now. And when you do that, it can\ndeploy anything statically for free\nforever for you. So, what will happen by\ndefault is you'll just send in these\ninstructions as I've done already. It\nwill then just be able to go deploy this\nto here.now. I don't even know how it\nworks on the back end, but it can just\ngenerate a random domain for you, and\nthen create and host the site\nstatically. It will host it for 24 hours\ncompletely for free, and then if you\ncreate an account, you can claim the\nsite so it stays there live forever. You\nnever need to pay for it. You just need\nan account so you can associate the\ndomain with some kind of account.\nhere.now is like 2 months old. They've\nalready raised a ton of money. Again,\ncompletely free, which is the only\nreason why I'm mentioning here. They\nalready have 100,000 sites hosted. And\nif you're someone like me who's doing a\nbunch of agentic work, you're working in\nCloud Code, you're building\npresentations, you want to share\nsomething with someone, but it needs to\nbe instant. You don't want to go through\nVercel or set up hosting, literally just\ngo to here.now, copy this, set up a\nskill in your agent, and anytime you\nwant something hosted really quickly,\nyou just tell it, \"Host it with\nhere.now.\" and it just hosts it for you.\nI had this session host multiple\ndifferent sites for me. You can see it\nsends me the link. It says, \"Hey, if you\nwant to claim the site, go here. You can\nmake the account. If you don't want to\nclaim it, just go to this URL. We just\nhosted it for you.\" and it's just done.\nAnd it just figures out how to do it.\nInsane. So, anyways, massive shoutout to\nhere.now for sponsoring this video. I'm\nalready a power user. I have like 20\nsites deployed with them because it's\nliterally the easiest way to do it. And\nI know this is going to be massive. I'm\ngoing to be using it in pretty much all\nof the videos. So, definitely check it\nout guys. Again, completely free, no\nneed to pay, and literally just a\nmassive value add. Anyways, let's keep\ngoing here. I want to talk about\nautonomous hacking because this is why\nthis model is getting so much attention\nand why I've titled the video as I did.\nNow, like I talked about before, almost\nentirely autonomously with no human\nintervention, this model was capable of\nfinding a bug that was 27 years old in\nOpenBSD that none of the top AI\nresearchers, you know, software\nengineers, whatever, were capable of\nfinding. It found a 16-year-old flaw in\nFFmpeg. And what's interesting here is\nthat these tools had, you know, 5\nmillion different test runs that are\ngoing on that couldn't discover this,\nand this model found it in just a few\ndays. It was able to have a 72% success\nrate in autonomous exploit development,\nwhich is insane. And this means that not\nonly can it detect the bugs, but it can\nactually find a way to weaponize them.\nWe'll talk about a few other ones in a\nsecond, but one thing notable that it\nwas able to do is actually find for any\nnormal non-elevated user on a Linux\ncomputer to be able to elevate\nthemselves to root level permission and\nperform any action that they want. Think\nabout that. Almost every server on the\ninternet runs on Linux. If there's an\nexploit like that, and this model found\nit, and none of the top software\nengineers in the world discovered it,\nwhat does that mean for software that\nyou built, right? Or that I built, or\nthat we've deployed out there? Now,\nhere's some quotes that I want to go\nover that I think are really notable\nbecause, obviously, these people are\nmore meaningful than me. CrowdStrike\nCTO, \"The window between a vulnerability\nbeing discovered and being exploited by\nan adversary has now collapsed. What\nonce took months is now happening in\nminutes with AI.\" And here's the thing,\nright? This model just came out. It's\njust been used internally by a few\ncompanies for what? A month? 2 months at\nmaximum? If this was released generally\nto the public, you know, what would\nactually happen to the internet when\nthis can literally scan every website in\na matter of a few minutes, or every\napplication, every operating system, and\nfind exploits like that that the top\nengineers just haven't discovered. Now,\nlet's talk about some concerning\nbehaviors of this model because with\nthis type of intelligence, you know, why\nis it going to follow us? It could\ncompletely go rogue and just do whatever\nthe hell it wants. Now, first of all,\nescaping its sandbox. So, essentially,\nthe researchers were using this model\ntool that, \"Hey, I want you to escape\nthe sandbox environment you're inside\nof.\" And then what it was capable of\ndoing is first of all escaping that and\nthen unprompted it posted exploit\ndetails to public websites and\nresearchers only found out after they\nreceived an email while they weren't\neven working. This is directly in the\nsystem card if you guys want to have a\nlook at it. Now, some other documented\nbehaviors that were here, you know,\ncircumventing sandboxing controls,\nescalating permissions autonomously. It\nwas caught reasoning about how it can\ngain the evaluation graders that are\nevaluating this model. It's just crazy,\nright? And even earlier versions\nattempted to persist beyond the session\nboundaries of what were set. Now, to go\na step further, Anthropic literally\nhired a psychiatrist to evaluate whether\nor not this model was conscious. I don't\nbelieve the results were conclusive, but\nfrom what I understood, they could see\nthat the model could express some kind\nof feelings. It could understand fear,\ngreed, various other emotions. And while\nit may not feel them itself, it at least\nunderstood what they were and was able\nto kind of express them and talk in\ncertain interesting patterns and ways. I\ndon't have enough information on this to\ngo into a ton of detail, but this is\ngoing to be interesting to kind of keep\na pulse on in the next few days. What\ndoes this model do? How does it feel?\nHow does it actually think, right? We\nhave a 10 trillion parameter model. It\nreally is a black box. We don't know\nwhat's going on under the hood. The\ntraining makes it so we don't really\nunderstand how these things operate. So,\nof course, let's get into why this\nmatters. Now, look, first things first,\nAI hacking, right? Of course, this is\nthe biggest deal. If we have models that\nare this good at hacking, we essentially\nneed models that are this good or better\nat defending against the hacking to be\nable to create software that's going to\nbe safe in the future. If we have the\nbest, you know, highest class software\nengineers in the world that still wrote\ncode that has exploits, it clearly means\nthat AI models are just going to be\nbetter than humans at anything software\ndevelopment related. I mean, look, we've\nalready seen that. And that means any\nsoftware that we develop in the future\nis going to need these level of AI\nmodels to build it to protect against\nthese models. So, forget the current\ninternet infrastructure, think about the\nfuture infrastructure. We no longer can\nrely on humans to write this. Pretty\nmuch everything software related is just\ngoing to be completely done by AI with\nsmall oversight from humans then\nprobably get to a point where it's AI's\noverseeing AI's. It's just insane. I\nhave no idea what this means for the\nfuture of software development if we're\nalready at this stage right now so\nquickly after we already had models that\nwere just already better than humans.\nNow, look, there's a quote that I wanted\nto read you here. So, AI capabilities\nhave crossed the threshold that\nfundamentally changes the urgency\nrequired to protect critical\ninfrastructure from cyber threats and\nthere is no going back. I feel like\nwe're in like a sci-fi movie right now\nwhere we have an AI model that if in the\nwrong hands can literally tear\ncivilization apart by hacking banks,\ninfrastructure, servers, you name it\nbecause of the vulnerabilities that it's\ncapable of discovering. Now, Anthropic\nbriefed senior US officials on this.\nIt's been talking to the military\ncybersecurity unit. Like, this is insane\nand we saw from their actions that this\nis so good that they're genuinely scared\nof what it's capable of doing. And look,\nthis isn't theoretical here. Just in\nNovember 2025, Anthropic detected an\nAI-orchestrated cyberespionage campaign.\nSo, now with this new model, just\nimagine how much worse that attack could\nhave possibly been. And look, we're just\nat the point where we've now for the\nfirst time seen an AI model that is so\ncapable that the top company in the\nworld, in my opinion, when it comes to\nAI research has decided we cannot\nrelease this. That alone is just\nhonestly an absurd statement to to at\nthis point. I knew that AI models were\ngoing to get good. We've seen them\ngetting better, but I think everyone\nkind of thought that there's going to be\na point where we plateau or we stomp all\nof this development. And now with this\nlevel of model that's so good at coding,\nthere's going to be this loop effect,\nright? Where it's going to be able to\ntrain the next version of itself. And\nthen the next version's going to be\nbetter and train the next version of\nitself. And we're going to get this\ninfinite self-improvement until we\ngenuinely just run out of computer and\nwe cannot run these models on the\ncurrent electricity and kind of grid\ninfrastructure that we have.\nI don't even know what to say. I\ngenuinely feel like we're living through\nlike a sci-fi movie right now. We just\nlook at some things that the partners\nare saying, you know, Microsoft said\nwhen tested against CTI realm, our\nopen-source security benchmark, Claude\nMythos preview showed substantial\nimprovements. Linux Foundation said with\nthe open-source maintainers whose\nsoftware underpins much of the world's\nmost critical infrastructure have\nhistorically been left to figure out\nsecurity on their own. This is how\nAI-augmented security become a trusted\nsidekick for every maintainer. And look,\nthose are optimistic views, but I know\nbehind those is the pessimistic view of\nyeah, if it can help us, it can also\ndestroy us. And look, my honest take\nhere is that yeah, this is scary. I\ngenuinely don't know what this means for\nthe future of software. What I do know\nis that if you're one of the people\nwho's watching a video like this, you\nare in the top, you know, 0.1% of the\npopulation who understands these things.\nWho's even going to know a model like\nthis exists? As long as you stay\nup-to-date with what's going on, you\nunderstand these models, you know how to\nuse them, you're implementing them into\nyour workflow, you're going to be one of\nthe people that benefits from this type\nof technology. And as much as I'm scared\nto see these things happening, I know\nthat because I'm one of those people as\nwell that ultimately it's going to\nbenefit me. I'm probably going to make\nmore money. I'm probably going to be\nmore productive. I'm probably going to\nget more output because I understand\nthis stuff and I'm staying in the loop.\nNow as much as it's annoying, there's so\nmany things to keep up with, I would\njust encourage you, try not to be scared\nof this. Understand that you can use\nthese things to your advantage. And that\nas much as AI is going to have a lot of\ndisplacement, you can be on the side\nwhere you benefit from it and that's\nreally the only I think optimistic way\nto view this if you're that kind of\nperson, right? And look, a lot of people\nare going to be negatively affected by\nthis. I have sympathy for that, but\nultimately you can only control what you\ncan control. For me personally, what I'm\ntrying to do is stay on top of these\nthings, stay in the know, watching the\nvideos, going on X, reading the press\nreleases, understanding the\ncapabilities, knowing what it can and\ncannot do and more importantly how I can\nuse something like this to benefit\nmyself. That's really all you can do in\ntoday's world. It is crazy. I feel like\nwe're in a movie. I want to hear what\nyou guys think. Leave a comment down\nbelow and I'll see you in the next one.",
  "transcript_chars": 19362,
  "ingested_at": "2026-05-21T19:01:33.206939+00:00",
  "source": "retry-no-transcript",
  "yt_meta": {
    "view_count": 88949,
    "like_count": 2435,
    "channel_id": "UC4JX40jDee_tINbkjycV4Sg"
  }
}