{
  "video_id": "PQU9o_5rHC4",
  "transcript": "At Anthropik, the way we thought about it is we don't build for the model of today, we build for the model six months from now. That's actually still my advice to founders that are building on LLMs. Just try to think about what is that frontier where the model is not very good at today, because it's going to get good at it. All of QuadCode has just been written and rewritten and rewritten and rewritten over and over and over. There's no part of QuadCode that was around six months ago. You try a thing, you give it to users, you talk to users, you learn, and then eventually you might end up at a good idea. Sometimes you don't. Are you also in the back of your mind thinking that maybe in six months you won't need to prompt that explicitly? That the model would just be good enough to figure out its own? Maybe in a month. No more need for plan mode in a month? Oh my god. Welcome to another episode of The Light Cone. And today we have an extremely special guest, Boris Cherny, the creator, engineer of Claude Code. Boris, thanks for joining us. Thanks for having me. Thanks for creating a thing that has taken away my sleep for about three weeks straight. I'm very addicted to Cloud Code and it feels like rocket boosters. Has it felt like this for people for months at this point? I think it was like end of November is where a lot of my friends said something changed. I remember for me, I felt this way when I first created Cloud Code and I didn't yet know if I was onto something. I kind of felt like I was onto something. und dann das ist wenn ich nicht schlafen. Und das war drei straight Monate. Das war September 2024. Ja, es war drei straight Monate. Ich habe nicht ein einziges vacationiert. Ich habe durch die weekends gearbeitet, habe ich alle Nacht gearbeitet. Ich habe mich, oh mein Gott, ich denke, das ist ein bisschen zu sein. Ich weiß nicht, dass es nicht mehr ist, weil es nicht mehr ist, weil es nicht mehr code wird. Wenn du nachdenst, wenn du jetzt auf die Momenten nach jetzt bist, was würde die meisten surprise über das Momenten jetzt sein? Es ist unglaublich, dass wir noch nicht mehr nutzen. Das war supposed to be the starting point. Ich denke, das war die Ende der Punkt. Und dann der zweite ist, dass es ebenso ist. Weil, du kennst, es nicht wirklich gut ist. Even in February, wenn wir geäte es, es hat vielleicht 10% von meinen Code oder so. Ich habe nicht wirklich benutzt, es war nicht gut. Ich habe noch nicht gut gemacht. Ich habe noch nicht mit meinem Code. So, die Fakt ist, eigentlich, unsere bets paid off. Und es wurde gut, die wir uns gedacht haben, es war nicht gut, weil es nicht so obvious. At Anthropik, die wir thought about es, wir nicht für die Model von heute. we build for the model six months from now. And that's actually still my advice to founders that are building on LLMs. Just try to think about what is that frontier where the model is not very good at today? Because it's going to get good at it. And you just have to wait. Going back, do you remember when you first got the idea? Can you just talk us through that? Was there something like a spark? Or what was even the first version of it in your mind? You know, it's funny. It was so accidental that it just kind of evolved into this. As Anthropik, I think for Ant, the VET has been coding for a long time. And the VET has been the path to safe AGI is through coding. And this has kind of always been the idea. And the way you get there is you teach the model how to code, then you teach it how to use tools, then you teach it how to use computers. And you can kind of see that because the first team that I joined at Anthropik, it was called the Anthropik Labs team. And it produced three products. It was QuadCode, MCP, and the Desktop App. So you can kind of see how these weave together. The particular product that we built, no one asked me to build a CLI. We kind of knew maybe it was time to build some kind of coding product because it seemed like the model was ready, but no one had yet really built the product that harnessed this capability. So still there's this insane feeling of product overhang, but at the time it was just even crazier because no one had built this yet. Und so, ich habe es so, wie ich es so, wie wir die Coding-Produkte haben. Was haben wir zuerst? Ich habe zu verstehen, wie ich die API habe, weil ich nicht die Anthropic-API habe. Und so, ich habe es so, wie ein Terminal-App, um die API zu nutzen. Das ist alles, was ich es. Und es war ein Chat-App. Weil, wenn du dich über die AI-Applications-Aberhafte, und für non-Coders heute, was die meisten Menschen nutzen ist, ist es ein Chat-App. So, das ist das ich. Und es war in der Terminal. Ich kann Fragen, ich gebe Fragen. Dann, ich denke Tool Use kam. Ich wollte auch Tool Use, weil ich nicht wirklich entstanden was, was ich habe. Tool Use, das ist cool. Ist das eigentlich nicht? Probably nicht. Ich habe es versucht, dass es einfach nicht. Du hast es einfach nur, weil es einfach nicht so gut ist. Ja, weil ich nicht mehr habe, weil ich nicht mehr so gut habe. Okay. Es war einfach nur mir. Das war die Idee, die Cursor und WindSurf waren, die Dinge wirklich aufzunehmen. Die Idee war, die Dinge zu tun, die sich nicht mehr so gut auszultert. Sie haben sich nicht mehr so gut auszultert. Wie soll das ein Plugin oder ein Fully Featured IDE selbst? Es war kein Pressure, weil wir nicht mehr wissen, was wir wollten. Die Team war einfach in Explore mode. Wir wussten, dass wir etwas in Coding wollten, aber es war nicht so. No-one war high-confidence. Das war mein Job. Und so, ich gave die Model den Batch Tool. Das war die erste Tool, dass ich es. Ich glaube, das war das literally die Examen in den Docs. Ich habe die Examen in Python. Ich habe es in TypeScript, weil ich es so geschrieben habe. Ich wusste, was die Model mit Batch zu tun. Ich habe es gefragt, wie es die File-File-File. Das war cool. Und dann war ich, okay, was kann ich eigentlich machen? Und ich fragte sie, was ich im Spruch am ich listening? Ich habe eine Apple-Scrip, um ich Mac zu verwenden, und ich habe die Musik in meinem Musik Player. Oh, mein Gott. Und das war Sona 3.5. Ich glaube, ich glaube, ich glaube, die Model könnte das tun. Und das war mein erstes, ich glaube, Fue the AGI moment. Es ist mir, oh mein Gott, die Model, es einfach nur für die Tools zu tun. Das ist alles. Das ist kind of fascinating. I mean, it's very kind of contrarian that Clockwork works so well in such an elegant, simple form factor. I mean, terminals have been around for a really long time, and that seemed to be like a good design constraint that allowed a lot of interesting developer experiences. It doesn't feel like working. It just feels fun as a developer. I don't think about files, where everything is. And that came by accident, almost? Yeah, it was an accident. I remember after the terminal started to take off internally, and honestly, after building this thing, I think two days after the first prototype, I started giving it to my team just for dogfitting. Because if you come up with an idea and it seems useful, the first thing you want to do is you want to give it to people to see how they use it. And then I came in the next day, and then Robert, who sits across from me, he's another engineer, he just had quadcode on his computer, and he was using it to code. I was like, what? What are you doing? This thing isn't ready. It's just a prototype. But yeah, it was already useful in that form factor. And I remember when we did our launch review to kind of launch QuadCode externally, this was in December, November or something like that in 2024. Dario asked and he was like, the Ushis chart internally, like the DAO chart is like vertical. Are you like forcing engineers to use it? Like, why are you mandating them? And I was just like, no, no, we didn't. I just like posted about it and they'd just been like telling each other about it. Honestly, it was just accidental. Wir haben mit der CLA, weil es die cheapest ist. Und es hat sich da für ein bisschen. In das 2024 period, wie waren die Engenhez using it? Waren sie sie mit dem Schiffen code mit es yet? Oder waren sie es in eine andere Weise? Der Model war nicht gut. Ich war es persönlich für Automatis Git. Ich denke, at this point, ich habe wahrscheinlich vergessen, die Git haben. Cloud Code hat sich schon so lange. Automatis Bash commands, das war ein sehr early use-case. Operatis Kubernetes und Dinge like this. People were using it for coding. So there were some early signs of this. I think the first use case was actually writing unit tests because it's a little bit lower risk and the model was still pretty bad at it. But people were kind of figuring it out and they were figuring out how to use this thing. And one thing that we saw is people started writing these markdown files for themselves and then having the model read that markdown file. And this is where QuadMD came from. Probably the single, for me, biggest principle and product is weight and demand. And just every bit of this product is built through weight and demand after their initial CLI. And so QuadMD is an example of that. There's this other general principle that I think is maybe interesting where you can build for the model and then you can build scaffolding around the model in order to improve performance a little bit. And depending on the domain, you can improve performance maybe 10, 20%, something like that. And then essentially the gain is wiped out with the next model. So either you can build the scaffolding and then get some performance gain and then rebuild it again, oder you just wait for the next model and then you kind of get it for free. The CloudMD and kind of the scaffolding is an example of that. And really, I think that's why we stayed in the CLI is because we felt there is no UI we could build that would still be relevant in six months because the model was improving so quickly. Earlier, we were saying like we should compare CloudMDs, but you said something very profound, which is, you know, yours is actually very short, which is almost like the opposite of what, you know, people might expect. Why is that? What's in your CloudMD? Okay, so I checked this before we came. So my QuadMD has two things. One is, it's just two lines. So the first line is, whenever you put up a PR, enable auto-merge. So as soon as someone accepts it, it's merged. That's just so I can code and I don't have to go back and forth with CR or whatever. And then the second one is, whenever I put up a PR, post it in our internal team stamps channel. Just so someone can stamp it and I can get unblocked. Und die Idee ist, jeder andere Instruktion ist in der Quad MD, das ist in der Codebase, und es ist etwas, das unsere Team anbieten, multiple times a week. Und sehr oft, ich sehe jemanden's PR, und sie machen einen Fehler, und ich habe einen Fehler, und ich habe einen Tag Cloud auf den PR. Ich habe einfach, Add Cloud, und ich habe das viele Times a week. Do Sie haben die Quad MD? Ich habe die Zeit, die ich an der Topdecke habe, saying your CloudMD is like thousands of tokens now. What do you do when you guys hit that? So our CloudMD is actually pretty short. I think it's like a couple thousand tokens, something like that. If you hit this, my recommendation would be delete your CloudMD and just start fresh. Interesting. I think a lot of people, they try to over-engineer this, right? And really, the capability changes with every model. And so the thing that you want is do the minimal possible thing in order to get the model on track. And so if you delete your CloudMD and then the model is getting off track, it does the wrong thing, that's when you add back a little bit at a time. And what you're probably going to find is with every model, you have to add less and less. For me, I consider myself a pretty average engineer, to be honest. I don't use a lot of fancy tools. I don't use Vim. I use VS Code because it's somewhere. Wait, really? I would have assumed that because you built this in the terminal that you were sort of like a diehard terminal, like Vim-only person, you know, screw those VS Code people. Well, we have people like that on the team. Adam Wolf, for example, he's on the team. He's like, you will never take Vim for my cold dead hands. Yeah, so there's definitely a lot of people like that on the team. And this is one of the things that I learned early on is every engineer likes to hold their dev tools differently. They like to use different tools. There's just no one tool that works for everyone. But I think also this is one of the things that makes it possible for quad code to be so good. Because I kind of think about it as what is the product that I would use that makes sense to me. And so to use quad code, you don't have to understand Vim. You don't have to understand TeamOps. You don't have to know how to SSH. You don't have to know all this stuff. You just have to open up the tool and it will guide you. It will do all this stuff. How do you decide how verbose you want the terminal to be? Sometimes you have to go Control-O and check it out. Is it like internal bike shed battles around longer, shorter? I mean, every user probably has a different opinion. How do you make those sorts of decisions? What's your opinion? Is it too verbose right now? Oh, I love the verbosity. because basically sometimes it just goes off the deep end and I'm watching and then I can just read very quickly and it's like, oh no, no, it's not that. And then I escape and then just stop it. And then it just stops an entire bug farm as it's happening. I mean, that's usually when I didn't do plan mode properly. This is something that we probably change pretty often. I remember early on, this was maybe six months ago, I tried to get rid of bash output just internally just to summarize it because I was like, these giant long bash commands, I don't actually care. Und dann, ich gebe es an Anthropik für ein paar Tage und alle einfach revoltiert. Ich wollte sie meinen Dash. Es ist eigentlich sehr gut für etwas. F etwas Git ist es nicht gut Aber wenn du kubernetes oder etwas hast du kannst es sehen Wir haben die Filereads und File Searches gesehen Du kannst du nicht sagen wie red foo es sagt, red one file, search one pattern. Und das ist etwas, ich glaube, wir nicht haben, sechs Monate, weil die Model nicht bereit war. Es ist schon schon erwähnt, dass es oft als User noch immer wieder und es immer wieder zu sein. Aber jetzt, ich habe es auf der Reihe, almost every time. And because it's using tools so much, it's actually a lot better just to summarize it. But then we shipped it, we dogfooded it for like a month, and then people on GitHub didn't like it. So there was a big issue where people were like, no, I want to see the details. And that was really great feedback. And so we added a new verbose mode. And so that's just like in slash config, you can enable verbose mode. And if you want to see all the file outputs, you can continue to do that. And then I posted on the issue and people still didn't like it, which is again, awesome, Das war's für mich. Das ist insane. Bugfixing ist einfach ein bisschen zu Sentry, Copy Markdown. Pretty soon ist es einfach ein bisschen zu bekannte MCP. Es ist ein Auto-Bugfixing und Test-Making. Was ist die neue Term, sie kenne es? Machen ein Start-up Factory. Da sind alle diese Konzepte jetzt. Rather than having to review die Code, ich bin old-school, so ich liebe die verbosity. Ich liebe es, du bist du das, aber ich will dich. that, right? But there's a totally different school of thought now that says, like, anytime a real human being has to look at code, that's bad. Yeah, yeah, yeah. Which is fascinating. I think, like, Dan Chipper talks about this a lot as kind of whenever you see the model make a mistake, try to put in the Quad MD, try to put it in, like, skills or something like this so it's reusable. But I think there's this meta point that I actually struggle with a lot. And people talk about, like, agents can do this, agents can do that. But actually what agents can do, it changes with every single model. And so sometimes there's a new person that joins the team and they actually use quad code more than I would have used it. And I'm just constantly surprised by this. Like, for example, there was a, we had like a memory leak and we were trying to debug it. And by the way, like Jared Sumner has just been on this crusade killing all the memory leaks and it's just been amazing. But before Jared was on the team, I had to do this. And there was this memory leak, I was trying to debug it. And so I took a heap dump, I opened it in DevTools, I was looking through the profile, then I was looking through the code and I was just trying to figure this out. And then another engineer on the team, Chris, he just asked QuadCode. He was like, hey, I think there's a memory leak. Can you learn this and then try to figure it out? And QuadCode took the heap dump. It wrote a little tool for itself to analyze the heap dump. And then it found the leak faster than I did. And this is just something I have to constantly relearn because my brain is still stuck somewhere six months ago at times. So what would be some advice for technical founders to really become maximalists at the latest model release? It sounds like people fresh off of school or that don't have any assumptions might be better suited than maybe sometimes engineers who have been working at it for a long time. And how do the experts get better? I think for yourself, it's kind of beginner mindset. And I don't know, maybe just like humility. Like I feel like engineers as a discipline, we've learned to have very strong opinions and senior engineers are kind of rewarded for this. In my old job at a big company, when I hired architects and this kind of type of engineer, you look for people that have a lot of experience and really strong opinions. But it actually turns out a lot of this stuff just isn't relevant anymore. And a lot of these opinions should change because the model is getting better. So I think actually the biggest skill is people that can think scientifically and can just think from first principles. How do you screen for that when you try to hire someone now for your team? I sometimes ask about what's an example of when you're wrong. Das ist ein wirklich guter. Some of these classic behavioral questions, not even coding questions, I think are quite useful. Because you can see if people can recognize their mistake in hindsight, if they can claim credit for the mistake, and if they learn something from it. And I think a lot of these very senior people, especially, there are some founder types like this, but I think founders in particular are actually quite good at it. But other people sometimes will never really take, they'll never take the blame for a mistake. But I don't know, for me personally, I'm wrong probably half the time. Half my ideas are bad. You just have to try stuff. You try a thing, you give it to users, you talk to users, you learn. Eventually, you might end up at a good idea. Sometimes you don't. This is the skill that I think in the past was very important for founders. But now I think it's very important for every engineer. Do you think you would ever hire someone based on the Cloud Code transcript of them working with the agent? Because we're actively doing that right now. We just added, just as a test, you can upload a transcript of you coding a feature with CloudCode or Codex or whatever it is. Personally, I think that it's going to work. I mean, you can figure out how someone thinks, whether they're looking at the logs or not. Can they correct the agent if it goes off the rails? Do they use plan mode? When they use plan mode, do they make sure that there are tests? All of these different things. Do they think about systems? Do they even understand systems? Like there's just so much that's sort of embedded in that, that I imagine. I just want like a spider web graph, you know, like in those video games, like NBA 2K. And it's like, oh, this person is really good at shooting or defense. It's like you can imagine a spider web graph of like, you know, someone's Claude Code skill level. Yeah, what would the skills be? I mean, I think it's like systems, testing, must be like user behavior. I mean, there's got to be a design part for sure. Like product sense. Maybe also just automating stuff. My favorite thing in CloudMD for me is I have a thing that says for every plan, decide whether it's over-engineered, under-engineered, or perfectly engineered and why. I think this is something that we're trying to figure out too. Because I think when I look at engineers on the team that I think are the most effective, there's essentially two. It's very bimodal. There's one side where it's extreme specialists. And so I named Jared before. He's a really good example of this. And kind of the BUN team is a really good example. just hyper-specialist. They understand DevTools better than anyone else. They understand JavaScript runtime systems better than anyone else. And then there's the flip side of kind of hyper-generalists and that's kind of the rest of the team. And a lot of people, they span like product and infra or product and design or, you know, like product and user research, product and business. I really like to see people that just do weird stuff. I think that's one of these things that was kind of a warning sign in the past because it's like, can these people actually build something useful? Das ist die Littmuss test. that that works. And then she put up that PR. And then she had Quad write its own tool instead of herself implementing it. And I think it's this kind of out-of-the-box thinking that is just so interesting because not a lot of people get it yet. You know, like we use the Quad Agent SDK to automate pretty much every part of development. It automates code review, security review. It labels all of our issues. It shepherds things to production. It does pretty much everything for us. But I think externally, I'm seeing a lot of people start to figure this out. But it's actually taken a while to figure out How do you use LMs in this way? How do you use this new kind of automation? So it's kind of a new skill. I guess one of the funnier things that I've been having office hours with various founders about is you have sort of the visionary founder who has the idea. They've built this crystal palace of the product that they want to build. They've totally loaded in their brain who the user is and what they feel and what they're motivated by. And then they're sitting in Cloud Code and they can do 50x work. But they have engineers who work for them who don't have the crystal memory palace of the platonic ideal of the product that the founder has. And they can only do 5x work. Are you hearing stories like that? There's usually a person who's the core designer of a thing and they're just trying to blast it out of their brain. What's the nature of teams like that? It seems like that's almost a stable configuration. Like you're going to have the visionary who like now is unleashed. But, you know, maybe going back to the top of it, like I'm experiencing this right now. It's like, oh, well, I'm only a solo person and, you know, I need to eat and sleep and I have, you know, a whole job. It's like, how am I going to do this? You know, you know, like we just launched quad teams and, you know, this is a way to do it. But you can also just build your own way to do it. It's pretty easy. What's the vision for quad teams? It's collaboration. Es ist ein ganzes Witz, der sich in den ganzen Fällen von Agenten zu lösen. Das ist ein ganzes Witz, wie man die Wäste kann, die sich in den Regen zu lösen. Es ist eine eine Sub-Idea, die ist un- und Correlated Kontext Windows. Und die Idee ist, dass es nur Multiple Agents haben, die sich in den gleichen Kontext zu lösen, mit den anderen Kontexten oder ihren eigenen Kontexten. Und wenn man mehr Kontexten hat, das ist ein Form von Test-Time-Compute. Und so, du hast mehr Witz, das ist. Und dann, wenn man die Witz auf dem Topolge auf dem, so die Agents kann kommunizieren, die Witz aus, in den richtigen Witz, then they can just build bigger stuff. And so Teams is kind of like one idea. There's a few more that are coming pretty soon. And the idea is just maybe it can build a little bit more. I think the first kind of big example where it worked is our plugins feature was entirely built by a swarm over a weekend. It just ran for like a few days. There wasn't really human intervention. And plugins is pretty much in the form that it was when it came out. How did you set that up? Did you spec out sort of the outcome that you were hoping for? and then let it sort of figure out the details and then like let it run? Yeah, an engineer on the team just gave Quad a spec and told Quad to use a Asana board. And then Quad just put up a bunch of tickets on Asana and then spawned a bunch of agents and the agents started picking up tasks. The main Quad just gave it instructions and they all just figured it out. The independent agents that didn't have the context of the bigger spec, right? Right. If you think about the way that, wie wir eigentlich starten nowadays. Und ich habe den Daten auf das, aber ich glaube, die majority of agents sind eigentlich prompted by Claude today in der Form von subagents. Because a subagent ist einfach ein recursive Claude code. Das ist all in der Code. Und es ist just prompted by, we call her Mama Claude. Und das ist all. Und ich glaube, if you look at most agents, they're launched in this way. My Claude insights just told me to do this more for debugging. I spend Ich habe ein paar Tage Zeit, und es wäre besser, zu haben, viele Subagents zu spüren und zu einem anderen Anwender etwas in parallel. Und dann habe ich das an, dass ich das mit CloudMD zu sagen, dass ich eine Agenten, die sich in den Logen oder eine Art, die in den CodePath ist. Das ist einfach nicht so gut. Für weird, scary bugs, ich versuche, die Bugs in Plan Mode zu versuchen, und dann es sich um die Agents zu suchen. Ja. Wenn du es einfach in der Linie machen, es ist einfach, okay, ich werde das eine Task instead of search wide. This is something I do all the time too. I just say, if the task seems kind of hard, this kind of research task, I'll calibrate the number of subagents I ask it to use based on the difficulty of the task. So if it's like really hard, I'll say like use three or maybe five or even 10 subagents, research in parallel, and then see what they come up with. I'm curious, so then why don't you put that in your clawed MD file? It's kind of case by case, you know, like clawed MD, like what is it? It's just a, it's a shortcut. Like if you find yourself repeating the same thing over and over, you put in the Quad MD. But otherwise, you don't have to put everything there. You can just prompt Quad. Are you also in the back of your mind thinking that maybe in six months you won need to prompt that explicitly The model would just be good enough to figure it out on its own Maybe in a month No more need for Plan Mode in a month Oh my god I think Plan Mode probably has a limited lifespan Interesting. That's some alpha for everyone here. What would the world look like without Plan Mode? Do you just describe it at the prompt level and it would just do it, one shot it? Yeah, we've started experimenting with this because Cloud Code can now enter Plan Mode by itself. I don't know if you guys have seen that. So we're trying to kind of get this experience really good. So it would enter plan mode the same point where a human would have wanted to enter it. So I think it's something like this. But actually plan mode, there's no big secret to it. All it does is it adds one sentence to the prompt that's like, please don't code. That's all it is. You can actually just say that. So it sounds like a lot of the feature development for ClockCode is very much what we talk about at YC. Talk to your users and then you come and implement it. It wasn't the other way that you had this master plan and then implemented all the features. Yeah, yeah. Ja, ich meine, das war es. Plan Mode war, wir sahen users, dass wir uns, hey, Claude, kommen mit einer Idee, plan das aus, aber nicht mehr schreiben. Und da waren verschiedene Versionen. Sometimes es war einfach nur ein Idee. Sometimes es war diese sehr sophisticated Specs das sie zu fragen, Claude zu schreiben. Aber die Common Dimension war, was das nicht ohne Kunde. Und so, literally, das war Sunday night at 10 p.m. Ich war einfach nur auf GitHub Issues und sehen, was die Leute reden. Und dann war die Internet Slack-Feedback-Channel. Und ich habe das Buch geschrieben, in 30 Minuten. und dann shipped it that night. It went out Monday morning. So do you mean that there will be no need for plan mode in the sense of I'm worried that the model is going to do the wrong thing or head off in the wrong direction, but there will still be a need for that? You need to think through the idea and figure out exactly what it is that you want and you have to do that somewhere. I kind of think about it in terms of increasing model capabilities. So maybe six months ago, a plan was insufficient. So you get Claude to make a plan, Let's say even with plan mode, you still have to sit there and babysit because it can go off track. Nowadays, what I do is probably 80% of my sessions, I say plan mode has a limited lifespan, but I'm a heavy plan mode user. Probably 80% of my sessions, I start in plan mode. And Claude will start, it'll start making a plan, I'll move on to my second terminal tab, and then I'll have it make another plan. And then when I run out of tabs, I open the desktop app, and then I go to the code tab, and then I just start a bunch of tabs there. And they all start in plan mode, probably 80% of the time. Once the plan is good and sometimes it takes a little back and forth, they just get Claude to execute. And nowadays what I find with Opus 4.5, I think it started with 4.6, it got really good. Once the plan is good, it just stays on track and it'll just do the thing exactly right almost every time. And so, you know, before you had to babysit after the plan and before the plan. Now it's just before the plan. So maybe the next thing is you just won't have to babysit. You can just kind of give a prompt and Claude will figure it out. The next step is Claude just speaks to your users directly. Ja, es ist ein bisschen Entire. Es ist ein bisschen, das ist eigentlich die Current-Sephraus. Wir quads, eigentlich, they talk to each other, they talk to our users on Slack, at least internally, pretty often. My quad will tweet once in a while. No way. But I actually delete it. It's a little cheesy. I don't love the tone. What does it want to tweet about? Sometimes it'll just respond to someone, because I always have co-work running in the background, and it's the co-work quad that really loves to do that, because it likes using a browser. That's funny. I really common pattern is I ask Quatt to build something, it'll look in the code base, it'll see some engineer touch something in the Git flame, and then it'll message that engineer on Slack, just like asking a clarifying question. And then once it gets the answer back, it'll keep going. What are some tips for founders now on how to build for the future? It sounds like everything is really changing. What are some principles that will stay on and what will change? So I think some of these are pretty basic, but I think they're even more important now than they were before. So one example is latent demand. I mentioned it a thousand times for me. It's just the single biggest idea in product. It's a thing that no one understands. It's a thing I certainly did not understand in my first few startups. And the idea is people will only do a thing that they already do. You can't get people to do a new thing. If people are trying to do a thing and you make it easier, that's a good idea. But if people are doing a thing and you try to make them do a different thing, they're not going to do that. And so you just have to make the thing that they're trying to do easier. And I think Quad is going to get increasingly good at kind of figuring out these kind of product ideas for you just because it can look at feedback, it can look at debug logs, kind of figure this out. That's what you mean by plan mode was latent demand, that people already had their clawed chat window open in the browser and were talking to it to figure out the spec and what it should do. And now plan mode just became that, you just do it in clawed code. Yeah, yeah. Sometimes what I'll do is I'll just walk around the office on our floor and I'll just kind of stand behind people. I'll say hi, so it's not good. And then I'll just see how they're using quadcode. And this is also just something I saw a lot. But it also came up in GitHub issues, people were talking about it. It seems like you're surprised how far the terminal has gone and how far it's been pushed. How far do you think it has left to go, just given with this world of SWAR, multiple agents? Do you think there's going to be a need for a different UI on top of it? It's funny, if you asked me this a year ago I would have said the terminal has like a 3 month lifespan and then we're going to move on to the next thing and you can see us experimenting with this right, because Quad Code started in a terminal but now it's in, you know, it's on web like Quad AS hash code, it's in the desktop app, you know, we've had that for like 3 months or 6 months or something just in the code tab it's in the iOS and Android apps, just like in the code tab, it's in Slack, it's in GitHub there's VS Code extensions, there's JetBrains extensions, so we're just like we're always experimenting with different form factors for this thing to figure out what's the next thing. I've been wrong so far about the lifespan of the CLI, so I'm probably not the person to forecast that. What about your advice to DevTool founders? Someone's building a DevTool company today. Should they just be building for engineers and humans, or should they be thinking more about what Claude's going to think and want and build for the agent? The way I would frame it is think about the thing that the model wants to do und zu setzen sich her. Und das ist das难 vonheier. Das dachte ich Creating многие Jäsen Brutke aus der Dette päin hat die to子in Tabakkonet certas maravilimm. Auch Serviette Dette Monster sind das steht an hier und esこれは zuやって die Welt版ınd ist dies die Tools ist zu tun das experts für your users. If you're Building Wenn ich eine DevTools starte habe, ich würde sagen, was die Problemen Sie wollen für die User wollen? Und dann, wenn Sie die Model zu lösen, was die Problemen die Model wollen? Und dann, was die Technik und Product-Solutionen die die Werte und Demand für beide? YC's Next Batch ist jetzt auf die App-Application. Got eine Start-up in Sie? Apply at ycombinator.com. It's never too early, and filling out the app will level up your idea. Okay, back to the video. Back in the day, more than 10 years ago, you were a very heavy user, and you wrote a book about TypeScript, right? Before TypeScript was cool. This is when everyone was deep in JavaScript. This is back in early 2010s, right? Yeah, something like that. Before TypeScript was a thing, because back then it's a very weird language. It's not supposed to do a lot of things with being typed in JavaScript. And now it's the right thing. And it feels like Cloud Code in the terminal has a lot of parallels with TypeScript at the beginning. TypeScript makes a lot of really weird language decisions. So if you look at the type system, pretty much anything can be a literal type, for example. And this is super weird because even Haskell doesn't even do this. It's just like it's too extreme. or it has conditional types, which I don't think any language thought of at all. It was very strongly typed. Yeah, it was very strongly typed. And the idea was when Joe Pamer and Anders and the early team was building this thing, the way they built it is, okay, we have these teams with these big untyped JavaScript code bases. We have to get types in there, but we're not going to get engineers to change the way that they code. You're not going to get JavaScript people to have 15 layers of class inheritance like you would a Java programmer. They're going to write code the way they're going to write it. They're going to use reflection and they're going to use mutation. And they're going to use all these features that traditionally are very, very difficult to type. They're a very unsafe type to any strong functional programmer. That's right. That's right. That's right. And so the thing that they did, instead of getting people to kind of change the way that they code, they built a type system around this. And it was just, it's brilliant because there's all these ideas that no one was thinking about. Even in academia, like no one thought of a bunch of these ideas. it purely came out of the practice of observing people and seeing how JavaScript programmers want to write code. And so for quad code, there are some ideas that are kind of similar in that you can use it like a Unix utility. You can pipe into it, you can pipe out of it. In some ways it is kind of rigorous in this way, but in almost every other way, it's just the tool that we wanted. Like I built a tool for myself and then the team built the tool for themselves and then for anthropic employees and then for users. And it just ends up being really useful. It's not this principled and academic thing. Which I think the proof is actually in the results. Now, fast forward more than 15 years later, not many codebases are in Haskell, which is more academic. And there's tons of them now in TypeScript, because it's way more practical. Right. Which is interesting. Yeah, it is interesting, right? It's like TypeScript solves a problem. I guess one thing that's cool, I don't know how many people know, but the Terminal is actually one of the most beautiful Terminal apps out there, und ist eigentlich mit React Terminal. Wenn ich erst begleitet, ich habe vor allem gemacht, und ich war auch ein Hybrid. Ich mache Design und User Research und RedCode und all das, und wir lieben Hiring Engineers das sind so wie das. Wir lieben Generalist. So für mich, es ist okay, ich bin ein Ding für die Terminal. Ich bin eigentlich ein shitty Vim-User. So wie ich buildet ein Ding für Leute wie mich das werden, die ich in einem Terminal arbeiten. Und ich denke, just the delight is so important and i feel like i see this as something you talk about a lot right it's like build a thing that people love if the product is useful but you don't fall in love with it that's not great um so it kind of has to do both designing for the terminal honestly has been hard right it's like a it's like 80 by 100 characters or whatever you have like 256 colors you have one font size you don't have like mouse interactions there's all this stuff you can't do and there's all these very hard trade-offs so like a little known thing for example is you can actually enable mouse interactions in a terminal. So you can enable clicking and stuff. Oh, how do you do that in Cloud Code? I've been trying to figure out how to do this. We don't have it in Cloud Code because we actually prototyped it a few times and it felt really bad because the trade-off is you have to virtualize scrolling. And so there's all these weird trade-offs because the way terminals work is like, there's no DOM, right? It's like there's anti-escape codes and these kind of weird organically evolved specs since like the 1960s or whatever. Oh yeah, it feels like BBSs. It's like a BBS door game. Yeah, yeah, yeah. Oh my gosh. That's like a great compliment. Ja. Ja, ja. Und du miejscest du lustig, endlich. Lord of the Red Dragons. Fantastisch. Oh, mein Gott. Ja. But we've had to just discover all these UX principles for building the terminal. Because no one really writes about this stuff. And if you look at the big terminal apps of the 80s or 90s or 2000s or whatever. They use Ed Ker cousin's and they have all these windows and things like this. And it just looks kind of janky by modern standards. It just looks too heavy and complicated. And so, we had to reinvent a lot. And, for example, something like the terminal spinner. just like the Spinner words, it's gone through probably I want to say like 50, maybe 100 iterations at this point and probably 80% of those didn't ship. So we tried it it didn't feel good, move on to the next one. Try it, didn't feel good, move on to the next one. And this was like sort of one of the amazing things about QuadCode, right? It's like you can write these prototypes and you can just do like 20 prototypes back to back, see which one you like and then ship that and the whole thing takes maybe a couple hours. Whereas in the past what you would have had to do is like weren't to use Origami or Framer oder so Let build a product that joyous and that people like to use Boris you had other advice for builders and we kept interrupting you because we have so many questions I would say maybe two pieces of advice that are kind of weird because it's about building for the model. So one is don't build for the model of today, build for the model of six months from now. This is sort of weird because you can't find PMF if the product doesn't work. But actually, this is the thing that you should do because otherwise what will happen is you spend a bunch of work. find PMF for the product right now, and then you're just going to get leapfrogged by someone else because they're building for the next model, and a new model comes out every few months. Use the model, feel out the boundary of what it can do, and then build for the model that you think will be the model maybe six months from now. I think the second thing is, you know, actually in the quad-code area where we sit, we have a framed copy of The Bitter Lesson on the wall. And this is this like Rich Sutton quad-code, like everyone should read it if you haven't. And the idea is the more general model will always beat the more specific model. And there's a lot of corollaries to this, but essentially what it boils down to is never bet against the model. And so this is just like a thing that we always think about, where we could build a feature into quad code, we could make it better as a product, and we call this scaffolding. It's all this code that's not the model itself. But we could also just wait like a couple months and the model can probably just do the thing instead. And there's always this trade-off, right? It's like engineering work now, and you can extend the capability a little bit, maybe 10-20% or whatever in whatever domain on this spider chart of what you're trying to extend. Or you can just wait and the next model will do it. So just always think in terms of this trade-off. Where do you actually want to invest? And assume that whatever the scaffolding is, it's just tech debt. How often do you rewrite the codeways of clock code? Is it every six months with this first visible? Is there scaffolding that you've deleted because you don't need it anymore because the model just improved? Oh, so much. Ja, all of QuadCode has just been written and rewritten and rewritten and rewritten over and over and over. We unship tools every couple of weeks. We add new tools every couple of weeks. There's no part of QuadCode that was around six months ago. It's just constantly rewritten. Would you say most of the code base for our current QuadCode is only, say 80% of it is only less than a couple of months old? Yeah, definitely. It might even be less than, yeah, maybe a couple of months. That feels about right. Just like the lifecycle of CodeNow, that's another alpha. expecting it to be the shelf life to be just a couple months for the best founders. Did you see Steve Yeggie's post about how awesome working at Anthropic is? And I think there's a line in there that says that an Anthropic engineer currently averages 1,000x more productivity than a Google engineer at Google's peak, which is really an insane number. Honestly, like 1,000x. Three years ago, we were still talking about 10x engineers. Now we're talking about 1,000x on top of a Google engineer in the prime. This is unbelievable, honestly. Yeah, I mean, internally, if you look at technical employees, they all use QuadCode every day. And even non-technical employees, I think half the sales team uses QuadCode. They've started switching to Co-Work because it's a little easier to use. It has a VM, so it's a little bit safer. But yeah, we just pulled the stat, and I think the team doubled in size last year, but productivity per engineer grew something like 70%. As measured by? Just like the simplest, stupidest measure, pull requests. But we also kind of cross-check that against commits and the lifetime of commits and things like this. And since QuadCode came out, productivity per engineer at Anthropic has grown 150%. Oh my God. And this is crazy because in my old life, I was responsible for code quality at Meta. And I was responsible for the quality of all of our code bases across every product, across Facebook, Instagram, WhatsApp, whatever. And one of the things that the team worked on was improving productivity. Und dann, ich habe ein Verwalt von 2% in Produktivität, das war ein Jahr von Arbeit bei 100% von Leuten. Und so 100% ist das unheard von, einfach unheard von. Was hat mich auf die Anthropik gemacht? Basically, als ein Builder, du kannst du irgendwo gehen. Was war die Moment, das macht dich sagen, das ist die Setup-Piefe oder das ist die Approche? Ich war in rural Japan, und ich war öffnig, Hacker News, und ich war öffnig, die News. und es war alles, es wurde ausgerichtet, wie ich AI-Stuff hatte, und ich habe diese Anerkennung zu nutzen, und ich habe die ersten paar Tage, ich habe es einfach ein bisschen mehr ausgedrückt. Das war sehr, aber das war eigentlich das Gefühl. Es war einfach das Gefühl, als ein Builder, ich habe das Gefühl, wie ich in diesen sehr, sehr, sehr early Produkten das in den Quad 2 Dase oder so, so ich ich habe zu sprechen, mit Ich habe einen Freund mit Lab, um zu sehen, was ich was. Und ich habe einen Mann, der ist ein Founder an Anthropik. Und er hat mich sofort über mich. Und als ich mit den Rest der Team, er hat mich über mich über. Und ich denke, es ist zwei Worte. Es ist ein, es ist ein Research Lab. Der Produktions-Tätig war es wirklich über die Safe Model. Das ist alles, was das alles. Und so diese Idee, die sich sehr close zu den Model und sehr close zu der Entwicklung, Und das ist nicht die wichtigste Sache, weil der Produkt nicht mehr ist. Das ist die Model ist die Sache, die das wichtigste. Das wirklich mit mir, nach dem Produkt für viele Jahre. Und dann die zweite Sache war, wie mission-driven es ist. Ich bin ein großer Sci-Fi-Reader. Mein Buchstuf ist einfach mit Sci-Fi. Und so ich weiß, wie bad das kann. Und wenn ich denke, was wird passieren, es wird total insane sein. Und in der worste, es kann sehr, sehr, sehr bad. und so I just wanted to be at a place that really understood that and kind of really internalized that and at Ant, you know, like if you overhear conversations in the lunchroom or in the hallway people are talking about AI safety this is really the thing that everyone cares about more than anything and so I just wanted to be in a place like that I know for me personally the mission is just so important. What is going to happen this year? Okay, so if you think back like six months ago and kind of what are the predictions that people are making, so Dario predicted that 90% of the code at Anthropic would be written by Quad. This is true. For me personally, it's been 100% since Opus 4.5. I uninstalled my IDE. I don't edit a single line of code by hand. It's just 100% Quad code in Opus. And I land 20 PRs a day, every day. If you look at Anthropic overall, it ranges between 70% to 90%. Depending on the team, for a lot of teams it's also 100%. For a lot of people, it's 100%. Und ich habe mich über das Prediktion in May, wenn wir GA-Di Cloud Code, dass man nicht mehr braucht, dass man nicht eine Idee zu code anymore. Und es war total crazy zu sagen. Ich fühle mich die Leute in der Audience gasped, weil es so eine silly Prediktion hat. Aber wirklich, all es ist, ist, dass du einfach den Exponential trast. Und das ist so deep in den DNA. Denn drei unserer Founders waren Co-Authors der Scaling Loss Paper. Sie sah das sehr early. Und so, das ist einfach Tracing Exponential. Das ist was, was wird passieren. Und, yes, das ist, das passiert. So continuing to trace the exponential, I think what will happen is coding will be generally solved for everyone. And I think today coding is practically solved for me, and I think it'll be the case for everyone, regardless of domain. I think we're going to start to see the title Software Engineer go away. And I think it's just going to be maybe Builder, maybe Product Manager. Maybe we'll keep the title as kind of a vestigial thing. But the work that people do, it's not just going to be coding. It's Software Engineers are also going to be writing specs. They're going to be talking to users. This thing that we're starting to see right now in our team where engineers are very much generalists and every single function on our team codes, like our PM's code, our designer's code, our EM codes, our finance guy codes, everyone on our team codes, we're going to start to see this everywhere. So this is kind of like the lower bound if we just continue the trend. The upper bound, I think, is a lot scarier. und das ist etwas wie wir hit ASL4 und an Anthropik, wir haben die Safety Levels ASL3 ist wo die Models sind jetzt ASL4 ist die Models ist recursively self-improving und so wenn das passiert, wir müssen wir ein paar Kriterien bevor wir können die Models release. Und so die Extreme ist, dass das passiert. Oder es gibt eine Katastrophic Misuse. People sind die Models zu designen, wir sind Zeroday, das ist das. Und das ist das wir wirklich, wirklich aktiv arbeiten, so das nicht passieren. I think it's just been, honestly, it's just been so exciting and humbling, seeing how people are using QuadCode. I just wanted to build a cool thing, and it ended up being really useful. And that was so surprising and so exciting. My impression from Twitter or just the outside is basically everyone went away over the holidays and then found out about QuadCode, and it's just been crazy ever since. Is that how it was for you at internment? Were you having a nice Christmas break and then came back and you're like, what happened? Well, actually, for all of December, I was traveling around, and I took a coding vacation. So we were kind of traveling around, and I was just coding every day. So that was really nice. And then I also started to use Twitter at the time, because I worked on Threads back then, way back when. So I've been a Threads user for a while. So I just tried to see other platforms where people are. Yeah, I think for a lot of people, that was the moment where they discovered Opus 4.5. I kind of already knew. And internally, Cloud Code's just been on this exponential tear for many, many months now. So that just like it became even more steep. That's what we saw. And if you look at QuadCode now, you know, there was some stat from Mercury that like 70% of startups are, you know, choosing Quad as their model of choice. There were some other stat from like Semi Analysis that 4% of all public commits are made by QuadCode, like of all code written everywhere. All the companies, you know, use QuadCode from like the biggest companies to kind of, you know, smallest startups. It plotted the course for Perseverance, for the Mars rover. This is the coolest thing for me. We even printed posters because the team was like, wow, this is just so cool that NASA chooses to use this thing. So yeah, it's humbling, but it also feels like the very beginning. What's the sort of interaction between CloudCode and then CoWork? Was it a fork of CloudCode? Was it like you had CloudCode, look at the CloudCode code and say, Let's make a new spec for non-technical people that keeps all the lessons. And then it sort of went off for a couple of days and did that. What's the genesis of that? And where do you think that goes? This is going to be my fifth time using the word wait and demand. It was just that. I mean, we were looking at Twitter and there was that one guy that was using QuadCode to monitor his tomato plants. There was this other person that was using it to recover wedding photos off of a corrupted hard drive. There are people that are using it for finance. When we looked internally at Anthropic, every designer is using it. The entire finance team at this point is using it. The entire data science team is using it, not for coding. People are jumping over hoops to install a thing in the terminal so that they can use this. So we knew for a while that we wanted to build something. And so we're experimenting with a bunch of different ideas. And the thing that kind of took off was just a little quadcode wrapper in a GUI in the desktop app. That's all it is. It's just quadcode under the hood. It's the same agent. Oh, wow. und Felix und die Team, und Felix war der Early Electron contributor. Er kennt das Stack wirklich gut, und er hat sich auf verschiedene Ideen. They built es in, ich denke, 10 days. Es war 100% written by QuadCode, und es war bereit, und es war bereit, zu release. Es war ein paar Sachen, wir hatten für nicht-technicale User, so es ein bisschen anders als Technikall Audience. All die Code runs in eine Virtual Machine. Es gibt viele protections für Deletion und Dinge wie das. There's a lot of permission prompting and other guardrails for users. But yeah, it was honestly pretty obvious. Boris, thank you so much for making something that is taking away all my sleep. But in return, it's making me feel creator mode again, sort of founder mode again. It's been an exhilarating three weeks. I can't believe I waited that long since November to actually get into it. Thank you so much for being with us. Thank you for building what you're building. Yeah, thanks for having me. Und send bugs. Sounds good.",
  "transcript_chars": 54620,
  "transcript_filled_at": "2026-05-23T11:20:37.749981+00:00",
  "transcript_filled_by": "groq"
}