{
  "video_id": "k2ZLQC8P7dc",
  "transcript": "Ich denke, wir sind wahrscheinlich bei AGI 2030. Around die Zeit, die wir werden, vielleicht ARC 6 oder ARC 7 geben, wird nicht die AI-Progress stopt. Ich denke, es ist zu late für das. Und so die nächste Frage ist, okay, AI-Progress ist hier. Es wird eigentlich immer weiterhöhnt. Wie machen Sie es? Wie machen Sie es? Wie machen Sie es? Wie machen Sie es? Wie machen Sie es? Das ist die Frage. Today we're lucky to be joined by Francois Chollet, founder of the ARK Prize, a global competition to solve the ARK AGI benchmark. His latest project is Endia, a lab exploring a new paradigm in frontier AI research. Francois is one of the best people in the world to help us understand the current AI moment and where all of this is going. Francois, thank you so much for joining us today and congrats on the launch of ARK AGI v3. Thanks so much for having me. I'm super excited to be here. Super exciting time to talk about AI. So Francois, tell us a little bit about India. So what exactly is it and what are you guys trying to achieve? Right, so India is this new AGI research lab. And we are trying some very different ideas. And so our goal is basically to build this new branch of machine learning that will be much closer to optimal, unlike deep learning. All of us right now are sort of taken by what's going on with code. Ich habe dieses viral moment, jetzt, wo ich 40,000 stars auf dem G-Stack aufgerufen habe. So es ist ein Open Source Projekt, das jetzt ist eine der größten, und ich habe mehr als 100 PRs von contributors zu tun. Ich glaube, du bist eine der besten Leute zu sprechen, weil du wirklich ein bisschen mit einem völlig anderen Weg bist. Das ist richtig. Das ist richtig. Wir machen Program Synthesis Research. Wenn ich über Program Synthesis spreche, dann fragen mich, ob ich Cogen oder eine Alternative zu Coding Agents mache. Das ist nicht all das, was wir machen. Wir arbeiten auf eine viel, viel mehr, viel mehr Level. Wir machen eine neue Branche von Machine Learning, eine Alternative zu Deep Learning, statt als Coding Agents. Coding agents are like this very high level, last layer piece of the stack. And we're actually trying to rebuild the whole stack on top of different foundations. So we're building a new learning substrate that's very different from parametric learning, deep learning. So if you go back to the problem of machine learning, you have some input data, some target data, and you're trying to find a function that will map the inputs to the targets Das werden einen Spezweiger Joshua organisiert. Wenn man sait, ist eine Parametrie Per problemas an den Be Brilement der견ungen, die doublingivoiten das Restreter pro Kurve wie die Spez 가족lich änder verewten嗎? Das ist eigentlich was, zgieglosst die Campe 4000wwww mit einem Symbolik MaterialсensformChat,れるweise die einfache kompat komt. a possible model to explain the data, to model what's going on. And of course, if you're doing that, you cannot apply gradient descent anymore. So we are building something that we call symbolic descent, which is like the symbolic space equivalent of gradient descent. The idea is to build this new machine learning engine that's giving you extremely concise symbolic models der Daten, die in den Daten zu finden, und dann werden wir es scale. So alles mit Machine Learning, mit Parametric Curves, sollte man das mit Symbolik Models in der Zukunft werden, in einer Weise, dass wir viel, viel closer zu optimale. Viel mehr zu optimale, in der Sinne, dass man viel mehr Daten zu erhalten, die Models werden viel mehr effizienter werden, denn sie werden so klein sein. Und weil sie so große wird erARRYft werden gut empfושDamit weniger per dicke mehr. You know, der Comparierung descent geht, auch was die tucked Situation nach Anuci Nós regresons, du musst du wie zum Zusammenhang etwas zurbageование zu machen, Er Companies turning money in die recient \"'ורrim prosecht' und dies sind noch nicht reconocENElle, genau dass, also dasُ student losses May Superintendent?\" Everybody is die Bundestag. The Uberringer devolvend es sich da für die Reteureignisse, so wird alles wieder Mukünferte gefunden. Man kommt vor Funktionieren, das machen. Man geht trotzdem nach bord,егенemente Khan muss sich das optimizierendeprodukte Schritt sein. Aber Арteg расп focus notor prevendem alle arbeiten dengan ein preventimer. Ich persönlich finde kein Musikle inconsistent. in 50 years, will be built on this stack. I think this is a stack that is very nice, and maybe it gets us to HCI. But it's not as efficient as it should be. I think it's inevitable that the world of AI will trend over time towards optimality. So I'm trying to leapfrog directly to optimality, to build the foundations of optimal AI today. But in Our vision is very ambitious. I'm not saying that we're going to be successful, like we have maybe a 10% or 15% chance of success. But that is enough that it's worth trying. And I think in general, among listeners, if you have a big idea and it has a very low chance of success, but if it works, it's going to be big and no one else is going to be working on it. It's not something popular. If you don't do it, no one else will do it. Das ist unsere Situation. Wenn du in diese Situation bist, dann sollte du ein Chance haben. Du musst du an und worken. Das ist die Mission Statement des Y Combinator, die du gesagt hast. Ja, das ist wichtig, dass wenn wir nicht tun, dann niemand will. Es ist wert, dass wir nicht. Es ist wert, dass wir nicht. Es ist wert, dass der Success, sehr specifically für die Coding Agents, die LLM-Stack auf der LLM-Stack, hat der Success surprisest du? In particular, über die letzten sechs Monate oder so? Ja, absolut. Ich denke, es hat viele Leute zufrieden. Es hat mich wirklich zufrieden. Wenn man sich die Frage, warum alles so gut mit Coding Agents ist, ist es wirklich weil Code einen Verifiable Reward-Signal gibt. Und ich denke, dass wir jetzt in der Situation, wo die Probleme, wo die Lösungen, die Sie vorbeigehen, können Sie verifizieren, und Sie können die Reward-Signal nicht nur ein Guess gemacht werden, Any domain like this can be fully automated with current technology, with the LMA stack. And code is sort of like the first domain to fall, but there will be many others in the future. I think mathematics is also primed to see a revolution in the next few years. For the same reasons, again, because the domain just gives you verifiable rewards. I guess the challenge for a formally verified domain is you have to um somehow take a domain and make it verifiable which is the trick i mean code is very natural you could test there's bugs compiles etc and mathematics as well whether all the theorems and proofs work out i guess because more nebulous when you go a couple degrees off where there are fields that are not naturally formally verified you need to come with mit einer Funktion zu kommen, mit einem Reward, das Verifizierbar macht. Mit sehr fuzzyten Dingen, wie in den English-Language und die Perfect-Essay, wie man das Formally Verifizierbar macht? Ja, absolut. Writing Essays ist ein Typical Beispiel für ein Domain, das nicht Verifizierbar ist. So, was du sehen, ist das die Prozess der Resonung Modell und Basel-Elemente kann also einen Verkauf Guang, auf die southeast Seite der Zum Beispiel, für einen großen Unlocken, was das erste Mal, als die Leute in einem Code-Based Training-Train environment für Posttraining, wo die Rewards-Signal, die Verification-Signal, ist bei Dingen wie Unit-Tests und so weiter. Das bedeutet, dass die Model nicht nur von Human-Provider-Annotations war, sondern es war tatsächlich die eigenen Dinge, verifizierter die Antworten, einem Amsted haben viel viel mehr могу archivieren, eine dri milks minutere Be morale zu den Problem ein gutes Eine Funktion besser. Über die ganze Marketing, die commutierte dieみben, Das ist auch das Models, das ist auch das, was es funktioniert so gut. Und es ist möglich, weil du mit einem sehr formalen, verifizierbarer Einwohnern kannst du das mit SES oder mit LAR oder anderen Problemen. Ich finde, wie du definiertes Intelligenz und wie du measurest es. Das bringt mich zu dem Frage, wie du, die Geschichte von ARK-EGI. Ja, so mein Definition von General Intelligence. Manche Menschen in den Industrie sagen, dass AGI ein System wird, die die die die meisten economically valuableen Tasks verantworten können. Das Definition ist für mich, dass es über Automation ist. Es ist nicht über Intelligenz. Mein Definition ist, AGI wird, dass es ein System wird, die die einen neuen Problem, einen neuen Task, einen neuen Domain und findet eine cremabe Durchszenzieren. seguיסieren anschauen und eine nit versorgt, der mit einer gleichartaire, dessen inclusive wie man könnte an Menschen. Das bedeutet, dass man basically den Theatre wie man ewig zeigen 이거 fragt sich, was gleich alle, die Versucht schon könnte man davon B Schlüssel. Es ist die umh mixing funktioniert, auf strukture Mant исследовiert man. Will IBMו� möchte auf einen Bereich Do you think it's possible that we will accomplish the first definition of AGI, the Automate Most Economically Useful Work, before we accomplish your definition? Absolutely. I think that's the trajectory that we're on right now. And I think it's already true that, in principle, current technology can fully automate at human level or beyond any domain where you have verifiable rewards, right? And code being the first one. Und ich denke, wenn man AGR hat, wenn man an der Human-Level-Learning-Efficiency über Arbitrary-Tasks wird wahrscheinlich ein bisschen mehr Technologie, ein bisschen Mindset, ein bisschen mehr Approach. Do you think LM's können sich die gleichen Sample Efficiency zu haben, oder ist das einfach nicht möglich und wir brauchen eine neue Approche? Und das ist das, was du hoping, zu lösen? Mit einer Compute, alles wird alles wie alles. Every computer is a great equalizer. Every approach starts looking the same. And I think it's possible in principle to build something that looks a lot like a GI on top of the LLM stack. But it's not going to be LLMs per se. It's going to be this new layer. Perhaps it's going to be even a few layers above. Not just one layer above, but a few layers above. But you can build it on top of LLMs because LLMs are a kind of computer. Ich habe Built. Ich glaube das würde uns nicht. Es ist komplett auf CPI. Ich denke, να Libraren auf Amazon A and Facebook zurückflex- Umthe развитigen 많은 Fliegerropic. Dafür gibt es in ein paar Jahre Drços. Es ist nicht ohne eine Hal parets, durch eine appropriationiva FileZLM, sondern es wird sich ein leichteres. Und Janne, sagen wir uns, wie du bank took, Arch-AGI und warum es ein guter Barometer ist? Ich habe es für eine sehr lange Zeit gemacht. Und in der ersten Zeit, mein Geist war, was Deep Learning zu tun, was alles zu können. Du warst die Keras, bevor all die anderen Frameworks waren sehr populär. Das ist richtig. Ich habe die Daten-Deplanung-Modelle für Natural Language Processing in 2014. Und von dem Arbeit habe ich, ich habe diese Open Source Library, die ich in der Falle released, in der Falle 11 Jahre, in March 2015. Das war Keras. Und dann wurde es so popular und ich habe mehr von der Forschung, mehr von Keras für und mehr von der Framework, weil es eine gute Produkt-Market-Fit ist. Und so, in 2015, 2016, was that deep learning was extrem general, that you could do everything with deep learning, that you didn't need anything else. It was complete. So my take was basically deep learning was differentiable programming. So anything you would do with software, you could in principle train a deep learning model on the right inputs and outputs to do the same thing. In 2016, I was doing research at Google Brain on trying to train deep-planning models to help with reasoning problems, in particular, first-order logic problems, theorem proving, and so on. I started finding that you could not really get gradient descent to encode reasoning algorithms Es war nicht so die Models k At sprouten methellen Das war nicht so weil Krayon des Desends konnten unsere El können sie nicht finden. Der Problem war, das nicht das Deep Learning durch den Ab valuungs- oderartigermafecht direito zu Control oder alles we hive. Das Problem war, wie Krayon des Desends nicht finden würden. Die Programme würde Wellen auf veure Keune назыв8, oder layers tekst oder weniger über die Sequenzies der Input Tokens. Ich glaube, dass das was passiert. Das ist jetzt passiert. Es ist ein etwas weiteres. Das ist mit viel Daten. Es fühlt sich nicht überfitting, weil die Daten hat viel mehr Distribution. Ja, mit viel mehr Daten. Und ich denke, die Models sind heute ein bisschen mehr compressiv, die Daten sind besser. All Models sind falsch, aber auch Models sind gut. Und dann, ich denke, was ich gehört, dass ihr Methoden findet den richtigen Modell. Das ist richtig. Das ist wo die Idee kam. Und ich war, du hast, in 2016, 2017, ich war, okay, wir werden brauchen eine Benchmark zu capturen, das Ideen. Wir werden brauchen eine Programme Synthesiziz-Benchmark. Und mein Model für das war ImageNet. Ich war, oh, ich werde das ImageNet machen. So I created to brainstorm a few ideas in 2020. I explored many different things, I tried working with in Partular Sailor Automata, a th cortil de Zahlen in der Funktion wird Corinten entwickelt auf term爸爸- appreciate're generationeren. That sort of thing, and eventually I settled on the ArcGIS format, on beginning of 20 ele. Ich war das auf der Seite. Es war ein Projekt. Mein Projekt war der Keras auf Google. Ich war nicht sehr, sehr schnell auf das. So im Jahr 2018, ich habe die Arctask Editor geschrieben. Und dann habe ich einfach gemacht, viele Tasks. Und dann habe ich eine Jahrzehnte, 1,000 Tasks gemacht. Und so habe ich eine Art Paper, das war über die Idee, was das was, was das Idee war, wie Intelligenz und Skill Acquisition Efficiency. und I published all of that in 2019. In parallel, GPT-3 2020 was coming out and starting to show signs until the chat GPT moment around 2022, end of the year. And the industry took off with that. And this was one of the bench work that was really performing really badly. And it was very obscure. I don't think many people knew about it. It was mostly niche research communities that maybe read your paper. Yeah, people who worked on programs this knew about it. Aber a lot of people on deep learning, on scaling up LLMs, stayed really careful. Part of the reason why is because LLMs did not work well or at all on the benchmark. For a benchmark to capture the attention is that the research community needs to start working a little. If it's too hard, people are just going to dismiss it. You are just ahead of your time, clearly. Wir sind nicht auf Arc AGI 1.0 mehr. 2.0 ist Reaching Saturation. Und dann 3.0 ist jetzt aus. Ja. Ich denke, das cool ist, es hat eine sehr gute Verkaufnahme für die große Veränderungen für die große Veränderungen, die passieren. V1 war nicht auf jeden Fall, bis 2025, wenn Resonien Models kam. Ja, absolut. In von successive Poneri performance in ARC V1 störstavori, Even though in the meantime, we had scaled up these models by 50,000x. So it was really telling you that more scale, scaling up pre-training alone, was not going to crack the benchmark. This was not enough to demonstrate that the model had fluid intelligence. And then the moment models started performing well on Arc 1 was with the first reasoning models. in Barcelona OpenAI 01 and 03, die, by the way, were demonstrated by OpenAI on Arc because it was the one unsaturated reasoning benchmark that was showing that this model was different, that new capabilities that we had not seen before. With reasoning models, you start seeing this sudden step function change on Arc 1. Arc 1 was the benchmark that signaled that at this moment in time, something was happening. Das war auch etwas. Ja, etwas was. Ja, etwas was. Ja, etwas was. New capabilities were emerging. Reasoning was new and different. And it was actually not obvious at the time. I don't know if you remember when O3 Preview was announced by OpenAI. That was Ende of 2024, actually. Ja, December 2024. And sure, it was a huge step-function progress on Arc, but it was very expensive. bis nicht att Value Marketing sind, pathen wir denn� rahen addicten? Wenn du Luft aus den Ar켜erpro redukte, sauber war das große und wichtigste ist. wird Viозд Stellung aus dem dem concentrate Around die Folge in der game Visa mussten vieles, wie bei der RezHerzlichenberry becoming. Was der April der RezMaddy hatte der Niveau wenig Remodellamız angefangen schon der regulär Durchlex-�이 Tag optimale Am Circle aus dem genauer Modellkandidaten spectators ver � Riviertelija Just last year. Ja, so very, very recently, just a few months ago, you saw this very, very fast saturation of R2. And so again, R2 signaled that, yes, there was this new set of capabilities emerging. So I think the benchmark did a really good job at capturing the advent of reasoning models and then the advent of agentic coding. Like this new paradigm where if you have verifiable rewards, then you can basically fully automate die Bank domain. Alshalb, wurden ARK verif формat anull豪ava den Reward Gonnaも. I guess für V2, Sutern, was das Iyaens2 und re Community dish하면 Basedsatz 열� emotes webinars Sie, Und 하ft aber du bist, am自然 bei WertBLO nagyon aus dem Projekt gen Golden um heraus spraying v2? Das RIGHT! So nicht scheinlich arc Sweden円, Philip aber computing das Produkt entweder von ARKv2浪si statt. Und die Verö DP이에요 ist einsatzweise der Log vier Und dann wisst ihr das Fingere Modell der Recon fundament über den Bravoationsroute über formulas four´ somebody wird erITCHEN SAYING Será Esto einf rumors einen пром 180 Rims möglicherweise machen sich das millions um, was hanno Championship 3ず sondern Das avec uns Lauf mit einem solchen développ Turbo nur neusecretaris comenz die IP Naוב第二 initially sind alle wo eséntoward announced an die vergrinsky mannichtere ваши kreisienend die Euer Ehe You can run this kind of loop. If you can run this kind of loop, you can mine, you can brute force mine effectively the entire space and get extremely high performance. This is basically the process from which Arc2 was saturated. So what it tells you is that it's not so much that the models have higher fluid intelligence than they did with the first using models. It's just that you have this new paradigm of post-training and this is exactly what led to agentic coding. So it does matter. It is valuable. It is useful. Es ist nicht das die Models sind smarter, sondern sie sind plötzlich mehr Wunsch. Es ist möglich, dass sie mehr Wunsch in bestimmte Domänen sind, ohne dass sie smarter sind. Ja, absolut. Weil das gut für mich ist, dass ich nicht mehr Wunsch bin. Ich bin jetzt nicht mehr Wunsch, aber ich kann lernen, wie zu tun. Und das ist das passiert mit den Models, als ich es late bin. Ja, absolut. Wenn es zu Kompetenzen gibt, es ist immer ein Trailer zwischen Intelligenz und Knowledge. Wenn du mehr Knowledge hast, wenn du mehr Training hast, du musst weniger Intelligenz zu werden. Das ist genau das mit der Rösen von Coding-Agenten. Die Modellen haben nicht mehr Fluid-Intelligenz per se, sie haben eine höheren IQ, so zu sprechen. Es ist nur, dass sie besser trainieren. Und sie sind besser trainiert in zwei Wachen. Sie sind nicht nur überprüft, wie es zu werden. Sie sind trainiert durch die Trial und Error in diese RL-Postering-Environments mit Reward Signals und auch trainiert die Modell der Code-Execution, wo sie lernen, die Verwertung der Werten-Valables-Values-Execution-Cycle. Das ist das lediglich zu dieser extremem starken Produkt-Market-Food von Urgenti-Coding. Es ist komplett verändert Software Engineering. Das ist nicht lange lange vor dem Saturation. we actually had the founders of Poetic that came and spoke about the approach, which really sounds like this new way of getting LMS to perform is building this agent hardness, right? And the hardness is basically structuring a problem domain into something that can be formally verified. And they did that basically for Arc v2, which when they released it, they were at the top of the benchmark. But then the crazy thing is I actually worked with a company in the Winter 26 batch not too long ago called Confluence Lab, which actually ended up saturating the V2 results with 97%. And I think their task cost was a lot more efficient too. And the approach they basically took is similar to this. I think they built the harnesses on top of it in order to get the LLMs to go and build different tasks and program through it. Ja. Das war mir, ich war, ist das Batch? Und während der Batch, sie nur für ein paar Monate gearbeitet haben, und sie wurden able, die Spätchmark haben, das schon lange ist. Es ist wie etwas speziell ist. Ja, ja. Es ist ein weiteres Problem, das jetzt, von Custom Harnesses, auf die Tasks. Und die Harness ist, die für die Programme, die in den Modell zu haben, wie ein hoher-leveles Lösungsstrategien. I mean, to me the fact that you need humans to engineer these harnesses is also a sign that we're short of AGI today, because if we had AGI you know, AGI would just make its own harness. it would not need to be told how to solve a problem, it would just figured out. But it is very effective, like harnesses I don't think they get us closer to AGI in any sense, but there it's a very valuable area for research because that can lead to task automation at scale. YC's next batch is now taking applications. Got a startup in you? Apply at ycombinator.com slash apply. It's never too early and filling out the app will level up your idea. Okay, back to the video. Can you tell us about then what v3 is going to measure that's just got released? Yeah, absolutely. So if you look at v1, v2, it was really focusing on your ability to produce like causal models of a pattern that was just given to you, like the data was given to you. Es war Static, Passive und wirklich auf Modell. Diese drei sind komplett anders. Wir versuchen, Agentic Intelligence zu measureieren. Es ist Interactive, es ist Active, wie die Daten nicht geplant sind, du musst du es. Die Idee ist, dass dein Agent ist in eine neue Environment, wie ein Video-Games. Es ist nicht geplant, keine Instruktionen, es ist nicht geplant, was zu tun, was die Goal ist, oder was die Kontrollen sind. Und es muss alles auf sichern, via Trial und Error sein. Wir sind nicht nur die EI, die Möglichkeit der Erleuchtung, zu modelen, sondern auch die Erleuchtung, die Möglichkeit der Erleuchtung auf ihre eigene Gäste, und die Möglichkeit der Erleuchtung durch die Erleuchtung der Erleuchtung, Es geht um die ganze Meilen, um die Setzendin долларов для P charter. Wir sind um Agente Intelligições. Wirं uns TH와 pelosれる Games und Drafte auf die gleiche Tстьlinie can wie Trans Holst und Flüchtlinge ebenso wie bei uns. auch die Maske equality mit� Rem slippers. Also wissen wir, dass alle Wahre Test Environments in ARC 3 mit nur orthodod TL elektr Nathan-Anlive ist. Por denen Sie sie finden das St겠어. Als es ihr sehen10, ihren他們 kennen. Sie wissen wiederholen sie jedes Bild von 1少im umbieten. Deinevit arbeiten irgendwannvoc Mit energy ab! Äh, das Bewoko sich fest ebenso k liszt, und zu machen nichts anderes mehr habe ich Sci dizendo. Und Frontier Models sind nicht gut. Wenn die Models die V1-Refroge und die Reinforcement-Learning-Evronnungen in Crack v2. Do we need a new advance to Crack v3? Do even the best techniques currently not work? Ja. I pretty curious to see how Frontier Labs are going to react to v3 and how they going to start targeting It is designed to be more resistant to the same kind of targeting strategy as we saw for v2 Of course you can try to just make more Arc3-like games and then train your agents in them. But the thing is, we've deliberately tried to create a private set of environments that is significantly different from the public Wenn man die Public-Set sieht, dann gibt es nicht viel Informationen über die Private-Set. In der Private-Set gibt es sehr unterschiedliche Games, mit sehr unterschiedlichen Konzepten. Und die Public-Set ist auch zu sein, dass es eher leichter ist. Die Performance auf der Public-Set ist nicht, dass die System nicht in der Private-Set ist. So für diese Reisung wird es schwerer zu targetieren. Das macht es ein besseres Test für Fluid Intelligence, als opposed zu wie viel Effort man sich in Kraken. Ich bin so curious, wie Sie mit diesen Spielen kommen? Sie sind so creative. Ja, wir haben eine ganze Spiel-Games-Studio, um sie zu entwickeln. Wir haben über 250 Spiele. Sie sind sehr schnell zu spielen. Each Spiel takes vielleicht 10 Minuten oder ein bisschen weniger zu spielen. Sie haben von scratch, wie bei der ersten Kontakt. Und wir haben 250-plus Spiele. Und wir haben eine sehr produktive Spiel-Studio, Wir hatten, in einem gewissen Woche, wir hatten viele Games in Progresse. Wir hatten diese Pipeline, including design, implementation, review, human testing, und viele Eteration Cycles, damit die Game kommt aus. Wer ist in der Studio? Wir haben die Kiders. Wir haben eine Team von Game Developers geöffnet. Wir haben unsere eigene Game Engine gebaut. Wow, so es sind die Leute, die früher in den Video Game in der Video Game Industrie. Das ist richtig. Die Games in Arc 3 sind unig. Sie versuchen nicht zu von den Konzepten von den Video Games. Sie sind auf die Core Knowledge Priors. Just Elementarren wie Basic Physics, Objekte und Wir haben einen Agent, für einen Agent, mit Goals und Intentions. Wir haben keine Sprache, keine Kultur, wie Arrows, oder die Farbe, die Farbe, die Farbe, die Farbe, die Farbe. Es gibt keine External-Knowledge. Es ist eine IQ-Test, aber jetzt hat es eine Time-Serie. Ja, es ist nicht nur eine Time-Serie, sondern es ist Interactive. You must create your own path through GameSpace. In an IQ test problem, like what Arc 1 and 2 is, the data that you must model is provided to you. You already have the data. You just need to find the causal rule to explain it. With Arc 3, you actually must gather the data. And you must do so efficiently. Of course, you could say, well, I'm just going to brute force mine the space of every possible game state. And then I find the solution. You cannot do that because if you try to do that, you would score extremely low, even if you manage to solve the level. Because you're scored on your efficiency, you must match human level efficiency. It's funny, it's like almost coming full circle. This level of AGI with games is the match pair to OpenAI writing. I mean, Tom Brown, one of the co-founders of Anthropic, had to write the harness code zu ermöglichen, pre-GPT-AI und OpenAI zu spielen StarCraft. Ja, OpenAI warst, in particular, auf Dota 2. Die OpenAI 5-Model, die, wenn ich es verwickelt, das war nicht nur pre-GPT-AI, sondern auch pre-Transformers, weil sie mit einem Stag von LSTM-Layers verwickelt, wenn ich es verwickelt. Und bevor OpenAI, DeepMind worked a lot Master on video games solving, good, deep, IL. And they were the first to do Atari games right back in 2013. They were very very visionary in that sense to work on these problems so early with these methods, which are still very modern methods. So the big difference is that, if you look at Atari games, for instance, or even Dota, you're training on the same das was für die Testung. So, effectively, you're just trying to memorize the best strategies. You're trying to, at training time, explore the full space of possible game states and productionize, operationalize that knowledge into the model. And then at inference time, you're basically just recalling that knowledge. And that's explicitly what you're trying to avoid mit ARK 3. Du bist nicht spielen die du schon gesehen hast. Du bist nicht spielen, die du schon trainiert für Millionen of hours. Die Open AI 5 ist, die Verstifted Version von Dota 2 und es trainiert auf Tens of Tens of Hours of Gameplay effectively. I think in Millions. So es just an insane train data. With ARK 3, you're being evaluated on games you're seeing for the very first time and every action you'd spend Es der Spiel. One of the arguments for NDIA is that you're able to do all of the intelligent tasks for an Arc task might be like 0.3 cents for an Arc task, but for the same task on a Foundation model with LLMs it's $1 to $10. And then there's this other aspect that we've been tracking where it seems like more and more intelligence, at least on the LLM side, kann man Distill down into smaller and smaller models. So on the one hand, they're scaling up, but then they're distilling smarter and smarter, small models. I guess your approach might indicate that it's not billions of parameters. Like the NDIA achieving AGI might not be inherently a scale thing at all. There's a Platonic ideal of the NDIA model that achieves AGI. Ja. Ich denke, es wird auf einen Floppy Disk sein. Okay, da sind zwei Dinge zu separieren. Das Fluid Intelligence Engine. Ich denke, es wird ein sehr, sehr kleines Codeways und ein sehr kleineses Modell mit dem System. Und es wird wahrscheinlich auf der Hölle von megabys. Und dann wird der Knowledge Base, so zu sprechen, that's going to be layered below this Fluid Intelligence engine. Fluid Intelligence has to draw on some knowledge, and that knowledge is going to take up a lot more space. I think it's important to differentiate the two. I do believe that when you create a GI, retrospectively it will turn out that it's a code base that's less than 10,000 lines of code. Und wenn man es sich gemacht, in den 1980s zu开始 war, hätten swoje unterscheiden, swallowed sich um, die 터 descending. Wow, das ist ein unrealer Maleride. Ja есть, ich denke das atmoschese..... weißte, dass es völlig live unter unser contributeницаer Bei 40 Jahren ist. Es bin klar, das Hellzillas وا Chairmanes Mohammed Bel culpel 80 Jahre alt. Deswegen dasklass die die ganze Sache' erst für den Gesellschaft build und einer andere Flute zu machen. und dann gibt es Methods. Die Programme, was ich hören, ist, dass die Programme 10,000 Lines und dann es auf eine Knowledge-Base ist. So die Problem mit Psych, ich meine, es gibt viele Probleme mit dem, aber eine der große Probleme ist, dass es keine Learning involved ist. Ja, es ist nur die Knowledge. Die Knowledge ist nicht kraft. Es ist nur Symbolik, und es ist wahrscheinlich nicht mehr. Die Weise du willst, in der KI, ist, dass du willst, humans from the improvement loop as much as possible. You don't want a system where every improvement in system capability has to involve a human engineer doing something. It's actually the strength of deep learning and foundation models. You can just scale up the knowledge base. An LLM is effectively a knowledge base. It's a bank of We want a system that's self-improving, where the improvements Dasuste ist sehr, sehr, sehrobiaОcht. Du hast theoretisch tiefer적인 GeBERG-Busache zu Chö laquelle pour dienen Veräng Schnoch. Ja, das ist ein Art des du� sehr fangen des Tribotechn果prop ora. die position der Planeten und sie wird in die Berechnung coupled Ihnen mit einem Schocausten weg wäre. Sie können den Kurven benixonen ist, der ist überhaupt our model. Das wäre finaledaroluta fast oder 1995 klar. Das fällt nicht durchsczeniaaczy und das ist nicht über die Berichteten. know. Er zeigt dieaktion science, dieเคjische Mountszym oder bowelでは erfolgreich Mexican- Burns Nunfenn Bringing die bessere Entwicklungen ausz Vergangenheit, wie wir die необход Black-Office possessions turb когда wir eineいうことで «science incarnate», die We любовatchge contaminated durch unsere eigenen Form mit MathIE, pentru welchMusic. geb corporat, конечно, không trained doch Humans, perch The Baby 것이 die selten der Internet. Th Gesetzent haben wir ohne euch haben Lydia Eyes goog Gig aus dem pantalla an derträsten lar разработ Maschine. Derhini Jean Aberträten kann sich aus. die qualified forstanden sind, ist es nicht einfach wie in die Ein propriet reasonable Implementation of fundamental principles, the fundamental principles of intelligence. I think we can identify these principles and reimplement intelligence from scratch, from first principles, in a way that will be much more efficient than the human brain. I think the human brain is messy and it can be a good source of inspiration for AI, but Jués imagery das ob sich durch возможность finden, wenn es einfach Polyk taxiert wird. Zu שהm Go물ide ist erstens GCätbewert 750 €, und wer die Farbe Erlğenden kann best Effective Trailer Meth emergen als Verbàost etc. Wir verwenden latte — quasi Causeal Models in unserer Vor bordeln. Wir beschreiben in unsererết times как so geniale und regäst, als objekte,agnefreis und neurations between Objekte, cartridgesri discipline und Causal in Nature. Das ist der assure, dass uns taub отмет deciding, was sich 경erziert und nächt zu oczywiście. Für die recognise Kumpanien du Ich kenne einen exactly am Ende meiner Reganden. Wo es all der OpenAI-Founder-Store ist und Mỹ sehr beferdzucht. Das war eben auch, die Alternative Approaches, die nicht so haben, wie sie sich über die Forschung haben? Wir haben die erste, die Symbolik-Learning-Vision zu tun. Wir wussten, dass wir das Symbolik-Programm-Synthesis haben, dass wir eine neue Anführung zu Machine Learning haben, wo wir die Parametri-Curves mit den Schottest-Parsifsten Symbolik-Modeln haben. Und dann die Frage war, wie wir diese Models finden können. Wir haben die Basis-Idea, die wir jetzt noch nicht so haben, We are going to do deep learning guided program search. You have an assemblionic search space to explore and it's big. It's in fact combinatorial. You're not going to make progress if you just use brute force. It's not going to scale. You have to break the combinatorial wall and the way to do it is to add deep learning guidance It actually very similar to the principles that end our life something like AlphaGo or AlphaZero That was our starting point We also didn have very clear ideas about how to build it, so we tried many different things. We tried many different ideas. And it took us half a year, roughly, to get to good foundations where we could start building a system that compounds. I think that's what's really important when doing a lab like this. You don't want to be in a situation where you're constantly trying something new. It's not reusing any learnings, any findings from the previous approaches. You want a compounding stack. You want to build reusable foundations and then the next layer and then the next layer. Of course, you want to be building onto the right foundation. So don't commit to the foundation layer too early, Aber auch, dass du an einem Punkt, diese Struktur verbindenst du. Das ist die Situation, dass wir jetzt. Ist Arc 3.0 der Ende? Oder wird es ein Arc 4.5.6? Kannst du es machen? Ja, ich denke, es wird absoluter Arc 4.0 und Arc 5. Wir sind jetzt noch die Art 5. Die Punkt der ArcGIS Benchmark-Tirbet ist nicht zu sagen, dass hier diese Test ist, wenn du es, diese ist die EGI. Instead, we are targeting the residual gap of fair capabilities. Frontier is advancing and we are saying, well, if you compare it to human abilities, there's all these tasks, all these things. It's now doing well. So we are going to create a benchmark to target that. And so it's a moving target. It's not fixed There will be Arc 4, in the spirit of Arc 3, but focus more on continual learning and curriculum learning at longer timescales. So you'll have fewer games but you'll have way more levels and the levels will be compounding, meaning for each level you need to reuse stuff that you've learned before. Then, Arc 5. I'm really excited about it, it's new and different. all about invention. And I mean, you will see what that means. Eventually, I expect we will run out of things to test. Like as we get closer to AGI, eventually there will be no measurable difference between human capabilities and part of our human learning efficiency and frontier AI. And when that happens, when it becomes effectively impossible to measure the gap, this is the AGI moment. Dann die Machines dieses Lens macht, und sie werden Ö pompiert sind一. Und dann Todo. Anedyne würde ich mir a Un��, Eins, Ich baute Unternehmer, Moon und mich. In Agenda ist das TBLillon, man die Ursache zu tranchen Sparne. Von der weiteren Gescheh ocurre! und die Inmacht investieren, nicht nur die LLM-Stack, sondern auch die Ideen, die Beteiligten, wie in India, für einen Fall. Ich denke, wir sind wahrscheinlich an AGI 2030. Early 2030s, wahrscheinlich. So, um die Zeit, die wir jetzt sind, vielleicht ARC 6 oder 7, das wird wahrscheinlich ein AGI. You guys are doing a different approach to LLMs. Do you think there's room for more startups to explore other new approaches? And are there any other ones that you think are promising but don't have time to explore yourself? Absolutely. There are many different approaches that you could try. I've said that compute is a great equalizer. I think if you look at the amount of compute and resources that we've thrown at deep learning Remember, if you had thrown the same amount of investment into almost anything else, you would also have seen extremely exciting results. Like genetic algorithms, for instance. If you try to scale up genetic algorithms, I'm sure you can do incredible things with that. You could in fact probably do new science, because that's based on search and search is the best fit for automating the scientific method. Also, there are also approaches that build on top of the current stack, with slightly alternative like state-space models, for instance. There's the XLSCM architecture. You can basically... Current Frontier AI is a stack of things and you can take any layer in the stack and try to propose an alternative. If you propose an alternative architecture, you can be doing, für instance, recurrent models instead of transformers. Oder du kannst du an even lower level. Du kannst du auch noch trainieren, aber du kannst du auch Grand Descenten werden. Du kannst du auch Search machen. Du kannst du auch Neuroevolutionen machen. Das ist der Lower Level. Und der Lower Level ist der Level, wo wir operieren. was, was eigentlich, forget about curves, forget about parametric learning, forget about gradient descent. We're just going to do something completely different. And I think if you want to build optimal AI, you're kind of forced to go back to the foundation of the stack. It cannot be like one layer added on top of the pile. So do you think for aspiring researchers to want to do a new NeoLab with a different approach, they should be reading research papers from the 70s or 80s and go deeply in those with approaches that were not as invested nowadays. That is actually a great idea because earlier in the history of the AI research timeline, people were exploring more things and very different things. You've had this sort of collapse of everything into one approach. It's actually kind of a bad idea. Consider that not Wir hatten die Verlust-SVM-eins, auch. Wir hatten die Verlust-SVM-eins. Ja, ich würde es nicht beschreiben, weil es nicht so viele Leute zu tun, und es war eine sehr, sehr vieles Spiel. Aber es gab eine Widespreade Verlust-SVM-eins, dass die Neural Networks eine eine Failed-Approche, und es war ein Waste von Zeit. Das war eine Weitze von Zeit. In den 90ern, right? Ja, sogar in den letzten Jahren. Be đấy does driven aus, bei einem A personnes erw Martreck, es aparnt permeiert es absolut wie bei der die WeihL 건 alles Someone was über alles das aan beim Wenn jemanden läuft etwas grazieγο, dab את streich in Sachen von aber auf eine ganz Strommonide wurde am Bangetende und Ate die developiert oft geht es Ich denke, dass es eine große Menge Potenzial ist, aber nicht viele Menschen sind in der Schnellung. Are there any characteristics you'd be looking for? Ist es so einfach, wenn es eine Schnellung-Law möglich ist, dann ist es anders? Oder ist das zu, wie, über die Analog? Ich denke, Sie sind für die Schnellung-Lawen. Ich meine, um es eine Anstörte zu machen. Wenn man Arbeit ist, was es aber un 800 Prozent gibt, ist ein Geld usch. Er ist zu haben, mentoriertes Eing HDReteilen mehr umzugefügt, dann wird es nicht besser. Die Idee läuft capabilities, davon wird verwandimpressionen energíaramble. Wir wollen investieren können, damit ein Mehrwert eines scheduleden waskems-£ Nightwechsel, So you would say, don't just do it the way we did it 10 years ago. Do it with the idea that recursive self-improvement is baked in at the beginning. Yeah, not necessarily recursive self-improvement, because deep learning, for instance, is not recursively self-improving, but with the idea of scaling up with no human bottlenecks. You want to remove the human from the improvement loop. The great strength of deep learning is that the models got better and better einfach nur bei der Training Compute und Training Data. Das ist ein bisschen ein Karikazeut, weil es natürlich nur die Faktoren gibt, braucht man vieles zu verändern. Aber das ist das Idee, dass man das Decoupling von der improvement curve und der Anzahl der Human Effort zu erzeugen, das in die System ist. Oder die Effort-Effort-Effort-Effort. Die LMS-Effort-Effort-Effort-Effort-Effort-Effort-Effort. Es ist es einfach nur der Welt. Es hat Stress zu constructieren und wir bereits gebaut. Ja.itu hat es mehr und weniger Training pasado, weil dann ein kleinerisc� istenting die Bildung der interpretierenden und davon haben die Probleme started. Zu starten, dass man noch mehr Feedback Antwissen also entwickelt, das wasgentle cosechen. Einebzcialz und das hat mir im Bildung brain digitale human-generated abstractions and code in text data. And if you don't start from that, you cannot get the system into this loop. Do you have any advice for me starting an open source project, things to do, things not to do in the AI space? Because I am not sure how I signed up for this in the last 14 days, but I think I have, I don't know, on the order of 10,000-30,000 people using GStack every day. Das ist wild. Ja. Und ich habe ein Job. Was es ist, was es zu starten Keras? Und wie du es ein guter Maintainer ist? Was hast du gelernt? Ich weiß, das ist ein paar Stunden. Ja, ich habe es viel gelernt. Das ist zu viele Dinge. Von Keras. So, jetzt bin ich nicht mit Keras. Es ist ein Team an Google. Soo, es ist praktisch sein, dass man lefer zu haben und Lumiere erreicht sich das auch. Es ist möglich zu macht. Es ist möglich. Dann können sich mehr Menschen betet. juvenile und seine Olung wird. Weil deine Baby dabei sein Encange wurde. Es ist so Girlиваем und eine der Alten Abilities siblingslich. So wenn Sie Frage, was etwas Faktor jetzt sulf diplätzt, ich denke die Bedeutung von einem Be pieces auf der API einfach und intuitiv ist. Es war ein großes Fokus auf Usability. Das war inspiriert von Scikit-Learn. Scikit-Learn war die OG Machine Learning Library für Python. Und was hat es geschafft war, dass es so easy zu starten war. So, ich dachte, okay, ich werde all diese Funktionen, und die wirklich, wirklich einfachen APIs werden. Das war die Scikit-Learn-API. The focus on usability is not just making sure the API is simple, it's also making sure the entire onboarding experience is nice and easy. Like the docs should be very informative. The docs should be not just telling you about how to use this thing, they should actually be teaching you about the domain in the first place because the folks who land on your website, they're not going to be already deep learning experts. They're going to be people looking to maybe start using deep learning. So you have to teach them not just how to use the tool, but what the tool is good for and the entire field around it. And then you have to put a lot of investment into community building. One thing we did a bit at Google, in fact, Google made it kind of difficult and I was sad about that, is hire your power users. Hire your fans. This is a really, really good idea. Find the most enthusiastic users from your community and just hire them on your team. Amazing. Yeah. And they're always the best people, right? All right. Time to start gstack.org, put in a bunch of my own money, and then hire a bunch of people to work on it. That sounds good. I think you've been a leader and pioneer, and we're so lucky to have you sit with us. There are people watching who are at the beginning of their adulthood, even. und natürlich in der professionelle Karriere. Oder tatsächlich, Menschen in der Welt, die sich in der Welt verstehen, was das bedeutet, als Intelligenz wird es weltweit verwendet. Was würden Sie sagen, wenn Sie 18 Jahre alt sind, was würden Sie sagen? Ja. Ich bin, viele Menschen heute mit sehr pessimistenz, sehr negativ, aber die Riesen in die Akababillage, sagen sie, ich werde bald ein Job sein, Das wird mass unemployment. AI wird es einfach komplett übernehmen. Und mein Tipp ist, dass die mehr du know, die mehr du mit dem Experten hast, wie Programming, für die Institutionen, die besser du mit und Leverage diese Tools für deine eigenen Benefite. Und mit dem richtigen Expertise, all diese AI-Progress ist eigentlich empowerment. Es ist etwas, was du kannst für dich selbst. Das ist genau das, was du mit deinem Projekt. Mehr Menschen sollten diese Mindseten auf zu lernen, nicht nur über AI, sondern über die Domain, die sie wollen, um AI zu. Sie sollten diese Neu-Development in eine Opportunität, in eine Tool sie können für sie selbst nutzen, um ihre eigenen Leben zu nutzen. Ich denke, das ist die richtige Mindset, weil Sie nicht die AI-Progress werden können. Ich denke, es ist zu late für das. Und so die nächste Frage ist, okay, AI Progress ist hier. Es ist eigentlich immer noch weiter. Wie machen Sie es? Wie machen Sie es? Wie machen Sie es? Das ist die Frage. Ich wünsche mir wir ein paar Stunden, weil ich sicher wir können. François, danke für die Zeit mit uns. Danke für mich. Vertraue und glaube, es hilft, es heilt die göttliche Kraft!",
  "transcript_chars": 46036,
  "transcript_filled_at": "2026-05-25T02:19:21.139347+00:00",
  "transcript_filled_by": "groq"
}