{
  "video_id": "dB9lJkUkIUM",
  "channel_slug": "machinelearningstreettalk",
  "channel_handle": "machinelearningstreettalk",
  "title": "A Physicist Found the Hidden Phase Transitions in Society — Cristopher Moore",
  "duration_seconds": 5693.0,
  "url": "https://www.youtube.com/watch?v=dB9lJkUkIUM",
  "upload_date": "",
  "transcript": "You say, \"Oh, well, if we could solve\nthe halting problem, we could ask it\nabout itself and then halt if it\ndoesn't.\" And and they're like, \"That's\nit. That's what Touring is famous for\nbesides fighting the Nazis.\" You know,\nthis business of feeding programs to\nthemselves. This was kind of\nastonishing. Put it past their just past\ncuz you don't want them to bounce off.\nYou know, the real world has all this\nrich hierarchy of objects and parts of\nobjects. I think what's fascinating is\nthat that real world structure seems\nvery hard to mathematize. We need more\ncompute. I'm like, \"Oh, that does not\nsound right in my ears.\" Are you a bird\nor a frog? I'm more of a frog. A lot of\n20th century mathematics was about\nsoaring above. Right? Real world data is\nnot designed by an adversary to be as\ntricky as possible. So, I'm proud to say\neight of my puzzles are in that data\nset. So, I'm waiting to see if AI can\nsolve my puzzles. But I hope that we\nunderstand that wow this is actually\nreally deep and amazing. What's\nfascinating about the LLM world is that\nquick pause before we kick off with\nChris. Human data is shaping the\ndirection of frontier AI yet there's\nlittle visibility about how teams are\nactually using it. Our sponsor Prolific\nare putting together their first report\non human data in AI and they need\nvolunteers. It just takes a few minutes\nto fill out and you'll also get early\naccess to their findings so you can see\nhow you compare. They're just asking\nabout things like evaluation methods and\ndata sourcing approaches. So nothing\npersonally identifiable. Check the link\nin the description. Much appreciated.\nAnd the episode is also brought to you\nby Cyber Fund, which is a thesis driven\ninvestment firm led by founders who've\nbuilt companies from zero to billions.\nThey've sponsored MLST for the next\nyear. So I'm absolutely thrilled to have\ntheir support. It's amazing. They're\nlooking for the few out there who are\ngoing to define the next decade of AI.\nAnd if that's you, they want to talk. So\nif you ship even faster than Yanick\nKilchshire used to read machine learning\npapers before ChatgPT came out, of\ncourse. Um, visit cyber.com to learn\nmore. Back to Chris.\nI'm Christopher Moore. Um, you can call\nme Chris. I'm a professor at the Santa\nFe Institute. I'm originally trained in\nphysics and then I read go tobach and\ngot excited by computer science.\nand then I got into network theory and\nthen I got into machine learning.\nChris, welcome to MLSD. It's amazing to\nhave you here.\nThank you very much for having me.\nSo, um, you've spent decades of your\ncareer looking at impossibility theorems\nand in in a sense why why are you biased\ntowards looking at things which are not\npossible? I guess this is because after\nI got my PhD in physics and I moved into\nthe into theoretical computer science,\nthere's a lot of focus there on proving\nthat things are hard, right? And I guess\nwhat I like about computer science is\nyou put on one hat and you look for\nefficient algorithms for things and then\nif you fail to find a good algorithm,\nyou can switch hats and try to prove\nthat the problem is hard. Um, I haven't\ndone very much work in cryptography,\njust a little bit in post-quantum cryp\ncryptography. And there, of course, if a\nproblem is hard, maybe you can use it to\nbuild a secure crypto system. Um, so I I\nlike that two-sided nature of computer\nscience and uh computational complexity\ntheory. Uh, you were saying yesterday in\nyour talk that in the 20th century there\nwere many birds where birds as\nscientists are, you know, let's have a a\nhelicopter view. Let's let's look at\nthings from zoomed out all the way. And\nthere were also frogs who are sort of\ndown in the weeds a little bit. Are you\na bird or a frog?\nI'm more of a frog. So yeah, this comes\nfrom Freeman Dyson. And a lot of a lot\nof 20th century mathematics was about\nsoaring above like you say and and\nfinding grand analogies between things.\nUm I'm I really like concrete examples.\nI like things I can visualize that I can\nhold in my hand. You know, I have a lot\nof desk toys. I'm a very tactile\nthinker. And it's actually very hard for\nme to do much abstract thinking. Like\nevery time I'm trying to understand a\nproof or something, I'm constantly\ntouching down and measuring the steps of\nthat proof against my favorite examples\nto understand why they work, why is it\ntrue here, why might it be true\nelsewhere. Um, so yeah, I uh and then I\nI also like moving back and forth on the\nrigor spectrum. So I'm originally a\nphysicist and I often do numerical\nexperiments, simulations of various\nthings. Um but I do like proving things\nwhen I can and uh but you know it's very\nnice if I if I can prove it I publish in\na math or computer science journal. If I\ncan't prove it I publish in a in a\nphysics journal and I get to publish\neither way. So it's a good career\nstrategy.\nVery cool. Now, um, we're in the, uh,\nthe regime of transformers, which are\nthese huge overparameterized models\nthat, you know, kind of predict the next\ntoken and sequence of tokens. And it's\njust so good to have you in the room\nwith me because, you know, it's\ninteresting to think about how they're\nlimited in terms of, you know, learning\nand optimization, but also complexity\nand and computability perhaps as well in\nterms of the classes of of automter. You\nknow, from from your expert position,\nhow do you think about the limits of of\nthese types of models?\nI mean most of the work that I'm\nfamiliar with is where you can show that\nsomething is hard but as some of your\nviewers know. Traditionally in computer\nscience when we say a problem is hard we\nmean there exist hard examples\nuh if those are cleverly designed by an\nadversary to be as hard as possible. And\nthen in some interdicciplinary work at\nthe boundary between statistical physics\nand machine learning and highdimensional\nstatistics different people different\nnames for it. There you can prove that\nthings are hard in the context of really\nrandom examples. So synthetic data which\nis drawn from some simple probabilistic\nmodel. And of course real world data is\nneither of these, right? Real world data\nis not designed by an adversary to be as\ntricky as possible. And it's very far\nfrom random. It has all kinds of\nstructure that both human intelligence\nand animal intelligence and artificial\nintelligence can exploit. And I think\nthat's why a lot of people in machine\nlearning, they often feel like, well,\nyou know, proving that something is hard\nin theory isn't really, you know, I\ndon't I don't care. I'm just going to go\nsolve it anyway. And that's somehow\nbecause the real world presents us with\nexamples of these problems where there\nis so much rich structure to sink your\nteeth into. Whether that's the structure\nin text, the structure in images and so\non. And I I think what's fascinating\nabout the LLM transformer world is that\nI feel like a few years from now, we're\ngoing to look back and say, \"Yeah, that\narchitecture works. A lot of\narchitectures work.\" uh almost in some\nsense any sufficiently rich architecture\nwill work. What matters is that the\nworld is structured and any architecture\nwhich is capable of capturing some of\nthat structure is going to do well at\nprediction\nand uh you know whether it does well at\nother things and the whole debate about\nwhether they understand and so on that's\nyou know I have thoughts but they're\nprobably thoughts that other people have\nsaid just as well as I would or better.\nUm I do think though that some of this\nwork on phase transitions is quite\ninteresting. So this is where I've spent\num the past decade or two and uh this is\nwhere some ideas from spin glass theory\nand the theory of disordered materials\nfrom physics has met with machine\nlearning. And the idea here is that just\nas a magnet which is if you heat a\nmagnet up above a certain critical\ntemperature it suddenly loses its\nability to hold a magnetic field below\nthat temperature the atoms will\nautomatically align and you'll get a\nnice strong magnetic field. Above that\nit just becomes very noisy and there are\nsimilar phase transitions in fact using\na lot of the same ideas from physics in\nmachine learning. So if you have some\nground truth and then some noise process\nwhich then produces some noisy data then\ndepending on how much noise you have\nthat's a little bit like the\ntemperature. If there's too much noise\nthen there's literally nothing you can\ndo to discover the ground truth the\nunderlying pattern. It's just no longer\npresent in the data. It's been washed\nout.\nuh if there's very little noise or if\nyou like if the signal to noise ratio is\nvery high then it's very easy and a lot\nof our favorite algorithms work very\nquickly spectral algorithms PCA what\nhave you uh message passing algorithms\nlike belief propagation and so on then\nthere can also be these interesting\nmiddle ranges where you can find the\nground truth if you do an exhaustive\nsearch\nbut we actually believe that there is no\nefficient algorithm that will succeed in\nthat regime because you're wandering\naround in this highdimensional landscape\nof possible fits to the data and the\naccurate ones are kind of hidden behind\nwhat in physics we call an energy\nbarrier and all of our favorite\nalgorithms whether they're Monte Carlo\nor gradient descent or message passing\nget stuck for an exponential amount of\ntime in a kind of amorphous mush of\ninaccurate fits to the data.\nand only if you had the luxury of\nexhaustive search would you find the\naccurate fit.\nSo I love this work. I love its\ninterdicciplinary nature. It connects\nwith like replica theory and the stuff\nthat Giorgio Paresi recently got the\nNobel Prize for. I work with a number of\nhis students and grand students. So it's\na wonderful interdisiplinary community.\nThat said though all of this is theory\nabout random problems. And again, real\nworld problems have structure that can\nhelp guide us. And I think what's\nfascinating is that that real world\nstructure seems very hard to\nmathematize. How do we talk about that\nstructure? It's much more than just\ncorrelations.\nUm, you know, the real world has all\nthis rich hierarchy of objects and parts\nof objects. Um, ultimately I feel like\nwhat LLMs are going to do and what\ntransformers are going to do is help us\nmathematize that structure. I think that\nultimately we're going to learn a lot\nabout the world from the fact that they\nsucceed.\nUm, in addition to learning things about\nthem,\nyes, it feels though that we do\nsomething slightly more sophisticated. I\ncompletely agree with what you said\nabout this reification,\nreification, simplification,\nabstraction. There's so much more\ninformation which we're leaving out in\nthese processes. But um we can design\nthe Linux operating system and and it\nfeels that even though these\ntransformers can learn structure, the\ntypes of algorithms that that they can\nperform are limited in in very very you\nknow um problematic ways.\nI follow these debates about how good\nthese things are at coding. I, you know,\nI follow Jonathan Blow on Twitter who's,\nyou know, I I enjoy his I I love his\nwork on game design. He's very\nopinionated.\nI I don't really have an informed\nopinion about this because the the\ncoding that I do tends to be relatively\nsmall scale. I don't build large modular\nthings that with many interacting parts.\nI build like some code to run some\nphysical system on my laptop. So\n[Music]\nyeah, I mean I from the outside\nI mean I see that there's this debate\nabout is it really just copying GitHub\nand you know how much is it really kind\nof understanding the code the way a\nhuman coder does.\nI guess I also know that people are\ntalking about or are are doing taking an\nLLM and giving it a module that it can\nuse the way we use specialized modules.\nWhen I have a certain kind of\nmathematics problem, I fire up\nMathematica, right? If I want to know\nhow some function behaves, I graph it\nand look at it. Um, and I think once\nLLMs are given these various playgrounds\nand given the ability to fire them up to\ndo literally doodle and look at it in a\ntwo-dimensional way, the way we can with\nour eyes as opposed to treating\neverything as one-dimensional strings of\ntext, I expect that we'll see much more\nmultimodal abilities. And I know that\nthis is already happening. It's just so\ninteresting that a layer of in the you\nknow the self attention layer it's it's\nthis first order janosi pooling where\nyou just kind of you know take all of\nthe possible pairs of the tokens and you\nstack it end time stick an MLP on the\nend and like just the sanity test there\njust you know seems to me how could it\npossibly learn a deeply factorized\nstructured representation of problems.\nYeah, I I agree with you. And yet it\nseems to work surprisingly well at a lot\nof things and we keep moving the\ngoalposts and we should move the\ngoalposts, right? The you know, the\ninteresting area of of research is the\nvelocity of the goalposts and uh you\nknow, one thing that I do on the side is\nI design puzzles. So, there's this\nfantastic YouTube channel called\nCracking the Cryptic, where these two\npuzzle champions from England, Mark and\nSimon, do uh pencil puzzles, and many of\nthem are uh well, they're online\nnowadays. Many of these are modern\nvariants of Sudoku.\nAnd so you know people take traditional\nSudoku and they invent all these cool\nnew rules like there are therm\nthermometers which are paths along which\nthe digits have to increase or there'll\nbe a box within which you're told what\nthe total of the digits in that box is\nor there are additional constraints like\ncells and knights move apart have to be\ndifferent. So there's an AI company I\nthink Sakana.\nOh yes I know them. Yeah.\nYeah. So they worked with these guys to\ncompile a lot of these. And so the goal\nis to can you get an AI to read these\nrules in English and then solve some of\nthese sudokus.\nLast time I looked the behavior so far\nwas pitiful. It was like you know they\nhad done a couple of 4x4 sedokus, right?\nI maybe a 6x6. I can't remember. And of\ncourse these these are puzzles that are\ndesigned by humans to have interesting\ninsights about them and uh you know cool\nglobal constraints that cause various\nkinds of logic which is just not present\nin traditional sudoku.\nAnd so so far the ability of AI to\nabsorb these rules and then use that to\ndo some kind of intelligent search that\nhasn't happened yet. Now I'm sure that\nit will I'm sure that it will improve. I\nthink one of the reasons why LLMs do\npoorly on these things I think is again\nthis basis on of one-dimensional text.\nAnd at least a year or two ago when I\ntried chat GPT on very simple tasks\ninvolving two-dimensional arrays of\nlike, you know, like the the classic\nqueen's problem and things like that, it\nit really couldn't do it. Um whereas we\nhave this sensorium, right? We're used\nto being able to look at a\ntwo-dimensional image. Our eyes, you\nknow, our pupils can sade around very\neasily. One of the reasons why we like\nSudoku is it's very easy to scan a row,\nscan a column, and scan a little 3x3\nbox. So, it fits with how we can address\nthat data structure, if you will. And uh\nthat lets us do I think much more uh\ndirected kinds of search. I mean, the\nlast thing you would want to do is\ntranslate it onto a big boolean\nsatisfiability problem and then use your\nfavorite boolean satisfiability solver.\nYou could do that, but that's certainly\nnot what Mark and Simon do on their\nYouTube channel. Um, they sort of sit\nthere and think about the rules and\nderive from them some heristic or some\nhighlevel logical constraint and then\nuse that. And I don't know, for me\nthat's a really interesting uh benchmark\nand I'll be very excited and a little\nannoyed if AI start solving those\nproblems. I'm proud to say eight of my\npuzzles are in that data set and so I'm\nwaiting to see if AI can solve my\npuzzles.\nYes. I suppose the paradox is that even\nthough they they are kind of compute\nrestricted, you made a a wonderful\nobservation yesterday that um you know\nwe have these hard problems and the art\nis transforming them into simpler\nproblems with with huristics. So in a\nsense the the intelligence is about\ndoing more with less. It's about making\nhard problems simple. And if only it\nwere possible just to to make that\ntransformation then the language models\nwould be able to do it. But what kind of\nintelligent process do you need? I mean\nwhat goes through your mind when you\ncome up with these incre you know these\ncreative flashes of insight?\nYeah. What one thing I like about the\npuzzle design community it's like there\nare like 10,000 people on this discord\nsort of built up around this channel is\nthat they talk a lot uh not typically in\na formal mathematical way but they talk\na lot about the art and science of\ndesigning insights and then finding\ninsights. And so acting not as an\nadversary but as a uh a challenging but\nultimately compassionate teacher who's\ntrying to create fun insights for the\nsolver to have. Um, and then you know a\nlot of talk about like because I I I\nthink the sensation you want to have as\na puzzle solver whether it's a wooden\npuzzle or you know fitting uh little\ntiles into some tray\nuh or or a sudoku\nyou want to have this at the first this\nvertigenous sense when you look at it\nlike oh my god I'm in this exponentially\nlarge search space op priori the last\nthing I want to have to do is exhaust\nhis search. It's boring. Humans are bad\nat it. That's the last thing anyone\nwants. Um, and so you want to sort of\nfeel that that sort of looking over this\nvast forbidding landscape\nand then you see an insight and you\nstart realizing things. I think one\nthing which is really fascinating is\nthat humans are quite good at uh\ndesigning on the fly different kinds of\npartial knowledge or partial solution to\na problem. So, you know, if you go back\nto the days of good oldfashioned AI\nwhere people were doing different\nbranching rules for backtracking search,\nDavis Putnham search, um certainly there\nwas a lot of clever ideas about if you\nhave some big boolean problem, which\nvariable should you try setting first?\nAnd people came up with these sensible\nheristics like if a variable occurs in\nmany different constraints, well, we\nshould set it first because that way\nwhichever way we set it, we'll satisfy a\nbunch of constraints, will make a bunch\nof other constraints more upset and that\nwill narrow the search space and that's\ngreat. Okay. Um, but humans do something\nricher than that. So like imagine you're\nsolving one of these wooden puzzles\nwhere you have tiles with different\nshapes. Um, penttoinos are my favorite.\nyou're trying to fit them together. Uh\nhumans will very fluidly switch from\nasking which piece can fit here and\nwhere can this piece go and which are\ntwo different kinds of variables in\nthese modern Sudoku variants powered\npartly by these really awesome apps.\nThere are you know traditionally uh\nSudoku fans had invented these different\nkinds of pencil marks. One of which\nmeans the thing here is either a two or\na seven which is one kind of partial\nknowledge and another is the three in\nthis box is either here there or there\nwhich is a different kind of partial\nknowledge. Now people are inventing new\nkinds of partial knowledge like uh these\ntwo cells I don't know what they are but\nthey have to be the same so I'll color\nthem both blue and then figure out what\ntheir numerical value is later or these\nthree cells I don't know what they are\nbut they all have to be different and so\nI don't know this this to me is a really\ninteresting frontier for AI where you\ntake the problem and you invent on the\nly what the what kind of variable you\nshould use to address the problem,\nright? Which is very different from\nbeing told here are the variables, here\nare the constraints. There's already a\nlot of interesting questions there. But\nhere it's more like okay you know fit\nthese things in you formalize the\nproblem you mathematize the problem and\nthen make some progress and maybe even\nfluidly jump from one mathematization to\nanother during the solving process and\nit's a lot I to me this is a lot like\nscience and when you're doing\nmathematical modeling right in many\ncases the challenge if you're working\nwith a social scientist or a biologist\nor whatever, 90% of the work is the\nmathematization,\nfiguring out what kind of mathematical\nstructure could fit here. And often once\nyou do that, it's relatively easy to\nsimulate the model or solve the model or\nprove something about it, whatever kind\nof work you're trying to do. Um that\nformalization process is something that\nI think is a really interesting kind of\ntask for an AI to do.\nYes. Yes. There's always this lingering\nproblem of resid, you know, residuals.\nWhat happens when we when we leave\nthings out? Um, but this process of\nepistemic foraging fascinates me, right?\nAnd also from a\nphrase, very good phrase.\nIt's wonderful. I got it from my friend\nCole Friston.\nOh, okay. The free energy guy.\nYeah, he's he's a great guy. And um so\nso there there's also this phoggenetic\nlocking in which I I I think is is good\nas a form of constraints but it's it's\nalso interesting from a flexibility\npoint of view. But we're trying to\nexplore this this space right to to\nforge a path. And the other thing is I'm\nnot sure whether you would call yourself\na Platonist or not, but there's this\nkind of interesting juxtaposition\nbetween are we converging on the real\nthing or are we constructing our own\nreality and where does culture and all\nof these different things come into it\nbecause we are very much just kind of\nlaying down the you called it partial\nknowledge. We're laying down the\nstepping stones and we're trying to move\nforwards,\nright? I mean, I guess when it's a\npuzzle, there is a ground truth. You've\nbeen promised there's a ground truth.\nUm, in the sudoku world, you've been\npromised that it's a unique solution,\nand when you found it, you know you\nfound it. Um, in real world problems, as\nyou say, it's very hard to know when\nwe've found the real thing and whether\nwe've failed to see something else. And\nand I guess right so I mean if what\nthese things know is what's on the\ninternet well that is a world it is not\nthe same as the physical world and uh\nyou know the so the grounding and\nmeaning so my friend Henry Ferrell who\nis a historian\num did a one of you know he tried out\none of these things where he wrote an\nessay and and in his essay he you know I\ncan't actually remember what the topic\nwas but he made a kind of subtle point\nthat was really a little bit sideways to\nthe various points that various people\nhad made and then he asked I forget\nwhich system to summarize his essay and\nit sort of blandified it right it it\nkind of lowest common denominator it.\nAnd it did kind of what people at a\ncocktail party might do if they're\nthinking pretty informally, maybe trying\nto impress each other a little bit. And\nit basically it saw it saw what he was\nwriting about. And then it produced a\nsummary based on the most common things\nthat people say about that. and it\ntotally missed the unique thing he was\ntrying to say that was different from\nthe common arguments on either side. And\nthis is interesting and I think you know\nfor him this\nthis was an indication again that these\nsystems are not grounded in meaning.\nThey don't really catch the oh that's an\ninteresting point. Now, you know, you\ncould say, oh, well, given a more\nsophisticated use of the statistics of\ntext,\neven if that's all they have, and then\nyou can argue about whether, you know,\ncompressing them forces them to build\nworld models, etc., etc., uh, maybe a\nbetter summarizer\nwould do, you know, would catch the cool\nthing, right? Just as maybe a better\nmusic uh maybe a better music or or book\nrecommendation system would challenge\nyou the way a friend challenges you in\nthat wonderful kind of directed way that\nfriends do. Like I know you don't think\nscience fiction is good literature, but\nyou have to check out Gene Wolf because\nthe pros is amazing and the characters\nare amazing and and I think it will meet\nyour literary needs and uh but I want to\nbring you over to science fiction. You\nknow, that sort of thing that our\nfriends do for us that I don't think any\nrecommendation system really does for\nus. It's like, you like this music?\nHere's some more music like that. Oh,\nyou kind of like that. Here's some more\nlike that. It's like, well, but give me\nsomething different, you know, challenge\nme. Um, and I think one source of those\nchallenges is the meaning the real world\nlike look at this cool thing or you know\nthis essay is actually about real\nthings. Think about those real things.\nDon't just look at the text and you know\nof course these are again like I I\npromised you that I I would say things\nthat other people have said better. Um\nso\nright.\nYeah. But this this is the this seems\nlike the debate and\num\non the other hand\nI do think like I said just just as I\nthink that once these things\nas they're already doing cannot just\nwrite code but run it and see whether it\nworks\nand then debug the code if it doesn't\nwork.\nuh or if it is a question about a\nthree-dimensional object, they could\nfire up a three-dimensional workspace or\na seven-dimensional workspace where they\ncan doodle and then kind of perceive the\nway we perceive, right? Um, I I am a bit\nof a platonist\nbecause\nif you and I close our eyes and we each\nthink of a cube,\nadmittedly, we live in a society with a\nlot of right angles and we've seen\nwireframes rotating on screen. So, we've\nhad a lot of practice with this. Um but\nboth of us can see in our minds a cube\nand we can count the fact we can just by\ncounting just by perception see that it\nhas eight corners and see that it has 12\nedges and if one of us thought it had 14\nedges the other would say no it's 12 and\nthe other one would look again at the\ncube in their mind and say oh yeah\nyou're right it's 12 right so we're\nreally perceiving something there\nand\nuh the fact that we can have that shared\nperception gives me and a lot of other\nmathematicians a sense that there is\nsome reality to these things. These are\nnot just subjective objects.\nAnd so I do think that um\nonce these systems\ncan\nswitch on the fly what kind of\nworkspaces they have and what kind of\nreasoning they do then I think that\nthey'll I think that they'll be much\ncloser to what we do right\nI mean even even if you ask them to do\nproofs\nof course there are proof finding\nsystems that are very formalized\nsystems. If you ask an LLM to construct\na proof, it will often construct some\nBS. It will be stylistically similar to\nproofs. It's read, but so far it doesn't\nseem to be able to do that reflection\nprocess and really check the steps in\nthe proof and see if it works. But of\ncourse, that is also a very specific\nthing that humans don't do very often,\nright? specific humans in specific\ncultures do this and have tools for\ndoing this. And we'll we might whip out\na sheet of paper and we might start\nwriting things with formal symbols and\nformal logic to see if our informal\nproof written in English or whatever\nactually holds.\nBut when we do that, we're firing up\nsome special mental models and we're\nusing some external tools, paper,\npencil, blackboards, computers to help\nus with them because actually formal\nlogic is not something that we're built\nto do. Um, so I guess I I expect\nI expect these systems to once they can\nreally play with all of these modules,\nincluding ones that we don't have like\nvisualizing things in seven dimensions,\nthey'll be able to do a lot. Just as\nwhen they start I'm not sure if we\nshould do this but when we give them\naccess to 3D printers and Fab Labs so\nthat they can start building things and\nseeing whether they work. Um well maybe\nwe should solve the alignment problem\nfirst. But uh yeah\nyes what what what you were saying about\nthe paste uh you know like the kind of\nGPT generated text was was very\ninteresting to me. Um so so they they\nmodel this statistical distribution and\nand they they they're greedily sampled\nand they just give you tokens from the\nbulk and you know one school of thought\nis well we'll make them more creative.\nWe'll just turn up the temperature and\nwe'll just sample tokens from the tail\nand and you you really get garbage there\nbecause you're kind of you know you're a\nlittle bit out of distribution now. And\none thing that would lead me to think\nthey were learning these kind of you\nknow factored representations of the\nworld is when when you did sample the\ntailor actually gave you something\ncreative and useful. But but I did I did\nwant to kind of like say that I'm not\nentirely sure whether it's that whether\nit is about learning meaningful\nstructured representations about things\nwhich are grounded in the world or um\nI'm a creative professional so I've\nlearned about video editing and I hire\nscript writers and so on and and I've\nnoticed that there might be something\nelse at play. I've noticed you know you\nwere saying you can add noise to\nproblems and and that actually makes\nthem more tractable. And in in audio and\nvideo, if you add noise, then you're\nactually training, you know, just\ntextures, high frequency patterns,\nyou're training human perception to to\nlook away. So you could blur, you could\nadd a texture and and and so on. And\nit's the same in writing that there are\nso many creative motifs just using\nslightly different language that\ndeliberately taking it away from the\nhead of the distribution, but such that\nit respects the epistemic fogyny. So it\nit needs to be meaningful but but still\ncreative and and it doesn't necessarily\nhave to have any grand meaning or be\ngrounded in in the world. So I guess the\nquestion is is it just kind of aesthetic\ncreativity or does it really need to\nrespect the rules? So this reminds me of\nuh Martin Amos and he has this book\ncalled The War Against Cliche and uh and\nhe makes some of the same points in a\nmemoir about his friendship with\nChristopher Hitchens. Um and his feeling\nas a pros writer is that any string of\nthree words which other people have used\nshould be avoided basically. I mean, I'm\nI'm paraphrasing him, but you know, if\nif you say um uh Okay, now I I'm having\ntrouble thinking of strings of three\nwords that other people put together,\nyou know. Um but you know, if you say\nanything which is a visible reference to\nsomething else,\nyou better be doing it on purpose, but\nyou shouldn't just be doing it because\nyou've heard it before and because it\nhas a high probability in the\ndistribution.\nSo I think his goal as a pros writer was\nto constantly produce new juxipositions\nand in order to well to intrigue the\nreader and to access a space in the you\nknow a to access a region of the\ncreative space the writing space which\nhadn't been accessed before\nand I guess and you know I'm sure you\nknow this as a creative person as Well,\njust as a just as a mathematician\nmight look at a proof\nand say, you know, hold it to the fire\nand say, okay, does this proof really\nwork? And we have lots of processes,\nboth individual and collective, to do\nthat.\nArtists make something and then hold it\nto the fire. Is this really good? Right?\nAnd of course, the pain of artists and\nmathematicians is that we crumple up a\nlot of pieces of paper and throw them in\nthe trash. Um, and it can be emotionally\nexhausting, but we have this very uh,\nyou know, we have this very high\nstandard for our own work. We don't just\nproduce things. We then reflect on them\nand show them to our friends and perhaps\nshow them to our critics and then try to\nmodify them or improve them or abandon\nthem. And\nuh and this this loop right I guess it's\na little bit like if you're a physicist\nyou do an experiment and see the\nexperiment works out. If you're if\nyou're a mathematician you do the quote\nexperiment but in a formal space of\nwhether it's logically sound and if\nyou're an artist you do the experiment\nof looking at it and judging it in the\nways we do. Right? And you don't\nOne thing I've learned from artists is\nthat art is not this kind of floppy\nthing, right? It is a very exacting\nthing. My PhD uh adviser, Philip Holmes,\nwrote poetry and he said this is much\nharder than doing math. You know, I I\ncompletely agree. Um on on that note, so\nwhen I look at um a video someone else\nhas edited because you know I'm very\nexperienced now and um it's it's very\nsimilar to mathematics or even the arc\nchallenge or something like that. You\nknow intelligence is the process of\ndecomposing you know something into the\nconstituent parts which made it. And as\na video editor, as any creative\nprofessional, you're trying to create\nthis progressive disclosure of\ncomplexity. So you're you know you're\ndelivering a sequence of artifacts to\nthe reader or the consumer which just\nincrease in complexity and you have to\njust put it past their prediction or\ntheir their cognitive horizon. So so so\nyou know\njust past\njust pass so you know\nyou don't want them to bounce off and\nyou don't want them to be bored. There's\na sweet spot.\nExactly. Even with audio production you\nhave these sound effects and they can\nhear the transitions. So you add\ntexturing, you add noise. But I can\nstill hear it. I'm going to add a little\nbit more noise. And at some point, the\nwhole thing just becomes more than the\nsum of its parts. And and the art the\nthe process of art is just building this\nup layer by layer and just having this\nalmost synchrony with the audience of\nknowing what their prediction horizon\nis,\nright? And you're and in order to do\nthat, you're\nyou're doing a great deal of mental\nmodeling of the viewer. You're\nconstantly putting yourself in their\nshoes. And uh\nand if I go back if I can jump back to\npuzzles for a little bit, right? So so\nlike when you design a puzzle, you're\nalso constantly putting yourself in the\nshoes of the solver and like would they\nget this? Are they going to see this? Is\nit,\nyou know, is it going to be uh visible\nbut just hard enough to see that it will\nbe a wonderful aha moment? And I think\none interesting thing philosophically is\nI think there are both subjective and\nobjective aspects to this, right? So\nsubjectively\num of course you're designing things for\nhumans and you as uh an editor and a\ncreator are designing things for humans\nwith well a certain level a certain\nlevel of literacy and familiarity with\nuh for instance the things that are\ntalked about on your channel and\nsimilarly you know if you're designing a\npuzzle well you're designing it for a\nhuman who has a certain a certain\ntolerance for search but not much more\num a uh again a certain\nit it's like if if you're in a chess\nplaying society, you kind of know about\nthe knight's move, right? So you so\nthere are certain things that you're\nfamiliar with. Um,\nand\nlike there there's a there's a variant\nmove in there's a variant rule in Sudoku\nwhich for some reason is called disjoint\ngroups that for me is very\nheadacheinducing which is that if\nthere's a seven in the top middle of\nthis 3x3 box, there cannot be a seven in\nthe top middle of any of the other 3x3\nboxes. This does not fit with my\nsensorium. I find it on a subjective\nlevel. I know that mathematically and\nlogically it's a it's a very nice\nextension of rows, columns, and boxes.\nIt's sort of like a three or\nfourdimensional extension treating the\nthing more like in a more hyperQB way.\nUm, but I hate it because I like have to\nlook over here and then look over there\nand then kind of painstakingly look over\nthere. I can't scan it in the nice way I\ncan scan rows, columns, and boxes. Um,\nso that's a qu that's an area where\nsubjectively I find puzzles involving\nthat constraint both harder and less\nfun. Um, it's also the case that if I\nwere a much more cognitively powerful\ncreature, then I maybe I would uh\nexperience just as much pleasure out of\n100 by 100 sudoku as I do out of 9\nby9's. Right? So, it's true that, you\nknow, I mean, I'm just 2 lbs of meat\nwith a one herz processor. I can, you\nknow, I can only handle the 9 by9\nthings. Is it two pounds? I'm not I\nhaven't weighed my brain. Um, you know,\non the other hand, I also I can't help\nbut feel that there are almost\nmathematically objective\naspects to aha moments. that maybe there\nare big aha moments and little aha\nmoments, but that we can all agree, we\ncan all sort of recognize them as\ninsights. Like when you're designing a\nvideo, you can agree that okay, at this\npoint now this concept is being like you\nhave a cognitive map of what concepts\nare being gained at each step and then\nused to build the next step. And maybe\nfor some viewers, some steps would be\nvery challenging, others they'd be kind\nof obvious, but they would all kind of\nunderstand that that's a step.\nSo I and in the puzzle world, there's a\nlot of recognition that a good puzzle\nand a hard puzzle, these are orthogonal\naxes.\nAnd there are simple there are simple\nbut beautiful puzzles. There are hard\nand beautiful puzzles. There's also\nsimple boring and hard and boring. These\nare really very different things. And uh\nyeah, I wish so. Theoretical computer\nscience supposedly helps us figure out\nwhat problems are easy and what problems\nare hard and what qualitatively makes\nthem easier or harder. What is it about\ntheir structure that makes them easy or\nhard? Why is this problem a smooth\nlandscape that a greedy algorithm can\njust find the optimum? And this problem\nis a very rugged landscape in physics.\nWe would say a very glassy landscape\nwhere there are many local optimates\nhard to navigate blah blah blah.\nI've tried a little bit to formalize\nwhat it is about these aha moments and I\nhaven't succeeded. It's a little bit\nlike public key public key cryptography\nwhere you have a function\nuh everyone can run the function\nforward. The challenge is inverting the\nfunction\nand um\nif you're given the public if you're\ngiven the private key then inverting the\nfunction becomes very easy.\nBut this is different. You have to find\nthe key yourself. You have to find the\ninsight yourself. Or there's this notion\nin computational complexity called\ncomputation with advice where again\nyou're given a big string of advice.\nWell, but again this is about finding\nthe advice. I feel like it's more like\nthe meta problem of designing an\nalgorithm.\nSo imagine that I show you an example of\na potentially hard problem, a you know\nan NP hard problem. But I promise you\nactually this example is easy.\nI promise you this example belongs to a\nlarge subclass of problems for which\nthere is an efficient algorithm. Now you\nhave to go find the algorithm.\nYeah,\nthat seems a little closer to\nthis puzzle design and maybe also a\nlittle closer to you're an intelligent\nentity dealing with a very structured\nworld. You're not having to parse\narbitrary images. You're parsing natural\nimages which through the processes of\nnatural selection\nthrough the structure of the built\nenvironment which is made by systems not\nentirely unlike you that are building an\nenvironment that they can understand and\nnavigate. Uh now your task is to\nuh understand, navigate, predict\nuh segment this data. Um\nand\nand what's really fascinating, right, is\nit's not so humans solving puzzles that\nwere invented by humans. Well, of course\nthat's a kind of ultimately it's a is\nyou're being communicated to by\nsomething with cognitive capacities and\ncognitive tastes, right? Enjoyments\nsimilar to yours and then you know you\ncan grab onto that. What's amazing is\nthat even the non-living world and even\nthe natural non-human world has all\nsorts of stuff that we can grab onto.\nyou can share it with other folks\nbecause I'm interested in in creativity\nand whether it's socially constructed or\nwhether it's grounded and as we were\nsaying the the other artists or the\nother mathematicians they can they can\nuh decompose the structure and they can\nsee if it fits the fogyny and they can\nidentify what the creative steps are. So\nthere's there's a kind of intrinsic\nvalue to it which is which is\nfascinating. And on on the other point\nabout you know actually forging paths in\nin this space I I wonder whether you\nwould agree it's related to\nundecidability and and even Wolram's um\ncomputational irreducibility. I mean,\nyou know, just imagine Wolffram would\ntalk about you have a cellular\nautomaton. Yeah. You know, like um\nwhat's one of his famous ones? Rule 134.\n110 110. Sorry, my my bad. But but you\nknow, or I know you you've um studied um\nyou know, like the the threebody\nproblem, right? So, so you you you do\nthis this um computation step by step\nand there are no analytical shortcuts,\nright? you just you just have to do this\nthis wide ranging divergent, you know,\nand then and then you find something and\nand that's amazing. You you've hit this\nstepping stone, but that there was no\nshortcut to get there,\nright? Or like a chaotic dynamical\nsystem where there's no closed form\nsolution, if you want to know the state\nit will be in at some future time, you\ncan't just plug t into some formula. Uh\nyou have to numerically integrate it and\num you have to do the work. you have to\nyou can't skip over its intervening\nhistory and so yeah I mean cellular\nautoma are a great playground for this\nthere are some the for the geeks like\nrule 150 which are kind of linear\nand uh they're like linear mod 2 or\nsomething and so if you want to if if I\ngive you the initial state and you want\nto know the state at some future time,\nyou can almost just plug it into a\nformula. You can you can create that\nfuture state with much less computation\nthan it would take you to actually\nsimulate.\nAnd then there are others where we\nstrongly believe that you have to go\nstep by step. uh the irreducibility as\nyou said that Wolffrin likes. One of the\none of the fascinating things about this\nis that our only techniques for proving\nthat prediction is hard that you have to\ndo the simulation is to build a computer\nout of the thing.\nSo, you know what happened with rule 110\nwas that Wolram observed all these cool\nparticles and thought, gee, you know,\nthese particles are doing all this cool\ncollisions almost like chemistry or\nalmost like almost like reading and\nwriting symbols on a touring machine's\ntape. And so um and then Matt Cook uh\ncame along and you know with Wolffrram\ncompleted this proof and proved\nwolffram's conjecture and uh then my\nfriend Damian Woods came along and and\ndid it more efficiently and so on. So\nthis is a lot like npmpleteness right we\nprove that problems are hard because\nthey have some kind of universal ability\nto encode or express other problems.\nTherefore, if they were easy, these\nother problems would be easy, too. And a\nfunny thing though is there are a lot of\nsystems\nwhere we don't know how to build a\ncomputer in them, but they still look\nreally irreducible.\nThey're still doing all kinds of stuff\nthat looks really nonlinear and it\nreally doesn't look like you could jump\nforward in time,\nbut the stuff is so uncontrolled that we\ndon't see how to build a computer out of\nit. So, we can't prove that we can't\nskip over the simulation.\nSo, you know, like imagine that you were\nwalking around in a world, imagine that\nyou were thinking about computational\ncomplexity several thousands of years\nago, which I guess you could have done.\nUm, and maybe in some philosophical\nsense some people did. But suppose you\ndon't yet have like wires or pipes or in\ngeneral things that can transmit\nsome information a bit or whatever very\ncleanly from here to there and you\ndidn't have little gates that take these\nclean wires and then produce something\nelse and send it out along another wire.\nSuppose you just had this kind of what a\nfriend of mine calls lava of just this\nchaotic stuff going all over the place\nor you know like imagine looking at the\nyou know the the the flow of plasma in\nthe sun, right? These amazing\num videos we have now from these solar\ntelescopes.\nYou see things briefly forming and then\nbreaking apart and it's very chaotic.\nIt's sort of like um\nit's like the planet Solaris or\nsomething, right? There's the uh what\nyou don't see is stuff that's out of\nwhich you could say, \"Oh, that is a nice\ncontrolled building block. I could use\nthat to store a bit that I could then\nwrite to later or read from later or\nlike combine to.\nSo, some cellar automa\nhave this kind of very chaotic, very\nnonlinear-look structure, but what they\ndon't have that we know of are these\nnice particles that we can use to\ntransmit and modify information and\nsimulate a touring machine or whatever.\nUm, and I wonder if a lot of natural\nsystems are in this weird middle ground,\nright? You can build hydrodnamic\ncomputers if you have pipes and valves.\nAnd before transistors were coming\nalong, people were trying to build\nmicrfluidic computers, right? There's\nsome wonderful alternate history in\nwhich we don't have transistors and what\nwe have is micrfluidics everywhere. Uh\nand that would be fun to think about.\nIt's a little bit like the difference\nengine uh the stuff with uh right Bruce\nSterling and William Gibson have a novel\nabout that where the kind of the Babbage\nsucceeded in building these mechanical\ncomputers and that's the technology we\nhave. Um but you uh but can you build a\ncomputer just out of water just the\nflows of water?\nMaybe, you know, just out of the Navier\nStokes equations using little\nflux donuts to travel from here to\nthere. Maybe um and some people say yes.\nAnd uh but it seems harder because you\nthings are not channeled.\nYes. So there's a difference between the\ncomplexity of a system\nand whether we can get it to do the\ncomputations we want it to do. Right? It\nmight be doing very complicated\ncomputations internally\nthat are indigenous to its own dynamics.\nThat doesn't mean that we could say oh\ngood now I can use it to build a\ncomputer. Formally we know that there\nare problems which are undecidable\nbut which are not touring complete in\nthe sense that if I gave you a box that\nsolves this problem an oracle for this\nproblem that you could then solve the\nhalting problem. So they're undecidable\nbut not because\nthe halting problem can be reduced to\nthem. Similarly, we know that if P and\nNP are different, which we believe, um\nthe academic we almost everyone I know\nbelieves that not everyone. Uh but if P\nand NP are different, then there are\nproblems in the middle ground which are\noutside P they cannot be solved in\npolinomial time. They cannot be solved\nefficiently and yet they're not\nNPcomplete. they don't have the ability\nto capture other things and you know but\nthe annoying thing is the only way we\ncan prove that a problem is hard is by\nshowing that it is complete that\nessentially by building a computer out\nof it.\nMy my co-host Dr. Doug, he's um a big\nfan of of you know touring machines\nbasically and and and he thinks that\ncurrent AI is is limited because\ntransformers are not touring complete\nand he thinks the reason for that is\nthat they're finite state automter and\nthey're trained in in such a way that\nmeans you know you can't really have a\nrecursive thing when you're doing\nstochastic radio descent because it\nwould just go on forever. But his\nfundamental hypothesis is that he he\nkind of thinks of GIS as being touring\nmachines and and even um the you know\nyou can have different strengths of\nagency. So a strong agent is a touring\nmachine um you know a thing which does\nsome computation you know it's it has an\nenvironment signal it takes an action\nbut if the C if the block in the middle\nis a touring machine then it it's\ncapable of strong agency. So I guess my\nquestion\nGI is general intelligences. Generative\nwhat's the G?\nOh, general. Yeah. But I mean as in I\ndon't know whether you would I mean I'm\nyou're the perfect person to ask about\nthis out of all the people we've ever\ninterviewed. And you know do you do you\nthink he's right to think about like you\nknow touring completeness as as being a\nway to demarcate different forms of of\nintelligence? And do you agree with his\ntheory that if it were possible for us\nto train a touring machine, you know,\nrather than a finite state or Thomas,\nyou know, so imagine we could\nempirically train it. Would would that\nlead to amazing things?\nWell,\nyeah. I mean, of course, Turing in his\n1951 paper is this classic paper where\nhe says that uh\nartificial intelligences, although I\ndon't remember him, I'm not even sure if\nhe used that phrase, artificial minds or\nsomething, would be trained or almost\nraised like children are rather than\nprogrammed. Touring machines as an\narchitecture\nI think are rather brittle\nand um\nI mean I think that the partly analog\nnature of neural networks and uh and\nLLMs this ability to\nuh I mean I know that they can be made\ndiscrete and so on but somehow that\ntheir ability to work in a continuous\nway with highdimensional vector spaces\nand embeddings. Um I I think that that\nis important to their trainability even\nif it's not ultimately important to\ntheir cognitive abilities. I mean I\nguess an easy repost to your question is\nthat I am also a finite state machine. I\nmean I have a very large number of\nstates. I will I will not have within my\nlifetime the ability to explore more\nthan a few of them. But I'm composed of\na finite number of neurons, a finite\nnumber of elementary particles. So I\nhave a very large but finite number of\nstates. Now\nI think the the difference is that\nuh because I am also a tool using and\ntool making entity.\nIf I realize that there's a problem\nwhich is difficult for me to do in my\nhead, which is most problems of any size\nat all, I can then build things, whether\nthat's uh uh a clay tablet or an abacus\nor a computer that extends my workspace,\nextends, if you will, the tape of my\ntouring machine. And and that gives me\nin principle recursion. I mean then we\ncan get into oh well is the universe\nactually finite blah blah blah that\nthat's not very interesting to me\nbecause we can reach fairly far into the\nkind of asintopia of recursion. Um\nfamously humans uh people joke about\nGerman speakers having a a stack depth\nof three or four and English speakers\nhaving a stack depth of one or two. I\nI'm not sure that I contain a stack. Um\nunlike what Nam Chosky supposedly said.\nI mean, if I had a stack in me, then it\nwould be a lot easier for me to repeat a\nstring of words backwards.\nYes.\nAnd that's very hard. If you give me a\nshort string of words, it'll be a lot\neasier for me to repeat it in the\noriginal order than backwards. So, I\ndon't think I'm very good at pushing and\npopping. I don't seem to have that kind\nof data structure in my mind. But if I\nneed it, I can build it with pencil and\npaper or a stack of plates on a table.\nSo I think it's that extensibility.\nYes. Yes.\nWhich gives us access to recursion\nand and universality.\nAnd that's partly why I guess I'm I'm\nexcited by the idea of AIS that can say,\ngee, this problem has a recursive\nnature. I cannot do this just in my own\ncontext window or my own embedding. I\nneed a data structure which lets me push\nand pop easily. And I know that before\ntransformers came along, people were\nalready working on these hybrids hybrid\nstructures where you have a deep\nnetwork. Rather than asking it to create\na stack in its own state space or like\ntrain a stack part of it which would be\nvery challenging give it a stack\nand let it take actions on that stack as\na data structure and let it learn how to\nuse it and play with it.\nCould I just refine this because I don't\nwant to misrepresent Keith. So\neverything you've said is absolutely\ntrue and when we have this discussion\nthere are many folks who say exactly as\nyou have done that um the the brain is\nan FSA and and we we extend you know\nlike we we can expand our memory by\nwriting things down on on on a hard you\nknow I can get another whiteboard I can\nget another whiteboard and so on but his\nargument is slightly more nuanced he's\nsaying that yes our brain is a finite\nstate automter but if you look at all of\nthe algorithms that are inside that\nclass there are a subset of algorithms\nwhich are those that can control a\ntouring machine and expand memory and\nand so on and those algorithms are not\ntraversible with stochcastic gradient\ndescent. So he's roughly saying that we\nwere you know maybe the chsky argument\nmaybe we've got the merge operation or\nsomething you know somehow our brains\nhave learned the special class of FSA\nalgorithms that can expand our memory.\nI see that's that's an interesting\nclaim.\nI mean I think you know we're at this\nworkshop this week where there was a\nwhole discussion yesterday\nuh about what do we actually need\nlanguage for\nand\nuh you know and what do we actually need\nsymbolic thinking for because that's\nwhere recursion seems to start right and\nuh there are plenty of intelligent\nentities out there like our close\nrelatives the great apes and possibly\nour ancestors who were already making\nstone tools and teaching each other to\nmake stone tools using gestures. They\ndidn't need, you know, the full-on\nmodular structure of language that we\nhave. And you can do a lot to navigate\nthe world uh without, if you will,\nTuring completeness. And well assuming\nthat what we mean by touring\ncompleteness is kind of the ability to\ndo symbolic recursion and so on. And\nthe funny thing is you know then we're\nwe're starting LLMs as language first\nthings. They're not tactile first things\nor visual first things or find food\nfirst things the way we were. They're\nlanguage first things. And because\nlanguage is the medium in which we do\nsymbolic thinking and recursion, then\nwe're like, \"Oh, good. They should be\nable to leap to all this formal stuff in\nmathematics.\"\nBut they're not formal systems, right?\nThey're they're token producing systems.\nThe way same way most human speech is a\ntoken producing system, right?\nformal reasoning is something that we is\nkind of a thin veneer that we do in\nspecific settings on top of token\nproducing, right? You know, when we're\nchatting with each other or even talking\nabout topics that we've had\nconversations with with before, we're\nacting very much like an LLM. You know,\nwe're we're cheerfully in a distribution\nwe're pretty familiar with. We're\ncheerfully emitting tokens. We're doing\nit. we don't really need to do that much\nself-reflection about it. It's when we\nhit some edge that we're forced to do\nwell kind of the self-reflection we were\ntalking about before about okay is what\nI'm about to say is it actually does\nthis actually make sense and and that's\nsomething most of us don't do that most\nof the time right you know\nand\nso\nyeah\nI mean I feel like the touring machine\nitself\nso for instance is when I when I teach\ntheoretical computer science I don't do\nit in a touring machine ccentric way\nright and I think if if you look at some\nmore recent textbooks uh they don't do\nwhat the older textbooks did where the\nfirst thing you do is here is a touring\nmachine and you know the touring machine\nis partly of historical interest now I\nmean it's a cool minimal thing that's\nuniversal but there are other many very\nsmall things that are universal\nand uh\nwhether those are families of boolean\ncircuits\nof increasing size with yes admittedly\nsome kind of uniformity to them um or\nwhether it's counter machines or finite\nstate automa with two stacks you know\npeople like Minsky and others had a lot\nof fun in the 60s and 70s is finding\nthese smallest possible machines that\ncan do that or solar automa or whatever.\nUm so for me the terrain machine isn't\ncentral actually it was the first thing\nlike it which had this ability to uh\nsimulate\nhad this universal ability to simulate\nother machines of its own kind. um had\nthis paradoxical ability to simulate\nitself and therefore the halting problem\nand so on. Um but I don't view that\narchitecture as central. It's very von\nnuminy, right? You have a CPU, you have\na memory. Um it's very magnetic tapy.\nYou know, you roll the tape over to this\npart. So yeah, I I feel I guess for me\nmathematically when I think about\ncomputational universality, I think\nabout\nthings like our favorite programming\nlanguages and their relationship with\nthe uh the theory of partial recursive\nfunctions, right? So, as I'm sure a lot\nof your viewers know, these basic\nnotions of recursion that were invented\nbefore terrain came along. Primitive\nrecursion is basically a for loop.\nThere's this other operator called\nminimization, don't worry about it,\nwhich is basically a while loop. And\nthen function composition is basically\nwell function composition.\nSo these tools can generate all of what\nare called the partial recursive\nfunctions which are more familiarly now\ncalled the computable things right these\nare the same things a touring machine\ncan do and then church you know has his\nwonderful lambda calculus and that shows\nup in Haskell and list and so on. So to\nme like the wonderful thing is that\nthese rather different architectures\nand the touring machine,\nthe grand unification which occurred in\nlike 1936.\nWhat is marvelous is that they can all\ndo the same thing.\nYes, maybe I should clarify that Keith\nwasn't talking about a physical touring\nmachine. He was talking about the\nstrength of the computation. I think in\nin a practical sense he was saying\nexactly as you were just saying that\nbeing able to do arbitrary uh loops and\nrecursion in an algorithm,\nright? Yeah. So ultimately, if you can\nif you can take building blocks and use\nthem to make more complicated things and\nthen use those things as building blocks\nand then wire them to each other and\nwire them to themselves, you can do\neverything we're talking about\ncomputationally.\nAnd that is what we do as technological\nbeings.\nyou know, we build these incredible\nscaffolded technologies\nand uh which, you know, we haven't seen\nthe end of yet. And if you can do that\nin a virtual space, then you are an\nunbounded technological being in the\nworld of mathematics. You can build\nthese arbitrary computable functions.\nYou can compute anything which any other\nuh reasonable architecture can compute.\nAnd\nnow right so\nyes touring machines can do this\nalthough they're very close to the\nhardware as it were right and of course\nTuring's great achievement was showing\nyou know that we can do software on top\nof a touring machine in a in a sense\nright and just as good's great achieve\nachievement was showing that\nmathematical formulas can talk about\nthemselves.\nHe showed how to do that, right? He he\ndid he if they built compilers, if you\nwill.\nUm, and it's funny like when you teach\nstudents nowadays, the halting the\nundecidability of the halting problem,\nyou say, \"Oh, well, if we could solve\nthe halting problem, we could ask it\nabout itself and then halt if it doesn't\nand and not halt if it does and there\nwe're done.\" And they're like, \"That's\nit. That's what touring is famous for\nbesides fighting the Nazis. It's like\nwell you got to understand you know this\nbusiness of feeding programs to\nthemselves.\nThis was kind of astonishing right? I\nmean\nin the early even in the early 20th\ncentury mathematics had this very\nstratified structure. There were numbers\nthen there were functions which act on\nnumbers and produce other numbers. Then\nthere were sort of meta functions or\nfunctionals like taking the derivative\nwhich takes a function and produces\nanother function. You couldn't feed\nthings to themselves that what does that\neven mean? That's nonsense. And then\nalong come tour Touring and Church and\nGoodel and show that you can that's\namazing right and like Doug Hashtder\ntalks about in good lesserbach it's sort\nof like enzymes and proteins are both\nprograms that can act on each other and\ndata like strings of amino acids and\nnowadays this is the air we breathe. A\ntext editor is a program that works on\nother programs. a compiler, an operating\nsystem, people in compiler class compile\ntheir own compilers, right? But I hope\nthat we understand that wow, this is\nactually really deep and amazing.\nSo yeah, that self-re reflexivity, that\nability to build things on top of other\nthings. You know what I like about the\nChsky hierarchy? If you look at a\ntouring machine, it is kind of a finite\nstate machine with an infinite number of\nstates, right?\nYeah.\nAnd and you know, if you're used to\nhaving these little graphs of like\nstates, well, it is one of those. It's\njust an infinite graph.\nThen you go to this higher level of\ndescription and say actually this thing\nhas a finite description, right? Just as\nif you have a stack of course you could\ndraw an infinite series of states where\nyou push push push push and then pop pop\npop and it would be a big binary tree if\nyou're pushing and popping binary\nsymbols and so on. um it's an infinite\nthing but then you move to a higher\nlevel and it has a finite description\nthat move\nI guess\nso let's\nthat move is something that\nyeah I don't saying touring machines can\ndo it I'm not sure\nyou can use touring machines to build it\nseeing the ability to do that,\nrecognizing this next level,\nuh, which makes a previous infinite\nthing finite,\nthat's a very cool thing for an\nintelligent entity to do.\nSo, I know this is a little sideways to\nyour friend's question. Um,\nbut\nyeah, that jump\num that jump to me is really\nfascinating.\nYeah.\nAnd\nI think that that's, you know,\nit's a little bit like, oh, recognizing\na statistical regularity and then being\nable to predict. Yeah. Well, but it\nseems like more. It seems like more. Um,\nand you've really recognized and\ncaptured an infinite set of objects all\nat once. You you've achieve you've done\na step of abstraction\nand um I would love to have artificial\npartners that can do that.\nIt also makes me think of that there's a\nbit of a a kind of a sandwich here. So\npeople like Wolram believe in digital\nphysics. So that the you know\nontologically the universe is is made\nout of computation. As a quick aside you\nsaid yesterday that you you you uh you\nlambasted you know folks for using\ncomputation as a noun which I thought\nwas was brilliant.\nCompute.\nSorry compute. Compute. Sorry. I\nYes.\nanyway like we need more compute and\nlike I'm like oh that does not sound\nright in my ears but okay I know\nverbing weirds language and as Calvin\nand Hobbs said and or and nouning does\ntoo but yeah\nI I know but um\nwe need more weird I guess yeah\nso so I guess you know like the the\nfirst part of the question is are you a\npan computationalist or you know maybe\nif you're not but the one step up from\nthat is do you think it's appropriate to\nuse the computation metaphor to talk\nabout a vector computations that the\nuniverse is doing and then you were\ngoing in an interesting direction a\nlittle while ago you know when you were\njust talking you know like Yosha Bark\nfor example he he talks about this kind\nof mimemetic virtual computation so our\nbrains are simulators and and we do this\nmeta programming and we share programs\naround and it's almost like the programs\nare the agents right so where's the\nlocus of agent\nselfish meme\nexactly yeah so so so you you've got the\nstack there and coming from the Santa Fe\nInstitute Of course, you know, we were\nsaying that I mean, my co-host Keith, of\ncourse, he's a he's a big fan of this\ntouring machine thing because he's an\ninternalist, but you're you're\nsurrounded by so many fascinating\nprofessors who have this very like\nexternalist um you know, complex systems\nview of things. So, there are so many\nideas of thinking about, you know,\nintelligence like what what does it mean\nto you?\nAll right. So, I like to talk about the\ncomputational lens.\nSo to me like as a kind of general\nscientist to the extent I am one\nI I like to be agnostic about what I\nshould focus on or if you will what lens\nI should look through when I look at a\nsystem and computation is one such lens.\nSo like I mean uh and to me that's the\nlens which focuses on the storage and\ntransmission and transformation of\ninformation in a system. So in the cell\nright I have friends who are studying\ncells and um who study the origin of\nlife and I mean the ribosome is clearly\nin part a computational device uh which\nis transforming information from one\nform into another. um the error\ncorrection mechanisms in DNA replication\nand so on are clearly in essence\ncomputational\nand so\nyou certainly learn a lot by looking at\nthat system looking for computation.\nUm\non the other hand, so like people who\ntried to build artificial life and who\npartly because they want to understand\nhow life began in the physical world.\nI've heard some people say that we move\ntoo far in a purely computational\norformational direction. Right? So like\none idea about the origin of life is\nsome kind of ultimately formal system\nwhere strings make more copies of\nthemselves, right? And this brings us to\nlike lambda expressions which make\ncopies of themselves or you know uh\ntouring machine or like in core wars,\nright? A little a little bit of assembly\ncode which copies itself elsewhere and\nmakes more copies of itself. Um and uh\nand that is one approach the sort of\nreplicator first approach to life. But\nsome people think that the problem with\nthat is that it doesn't recognize that\nuh organisms are really dealing with\nthermodynamic constraints. They really\nhave to get energy. They have to\nuh manage chemical gradients and and\nextract free energy from these chemical\ngradients. So there's a lot of physics\nthat they have to do and chemistry that\nthey have to do. And here things like\nthe abundance of of different uh\nelements or electrical charge or uh\nlight thermodynamic stuff really\nmatters. And so from this point of view\nthe fundamental thing is not the\nreplicator. It's more like the\nmetabolism. the thing which channels\nfree energy the way uh a river channels\nwater or the way a lightning bolt\nchannels electrical charge. And so\nyou know this is this is a caricature\nbut so you could say oh all\nthisformational stuff the genome um all\nthese wonderful strings of symbols that\nlook very much like touring machines to\nus. This is just stuff that the selfish\nmetabolism built\nto better channel free energy,\nright? As opposed to metabolisms are\nthings that replicators built to get the\nenergy we need to replicate.\nAnd maybe, you know, who knows which\nthing is the tail and which thing is the\ndog? Um, which came first. And maybe\nthere's some truth to all of this.\nSimilarly like you know\nyou could say that even the orbits of\nplanets in the solar system are\ncomputing they're computing their own\nfuture positions and yes you can say\nthat I'm not sure what we learn about\nplanets by saying that you know so for\nme I I mean as a pan computationalist do\nI think everything is computing yeah but\nI think in some cases that's an\ninformative thing to say and in other\ncases a less informative thing to say.\nUm so I think that focusing on that just\nas other sorts of lenses like another\nlens is adaptation. Are things evolving?\nAre they adapting? Are they learning?\nEither learning within a lifetime or\nover evolutionary time. Yes, that's\nanother thing that a lot of things are\ndoing. Um sometimes that's really\nimportant understanding them. Um, other\ntimes it might be less, you know? So,\nlike,\nuh, for instance, I I have a strong\nallergy to evolutionary psychology,\nright? Like I know maybe some of the\nways we treat each other and think and\nfeel might be because it's adaptive when\nbecause we're social primates, blah blah\nblah. But I don't really find that\nhelpful to thinking about I certainly\ndon't find it helpful to think about\nethics, right? And um it's like the\norigins of things are not always the\nimportant thing about them. Um the\nconstitution was written by slave\nowners. Yes, that's historically\nimportant. It's also this system that we\ncan use and\nyou know call upon now to try to do good\nthings in society.\nyou know, so each of these lenses are\ninteresting and they reveal different\nthings and it depends on what you're\ntrying to do and what kind of phenomenon\nyou're trying to understand. Um, and I\nthink we should kind of freely and\nfluidly switch back and forth between\nthem when we're trying to understand\ndifferent things.\nI didn't quite get whether you would\nagree or disagree that the universe\nonlogically, you know, in its primacy\ncould be thought of as computational. So\nI guess I translate that into is it\nsimulatable by a computer?\nOh, is there is there a distinction\nthough? Because I I I would think it\nwould be possible for the stuff the\nuniverse is made out of to not be\nsimulated by a computer, but we could\nsimulate it. Maybe I've just said\nsomething very stupid there. Maybe what\nyou said was correct. Is there a\ndistinction there?\nI don't know. I mean, is a computer\ncomputational? Right. So a computer is\nthis thing made out of elementary\nparticles which are doing all sorts of\ncrazy things. Um we exploit a small\nfraction of their dynamics to do to make\npixels and do you know do anyway right I\nmean obviously the laptop is doing all\nsorts of things other than the\ncomputation we want it to do. Um,\nso\nis it at the fundamental level a\ncomputer? I mean, I don't know. I'm not\ntrying to slip out of the question.\nI like Fineman's question of if I have a\nspace-time box, a single cubic meter\nsecond of spacetime,\num is the amount of information\nprocessing or shall we say computation\nin there finite?\nAnd you know, I'm inclined to think so,\nbut I don't really know. I mean I don't\nthink that at the fundamental level\nthings are solar automa because I think\nthat doesn't really work with quantum\nmechanics. Um\nI like this picture that at the plank\nscale something funny happens to space\ntime so that you don't really have an\ninfinite infinitely divisible continuum\nof space and time at those smallest\nscales. I don't think it's a lattice,\nbut maybe it's something more amorphous.\nAnd you know, people talk about causal\nnetworks and so on. And my first paper\nwas about causal networks. And anyway, I\nmean,\nI'm inclined to think that at the end of\nthe day that the physical church touring\nthesis is true\nthat uh any\nthat that the universe\nright there there one form of that says\nthat any device that we could actually\nbuild\nwould be simulatable by say a quantum\ncomputer with finite resources.\nThen there's also the question about\nwhat does it do by itself, right? Like\neven in this box,\nthere could be analog degrees of freedom\nthat go all the way out to infinity.\nLike you know the states of this box\ncould be real numbers that really have\nan infinite number of digits. And I\nspent some time in my in my career\nthinking about analog computation which\nby the way is a wonderful cool history\nwith Claude Shannon building mechanical\ncomputers and so on. So if you have real\nnumber computation then in theory there\nare an infinite number of bits there\nthat you could call upon. The question\nis can you read to them? Can you write\nfrom them? But even if we couldn't\naccess them as engineers to do an\ninfinite amount of computation, maybe\nit's still doing an infinite amount of\ncomputation itself, if you know what I\nmean. Um\nI'm inclined to think that's not the\ncase. Uh so yes I'm inclined to think\nthat there is a finite amount of\ncomputation happening\nand if you want to say that that\nmeans that it's computational\nalthough if there were infinite amount\nwe could say it's computational. It's\njust a really awesome kind of infinite\nhypercomputation. Yeah, some of the work\non hypermp computation is a little\nsilly. Um, but you know, I mean, what\nhappens in black holes and can you uh\nyou can set uh set your grad students up\nin orbit around a black hole. Um, make\nsure that they have a hereditary\nuh monkhood which will keep working on a\nproblem. Then you wave goodbye and fall\ninto the event horizon. And if they ever\nif their touring machine ever halts, if\nthey ever solve the problem, then they\nsend you a signal which of course will\nvaporize you because it will be blue\nshifted into gamma rays. But then, you\nknow, in theory, maybe you could learn\nsomething about, you know, closed if you\nhave a closed timelike curve, you can do\nreally awesome things. Um, yeah. I mean,\nI don't know. Part of it. So,\nphysicists, which is my original\nculture,\nhave an allergy have an have an an\nallergy, why did I say that? Have a oh,\nhave an allergic reaction. That's why I\nstressed that syllable to infinity.\nSo for us when something is blowing up\num and like when an integral is\ndiverging say or some infinite sum is\ndiverging instead of converging\nlike you seem to be able to do an\ninfinite amount of computation in finite\ntime\nfor us this is a sign that something is\nbreaking down right and so the attitude\nis oh well the problem with this black\nhole idea is that\nthere's cosmic censorship which will\nactually prevent us from making uh\nclosed timelike curves or there's going\nto be some sort of noise or firewall\nwhatever at the event horizon which will\nblow up our ability to do this or\nor you know as Sean Carroll says the\nuniverse is expanding so fast that uh\nit's rather grim in my opinion that we\ncan only do a finite amount of\ncomputation before uh the stars go out\nfrom the accelerating expansion. And you\nknow, this just pisses me off. I'm like,\nyou know, we'll do something about it.\nWe should, you know, the point is not to\nstudy the world, but to change it. So,\nlet you know, do we should do something\nabout that. Um, you know, which gets us\ninto science fiction. Uh\nso but for physicists right I mean every\ntime in particle physics there's\nsomething which\nseems to give an infinite answer we\nthink that means that our theory is\nbreaking down somewhere and historically\nthat's been true.\nSo that that gives us this sense that\nthere aren't any real infinities.\nAnd in particular, that sort of fits\nwith the idea that we're never going to\nbe able to build a box or even find a\nbox out there made of black holes or\nwhatever that can solve undecidable\nproblems. Um, but we don't really know,\nright? Ultimately, this is a a claim\nabout the world which may or may not be\ntrue. Um, I'm inclined to think it is.\nAnd I guess yeah I mean like I I grew up\non like Fredken and Tfulli's uh digital\nphysics and and reading Wolfrrim and\nplaying with cellar. So yeah I kind of\nthink something\ndiscreetish is happening at the finest\nscales of space and time.\nFascinating. And and just before we go\nChris we haven't really spoken about the\nalgorithmic justice. So this is this is\nsomething that you've been spending a\nlot of a lot of time looking at recently\nand I suppose it's difficult because we\nwe we're building these inscrutable um\nneural network models and I I think you\nknow certainly in common parliament\nthere's this intuition that it needs to\nbe inscrutable because if we make them\ninterpretable and if we kind of dumb\nthem down to be understandable then they\ndon't work as well. But we now have\nthese unbelievable illeible black boxes\nthat are making consequential decisions\nin our society.\nRight. Yeah. So I I I have thoughts\nabout this and maybe this is another\nconversation. Um I don't think these\nthings should be inscrable. I mean,\nI I I think that there is a range of\napplications, right? If you recommend\nmovies to me using a blackbox and I like\nthe movie, everybody's happy. That\ndoesn't bother me, right? Um maybe if I\nwere a filmmaker, I would want to know\nmore, but it doesn't really bother me as\na consumer. Uh, at the other extreme, if\nyou are putting me in jail even though\nI've not yet been found guilty of a\ncrime, or if you're using AI to help\nfind me guilty of a crime, you know, we\nhave these things in the Bill of Rights\nthat say I I should be able to confront\nmy accuser. I should be able to\ncross-examine witnesses. I should be\nable to contest evidence. And the\ninteresting thing about these things,\nthis is what people call procedural\nfairness.\nUm,\nand the interesting thing about the\ncriminal justice system is we explicitly\ncare about things other than accuracy,\nright? So, for instance, we've all\nwatched TV shows where the guy actually\ndid the deed, but the police planted the\nevidence.\nThey violated the rules of evidence.\nThey knew he was guilty. They wanted to\nput him away. And then they crossed the\nline and because of that, he got to\nwalk. And in our society, we think\nthat's how it ought to work, right?\nBecause we don't just want to be\naccurate\nin putting away guilty people and\nreleasing innocent people. We want to\nhave a certain relationship between\ngovernment and its citizens. We want to\nhave rules about how can the government\nuh surveil you, investigate you and\nthat's really profound, right? And\nthat's just, you know, how do you\noptimize for that? How do you even\nmathematize that? In some of the work on\nfair machine learning, you know, people\nlook at statistical notions of fairness.\nOh, we have this group of people, we\nhave that group of people. Um I'm a\nlittle bit disturbed by the assumption\nthat everybody belongs cleanly to one of\nthese two groups. I think that's part of\nthe problem. But to the extent that we\ncan divide the world into subopuls and\nit's like well we want the false\npositive rate to be equal or whatever.\nWell that's a constraint. We can add\nthat to the model. We can tack that onto\nthe algorithm and people have done lots\nof good work in that direction. But I'm\nreally fascinated by these other harder\nto mathematize\nnotions not just of fairness but yeah I\nmean what do we really want these\nsystems to do? Um one interesting fact\nwhich I recently learned from a guy\nnamed Mark Kaneus who has a PhD in\naerospace engineering and then went to\nlaw school and became a public defender.\nI mean is that a lot of the software\nproducts which are being used to do DNA\ntesting\nspecifically this thing called\nprobabilistic genotyping where it's been\na couple of days there are multiple\npeople who passed through the scene the\nDNA has fallen apart into pieces you\nknow how do you then there's there are\nsome choices to be made here about how\nthis is not a perfectly clean math\nproblem about was the defendant at the\nscene of the crime in that setting. Um,\nmany of the software tools that are used\nfor this, there are kind of two or three\npopular ones. Some of them have never\nbeen or at least until recently were not\nindependently tested by anyone.\nMany of them were not open- source. They\nwere proprietary products and they\nsometimes disagreed with each other.\nRight? So,\nwhat is the right metaphor here? Is this\num are these things expert witnesses\nthat you can cross-examine? Not really.\nAre their designers the witnesses? Do\nyou cross-examine their coders or the\nbioinformatics behind them? Um you know,\nif they disagree with each other, how\nare judges and juries supposed to\nevaluate which one is better? Um,\nso for me I I like the idea of\ntransparency\nwhich\nfor me is a stronger word than\nexplanability or interpretability.\nI agree transparency is a moving target.\nIn some settings, it might just be, has\nsome independent agency, consumer\nreports or underwriters laboratories\ntested this thing and can they verify\nthe vendor's claims that it works. In\nsome in some settings that might be\nenough and in a lot of settings even\nthat is missing right from things that\nare being used right now to make\nimportant decisions about people. In\nsome other settings, I really want to be\nable to look under the hood. And\num yes, I know deep networks are hard to\ninterpret and but at least it's a start.\nIf I can look under the hood, I can do\nthese sort of fMRI experiments like you\nknow like the Athell paper where people\ntry to do the tomography and figure out\nwhat kind of model it's building. I\nthink that's a very interesting line of\nwork. Um, so I think that as humans it\nwould be very good for all of us,\nespecially if we want a democratic\nsociety where we're making kind of\ninformed collective decisions about when\nto use these things, if to use these\nthings, in what settings to use them. We\nshould all try to understand these\nthings as well as possible. There are\nmultiple sources of gaps in our\nunderstanding.\nSome gaps are there for\nhonestly good reasons like deep networks\nare hard to understand. Some gaps are\nthere because of intellectual property\nand because people don't want to reveal\nhow these things work because they want\nthem to be proprietary.\nI am not very sympathetic to that second\nkind of gap and I think that kind of gap\nshould be closed. I don't think we\nshould be using opaque proprietary tools\nto make decisions that affect people's\nfundamental human rights. I think it's a\ncontinuum. Like in health, it's\ninteresting. Like if you are using a\nproprietary tool to diagnose my cancer,\nwell, I mean, I'm a geeky guy. I'm\nreally curious how it works. I would,\nyou know, um, if it's been independently\ntested\nby people who are not paid by the vendor\nof this system and it's really led to\ngood outcomes, even if it's a black box,\nI might I might go along with it because\nI want to live, you know. Um,\nyeah, I don't know. I mean, I think it's\na continuum like what what level of\ntransparency we would demand, but when\nwe get into sort of constitutional\nrights, I think we should demand every\npossible form of transparency.\nUm, and uh, yeah,\nChris, it's been so lovely to have you\non the show. Thank you so much for\njoining us today.\nThank you very much. It's been a great\ntime.",
  "transcript_chars": 80875,
  "ingested_at": "2026-05-12T00:42:45.116289+00:00",
  "source": "channel",
  "yt_meta": {
    "view_count": 14219,
    "like_count": 433,
    "channel_id": "UCMLtBahI5DMrt0NPvDSoIRQ",
    "categories": [
      "Science & Technology"
    ],
    "tags": []
  }
}