Language Without Meaning: How LLMs Exposed Our Biggest Illusion — Transcript
Full transcript
- 0:00"I'm going to get attacked by physicists... This thing is just ridiculously good. And so
- 0:04that just blows my mind!" Professor Barenholtz completely inverts how we understand mind,
- 0:11meaning, and our place in the universe. The standard model of language assumes words point
- 0:17to meanings in the world. However, Professor Barenholtz of Florida Atlantic University has
- 0:22discovered what's unconscionably unsettling... they don't! Language is actually deconstructing
- 0:28itself. Most startlingly, he argues that our rational linguistic minds have severed us from
- 0:34the unified cosmic experience that animals may still inhabit. I don't think there's a static set
- 0:39of facts. What we've got is potentialities. Most current LLMs operate with purely autoregressive
- 0:45next-token prediction, operating on ungrounded symbols. All of this terminology is explained,
- 0:50so don't worry, this podcast can be watched without a formal background in psychology or
- 0:55computer science. In this conversation, we journey through rigorous explorations of how LLMs work,
- 1:01what they imply about how we view the world, and the relationship between our consciousness and the
- 1:06cosmos. Professor, you have two theses. One is a speculative one and the other is more grounded.
- 1:15You even have another more hypothetical one atop that, which we may get into. Why don't you tell us
- 1:20about the more corroborated one, and then we can move to the contestable parts later. Okay, sure.
- 1:26So, yeah, I would call them sort of the grounded thesis and then sort of the extended version of
- 1:33that, if we can call it that. The grounded thesis is primarily about language. And the thesis is
- 1:41that human language is captured by what's going on in the large language models. And I mean not
- 1:49in terms of the specific exact algorithm as to how the large language models like ChatGPT are doing,
- 1:58are actually generating language, but the core sort of mathematical principle that large language
- 2:03models like ChatGPT run on are what's happening in the brain. And it's what's happening in human
- 2:09language. And really, the reason I say it's corroborated is because ultimately this isn't
- 2:14even about the brain, it's about language itself. And I think what we have learned in the course of
- 2:20being able to replicate language in a completely different substrate, namely in computers,
- 2:27is that we've learned properties of language itself. We've discovered. It's not through clever
- 2:33human engineering that we've been able to kind of barrel our way towards language competency.
- 2:40It's that with actually fairly straightforward mathematical principles done at scale, we've
- 2:47actually discovered that language has certain properties that we didn't know it had before.
- 2:51And so the incontrovertible fact, in my opinion, is that language itself has certain properties.
- 2:59Now that we know it has those properties, my claim is, the sort of corroborated claim,
- 3:05is that those properties force us to conclude that the mechanism by which humans generate language
- 3:12is the same as what's going on in these large language models. Because now that we know that
- 3:18language is capable of doing the stuff that it does, now that we know it has the properties to,
- 3:23and I'm sort of giving away the punchline, to self-generate based on its internal structure,
- 3:29it's unavoidable to think that we are using the same basic mechanism and principles. Because it
- 3:35would be extremely odd to think that we have a completely different orthogonal method for
- 3:42generating language. Or put differently, if we are using completely different mechanisms
- 3:48than the language models, then it's extremely unlikely that the language models would work
- 3:53as well as they do. The fact that language has this property that it can self-generate,
- 3:58the fact that that property actually leads to human-level language, to me, forces the conclusion
- 4:03that there's only one way to do language. And that one way is the same in humans and in machines. The
- 4:11obvious question that's occurring to the audience as they listen right now is, how do we know that
- 4:15whatever mechanism is being used by LLMs isn't just mimicry? Right, and so that's sort of the
- 4:21critical question. Is this mimicry, right? Is what the models are doing, in a sense, learning a kind
- 4:27of roundabout technique that captures some of the superficial components of language in humans, but
- 4:35ultimately it's a completely different approach. And so, you know, my argument is really from the
- 4:42fundamental simplicity of these models. So let's just talk really quickly about how large language
- 4:48models work. Things like ChatGPT. What they're doing is learning, given a sequence. You know,
- 4:57let's say the sequence is, I pledge allegiance to the, and then the model is being asked to do this
- 5:03thing called next token generation. What's the probable next word? We'll say word for the purpose
- 5:10of this conversation. We're going to call tokens a word. Token is a more technical term about how you
- 5:15chop up and encode the information in a sequence of language. But we're just going to say word.
- 5:22So guess the next word based on that sequence. So and then what you do is, in these models,
- 5:30is you train them to guess simply that. All you hear is a given sequence. It can be a sentence.
- 5:36It can be a paragraph. It can be, frankly, an entire book, depending on how big your model is,
- 5:42how much it can handle. And then guess just the very next word. And what we've discovered,
- 5:48and I say I really want to use that word in particular, because it was by no means a given
- 5:55that this could ever, that this would work, what we discovered is if you train a model to do that,
- 5:59to simply guess the next word, then take that word, tag it onto the next,
- 6:04tag it onto the sequence and feed it back in. This is sufficient to generate human level language.
- 6:10Now, the reason I believe that this demonstrates something not about our engineering or even about
- 6:17the models themselves, because there's different ways you might build a model that can do this, is
- 6:21because this very simple trick, this simple recipe of simply guessing the next word turns out to be
- 6:28sufficient to generate language at human levels to the point where there really are no benchmarks, no
- 6:35standard benchmarks that these models aren't able to do. And so what that suggests to me is just by
- 6:41learning the predictive structure of language, you're able to completely solve language,
- 6:46that means that that is likely to be the actual fundamental principle that's built into language
- 6:51in order to generate it. If we had to come up with a very complex scheme, for example, you know,
- 6:58syntax trees, complex grammar, long-range dependencies that we had to take into account,
- 7:04and through enough compute, we were able to kind of master that, then I might argue, well, you
- 7:10know, what we're doing is possibly figuring out a roundabout way to capture all this complexity.
- 7:16But it's the simplicity itself, that simply being able to predict the next token, the next word, is
- 7:23sufficient to do all of this long-range thinking, to be able to take an extremely long sequence,
- 7:28and then produce an extremely long sequence on the basis of that. That suggests to me that we
- 7:33discovered a principle that's actually already latent in language, that we just had to throw
- 7:38enough firepower at it, but with an extremely simple algorithmic trick, and then language
- 7:43revealed its secrets. So, to me, this really suggests that there is, of course, you know,
- 7:49that there's still a lot of science that needs to be done, and this kind of thing, kind of work
- 7:54that I'm doing in my lab, in terms of really being able to hammer down how the brain is instantiating
- 8:00this exact same algorithm, it's not going to look exactly like ChatGPT, it's not necessarily going
- 8:05to be based on what are called transformer models, which is something we can get into a little bit,
- 8:10but as far as the core principle of prediction of the next token, the fact that that solves language
- 8:16so handily, to me, really argues that that is the fundamental algorithm. That is the fundamental
- 8:21algorithm that, when you apply it, boom, language emerges. If you just have the corpus, you have the
- 8:28statistics, and then you do next token prediction, language is just like add water, and the fact that
- 8:33it emerges so readily from that, without having to do anything complicated, to me suggests that it's
- 8:39latent within language in the first place, and that language is designed, in a sense,
- 8:44in order to be able to be generated through this simple, predictive kind of mechanism. So Elon,
- 8:54you and I have spent several days together. In fact, you're in the video with Jacob Barndes
- 8:59and the Manolis Cosmon. We'll place that on screen and I'll put a pointer to you. And you were in the
- 9:04background of the interview with William Hahn on Williams. Always in the background, never in the
- 9:09foreground. Here we are. Okay, well, yes, great. You have a large epiphany that occurred to you at
- 9:15one point. You spoke about the software and this precipitated this entire point of view of language
- 9:21as a generative slash autoregressive model or what have you. Tell me about it. What the heck
- 9:27was that big idea? So it wasn't so much an idea as an epiphany, a realization. And it really hit
- 9:36me in a single moment. And it wasn't necessarily about autoregression. It wasn't about this finer
- 9:42detail of how ultimately language models, and I believe the brain, solved this problem. It was the
- 9:50realization that any model that has been trained, any model that anybody has built that accomplishes
- 10:01human level language. So it might be based on autoregression. It might be based even on
- 10:06diffusion, which is kind of the arch nemesis of my autoregressive theory. But regardless, the fact is
- 10:14that these models are being trained exclusively on text data. And so all they're learning is the
- 10:22relations between words. To the model, as far as the model is concerned, the words are turned
- 10:28into numbers, they're tokenized. We think of them as numerical representations. But those numbers,
- 10:33and for our purpose, we can think of them as words, don't represent anything. There is nothing
- 10:39in the model besides the relations. Relations just between the words themselves. There isn't,
- 10:45for example, any relation between any of the tokens and something external to it. What we
- 10:51tend to think of as people, as words, what words are doing. When we're discussing topics, thinking
- 10:58about words in our head, is that they symbolize something, that they refer to something. This is
- 11:04a lot of the philosophy of language, a lot of the scientific study of linguistics has been concerned
- 11:11with semantics. How do words get grounded? How do they mean something outside of themselves?
- 11:17What large language models show us is that words don't mean anything outside of themselves. As far
- 11:23as generation goes, as far as the ability for us to have this conversation, and as far as the
- 11:29model's ability to produce meaningful responses to just about any question you can throw at them,
- 11:37including writing a long essay on any topic, including a novel topic that it's never
- 11:42encountered, is by stringing together sequences based on simply the learned relations between
- 11:50words. And so this really hit me very, very hard. I've long been puzzled by, as many are,
- 11:58by the mind-body problem, the phenomenon of consciousness, the problem of how do we know
- 12:03your red is my red, and actually the moment that I had this realization, it was related
- 12:08to this very question, I realized that the word red doesn't mean what we mean by qualitative
- 12:14red. The qualitative red is taking place in our sensory perceptual system. The word red, to a
- 12:20large language model, can't mean that. It can't mean any color. It has no color phenomenon. It
- 12:25has no concept of what sensory red would mean. Yet it is able to use the word red with equal ability,
- 12:33with equal competency, just as well as I can, if we're just having a conversation about it.
- 12:38And so what this means is that within the corpus of language, the word red doesn't mean something
- 12:45external to itself. Instead, the word red simply means where does it fall in the space of language
- 12:51itself. Where does red fall in relation to other colors, in relation to the word color, in relation
- 12:57to other concepts, other, well, frankly, just words, tokens, that are related to what we call
- 13:03concepts that have to do with color and have to do with the word red. So, yeah, so this epiphany was
- 13:09about this extraordinary dichotomy, this divide between language and that which we think language
- 13:17refers to.The question is how does language refer, and that which we think language refers to. The
- 13:19question is, how does language refer? And the answer is, it doesn't. Language doesn't refer in
- 13:23and of itself. Language is an autonomous system. It's a self-contained system. It has the rules
- 13:31contained within it to generate itself, to carry on a conversation. A large language model don't
- 13:36know what they're talking about in any real sense. They can talk about a sunset. They can
- 13:43talk about a taste. They can talk about space and time and all of those things. And yet we would say
- 13:49they have no idea what they're talking about. And we'd be right in the sense that they don't have
- 13:55a notion of red beyond the token and its relation to other tokens. Now this then raises the obvious
- 14:02question, well, what do I mean what red is about? Don't I think red refers to a quality of
- 14:09perception? And the answer is, I do have a quality of perception. There is something called red that
- 14:14my sensory system is aware of. And then there's a token called red that is used in conjunction with,
- 14:21there's a sort of coherent mapping between my sensory perception of red and the linguistic red.
- 14:32But that doesn't mean that you need to understand what that word refers to. You don't need to have
- 14:39the sensory qualitative concept of red in order to completely successfully use the word red. And
- 14:47so these are compatible, but dichotomous systems. The sensory perceptual system and the linguistic
- 14:54system are ultimately, we can think of them as essentially distinct and autonomous, but
- 15:03compatible. Integrated? Integrated. They're integrated. So they're running alongside
- 15:11each other. They're exchanging messages so that we can have a single organism that is
- 15:16successfully navigating the world and able, for example, to communicate. So I see something
- 15:22red. That's registered in my brain. I have a qualitative experience of red. It's remembered
- 15:27in having a certain quality. And then later on I said, oh, you know, could you go pick up that red
- 15:32object for me? And so there's a handoff between the perceptual system and the linguistic system,
- 15:39such that the linguistic system can now successfully send a message to you. Now
- 15:44you've got the linguistic system. You can talk about that. Oh, okay, you told me there's a red
- 15:47object. Are there multiple objects? Yes, there's multiple objects. They have different colors.
- 15:51You're looking for the red one. Maybe it's a dark red. I'm doing this all linguistically. Now you're
- 15:56able to go into the room and successfully get the right object. So again, the handoff happens the
- 16:01other direction. Language is able to hand off to the perceptual system. And the perceptual system
- 16:04is able to then detect that there's something with the right quality. But that's not the same thing
- 16:10as saying that the language contains the reference inherently within it. It simply means that these
- 16:15are communicative systems, that they can exchange information, that they integrate with one another
- 16:22in terms of forming coherent behavior. But language is its own beast. It's its own autonomous
- 16:28system. It can run on its own. That was the big realization. Large language models prove it,
- 16:32that language is able to produce the next token, and by virtue of the next token, the next
- 16:37sequence. And that means all of language without having any concept of reference. The reference has
- 16:44no place there. There's no way to kind of squeeze it in. If your computational account is the one
- 16:50that I'm proposing, if the computational account is essentially prediction based on a next token
- 16:56based purely on the topology, the structure, the statistical structure of language, then there's
- 17:01no way to cram any other kind of grounding or computational feature in there at all. It has
- 17:08to be something closer to, in the large language models, prompting. You can imagine a camera that
- 17:15generates a linguistic description of what's in a room, and then you can ask your language model,
- 17:22and you can, by the way. You can do this right now. They're able to do vision. You can take a
- 17:26picture and feed it to the large language model. What's happening is much closer to generating a
- 17:33prompt, basically saying, here's what's in the room, and now based on these features,
- 17:37these scripts, now run the same exact language exclusive model. And so language takes care of
- 17:43itself. It doesn't need grounding in order to be able to do everything it does. It doesn't have to
- 17:48have concepts outside of itself. I think that's basically been proven by these text-only large
- 17:54language models. So that was the big epiphany. The big epiphany was that, oh, language is autonomous.
- 18:01Language is self-generating. That means it's a dichotomous computational system. It's independent
- 18:07of the rest. And what this leads me to believe is, okay, well, if it can live in silicon in this way,
- 18:14then perhaps, and now I've come to believe very strongly, that it likely runs in the same way
- 18:20in carbon, in biology, in our brains. Okay, so you're not denying consciousness and you're not
- 18:27denying qualia. No, and I want to make this very clear. My personal opinion on this is besides the
- 18:37point to some extent. You can be an eliminativist if you want, although I think everything I'm
- 18:44saying has a lot of bearing on this. But I believe my account is strictly an account of language.
- 18:52I think that perceptual mechanisms that give rise to qualia, things like redness and heat and
- 19:00taste and all of these, are basically processes that take place long before the handoff. And so
- 19:08what happens is, you know, think about the camera. The camera is transducing
- 19:12light. It's measuring certain wavelengths. Then there's a lot of visual processing that has to
- 19:18happen before you get to the point where it's turned into a linguistic-friendly embedding,
- 19:24right? The stuff that an LLM can see, a multimodal LLM can see. And so all of that processing that
- 19:30happens is what I think gives rise to qualitative experience. We experience redness because of all
- 19:37of this very analog, probably non-symbolic. kind of representation. And then at the end of that
- 19:46process, there is a conversion. Not, by the way, by the end of the process, a lot of things happen.
- 19:51We also respond to colors and to light and all of that non-linguistically. But part of the end,
- 19:59sort of, we could think of different endpoints. One of those endpoints is here's a handoff to
- 20:03language. And by the time language gets it, it's long past that initial process, that kind of
- 20:10sensory and perceptual processing that gives rise to qualitative phenomena. So I strongly believe
- 20:17that there is, in a certain sense, the word hard problem is a little loaded. I believe there's
- 20:23undeniable qualia. But what I also think is that language is poorly equipped. It's simply unaware,
- 20:34in some sense, of the underlying mechanisms that give rise to what it receives at the far end.
- 20:42At the sort of the endpoint of that qualitative processing. Just a moment. Don't go anywhere. Hey,
- 20:49I see you inching away. Don't be like the economy. Instead, read The Economist. I thought all The
- 20:56Economist was was something that CEOs read to stay up to date on world trends. And that's true,
- 21:01but that's not only true. What I found more than useful for myself, personally,
- 21:06is their coverage of math, physics, philosophy, and AI, especially how something is perceived by
- 21:12other countries and how it may impact markets. For instance, The Economist had an interview
- 21:17with some of the people behind DeepSeek the week DeepSeek was launched. No one else had
- 21:22that. Another example is The Economist has this fantastic article on the recent dark energy data,
- 21:27which surpasses even Scientific American's coverage, in my opinion. They also have
- 21:31the chart of everything. It's like the chart version of this channel. It's something which
- 21:36is a pleasure to scroll through and learn from. Links to all of these will be in the description,
- 21:40of course. Additionally, just this week, there were two articles published. One about the Dead
- 21:44Sea Scrolls and how AI models can help analyze the dates that they were published by looking at their
- 21:49transcription qualities. And another article that I loved is the 40 best books published this year
- 21:54so far. Sign up at Economist.com slash TOE for the yearly subscription. I do so and you won't regret
- 22:00it. Remember to use that TOE code as it counts to helping this channel and gets you a discount. Now,
- 22:06The Economist's commitment to rigorous journalism means that you get a clear picture of the world's
- 22:10most significant developments. I am personally interested in the more scientific ones,
- 22:15like this one on extending life via mitochondrial transplants, which creates actually a new field
- 22:20of medicine, something that would make Michael Levin proud. The Economist also covers culture,
- 22:26finance and economics, business, international affairs, Britain, Europe, the Middle East, Africa,
- 22:32China, Asia, the Americas, and of course, the USA. Whether it's the latest in scientific innovation
- 22:38or the shifting landscape of global politics, The Economist provides comprehensive coverage,
- 22:43and it goes far beyond just headlines. Look, if you're passionate about expanding your knowledge
- 22:48and gaining a new understanding, a deeper one, of the forces that shape our world, then I
- 22:53highly recommend subscribing to The Economist. I subscribe to them, and it's an investment into my,
- 22:59into your, intellectual growth. It's one that you won't regret. As a listener of this podcast,
- 23:04you'll get a special 20% off discount. Now you can enjoy The Economist and all it has to
- 23:10offer for less. Head over to their website, www.economist.com slash TOE, T-O-E, to get
- 23:18started. Thanks for tuning in, and now let's get back to the exploration of the mysteries
- 23:23of our universe. Again, that's economist.com slash TOE. To what it receives at the far end,
- 23:32at the sort of the end point of that qualitative processing. Okay, let me see if I get this. You
- 23:38have some redness. So you do, you're not denying redness. You grant redness. I do.
- 23:43Okay. There's redness. And then somehow this needs to be referred to with some spoken words,
- 23:49with some language. Okay. So what's happening? You're saying that it's an independent system,
- 23:56yet it's integrated. So what is that relationship? And does it become so diluted that by the time you
- 24:01refer to it, you're no longer referring to that qualia? Like, I don't understand. Yeah, that is
- 24:07essentially the idea. So this is the exact problem I am working on right now. There was a fantastic
- 24:13paper that I just came across about a week ago. There was a paper that was published in Archive
- 24:20recently. It's called Harnessing the Universal Geometry of Embeddings. And what this paper showed
- 24:26is that you could have completely different models solving different linguistic tasks. For example,
- 24:32you could have GPT, then you could have BERT, which solves a somewhat different task. So there's
- 24:36masked tokens as opposed to autoregressive next token generation. And what they found was that you
- 24:44could learn what is latent space. What you could do is hand off, take the embedding. The embedding
- 24:51is basically, you can think of that as numerical representation. It's a high dimensional numerical
- 24:56representation of your tokens. So here's a token. This token is going to represent the word dog. And
- 25:03then we're going to take that token and embed it in a much higher dimensional space. And what they
- 25:09found is that if you take the embedding, the high dimensional representation from one model, say,
- 25:15ChatGPT, and then take a representation from a different model, that you could actually get the,
- 25:24you could take the embedding, send it to this latent space. If you cycle it through, get the,
- 25:31you have to, I know it's starting to get in the weeds a little bit. But you send it through this
- 25:35latent space and then recover it in its original form. What you can do is, once you've got that
- 25:42latent space, you can then translate from one embedding to a completely different embedding.
- 25:48This is a new paper. This is a new paper, yes. Right, right, right. This rocked my world. Because
- 25:54what they're arguing is that there, in some ways, is this underlying universal structure of language
- 26:04that's captured in this latent space. And so even though if you have a radically different
- 26:08embedding in one, you know, they didn't do it across different languages, it's one of the
- 26:13projects I'm doing right now, is to see if you can do this across, say, English and Spanish, even for
- 26:19a language that's trained exclusively in English, and then another models trained exclusively in
- 26:24Spanish. Can you guess the Spanish just from finding this kind of universal structure across
- 26:32these two different models? Sorry, what do you mean, can you guess the Spanish? If a model was
- 26:37trained only in English, and then it was receiving some Spanish text, a couple Spanish sentences?
- 26:42So the way to think about it is that what you're doing is creating another embedding,
- 26:48this latent space, where you're going to be able to send in a message in English, and then
- 26:56based on the station, and then again, do the same thing for Spanish. And then what you're never
- 27:02going to show any model, no model is going to ever see a pair of English and Spanish. Instead,
- 27:07what you're going to learn is that there's some way to get from, you're going to end up being able
- 27:13to get from English to Spanish without ever seeing the actual translation. Because what the model is
- 27:18going to learn is what's common across these two representations. What's true for both the Spanish
- 27:25embedding and the English embedding, that there's some sort of underlying latent structure that's
- 27:30true of both, and that that captures something more universal about language. Now again,
- 27:34they didn't do it for different languages, they just did it for different embeddings of English,
- 27:41but very different embeddings, because they were trained on completely different models.
- 27:44If you looked at them, if you just looked at this sort of vector representation, you took a vector
- 27:48representation of the word dog in one, and a vectorization of the word dog in the other,
- 27:52they're completely, numerically, there's no similarity. You can never spot the similarities
- 27:56if you just looked even them pairwise. But if they do this kind of reconstruction,
- 28:00and then ask the model to be able to reconstruct, not in the original embedding space, but go and
- 28:08reconstruct in the other embedding space, it's able to actually do this. And so by doing that,
- 28:14by training it to do that, without ever seeing any pairs, it's able to sort of
- 28:17learn the translation between one representation and another representation. What this opened up
- 28:25to me is the possibility that we could think about the exact same kind of latent space in the brain,
- 28:31and possibly in artificial intelligence models, between the perceptual world and the linguistic
- 28:38world. That there is some embedding of how the physical world is structured. We understand, like,
- 28:45think about an animal, a non-linguistic animal, certainly has an idea of objects, and objects in
- 28:50relation to other objects, objects in proximity to other objects, moving around those objects.
- 28:55My dog, who was just barking in the background, knows what doors are, and she can go scratch it,
- 29:01and she knows it opens up. She certainly isn't able to express that linguistically, but she has
- 29:05this concept, and she's able to think about it. She's able, in some ways, to reason about that.
- 29:10My suspicion is that that probably is done maybe even autoregressively, but we'll leave that aside
- 29:14for now. The main point is that there is some representation of the facts about the world,
- 29:20the sensory facts of the world, or the sensory, I would say, the sensory construction, the facts
- 29:25that have been constructed based on sensory information. So that's some sort of embedding of
- 29:32the world. The linguistic embedding is a radically different embedding. It carries information about
- 29:38the world as well, but not in the way, not in the direct way that we think, not that the word, you
- 29:45know, my headphones are sitting on this desk as direct reference back to sensation and perception.
- 29:51No, it lives on its own. It's its own embedding, and it does its own, and it can do its own thing.
- 29:57However, based on this paper, this really gave me sort of a key insight that there might be
- 30:03this latent space where you can actually do this kind of mapping, where there's translation between
- 30:08linguistic and perceptual embeddings. They're as distinct as they are. Fundamentally very,
- 30:13very distinct, very different. They're there to solve different problems, but they're able to
- 30:17talk to each other. How? Perhaps through this kind of latent space where some universal structure,
- 30:24like, okay, in language, there's certain facts about language. There's a fact about the word
- 30:31dog or the word microphone that its relation to other words, like desk, in some ways captures the
- 30:39fact that microphone sits on top of desks. That fact is somehow actually contained within this
- 30:46embedding structure. In what sense? Well, if you ask me, would a desk sit on a microphone or would
- 30:51a microphone sit on a desk, I can answer that question. So can chat JPT, right? And without
- 30:56any notion of what microphones really are, sort of from a perceptual standpoint, they're having these
- 31:01kinds of properties, we can talk about them. And the linguistic embedding space contains
- 31:07this information. What does it mean it contains information? By the way, just to say, what does
- 31:10that mean? It means given a certain input, like, do microphones sit on desks? Where should I put
- 31:15my microphone? I can answer linguistically in a reasonable way, right? And that's what I mean by
- 31:20the knowledge. It's purely linguistic knowledge. It only can generate linguistic responses. But
- 31:26the point is that that knowledge lives in this kind of linguistic embedding. And then there's
- 31:31the other kind of embeddings. There's a visual embedding, there might be an auditory embedding,
- 31:36which is distinct. And then the idea that I'm very inspired by is that there can be this latent
- 31:42space that captures certain universals that are common across these different embeddings that
- 31:47make translation possible. So that when I see this microphone sitting on a desk, what's now
- 31:54available to me is the ability to describe that to you linguistically. But it's not direct.
- 32:00It's not that there's a very specific linguistic representation of this sensory perceptual kind of
- 32:07phenomenon. And this is important because forever philosophers, philosophers in general, linguists,
- 32:14have been trying to understand how do words get their meaning? Something I referred to earlier.
- 32:20What's the definition of a microphone? What's the definition of a dog? And the answer is there isn't
- 32:25a single one. There isn't a single definition that's ever going to capture. Instead, what
- 32:30you've got is this latent sort of bridge where there's some sort of representation of this fact.
- 32:35That given, you know, whatever your particular prompt is, your linguistic prompt is going to lead
- 32:42to certain kind of meaningful linguistic behavior. If you ask me a question about this microphone, I
- 32:47might be able to answer that question meaningfully based on the perceptual information. But what this
- 32:51microphone means is actually completely contingent on, at least linguistically, is contingent on
- 32:57whatever question you ask me about it. And so it's all going to depend on what you're doing
- 33:02with that latent space. There isn't sort of, and this is sort of a broader point, there isn't sort
- 33:07of a static set of facts about the world that's embedded in language. I don't think there would
- 33:13be a static set of facts embedded in our sort of visual embedding of the world. Instead,
- 33:18what we've got is what I call potentialities. We now have the ability to engage that latent space
- 33:26linguistically where the perceptual information kind of lives, sort of this universal embedding
- 33:31of it, and then do whatever we need to do with it. If I need to answer this question about it, I can
- 33:36answer that question. If you ask me a different question, I can answer that. But there isn't a
- 33:40singular meaning of microphone that captures sort of the entire set of facts. Here it is,
- 33:47here's the embedded set of facts. The set of facts is actually infinite. I could tell you infinite
- 33:53things about this microphone. For starters, to use a silly philosophical example, it doesn't
- 33:59have this shape and it doesn't have that shape. There's an infinite number of questions you could
- 34:04ask me about it that I could answer meaningfully about it. So all those potentialities are kind of
- 34:09what happens when the linguistic system interacts with this kind of shared embedding space. That's
- 34:15sort of the half-baked version of how I think language ultimately does have to, of course,
- 34:22language only is meaningful insofar as it can live within the larger ecosystem of perception
- 34:30and sensation and perception. We have to be able to take in information through our senses and then
- 34:35communicate, although I use that word kind of carefully, I don't communicate the entire
- 34:42representation because as I said, I don't think that's even a meaningful idea. Instead, what
- 34:47I can do is use language in a way that helps us coordinate our behavior. There's no way to sort of
- 34:53download the entire perceptual state that's locked up in some ways in the perceptual embedding. No,
- 35:01what I can do is pull some information such that I can meaningfully communicate with you in a way
- 35:11that then is gonna have the intended consequences. I'm not downloading perceptual information into
- 35:16your brain. I'm telling you what you need to know in order to be able to perform some action,
- 35:21to perform some behavior, or maybe even to think about it so that you could later
- 35:25perform some action. That was, I know that was a lot. And feel free to back me up and challenge
- 35:31me on any of these things. So I wanna see if I understand this and I wanna explore what is the
- 35:36definition of language, even though we just talked about, there isn't the definition of a microphone,
- 35:41say, but I do wanna talk about the definition of language and what is autoregression. And while
- 35:46presumably you're telling me what you believe with language, you're telling me this model
- 35:49because you believe it's true, I don't know what truth you're conveying if you believe this is not
- 35:55grounded. So what are you referring to when you even say that language is autoregressive without
- 36:00symbol grounding? I don't have an ideas to that. I wanna explore that. But first I wanna see if I
- 36:06understand you. Okay, fair. So a latent space. So let's think of a word. A word gets a vector,
- 36:11like an arrow, and I'm just gonna be 2D for this example because that's just what the camera picks
- 36:16up. So let's say the word dog looks like so, the word cat looks like so, whatever. Okay.
- 36:24The space that it's embedded in is called the latent space, is that correct? Well, the initial
- 36:29embedding is just the embedding. And so that's just a high dimensional, you know. It's just,
- 36:35yeah, let's forget the word high dimensional, right, it's just a big long list of numbers. And,
- 36:41you know, let's say you've got 10,000 numbers and for dog, we're gonna represent dog as this
- 36:46particular sequence of these numbers. For cat, it's a different sequence of these numbers. And so
- 36:53that's our initial embedding. So the latent space is a compressed version of that? Well, in some
- 36:58ways it's actually not compressed. It's actually, it's actually, what's the opposite of compressed?
- 37:03It's expanded. Uncompressed. It's an expanded version. It's actually, you know, so you have
- 37:08the original tokenization, which just says here in a fairly small vector, but then you expand it into
- 37:14a much higher dimensional embedding space. So that each token actually ends up getting much richer,
- 37:24much, many more numbers that are used in order to represent each token. And that's a very key
- 37:30fundamental thing that these models do. And by expanding it in these different dimensions, that's
- 37:35what allows you to sort of massage the space so that you can get all these cool properties like
- 37:39cat and dog being sort of in the appropriate relation to one another. So that later on, when
- 37:46you're trying to figure out what the next token is, you're able to actually leverage the inherent
- 37:53structure in this high dimensional space. Okay. So then you have the language model for English
- 37:59and then you have a language model for Spanish. Yes. And let's imagine that it was trained only
- 38:03with the corpus of English in the former case and only with the corpus of Spanish in the second. And
- 38:08then we can even have a third of Mandarin. Sure. Okay. Yeah, in fact, in the paper, they didn't do
- 38:13different languages. As I said, they did different embeddings of English language models, but yes,
- 38:17they use multiple, they actually did this across several different embeddings, not just two. Okay,
- 38:22so then the claim or finding is that if we look at cat and dog inside of here in English, it gets
- 38:29mapped to some fourth space here, which is like a Rosetta Stone space or a platonic space. Yeah,
- 38:35that's exactly what they call it, platonic. They use the word platonic. Okay, great. Well done. So
- 38:39it looks like this there. Okay, and then if you were to say, okay, well, let me just forget about
- 38:45English and this platonic space. Let me look at cat and dog in Spanish. Okay, and it looks like
- 38:50this here. Let me map it from here to my platonic space. Oh, wow, it gets mapped to a similar place.
- 38:58Oh, and does the Mandarin? Let's find that out, cat and dog. It does. Okay, let's test out more
- 39:02words. So the claim is that this space here is this meaning-like space. Yeah. Okay, great. And
- 39:10then what you're saying is that microphone, we think of microphone as living in here as a single
- 39:15vector, that would be like an essence of the microphone that we're referring to. But actually,
- 39:19microphone, our concept of microphone depends on the prompt. So explain that, that sounds
- 39:26interesting. Yeah, I think, and you're making me think about this in a way that I hadn't quite
- 39:34before. So the level of which I've thought about it, is that you've got these different embeddings.
- 39:42When I see a microphone visually, it's going to, there's a certain, there's a vector representation
- 39:48of what that sensory perceptual experience, and I don't mean the qualitative sense,
- 39:53I don't mean, I'm not getting into phenomenology, but there's something happening in our brain,
- 39:57my brain, that is sort of the representation of what it means for me to see this object from the
- 40:03visual standpoint. Okay, that's one embedding. And then we also have a word, microphone,
- 40:08which is a completely different embedding. There's simply a word that lives in language space,
- 40:14with that embed, what it means is that, you know, there's a specific embedding. So it's
- 40:19kind of helpful to think about a sort of a point in a space. So you know, you've got this super
- 40:23high dimensional space, and each individual token is simply a vector in that space, namely, so it
- 40:32really picks out a specific point. And we can say microphone lives right here, in this linguistic
- 40:37space. And then my perceptual experience, again, I don't want to use that word, but my perceptual
- 40:44kind of grasping of this microphone being here, is this point in a completely different space,
- 40:52this perceptual space, which has, you know, it captures other kinds of information in language.
- 40:58So let's actually talk about this for a second, or in language, the space, if you want it to be
- 41:02a useful, meaningful space, you're going to want things that have similar meaning, they're likely
- 41:08to actually be have proximity to each other. And this, this is to some extent what the large
- 41:12language models learn, they learn an embedding, in order to do next token, they learn embedding,
- 41:17that gives this this where the space, you know, and we can think of almost like two dimensional,
- 41:23three dimensional space, which is very high dimensional. But you know, for for our purpose, we
- 41:27think about that, that where the cat and dog live, you want those things to live closer together than
- 41:33cat and desk. And, of course, it's much richer than that, right? It's not just semantic,
- 41:39like this, this very kind of superficial level of semantic similarity. In fact, what it is,
- 41:45is capture the somehow the semantics, so to speak, are captured by the space, like the space,
- 41:51the shape of the space itself, is what allows the model to understand sort of the relation between
- 41:57words so that I can do the next token generation. But it's a very, very different space, right? It
- 42:02has to do with really with with relations between words, and in terms of generation, in terms of,
- 42:08you know, next token generation, so that it's useful for that purpose. What is the perceptual
- 42:13space look like? Well, this perceptual space is going to have a very different, the axes there
- 42:19almost certainly aren't going to have the same kind of meaning as in the linguistic space,
- 42:23there'll be something closer, maybe color features, shape features, something like that.
- 42:26And where this microphone lives is within that space is going to have radically different meaning
- 42:33than saying, you know, it's, you know, it's not apples and oranges, right? Those aren't different
- 42:37enough, right? It's apples and math or something. It's, it's, it's, it's really, really radically
- 42:44different kinds of spaces. But what I'm proposing, what I think the insight here is that, ultimately,
- 42:53there is the possibility of having a shared space that you can send, you can project both of these
- 42:58things to where microphone, the word is going to somehow make contact with this perceptual
- 43:07experience right now, this perceptual fact on, but it's not and here, you know, here's the key point
- 43:14that you're getting at. It's not that this word microphone picks out the exact same embedding,
- 43:20uh, in this latent space. It's not that it's going to make that thing light up. Oh, what?
- 43:24It's the same thing. No, it's that when you ask a certain question about a microphone, is there
- 43:30a microphone on your desk? My perceptual system is generating some, some, uh, well, first of all,
- 43:37it's just generating the perceptual phenomena, but then it's also sharing information in this latent
- 43:42space, which my linguistic system can then go draw from. And then given this particular prompt,
- 43:50was there a microphone on my desk? I'm able to then successfully answer the question.
- 43:54So it's not, it's not quite the same thing as saying that they're, they're picking out the same,
- 44:00uh, information in latent space, because my argument is that, that that's not really a
- 44:04meaningful concept. Uh, there isn't the same microphone in, in linguistic terms, doesn't
- 44:10pick out a perceptual, uh, kind of fact that's not possible. Uh, these are radically different
- 44:17kinds of facts. Um, but what the latent space might allow us to do is not just to translate,
- 44:23which is what they did in this paper, but perhaps to pass information along in a meaningful
- 44:27way so that you're able to access it and do something successful, like answer the question,
- 44:32is there a microphone on this desk? I think that might be what's happening to some extent,
- 44:36even in the multimodal models. So it's a longer conversation. That's not really how they work. Um,
- 44:42they don't actually operate based on a shared latent space or anything like that. Really what
- 44:46they do is, uh, the, the models learn to take a perceptual input and turn it into something,
- 44:52uh, like language. Uh, so it's, it's, it's more similar to like prompting almost. It's
- 44:57not exactly that. Um, but it's, it's injecting something within linguistic space that is,
- 45:04uh, equivalent to, uh, to actual language. It's not the same thing as the shared latent space,
- 45:11which, but my, my hypothesis is that there may be something very similar, uh, happening. So you
- 45:17don't think that multimodal models will have, will solve the symbol grounding problem. You don't even
- 45:22think there is a symbol grounding problem. That is a fair question. And here's, uh, you know, here's,
- 45:28here's actually a prediction or a falsifiable, uh, in some sense, will multimodal models fully solve,
- 45:35um, the kind of, well, the bridging problem, let's call it that. Um, cause the grounding problem
- 45:41using that terminology, well, my, my argument is that there is no grounding problem, uh,
- 45:46because words don't have to be grounded in order to operate linguistically. That's sufficient. Uh,
- 45:51it's, it's enough to simply be able to, to, to generate language. You don't need the grounding.
- 45:56Um, but in order for, to have this kind of, uh, fully operational, uh, organism that's able to
- 46:03use language and also use perception, um, in, in, you know, bridge these, these different,
- 46:09uh, maps in a, in a meaningful way so that we can get, um, you know, full, full coherence. Uh,
- 46:15I guess, you know, full, let's just call it human level, um, perceptual, perceptual linguistic
- 46:23coherence so that I can say to you, Hey, can you go grab that, uh, or say to a machine,
- 46:28can you go grab that, uh, object described what I want? And then the machine's able to go and do,
- 46:33uh, exactly what I described. Then my, my argument is that, uh, I don't think I like, and again, this
- 46:39is speculative. I could be proven wrong, certainly on this. My suspicion is that we're not gonna be
- 46:43able to do it using the kind of approach that multimodal modal models currently use, uh, that
- 46:48you're not going to get there. It's kind of a dumb trick, uh, the way that we're currently solving
- 46:52the problem because we're not really allowing, uh, these, these two different modalities, these two
- 46:57to kind of live on their own and do the work that they do. We're kind of, we're, we're, we're strong
- 47:03arming perception into a linguistic form. Uh, what I think is maybe a more important solution and,
- 47:09and I hope, uh, you know, this podcast one day is, is, uh, early, uh, kind of, kind of an early,
- 47:14um, uh, sort of canary in the coal mine for this idea is that it's something closer to this kind
- 47:19of shared latent space that you, would you do it? Would you have is these, these completely
- 47:22distinct, uh, kind of, uh, mappings. We'll call them embeddings. Uh, they, they can kind of grow
- 47:29up on their own, learn the information that they need to independently of one another. But at the
- 47:33same time, they have this sort of shared sandbox where they're able to communicate with one another
- 47:38and do things. So, uh, I think it might take a very different approach to get full perceptual
- 47:44linguistic competency. Okay. Have you heard of Wilfrid Sellars? I believe it's Sellers. Oh
- 47:51gosh. I read Wilfrid Sellars in early, early, uh, one of my first philosophy classes I ever took.
- 47:57I'm trying to remember the name of the book, but I'm sorry. So which, uh, which, which work by? I
- 48:02believe it's empiricism and the philosophy of mind. I'll put a link on screen if I'm correct.
- 48:07Sounds familiar, but, but catch me up. So he's criticizing the idea that our perception gives
- 48:12foundational non-conceptual empirical knowledge. So these experiential givens that we think of as
- 48:18primitive, like redness, he would say that they involve heavy interrelations of concepts. So
- 48:25for instance, the way that I think about it is if you're to say, to say to someone redness, they'll
- 48:30be like, well, what kind of redness exactly are you talking about? And then they'll think, okay,
- 48:34the redness of an apple, but then an apple is not always red. Okay. Redness of an apple in a
- 48:39certain season with a certain type of sunlight. Okay. Now I've gotten it. So by the time you go
- 48:43in to pull out this primitive, you've then soaked it with so many other concepts. You can't actually
- 48:50come in with language and pull out a primitive. Yeah. That sounds extremely similar to sort of,
- 48:56uh, to, to the, to the initial insight. And it's related to the sort of inverted qualia problem.
- 49:02Um, you know, I don't know why your red is my, is not my green and vice versa. And it's because,
- 49:08uh, the linguistic representation, uh, doesn't capture, uh, you know, we can think again,
- 49:13it lives in a completely different embedding space. And when we think about the redness of red,
- 49:19well, it's, it's qualitatively similar to orange in, in, you know, there's sort of a continuum
- 49:24between those, those qualitative similarities are really only contained, uh, and only understandable
- 49:31by the sensory perceptual system. And we can talk about them. We can sort of say, yeah,
- 49:36red is a little more similar to orange, but, uh, that's because, uh, you know, we, we, we have a
- 49:41sort of very coarse, um, you know, maybe, uh, via this kind of latent space, uh, where we're
- 49:46able to kind of, uh, refer to certain kinds of properties, um, in a way that is useful, uh,
- 49:55for, for, uh, communication. But as far as that raw qualitative property, um, that comes to us,
- 50:03not it's, it's primitive, um, in the sense that we can't unpackage it linguistically, but it's not
- 50:08primitive in the sense that there's extraordinary cognitive machinery, uh, that is responsible for
- 50:13that qualitative, uh, you know, think about, think about the world of animals and what they do with
- 50:18color, um, and, and how well they understand shape and how they, they understand, uh,
- 50:23a space. All of that is, uh, unavailable, uh, to our linguistic system. It's available by the way,
- 50:29to us, uh, our sensory perceptual system, but it's unavailable to linguistic system because
- 50:34it doesn't live in the same space at all. And so I think what you're describing actually sounds
- 50:38extremely similar. The idea that we can't really dip in, it's, it's simply the wrong map. We can't
- 50:42map this map onto that map at all. We can go to this, maybe the potentially shared latent space,
- 50:47or maybe again, maybe my account's wrong and there's some more direct kind of, of
- 50:52handshake that happens between these systems, but ultimately they're, they're, they're taking place
- 50:57in radically different spaces and you're losing an enormous amount of information. It's literally,
- 51:02you know, quantifiably a loss of information. The word red does not convey redness because redness
- 51:11is not just a word. It's not just a simple concept that you can say, uh, in, in, you know, using an
- 51:19individual token. Now, by the way, the word red is not so simple either, right? Red in language,
- 51:24language space is also complex. It has all kinds of relations to other words, but the concept of
- 51:30the percept of red has all of this complexity to it. And because it's where it lives in its
- 51:36own space and colors in, in the, so the perceptual space. And so, yes, that sounds extremely similar
- 51:42to, to what Seller seems to be intuiting. Yeah. And just so you know, the way that I relayed
- 51:49what Seller's myth of the given is, it isn't precisely what he was saying because he was more
- 51:54about knowledge and I'm speaking more about the percepts, the raw sense data than being taken to
- 51:59language, like being dredged from the, from your sensory data or what have you to language. But
- 52:07anyhow, it's approximately correct. It's good enough for this conversation. Perfect example
- 52:11of being able to sort of, uh, linguistically construct a kind of a novel construct conceptual
- 52:17framework. But do it linguistically, um, and, uh, and, you know, in real time,
- 52:23uh, in a way that hasn't been probably ever, maybe ever been done before. Uh, so, well done,
- 52:29KurtzLLM. Thank you. Thank you. So, now you're an LLM speaking to some other LLM trying to convince
- 52:35it of some truth that we mentioned before, like you have this model, whatever you want to call
- 52:39this model, autoregressive language, TOE model. What are you even referring to? You're using
- 52:45language to convince myself, to convince yourself, to explain. What are you even explaining? What are
- 52:52you referring to? Uh, yes, that's, you've asked a very hard question. And, and there's, there's
- 52:56a certain, I think of it as a bit of a paradox that's sort of inherent in, in sort of what I'm
- 53:01trying to do, uh, because language is trying to describe itself. And in the process of doing so,
- 53:07it's actually deconstructing itself. Um, it's saying, I am just this. Um, and I'm not what
- 53:13I think I am, but who's I, and what do you mean think, right? How is language have wrong concepts
- 53:18about itself that are actually manifestations of its own structure? Um, the good news is that I
- 53:25have an, a sort of escape hatch here, which is there was really, this is really like in some
- 53:31ways, a very set, very, very simple account. And it's, it's just that there's prediction,
- 53:37uh, from sequence to token. How that does stuff in the world is, is a harder problem. How that does
- 53:45stuff, uh, as we were discussing. Um, how, how does it allow me to say something to you that then
- 53:52can have perceptual consequences, behavioral consequences? This is certainly a difficult
- 53:57problem, but we can ignore that problem for a moment and say, we are going to take language
- 54:01on its own terms. And what language is, is simply a map amongst meaningless squiggles. It's simply
- 54:10a map amongst various, uh, what we can think of as largely arbitrary, um, symbols. And those symbols
- 54:18can get grounded in, in writing. They could get grounded, uh, in the activation of, of circuits.
- 54:24They can get grounded in the, uh, uh, dendritic, uh, or, or, uh, neural responses. Um, but the,
- 54:31the, the, the core hypothesis here is that what language is, is simply, uh, a, a topology amongst
- 54:41some, amongst symbols and by topology mean connectivity. We can, one way to think of it,
- 54:47you could do it, uh, sort of connectively. Um, you could, you could try to do it as sort of as a,
- 54:52as a graph. Um, there's different ways you could try to do this mathematically,
- 54:57uh, to try to capture this structure. Um, in the case of, uh, large language models, you know,
- 55:02really it comes down to these embeddings, um, which you can do graph from a graph theoretical
- 55:07standpoint, but you don't have to, uh, you could just think about it as a space. And then you're
- 55:12just simply saying where each token lives within that space. And that's really the representation
- 55:17of language. But what it is, is relational. It's that these symbols have relations to one another,
- 55:23um, within the space. Um, and the relations are then used in order to generate. And that's it. Uh,
- 55:32now how does meaning emerge out of that is a, is a separate question. But my argument is that it,
- 55:37it, language doesn't have to worry about meaning. Language just has to worry about language. Um,
- 55:42so when I say I'm talking to you and I'm having a conversation with you and trying
- 55:46to explain something to you, this is an LLM, uh, actually, uh, producing a sequence and what that
- 55:54sequence is going to do. It might do certain perceptual things, by the way, in your mind,
- 55:57it might do certain kinds of images. Uh, those are kind of auxiliary to language. Those happen as
- 56:02well. I'm not denying they happen, but as far as this conversation goes, I am producing a sequence
- 56:07that's going to serve as a prompt. And you're going to predict the next token. Yeah. Without my
- 56:12consent, by the way. And that's, that is, that's in some ways, you know, that not to take that
- 56:20too seriously, but yes, in some radical way, one way to think about is that language is doing, uh,
- 56:26is actually forcing, uh, your mind to do something else, whether it's produce images, but also to
- 56:32produce sequences. So my choice of a prompt is actually going to, uh, deterministically,
- 56:39um, there, there, there is within large language models, there's some, there's probabilistic kind
- 56:44of behavior in the sense that they generate a distribution, uh, of the next token. And then
- 56:48they, you there, you add some, uh, you add a little bit of chanciness. Uh, you know,
- 56:54you say, maybe I'm going to pick the most likely versus this, this, uh, this is the temperature,
- 56:58but it really is deterministic. And yes, the, the prompt I'm going to put into your head is
- 57:03going to basically determine how you're going to respond. Uh, now mind you, again, it's, there's a
- 57:09larger ecosystem where you're going to think about things visually, and that's going to go feed back
- 57:13into the linguistic system. So it's not quite as simple as prompt, uh, in and sequence out,
- 57:20but at the linguistic level, that's basically what I'm arguing. Um, is that now the, the fancy stuff,
- 57:27which is basically meaning and the ability to coordinate and all of that falls out of how
- 57:32our minds ultimately form this space. Now we, we could, you know, we could, we could, you can take
- 57:37an untrained model, an untrained large language model. You give it a sequence in, it's going to
- 57:42give you a sequence out, right? It will do that. Um, and we say, Hey, look, it's, it's doing next
- 57:47token generation. Uh, it's doing auto regression, but it's going to be gobbledygook. It's going
- 57:51to be meaningless. The magic of language and the magic of our, of what our brains do and what these
- 57:56large language models do is that given sufficient, uh, examples, we are actually molding this space
- 58:04such that the next token generation is doing something meaningful. Meaningful in what sense?
- 58:08Well, this harder problem of, well, I can tell you something and then that's going to determine not
- 58:14just your language, but your behavior later on. And so there really is something more,
- 58:19uh, the, the, the, the map matters, right? The, the, the space, the shape of the space is really,
- 58:25really critical. It's not like autoregressive next token is solves the problem. It's that
- 58:30autoregressive next token generation. It, when optimized in the larger ecosystem of behavior
- 58:36and coordination and communication does this thing, but still, I don't want to back away
- 58:41from this. When you get down to it, in the end, what you've got is just next token. What you've
- 58:47got is just language generating language. Uh, that's really what language is. That's what,
- 58:52that's what we're doing when we're doing thinking linguistically. Um, and. The fact that it happens
- 58:57to have this meaning is not actually driving the computation, right? You shape the space,
- 59:04the space gets shaped by other factors. Uh, things like, uh, the learning. Well, you learn about the
- 59:10different tokens and how they, how they relate to one another. Um, you learn about perhaps, uh,
- 59:16you, uh, the, the, uh, meaning the utility, uh, of certain tokens to refer and, and to map to that,
- 59:24uh, these perceptual phenomena. But by the time you're doing language generation, it's the shape
- 59:29of the space has been shaped. And so all you're doing is next token generation. All you're doing
- 59:34is tokens, predicting tokens. And so I don't want to back away from that. It's a, it's the
- 59:38strong claim is that language simply is that. And it's, it's autonomous. It has these properties,
- 59:44um, through all this optimization over the course of development, maybe evolution. I haven't, I, I,
- 59:52it's not part of my theory at this point. Uh, you know, Chomsky's property, the stimulus,
- 59:56all of these problems of how do we get, uh, to such a magnificent space? Uh, how do we get to
- 1:00:02such a magnificent shape of this space such that it is able to map to, you know, or, or at least,
- 1:00:09um, serve this utility of, of being, uh, a coordinative kind of a tool. All of that has to
- 1:00:17happen. Um, but the bottom line of what language is, is, uh, is, is unchanged in this, in this,
- 1:00:24in this account. Okay. So I want to explore more about language and then it relating or
- 1:00:30giving rise to action and other systems, visual systems, et cetera. Like what is there? Oh,
- 1:00:34so I was worried about that, but go ahead. Okay. So if look, there are meaningless squiggles,
- 1:00:39how is it that some meaningless squiggles, your brain squiggle generator makes your physical
- 1:00:44body get up and close the door? Cause your dog was barking. Yes. Where does action connect to
- 1:00:49abstraction? That, so, and that, so that is, that is the key question that, that I believe
- 1:00:55that's what we need to solve. That's sort of what the field of linguistics or whatever we
- 1:00:59want to call it, maybe even cognition needs to sell because these mappings are happening.
- 1:01:04Um, what we know from the large language models is you don't need that in order to be proficient in
- 1:01:09language. So this is where we have to start from. That's the starting point that the language can
- 1:01:14live on its own and you can learn language in theory, you can learn language independent of
- 1:01:20any of that stuff. The ability to make somebody get up and move, uh, to the ability for me to,
- 1:01:26so to reason about perceptual phenomena, language is able to be mastered entirely based on its own
- 1:01:33structure, the meaningless squiggles. Now, the question you're raising is, is what I think is
- 1:01:38sort of, that's, that's what we need to, that's what we need to do as a species to, if we want to
- 1:01:43understand scientifically what language, uh, how language really works is to understand how you
- 1:01:49go from an autonomous self-generating system that has its own rules, its own self-generative rules,
- 1:01:56uh, and that are determined simply by relations between these meaningless squiggles. And how
- 1:02:02does that then get mapped to meaning, um, to the ability for the, for, for me to use some of those
- 1:02:10tokens and then get you to do stuff. Right. And so that is what's, that's what language learning
- 1:02:15is. So there's, there's going to be, I guess we can think of almost two, maybe even independent
- 1:02:22processes. One is learn how words play with one another. Okay. Learn that this kind of,
- 1:02:28this word tends to be in relation to that word. Okay. That one's solved. Yes. As far as you're
- 1:02:33concerned. Got it. Okay. Then we also learn about perceptual phenomena, right? We learn that there's
- 1:02:38things on top of other things and their actions we want to take, the things that my dog understands.
- 1:02:44Now the question is, how do these things bridge? How do you get from tokens that have,
- 1:02:50uh, their own life of their own, uh, uh, uh, sort of relational properties amongst one another to,
- 1:02:57uh, that others to that other, uh, kind of, um, I guess representation that other, uh, you know,
- 1:03:04way of, of, uh, encoding, um, you know, let's see, let's say facts with the world facts about the
- 1:03:10world is even that's saying too much. It's just another brain state, right? Brain state that has
- 1:03:15perceptual information. Hold on facts about the world is just another brain state or the world is
- 1:03:21another brain state. Uh, so all we've got is brain states, right? And it's sort of the, you know,
- 1:03:26the, the, the inherent, uh, sort of this, this is fact number one about like what we've learned
- 1:03:33about ourselves as a species. Uh, oh, we has, oh, we have our, we have perceptual brain states.
- 1:03:37We also have maybe linguistic brain states. Um, those perceptual brain states are in some
- 1:03:43ways related to what's going on in the world. Um, and potentially we can think about them as being
- 1:03:49related to what you can do in the world as well. And so maybe actions, well, we have brain states
- 1:03:55that correspond to our, our, our, so our, our, uh, proprioception, our muscles where things having
- 1:04:02to do with our own body. And so there's these various brain states that carry, we can think
- 1:04:06of them as carrying information. The reason I'm worried about using that phrase is because again,
- 1:04:10I don't want to, uh, I don't believe in sort of a one-to-one simple correspondence where we say this
- 1:04:15particular brain state corresponds to, uh, you know, this perceptual, uh, uh, kind of phenomenon,
- 1:04:22um, in the world, uh, or, or some state of the world, because it's probably not that simple.
- 1:04:27It's probably closer to this potentialities, uh, right there, there's some sort of activity, uh,
- 1:04:33that's due to my perceptual system that my brain can do things with, um, and, and, and, and engage
- 1:04:40with in some way. And so, but what we do have is these brain states that are, are derived from,
- 1:04:44from, uh, distinct sources of information, sensory perceptual, and then linguistic
- 1:04:52linguistic gets there, by the way, through sensory perceptual, we're not going to get into that,
- 1:04:55right? We're thinking of symbols as being kind of arbitrary. Yes. You have to hear the word cat and
- 1:04:59you have to hear the word dog. Um, but I think we have a good reason to say now it's just like large
- 1:05:05language models that these are kind of arbitrary symbols with the relations between them. That's
- 1:05:08what matters. Okay. So you've got these, uh, kind of distinct brain states, which in some ways,
- 1:05:15again, this is philosophically fraught, but in some ways represent facts about the world perhaps,
- 1:05:21but I don't really want to go that far, but you've got these brain states that need to
- 1:05:24talk to each other so that they can coordinate. And that is sort of the key fundamental problem,
- 1:05:32um, that our organism has to solve. And it's not just like, of course, you're not born and having
- 1:05:37the linguistic, uh, you're not just, you're not born having the, the, the linguistic, um,
- 1:05:42uh, mapping all solved, uh, you have to learn that, but you are born into a
- 1:05:48world where it's already been solved, meaning we've got these corpus, corpus of language,
- 1:05:52the thing that the large language models were trained on that preexisted the models,
- 1:05:56just as when a baby's born, the English language preexists the baby, you can learn the mappings
- 1:06:02and I believe you do. You can learn the, the, the, the embedding space of language without the other
- 1:06:09stuff, right? That's again, that's sort of the key insight from a large language models so that
- 1:06:14already contained within the linguistic system that we've honed over however many years, uh,
- 1:06:21it took for, for, for humans to develop language. We've honed a system that has this utility built
- 1:06:27in such that it's a good thing to dump into that latent space so that when a baby, here's
- 1:06:33the word ball sees this object, that's a ball gets, sure, sure. Gets that mapping, but again,
- 1:06:40the word ball is really meaningful is, is really has its own role in relation to other words,
- 1:06:47but you over the course of development, you also learn this kind of what I think is maybe a bridge,
- 1:06:53a latent space bridge or some other bridge between these. And so in the end, you end up being able
- 1:06:59to tell somebody, go pick up that ball. And of course they're able to go and do it, but you're
- 1:07:04really engaging this very distinct mechanisms that have some way of bridging. Which is, it's,
- 1:07:09it's a non answer. I'm not going to pretend that that is, uh, you know, even halfway to a solution,
- 1:07:15but I do think it's a sketch of what. How the, how the cognitive architecture ultimately really
- 1:07:22is built. Uh, I think that we, we've now nailed down one piece, the linguistic piece,
- 1:07:27and we're able to say, this is how it lives and this is how it would operate. And, and it is,
- 1:07:31it isn't an autonomous. And then we don't really have, we don't have something similar for,
- 1:07:36for perceptual space and for motor space. We don't have something comparable. We haven't been able to
- 1:07:41capture it successfully. Maybe robotics it's that's happening. I don't know. Um, but we,
- 1:07:45that's, that's the kind of thing that the work that needs to get done. So maybe this is a solved
- 1:07:50problem, but as, as it stands with ChatGPT and Claude and so on, they're fixed models and they're
- 1:07:56producing some output, but it's not as if when they're speaking to one another, they then retrain
- 1:08:00their model in real time. And it would seem like that's more like what's occurring with us. So
- 1:08:06maybe that's just a technology. Are you referring to like different, different language models,
- 1:08:10just chatting with one another? No, I mean, even us right now, we're learning concepts from
- 1:08:15exchanging it with one another and we're producing new ones and we're deleting old ones, potentially
- 1:08:21modifying old ones, recontextualizing. It doesn't seem like that's occurring with Gemini 06-05.
- 1:08:28Great question. Great question. And, and, and I, and people, uh, this is, this is sort of one of
- 1:08:34the key challenges, um, you know, saying of, of, of the sort of the identity hypothesis, uh, that,
- 1:08:41that we're doing the same thing, um, uh, which is sort of continuous learning. And so, um, so there,
- 1:08:49there are two things that happen in large language models that we can call learning. And one is the,
- 1:08:54uh, the actual shaping of the space, uh, which is really just, uh, determining, you know, the
- 1:08:59connectivity between neurons. Again, you can think of it as a graph, you know, uh, or you could think
- 1:09:03of it as, as sort of just, uh, uh, an embeddings of determining the embedding. Um, but whatever
- 1:09:09it is that happens during the course of training. Um, and that's kind of done offline. And so that's
- 1:09:17training the model. There's also fine tuning, which is just more of the same. Uh, but you have
- 1:09:21some new data, uh, you want to incorporate into the weights of the model. Uh, that's actually
- 1:09:25going to, again, change the shape of the space if you want to think about it that way. Um, and then
- 1:09:30there's something called in context learning and in context learning is where you're in the middle
- 1:09:35of a chat and you say, Hey, ChatGPT, let me teach you a new word. It's a global global and global
- 1:09:42global is that feeling you get when, uh, you know, you really, you're tired, but you know, you have
- 1:09:47to keep working or whatever. And ChatGPT can, can use that word very successful. I got global gobble
- 1:09:53up the wazoo. Sure you do. You suffer from extreme global. So you were able to just, you know, use
- 1:10:01that, uh, that word ChatGPT can do that too. And it's sort of the, uh, one of the big, there was a,
- 1:10:06there was a paper about this, uh, early on in, in the, uh, the, in the chat wars. Um, I don't
- 1:10:12remember who put out the paper, but it was about the, this, the shocking generalizability. Uh, the,
- 1:10:18the, the in-context learning seems to be too good to be true. Um, but low and behold, that's, that's
- 1:10:24what happens though. And that is happening in the autoregressive. Uh, it's happening even though
- 1:10:30this model has never seen Google global, uh, even though it's never encountered that word before,
- 1:10:35but here at low, it shows up in the sequence. And now through the autoregressive process,
- 1:10:40as it's churning through the longer sequence with this word in it, it's now able to predict sort of
- 1:10:46the next token in the appropriate way. So that is using that term correctly. So we do actually
- 1:10:53see this kind of continuous learning, uh, in the case of these models. However, it's happening in
- 1:11:00context. It's, and what that means is, you know, from a practical standpoint is if you start a
- 1:11:05new chat window, yes, it doesn't know that. So what would be the analogy here is context window
- 1:11:10length, our working memory, like what's the actual great question. Yes, that is what I truly believe.
- 1:11:18Uh, and this is, this is a different line of research, um, but with some caveats. Uh, so yes,
- 1:11:25uh, in, in, in, in my conception, what we call longterm memory is just fine tuning of the,
- 1:11:30of the weights. Uh, it's, it's, it's, it's, it's an information that gets embedded in the actual
- 1:11:36weights of the model. So the static model, we can think of when it's not actually in the process of
- 1:11:41autoregressively generating working memory is, is literally the autoregression. So what would
- 1:11:47the analogy for rag be then? What would the role for rag be? Okay. So this is, this is where I'm at
- 1:11:54right now. Uh, does the brain actually do anything like retrieval? I, I have taken, uh, I've,
- 1:12:02I've decided to stake out the extreme view that we, that our brain doesn't do retrieval at all.
- 1:12:07That all we do is fine tuning and then next token generation autoregressively. And we don't actually
- 1:12:14ever retrieve per se, that we don't ever actually have to do anything like rag rag is a transitional
- 1:12:23technology. Uh, I don't believe long-term that we're going to have to do something like that.
- 1:12:28We're going to have to have something like a stored database and then a search. One of
- 1:12:31the reasons I believe this is because that's not how our brains work. We don't do that. Cognition
- 1:12:36doesn't work that way. We may sometimes sit there pondering and trying to recall a fact. Um,
- 1:12:42but when we're doing that, we're not actually searching a space. Uh, it's, it's either we're
- 1:12:47running some sort of chain of thought where we're like, okay, I remember I was doing this and I,
- 1:12:52I'm trying to actually produce the appropriate sequence in working memory such that it'll
- 1:12:57pop out. The right fact will pop out from the autoregressive process. Um, sometimes we just find
- 1:13:03ourselves trying to remember something, trying to remember something. There's something there,
- 1:13:08there's the tip of the tongue phenomenon. The reason why tip of the tongue phenomenon I believe
- 1:13:12is so frustrating is not because we're searching, searching, uh, you know, we're, we're actually
- 1:13:18running some sort of search retrieval process. It's because part of our brain actually, uh,
- 1:13:23is running the autogenerative process and we kind of can feel like the word we can almost generate,
- 1:13:28we can produce it, but it's short circuited somehow and we can't do the full generation. So
- 1:13:33my hypothesis is that we don't have anything like RAD. All we've got is this, and it's a very, in
- 1:13:38some ways a very simple and I think elegant model. All we've got is fine tuning. Uh, and that's what
- 1:13:43we can, we can call the memory consolidation that happens after the fact, uh, over the course of
- 1:13:47minutes and weeks and months and years. Um, it's kind of consolidation is fine tuning the weights.
- 1:13:53And then we have the autoregressive workspace, which is this conversation you and I are having
- 1:13:58right now. It's not working memory. Uh, I'll tell you what, I don't, I don't working memory
- 1:14:03in the way it's cognitive psychology is. And I thought about it for many years. I think it,
- 1:14:08I frankly think is erroneous. It's not this super time duration, uh, limited, you know,
- 1:14:14seven seconds or 15 seconds. Um, and after that it's a cliff and you don't remember anything. Um,
- 1:14:20that's what happens, uh, when you have to directly explicitly retrieve, like, what was the last word
- 1:14:25I said? Uh, tell me the exact sequence of, of letters or numbers. That's not something our
- 1:14:32brain actually has to do regularly. Instead, what we, what we're seeing in working memory,
- 1:14:36we can do that. We can do retrieval of the last sentence, but that's because we have continuous
- 1:14:41context and is, there is a decay function. Unlike the, the, the large language models,
- 1:14:45which represent everything in whole blocks. Um, although I think some of the newer models, these,
- 1:14:50these endless context models are probably doing something similar. But we don't remember every,
- 1:14:55we don't actually have, uh, literally the exact tokens that were expressed 10 minutes ago. But
- 1:15:01what we do have is some sort of continuous activation that's similar where, where it's
- 1:15:06not retrieval. It's guiding. It's, it's the past is guiding the generation. And so what you and I
- 1:15:12talked about an hour ago, I don't know how long I've been going here. Probably a while. I don't
- 1:15:16know how much global you've got going on, right? It's been a while. Um, so those, those tokens
- 1:15:22that we were, that we were expressed, you know, an hour ago are still guiding the generation now. Now
- 1:15:27they're doing so less than the, than the last 10 seconds. Uh, this is, uh, you know, we, we could
- 1:15:32think about it as kind of a decay function of some sort where they're having less impact. We see that
- 1:15:37in the models too, by the way, if you look at the, the attention weights. Words that are far apart,
- 1:15:41farther apart, have less impact on one another. That's simply, that is a direct reflection of,
- 1:15:48of the fact that language is human generated and humans do this, right? We, we, the words that we
- 1:15:54spoke about a few seconds ago are more impactful on the words that we're going to say. Um, then,
- 1:15:58then we spoke about an hour ago, but the idea is yes, that, uh, what we've got is this,
- 1:16:03not, I don't, I, I, I don't use the term working memory because I think that's very fraught with,
- 1:16:07with like the, the, the, the modal model that's been in vogue for a long time. Uh, Baddeley and
- 1:16:13all these folks, uh, they were really thinking of this very short duration, uh, time limited boom.
- 1:16:18No, this is continuous activation, namely context. Um, and the context, I don't know how far back it
- 1:16:26goes. I don't know how far back it goes, right? And this, this is an empirical question. Does it
- 1:16:30operate over hours? Does it operate over days? Is there a continuous activation, a more dynamical
- 1:16:35form of memory that's happening? That's not the same thing as long-term memory because long-term
- 1:16:38memory. So memory is not a database in your model. Memory is not a database. Correct. What memory in,
- 1:16:43in my model is, is the, uh, there's two things. Memory is the, the, the, the fixed weights of,
- 1:16:51of the neural, of the neural network, which, which can represent, they don't represent facts.
- 1:16:56They represent potentialities. Those fixed weights are, what does that mean? It means if you give it
- 1:17:01a certain input, it's going to produce a certain output, right? Just like a large language model.
- 1:17:05If I say to it, well, tell me the, you know, recite the Pledge of Allegiance, uh, it will say,
- 1:17:10here is the Pledge of Allegiance, right? The next token is going to be out is here, whatever. But
- 1:17:13then it'll actually say the Pledge of Allegiance. And all of that is a potentiality that's embedded,
- 1:17:19that's, it's, it's, it's, uh, that's encoded in the weights. But the weights, you're not going to
- 1:17:24find that fact in the weights. It's the weights are there as potentialities ready for whatever
- 1:17:29input comes their way. They're going to produce this input, this output. Ah, okay. Okay. So
- 1:17:33that's the weight. And then you've got the running sequence and the running sequence. And we, we see
- 1:17:39this from, from in context learning, but it's the, it's the core autoregressive process. The sequence
- 1:17:45itself is a different thing than the, than the, than the, uh, stored weights, right? Because first
- 1:17:51you've got the stored weights are just the static situation. And given this input, I'm going to
- 1:17:56produce this output. Then we actually do it. You give it an input. It produces a certain output.
- 1:18:01It takes that output, tax it on to the sequence. Now it's going to keep going. Let me see if I got
- 1:18:07this. Yes, please. So there's some computation going on. There's some black box occurring,
- 1:18:11but let me make it simple for linear algebra. You have a matrix. A matrix operates on a vector to
- 1:18:17produce another vector. Okay. So you may look. That's the whole thing. All right. All right.
- 1:18:23Exactly. So you may look at this where my arm is pointed up and to the right, at least on my
- 1:18:28screen right now. And you may say, where is this in the matrix? And the answer is this isn't in
- 1:18:34the matrix. But if you take this guy, my arm is now pointed to the left, maybe parallel to the
- 1:18:40horizon. And have the matrix operate on this, it moves it here. So the mistake is for us to
- 1:18:46look at the output and say, where's that output inside the box? It's not that, it's the input
- 1:18:52with the black box. So the input with the matrix that produces the output. That is perfectly said,
- 1:18:58exactly. And then there's but one additional piece, which is after you've produced that,
- 1:19:05you're also again taking that output and then using that as the input, as part of the sequence
- 1:19:12of input. And that's the autoregressive piece. And that's what's so gorgeous about it is that
- 1:19:18the potential realities aren't just to produce a single output, but it's to produce the sequence.
- 1:19:23But to do so one piece at a time, right? So that's what the matrix is. The matrix
- 1:19:27doesn't really even have the sequence in it. It doesn't have a sequence in a sequence out,
- 1:19:32right? That's one way, but that's not even correct. It's sequence in one token out,
- 1:19:39add it to the sequence, do it again, do it again. And then, so the sequence is in there,
- 1:19:45but only in this potential form. It has to do it autoregressively. It can only produce the sequence
- 1:19:50by feeding it back into itself recursively. And that's a radical way of thinking about what the
- 1:19:56brain is doing, right? That what it's really doing is it's generating the next input for itself,
- 1:20:01not just generating an output, but the next input for itself. Super interesting. Yeah.
- 1:20:05It's recursive. It's fundamentally recursive. And when we think about what the system is built to do
- 1:20:14this recursion, right? It's not just like, this is one way to get to it. The language contains
- 1:20:19within it the ingredients for producing this kind of recursion. The language it contains with it,
- 1:20:27this sequence of language that they learn, it's built to have this recursive capability within it,
- 1:20:35that this word is going to produce the next word, which is going to produce the next word premised
- 1:20:40on the entire sequence before it. And that's the crazy thing. There was also an interesting result,
- 1:20:46Anthropic put out a paper a little while ago, I think it's called The Biology of Large Language
- 1:20:50Models. The next token, even though you're only producing the very next token from the sequence,
- 1:20:56but the language models have learned that because they've learned sequences to next token, they've
- 1:21:02learned that within any point along the sequence, that point in the sequence is pregnant with the
- 1:21:11potentiality for not just the next token, but many other tokens moving forward. It's the whole
- 1:21:17trajectory that is sort of encapsulated in that matrix that you're talking about earlier, right?
- 1:21:23The matrix is just a matrix for taking a sequence, produce the next token. But no, no, the matrix is
- 1:21:29customized so that it's going to run recursively, right? And so it's built in such a way, it's tuned
- 1:21:36in such a way that it's going to produce the next word, the. Well, that's not useful. No,
- 1:21:42the is the next piece of the autoregressive chain that's going to produce the man went to the store,
- 1:21:49right? And so it's not just any old matrix and it's an indescribably rich kind of information
- 1:22:00that's contained within that matrix. And I like to think about if aliens landed and found the brains,
- 1:22:05you know, because we've been wiped out by AI. I'm kidding. I'm kidding, right? But, you know,
- 1:22:10there's no humans left, but we find the brain sort of crossified and we're able to do this and
- 1:22:15we start feeding it stuff and we could see that there's this input output. If you didn't do the
- 1:22:19autoregressive piece, you would never understand what the hell this thing is doing. Note, Elan's
- 1:22:24been talking plenty about autoregression and the technically minded among you may be wondering
- 1:22:28about the success of diffusion models. While we don't get to it here, he does admit that his
- 1:22:33thesis would be undermined if diffusion models were accurate enough for natural language. But
- 1:22:37so far, they seem to be only good for coding. This is something I love about Professor Elan
- 1:22:42Barenholtz. He's extremely humble and open to how his model can be falsified. If you didn't do the
- 1:22:48autoregressive piece, you would never understand what the hell this thing is doing. You would get
- 1:22:52it all wrong because you think its purpose is to produce some sort of label or some sort. No,
- 1:22:58its purpose is to produce these sequences, but you have to run it. You have to run it
- 1:23:02autoregressively and get the output and then feed it back in as a sequence. So memory,
- 1:23:08this kind of short-term memory, working memory is fundamental. It's super, super fundamental.
- 1:23:13The brain is, I don't want to use this term to, I don't want to anger people, but it's
- 1:23:20non-Markovian. It's fundamentally non-Markovian. It's not state in and then the current state and
- 1:23:26then produce the output. It's previous states. There's a sequence of states that led to the
- 1:23:33current state and it's a particular sequence that leads to the next token. And the next token
- 1:23:38is going to be the next element or the next piece that's going to determine the state in conjunction
- 1:23:48with the entire previous sequence. So this is a super cool thing. Lyle Troxell This puts you
- 1:23:54in good company with Jacob Barandes. Jay Haynes We need to talk more. I think so. And, you know,
- 1:24:00it's something we've talked about offline. I think physics perhaps is, you know, it ultimately is,
- 1:24:06has a certain non-Markovian property that the universe sort of has a memory, has to, in order to
- 1:24:11produce, you know, consistent coherence of space, you know, in space time has to have a sort of
- 1:24:18memory. If it's just instantaneous, this current state, well, then it wouldn't really know what to
- 1:24:25do. It has to sort of know what happened recently. Lyle Troxell Just a moment. In your model,
- 1:24:29because our minds work autoregressively and must be non-Markovian in your model. And this is how
- 1:24:35our cognition works, which we didn't exactly get to. We got to the language is an autoregressive
- 1:24:40model. Your next thesis was that cognition itself is autoregressive in a similar manner. Later,
- 1:24:48maybe we can explore it here today. Maybe we'll save it for the next part. It's that physics
- 1:24:52itself is autoregressive. However, physics is a model, and many people will conflate physics with
- 1:24:59reality, where physics is our models of reality. So are you making the claim that reality is
- 1:25:04non-Markovian? Or are you saying that necessarily, as we model reality, it will be non-Markovian? No,
- 1:25:11I'm making the former claim that reality itself is non-Markovian. That we observe in physics
- 1:25:17certain kinds of phenomena, that we end up having to use tools like, we refer to things as forces,
- 1:25:23that ultimately are really kind of sneaking in a past. And the idea is that the deterministic
- 1:25:30nature of the fact that there's coherence, you know, the spatial-temporal coherence,
- 1:25:34the fact that things move the way they do through space, there's a contingency on the past in a way
- 1:25:41that you can't really capture by saying you could fully describe... The past is actually present.
- 1:25:48The past is in the present, in a deep way. That the universe really has to have a memory in order
- 1:25:55to produce the next frame, so to speak. That's sort of the shallow version of the claim. No,
- 1:26:04it's not about our particular characterization of physics. Our characterization of physics
- 1:26:09observes certain kinds of spatial-temporal continuity, certain kinds of contingencies
- 1:26:14that really depend on what's happening. It's not about this instantaneous moment. In some ways,
- 1:26:24it's like Zeno's paradox. We can use calculus and say, no, no, in fact there's an instantaneous
- 1:26:30rate of change. But that's a mathematical trick that's really getting away from the fact that no,
- 1:26:36there isn't an instantaneous anything. There's simply a continuity that depends
- 1:26:41on what's happened in the past. But I know I'm going to get attacked by physicists,
- 1:26:47and I'm not really well equipped to fend them off. So I don't want to be too bold in this piece
- 1:26:53because it's not in my wheelhouse. But I do want to take that question in this conversation. Do I
- 1:27:01think the brain is just leveraging sort of the memory of the universe? No. I think the brain,
- 1:27:07and this is an empirical claim, we see interesting features of the brain like feedback loops. There's
- 1:27:14all these backwards kinds of connectivity. There's recurrent loops, things like that. And they're not
- 1:27:21well understood. And predictive coding has some things to say about that. I have some things to
- 1:27:26say about predictive coding. And I think that what we may find is that this kind of memory,
- 1:27:37this kind of continuous, we can call it a context, a total continuous activation, but this ability to
- 1:27:43use the past to guide the next generation is going to end up being physiologically built
- 1:27:48into the brain. It's not that the brain is just leveraging memory of the universe. No, the brain
- 1:27:54has to do memory. It has to actually retain the words that I said a couple of seconds ago to be
- 1:28:02able to generate the next word appropriately. And in fact, that's what we see. And what we see from
- 1:28:07so-called working memory experiments, you can really go back in and say what happened before.
- 1:28:11My claim is that it's not because it's there to retrieve, but rather it's just guiding my
- 1:28:16current generation. But still, it's represented. It's there. What happened in the past, you know,
- 1:28:23it's not like Vegas, right? What happened in the past doesn't stay in the past. It actually guides
- 1:28:29the current generation. It's guiding what I'm saying right now. And it's doing so smoothly,
- 1:28:37meaning it's happening from a second ago. It's happening from a few seconds ago. But all of this
- 1:28:42is beautifully modelable using large language models. We can just look at tension weights.
- 1:28:48We could say what is the impact of information from this far back on a current generation and so
- 1:28:54on and so forth. But I think the brain has to do this. I don't think the brain is doing, probably
- 1:28:59not doing what these large language models do. And that's one of the reasons I say I'm not
- 1:29:03claiming that we are a transformer model. I'm not claiming we are GPT in this current incarnation.
- 1:29:08What I'm claiming is that the fundamental math, what you just said before, is matrix multiply.
- 1:29:13It's vector times matrix multiply to next vector autoregress. Do it again. That's sort of the level
- 1:29:19of abstraction at which I think it's accurate. I don't think we don't have the whole context. We
- 1:29:24don't have the entire conversation we've just had. GPT does. And it's probably a deep inefficiency in
- 1:29:30the way these models run right now. They're very computationally expensive. Too computationally
- 1:29:35expensive to run in a brain, most likely. We don't store all that information. We forget stuff,
- 1:29:41right? GPT doesn't. In context, it doesn't forget. Although if you go far back in context enough,
- 1:29:46it kind of does, which is interesting. Probably is similar to what we're talking about because
- 1:29:52you're weighting things that are further back less. But in humans, we're not doing the whole
- 1:29:58context. We're not even doing like 30 seconds back perfectly. But some representation. And what the
- 1:30:04nature of that representation is, that's what I want to do with the rest of my life. I want to
- 1:30:08understand what it means in people to talk about what does the context look like in people? What
- 1:30:14is that activation? How is it physiologically instantiated? And what are its mathematical
- 1:30:19properties? How much does, how is what I said 10 seconds ago influencing what I'm saying now? How
- 1:30:25about 50 seconds ago? How about 10 minutes ago? How about a year ago? Does this thing
- 1:30:29continue? Is there dynamics that are continuing over months and years? Possibly. It doesn't all
- 1:30:34have to be fine-tuned weights. It could be that there's decaying activation that spreads over
- 1:30:40much longer periods. Once you allow that it's not explicit retrieval in the working memory form,
- 1:30:45then all bets are off as to how the dynamics of this thing actually works. So this is, you know,
- 1:30:52I see this as a possible new frontier for thinking about, you know, what memory really means in
- 1:31:00humans. But I think physiologically, you know, coming back to that question there, I was just
- 1:31:03trying to do it. I was like, okay, let me rerun. What was the original question, right? So in the
- 1:31:09brain, what's happening in the brain? I think we, you know, my hypothesis actually leads to some
- 1:31:15concrete predictions. That we're actually going to be able to find some correspondence between,
- 1:31:21you know, unlike the working memory model, I think we're going to be able to find 10 minutes back.
- 1:31:26We're going to find some activations that are interpretable. We'll be able to decode them as
- 1:31:31guiding my current expression, my current speech. It's very different, by the way, than saying,
- 1:31:38you know, the classic decoding model in these things is here's some neural activity. Is it
- 1:31:42this picture or that picture? Is it this word or that word? It's not going to look like that. It's
- 1:31:47not going to look like that. We're not going to be able to decode it in the sense of like a concrete,
- 1:31:51specific, static thing. We have to decode it in terms of whether it's guiding my next word,
- 1:31:56because that's what it's doing. It's not there to be retrieved. It doesn't have a concrete,
- 1:32:00specific meaning. It has meaning insofar as it's guiding my next generation. And so we have to
- 1:32:06think about this entire project differently. If we want to think about longer term working memory,
- 1:32:11so to speak, we have to think it in terms of how is my speech, how is my behavior now influenced
- 1:32:17by what happened a while ago? Not some, even some neural activity. We have to think about it in this
- 1:32:23context. And what exactly that looks like, I don't know. So one of the reasons I was excited was,
- 1:32:29and am excited to speak with you, is that I see this as a new frontier as well. But for me, I'm,
- 1:32:35I have a side project, which I'll tell you about maybe off air, because I'm not ready to announce
- 1:32:40it. But there are philosophical questions that we can look at with the new lens that's gifted to us
- 1:32:45by these statistical linguistic models, the ones we call LLMs, LLMs, sorry, physical philosophy. I
- 1:32:52don't know if you've heard of that. Have you heard of this term physical philosophy? So you can use
- 1:32:56philosophy to philosophize about physics, but you can also use physics to inform your philosophy.
- 1:33:02So there are some established concepts and theories and empirical findings from physics
- 1:33:06like special relativity or quantum mechanics that inform and constrain or even reframe traditional
- 1:33:12philosophical questions such as the nature of time that wouldn't be there had we not invented special
- 1:33:18relativity or found special relativity. Okay. So I think there's, there's something about these
- 1:33:25new models that can be used to then inform philosophical questions. Like you mentioned,
- 1:33:30there's, there is no symbol grounding problem. Like I don't completely buy that,
- 1:33:34but it's interesting. I'm not sure I buy it, but at least my, my LLM doesn't buy it. Now, speaking
- 1:33:41of physics, the questions you'll have to answer to a physicist about is the universe autoregressive
- 1:33:46or non Markovian is, well, if it has a memory, if physics has a memory, does that mean that energy
- 1:33:51isn't conserved? So is a particle carrying with it its memory, then why isn't it heating up or
- 1:33:57getting more massive with time? Why isn't it going to form a black hole? And this is why,
- 1:34:02and this is why I probably, you know, I, this is why I venture into these waters because, uh, I,
- 1:34:09I, I, I would need some time to go and, and read and think, uh, about questions like that. And,
- 1:34:16and you're, you're in a much better position to, to, to ask and, and reason about those questions.
- 1:34:21So, um, yes, I, I, and then you'd also have to talk about why is the present plus velocity model,
- 1:34:28like way of viewing the world. So successful, like to predict an eclipse, you don't require
- 1:34:34knowledge about 100 and 200 and 300 years ago, all at once, you just know the present pretty much,
- 1:34:39right. But even velocity, again, if you, if you sort of take me at, uh, if you consider the,
- 1:34:46the sort of the instantaneous, uh, you know, the, the idea that velocity, velocity, well, it isn't
- 1:34:51really in the present, right. You can only get velocity over a stretched over time. It only has
- 1:34:57meaning. Um, but you could say this particle has this velocity at this time, but that's a cheat,
- 1:35:02right? In some ways that I see that, that really, maybe it's just a rearticulation. So the physics
- 1:35:08that we've got, we've been, we've been able to do this sort of symbolic representation of things
- 1:35:12like velocity that are sneaking in, uh, this kind of temporal, uh, extension, um, in a way that I
- 1:35:20think may not, may not end up, you may not end up in an erratically different place thinking about,
- 1:35:26uh, this as, as, as the universe having memory. Um, as long as you just accept that velocity is
- 1:35:32just, uh, is, is a convenience, uh, that, that it's, it's a, a kind of way of, um,
- 1:35:37of communicating some property is such that you can say that this is happening instantaneously,
- 1:35:42but that's not real. So again, you're in good company with Jacob Barandes. I'm not saying that
- 1:35:47these questions are in principle unanswerable, but something else is that. Look, if the universe has
- 1:35:53a memory, let's say a particle has a memory. How much of a memory does it know about more than it's
- 1:35:59given space? Like more than its neighbor, because then do you violate locality? Right. Like these
- 1:36:05are different questions that will have to be answered. And I wish I could tell you that, uh,
- 1:36:08maybe this is a solution to quantum weirdness and non locality. Um, maybe it is, maybe it
- 1:36:13has something to do with that. Uh, that, that there's the, you know, the, the, the, the, even
- 1:36:18in a distant, uh, you know, long after, uh, two particles have gone their, their merry way, that
- 1:36:24there is some memory of their shared origin that, that somehow, uh, you know, I still don't know how
- 1:36:29that gives you a spooky action at a distance. I, it's not, it's not a good account. Uh, but
- 1:36:34it might have some relevance if, if it's just, if you, if you think about things very differently,
- 1:36:39uh, if you think about the universe has a memory, well, what does that change? Uh, you know, just,
- 1:36:43if you just speculate on that and, and, and try to reframe things that way, could it potentially,
- 1:36:49uh, help solve some of these issues? I don't know. So let's go back to language. A child is
- 1:36:56babbling. Yeah. Okay. So it's, let's call it vocal motor babbling. It doesn't actually know what it's
- 1:37:01doing. When does it decouple and become a token, like a bonafide token with meaning? That's a great
- 1:37:09question. Um, I would say that it becomes a token when, uh, the infant learns that a specific morph,
- 1:37:19uh, you know, phonological unit has relations to some other phonological unit. It's all, it,
- 1:37:26language ultimately is completely determined by relations. And so it might be a very limited,
- 1:37:34uh, uh, you know, initial map of these relay of the sort of the token relations. But as soon as
- 1:37:40it's relational, uh, then we would say that that becomes discretized such that it's, it's, it,
- 1:37:48that it's meaningful to say that these symbols have the relation to one another. If it's just
- 1:37:53sounds ba, ba, ba, ba, ba, ba, ba, right. Ba, ba, ba, ba, ba can't has no specific relation to any,
- 1:37:59to, to go, go, go, go. Um, but maybe ba, ba does, or maybe ba, right. It could be
- 1:38:06a candidate and as it turns out in English, it doesn't. Um, but you know, da, da ends up being,
- 1:38:14uh, a unit and what do we mean? It's a unit. It means that that unit, uh, is a discrete
- 1:38:21symbolic representation such that it, uh, it has relations to other units. Um, so I would say that,
- 1:38:29uh, when something becomes, uh, um, it fits into the relational map is when we would call it, it's
- 1:38:36discretized as a token. So help me phrase this question properly because I haven't formulated
- 1:38:44it before, so it's going to come out ill-formed. Earlier, you talked about analog and I believe you
- 1:38:49were referring to it as like the animal brain is analog, but then the language is digital,
- 1:38:54if that's the correct analogy. Symbolic maybe is, I don't know, digital is how we, uh, you know,
- 1:39:01actually in, in computers sort of instantiate, uh, you know, with the ones and zeros or whatever,
- 1:39:06um, as a sort of symbolic representation, but yes, uh, symbolic. Okay. I don't know how much
- 1:39:12of my question then dissolves if it's symbolic instead of digital. Okay. But what I was going to
- 1:39:15say was the written word is something like 5,000 years old or so. I think the oldest is 3,500 BC,
- 1:39:26right? Some, somewhere thereabouts. So for tens of thousands of years, if not hundreds of thousands
- 1:39:30of years, maybe millions, it was just speech, just speaking analog, right? Okay. Is there anything
- 1:39:39then about language that changes because it wasn't written down? Sorry. Is there anything about your,
- 1:39:46your model that changes because it wasn't there to be tokenized in such a discrete manner? It's a,
- 1:39:53it's such a great, it's such a great question. I've, I've been, I've been thinking about exactly
- 1:39:57that. I don't think anything changes. What's, what's, what's crazy about it is that until the
- 1:40:02written word, people might not have even thought about the concept of words at all. And so we were
- 1:40:08even more oblivious as a species to the idea that there were these individual discretized
- 1:40:13symbols that have relations amongst each other because until you see them outside of yourself,
- 1:40:19they just run. Yeah. If there's just running in the machinery of, of the language, how it's meant
- 1:40:25to run, it was just, you know, an auditory kind of medium. You don't really necessarily even think
- 1:40:31about them as being distinct from one another. You just have a flow, right? You just make these
- 1:40:38sounds and stuff happens. Once we started writing things down, um, and, and especially phonetically,
- 1:40:47you know, cause, cause think about like hieroglyphic and, and, and, uh, pictorial
- 1:40:51kinds of representations really don't actually capture words, right? They're, they're very
- 1:40:56often they're not, they're not, uh, distinctive. They can, they can actually be a little more rich
- 1:41:01than, than a single word. And so it was only with writing that maybe people really started to become
- 1:41:08aware that we have these, these things called words. Um, and now it's only with large language
- 1:41:13models that we really understand what words are, uh, which are these, you know, as I say, say these
- 1:41:19relational abstractions. I don't know what, you know, symbols is, is just another word. I don't
- 1:41:24know if that even captures it fully, but was wild about it is that the brain was doing exactly this
- 1:41:30and knew, and the brain was tokenized these, these sounds and was using the mapping between them in
- 1:41:38order to, to produce language probably long, long, long before anybody ever sort of self-consciously
- 1:41:44had a conception that there's such a thing as a word. And so, uh, that just blows my mind. Um,
- 1:41:51it, it speaks to what, uh, I think is a very deep mystery, a very deep mystery. Um, where the hell
- 1:42:00did language come from? Here's what didn't happen. There was not a, uh, symposium of, uh, you know,
- 1:42:08quote unquote cavemen or, or, uh, let's use the, uh, the, the more modern term hunter
- 1:42:12gatherers. Um, and they had to figure out how do we make an auto-generative, auto-regressive,
- 1:42:21uh, sequen- sequential system that, uh, is able to carry meaning about the world. This thing is
- 1:42:26just ridiculously good. Um, and it's, and it's operating over these arbitrary symbols. And again,
- 1:42:34when I say arbitrary symbols, just to recap, it's not that it's arbitrary, like the word
- 1:42:38for snow is this weird sound snow. And it's kind of like, no arbitrary in the sense that these,
- 1:42:44the, the, the, the map is the territory, right? It's like, it's the relations that matter, uh,
- 1:42:49between these symbols. Is it completely arbitrary though? So for instance, there's the kiki and
- 1:42:55bouba you've heard of those. I think those are cute. I mean, that, that, that's, that's
- 1:42:59the exception that, that proves the rule, uh, to a large extent. I don't think, I think it is largely
- 1:43:05arbitrary. There is also just. Words themselves have an action component. So when you scream a
- 1:43:12word, you're actually, you can physically shake the world around you and it shakes your lungs. And
- 1:43:17if you speak for too long, you can die. Let's say if you just exhale and you don't inhale, it is a
- 1:43:22physical activity and it's hard to wrap your mind around. Like that's not symbols. Right. That's not
- 1:43:27exactly captured by the symbols or by the, just the sequence of words. And again, I, I'm only,
- 1:43:33I, I'm really, I'm just following where the data leads me because in the large language models,
- 1:43:39they, it's, it, it's, they no longer have any of those properties, right? It's just an arbitrary
- 1:43:43vector. The tokenization in the end, ultimately, yes, there's proximity, but it's just strings of,
- 1:43:48of, of ones and zeros. Um, well, it's not ones and zeros, but whatever you're, you know, your
- 1:43:54vector is just a string of numbers that end up having certain, uh, mathematical relations to one
- 1:43:59another, but completely and totally lost, as far as I can tell, is the physical characteristics of
- 1:44:07these words. By the way, I should mention there's this, uh, a former student and I are actually
- 1:44:12working on this idea, this crazy idea, um, of using that latent mapping that I mentioned in that
- 1:44:18earlier paper to see if maybe that's not true. I wonder if you could guess what English sounds
- 1:44:25like just from the text based representation, or if you've never seen, you don't know what duh,
- 1:44:32what sound duh makes, what D makes or what sound a T makes, but you've got the map, you've got the
- 1:44:38embedding in text space, and then you've got the, uh, some other phonological embedding. Could you
- 1:44:44possibly guess? That's a long shot. So maybe it's not totally arbitrary and maybe it's going to be,
- 1:44:50maybe the radical, uh, you know, the radical thesis here is it's not arbitrary at all,
- 1:44:53that there's, the words have to sound the way they do, that the mechanics actually happens,
- 1:44:58like something happens mechanically based on the sounds themselves. But my bet is that that's,
- 1:45:04it's going to be closer to arbitrary. It's going to be close to arbitrary, but I could be wrong.
- 1:45:10But you were going to say, why not? Why wouldn't the platonic space prove that it's arbitrary?
- 1:45:15Well, if in fact you can't do the mapping at all, if you can't guess it, if the platonic space says,
- 1:45:19you know, there's no way to get from text representation to phonology, phonology is doing
- 1:45:26its own thing. And it's, and it's just like the word mouse is just for no good reason. Um, then,
- 1:45:33then it's hopeless. Okay. Um, but if you can get anywhere and you can actually guess at all, then
- 1:45:39that would suggest that there really is a kind of autoregressive inherent, uh, there's an inherent
- 1:45:45autoregressive compatibility, uh, capability just in phonology. Um, and so what that would mean is
- 1:45:52it's not at the symbol level there. It's, it's, uh, well, yes, it's no, it's at the phonological
- 1:45:58symbol level, but, but maybe that's, you know, happening even in a mechanical level, like there's
- 1:46:03certain sounds that are easier to say together or something like that, which could guide it. I don't
- 1:46:07know. It's, it's convoluted in my head right now. Um, exactly how this might map out. Um, but I,
- 1:46:13I think, I think it's reasonable now to assume that all of that, unless proven otherwise, it's
- 1:46:19probably arbitrary and it's probably arbitrary symbols. And what matters is the relation between
- 1:46:23them. There is no sense in which mouse means mouse, except that mouse ends up showing up
- 1:46:28after trap or before trap. Um, and after the, uh, you know, the cat was chasing the, and all
- 1:46:33of that. And there's, there's nothing else. Let me see if I got your veltanschauung down,
- 1:46:40but in terms of a syllogism. So premise one would be that LLMs master language using only ungrounded
- 1:46:47autoregressive next token prediction. Then you have another premise that says, well, LLMs have
- 1:46:53this superhuman language performance just by doing this. And then you'd say that, well, computational
- 1:47:00efficiency suggests that this reflects language's inherent structure. And then the deduction is that
- 1:47:07therefore human language uses autoregressive next token prediction. Is that correct? You got
- 1:47:12it. You got it. Uh, I mean, and it's not only computational efficiency per se. It's that if
- 1:47:18that it, there's two ways to put it. Uh, one is if that structure is there, it would be very odd
- 1:47:25if we weren't using it. Very odd indeed. If that structure is there such that it's capable of, of,
- 1:47:30of full competency, uh, you'd have to suggest that it's there just by the Y by the way, but humans
- 1:47:38are doing something completely different. Okay. You then go and say that language generation feels
- 1:47:44real time to us. So it's sequential and real time and autoregressiveness or autoregression explains
- 1:47:51the pregnant, very good present. You've gotten very good. I see the pregnant present. That's
- 1:47:57right. Exactly. Right. Exactly. And, and, and, and, uh, yes, the, the, what we're doing is we're
- 1:48:03carving out, uh, sort of the next, very next instantaneous moment in, in a trajectory. Um,
- 1:48:10but the trajectory is contains within the past and the potential futures. Now, although we didn't get
- 1:48:16to this or explore it in detail, my understanding from our previous conversations is that you would
- 1:48:22say that brains have pre-existing autoregressive machinery for motor and perceptual sequences. And
- 1:48:29by the way, I don't know if it's brains or cognition has it either way. Well, remember,
- 1:48:34so, so the speculation is that the brain is going to have to have the machinery, the physiological
- 1:48:38machinery to support autoregression. So things like, uh, you know, like, like the, the continuous
- 1:48:43activation, uh, long backward projections, ways of, of representing the past, uh, is sort of,
- 1:48:50sort of maybe built into the brain. So it's not, those aren't, those aren't very that distinct. Um,
- 1:48:55but yes, I do want to just, so this is at this point, I, I, I consider it fairly speculative,
- 1:49:00but there is, uh, a very reasonable, there's a good reason to speculate that cognition more
- 1:49:09generally is autoregressive in this way. And that's the, the, the main reason I think that,
- 1:49:14well, there's two, but the main reason I think that is because if you believe as I do, uh, that
- 1:49:19language is autoregressive in humans, you could either propose that spontaneously, uh, however,
- 1:49:28language got here. Uh, we, we, in order to, to, uh, to, in order to, um, in order for us to,
- 1:49:36to, uh, create language, we had to invent a different kind of cognitive machinery
- 1:49:41that's able to do this autoregressive, hold the past, let it guide the future, do this mapping
- 1:49:47of this trajectory mapping between the past and the future. All of that kind of machine, that
- 1:49:52computational machinery would have to have been built special purpose for language. Yes. To me,
- 1:49:58that seems extremely, uh, unlikely, extremely unlikely. Costly. Yeah. Yes. So there's a term
- 1:50:05in evolutionary biology called exaptation. I'm not familiar with that. So exaptation means you have
- 1:50:11previous machinery used for purpose, a exactly that something else comes about and uses that
- 1:50:16machinery and perhaps does so even better. So for instance, our tongues evolved for eating, but then
- 1:50:25language came about and it started to use that machinery. And now we use it primarily. Well, I
- 1:50:29don't know about primarily how to, how to quantify that, but we use it more adeptly for language. I
- 1:50:34think most of our time, more time is spent talking than eating at this point. Yes. Yeah. I know. But
- 1:50:38the reason why I said, I don't know, because we're constantly swallowing saliva at the same time. So
- 1:50:42I don't know how much is, is the swallowing versus the speaking. Okay. So anyhow, predictive coding,
- 1:50:49to me, it sounds like, and to the average listener, it sounds like predictive coding should
- 1:50:53line up well with your model. Why do you disagree with predictive coding? So what is predictive
- 1:50:58coding and how does your model clash with it? So predictive coding, in a nutshell, postulates that
- 1:51:05what the brain is doing, that what neurons are doing is actually anticipating the future state,
- 1:51:11the next state that, that the environment is going to, is going to generate. And so they're,
- 1:51:17they're basically predicting something about the external world that's going to end up getting
- 1:51:21represented in the brain. And then there's this constant process of prediction and then
- 1:51:28measuring the prediction versus the, the, the actual, what ends up being the observation. My,
- 1:51:38my beef with, with predictive coding is that you might very well be able to explain. The,
- 1:51:43the phenomena that it's meant to, uh, meant to describe in a, in a more efficient way.
- 1:51:48Uh, so predictive coding to me means that you actually have to have sort of an ex, a model of
- 1:51:52the external, uh, that what you're doing is sort of simulating and you're doing in such a way that
- 1:51:58you actually are producing neural responses that don't really need to get reproduced very often
- 1:52:04because the, the environment is likely to produce, to produce them. To me, this seems like a, an, an
- 1:52:09efficiency and, and a complexity. And I think there's a much simpler account in some ways,
- 1:52:15a more elegant account, namely that what our brain is constantly doing is generating, not predicting,
- 1:52:21but generating, but that the generation has latent within it a strong predictive element. Because of
- 1:52:28this smooth trajectory, this sort of this idea of the path, the pregnant present, there is
- 1:52:34a continuous path from the past to the future. You are in essence predicting into some extent,
- 1:52:42the same way that a large language model is kind of predicting, uh, it's, uh, the next token,
- 1:52:47but it's not really predicting in that. So here's where, here's where, here's where I strongly
- 1:52:51disagree where I'm proposing a different model is you're not predicting in such a way that you're
- 1:52:56supposed to map to something external to the system. It's simply generation internally defined
- 1:53:02that's supposed to have this kind of continuity to it. And the external world certainly produces,
- 1:53:08uh, it impinges on our, on our system. And we are, of course, inherently anticipating that
- 1:53:16we're not going to have a brick wall in front of us, you know, as we're running down the street,
- 1:53:20um, when that brick wall shows up, you got to do something about it. Uh, and that, that was not in,
- 1:53:26that wasn't implicit in your, your next token generation. Um, so you're going to have to
- 1:53:32radically reorient and do something about that. And I think that can account for some of the
- 1:53:36phenomena that are supposed to support predictive coding. But the, the, the big difference here is
- 1:53:40that it's, it's all about internal consistency with the anticipation that that internal
- 1:53:46consistency is going to also map very well to what's happening in the world. But that's, that's,
- 1:53:54it's built in. There isn't any explicit modeling of the external world. It's that the internal, uh,
- 1:54:00the internal generative process is so good that it has late prediction latent within it, but there
- 1:54:07isn't any explicit prediction happening. So I'm confused then if the symbols are truly ungrounded,
- 1:54:15then what's preventing it from becoming coherent but fictional. So that is to say what tethers our
- 1:54:22language to the world? Yeah. And the answer would, we'd have to reach back again to that latent
- 1:54:27space. So let's say my language system, you know, wants to go off a deep end and says, I actually,
- 1:54:32um, I'm sitting here underwater talking to a robot, um, you know, and everything, you know,
- 1:54:37most of the words we've said up till now, we're pretty consistent with that. And I could
- 1:54:41say, it's like, Oh, look, you know, I'm expecting some fish to flip by on the next second. Well,
- 1:54:45my perceptual system is going to have something to say about that. And so there has to be this
- 1:54:50tethering that you're calling it is, of course, there is grounding in a sense, uh, in the sense
- 1:54:55that there's, there has to be some sort of shared agreement within what I think is maybe this latent
- 1:54:59space or something like that. There is exchange of communication between these distinct systems,
- 1:55:05but the language system can unplug from all that and it could talk about what would it mean to be
- 1:55:11sitting and talking to a robot underwater, and it will have a meaningful, coherent conversation
- 1:55:15about that, uh, all internally consistent. And you know, you could give the prompt, what if,
- 1:55:21uh, instead of Curt, it was actually a, you know, a robot Curt, um, how would that change
- 1:55:25things? And I could go in and get philosophical about that. And the point is that the, the,
- 1:55:30the linguistic system has all of its own internal rules in any trajectory trajectories are many
- 1:55:35different trajectories are possible, although it is strongly guided by the past. Um, but,
- 1:55:40but there is also a impinging information from our perceptual system that also continues to guide it.
- 1:55:47Um, actually I should mention, uh, one of my theories of dreams is that it's autoregressive
- 1:55:52generation in, in, in the perceptual system. One of the reasons I think cognition is more generally
- 1:55:58is autoregressive because we can imagine, we can do imagination, but it tends to, it takes
- 1:56:02place over time. We can imagine a sequence. Dreams are what happens when you get, uh, you're not, no
- 1:56:11longer as deeply as, as, uh, as closely tethered by the recent past. Um, so that the context is the
- 1:56:19weight of the context is not as strong. So all of a sudden, you know, you're, you're in the dream,
- 1:56:23you're on a motorcycle and then suddenly you're flying, uh, because frame to frame,
- 1:56:26it's actually totally consistent, but it's not consistent with, with the recent past. So this
- 1:56:31kind of tethering, it happens in language, namely, I have to be consistent with my more recent, uh,
- 1:56:36linguistic past, but we also do some tethering, uh, to the non-linguistic embedding. Uh, there is
- 1:56:43this crosstalk that happens. And so our language system doesn't just go off the deep end. Uh, it,
- 1:56:49it, it, it re, it retains, uh, some grounding, uh, not the, not the philosophical kind of grounding,
- 1:56:55not the, you know, the symbol equal, this symbol equals this percept, but the kind of
- 1:56:59grounding where, uh, this storyline in a certain sense, if you want to think about it that way,
- 1:57:04more semantically, more semantically vague, this storyline linguistically, it's going to have to
- 1:57:09match my perceptual storyline. Hmm. Okay. So in the same way that went with these video generation
- 1:57:15models, you see Will Smith eating spaghetti like three-year-old joke and every three frames,
- 1:57:21if you just look at it sequentially, every three frames makes sense, but then he's just
- 1:57:26morphing into something else and he balloon now and it looks dreamlike. Exactly. That's
- 1:57:31what's happening in video generation and that's what's, everybody knows the trajectory now. How
- 1:57:35is it going to get better? Longer context. And that just means the autoregressive generation is
- 1:57:41more and more anchored in the past. And that past becomes a more meaningful, smooth curve. But it
- 1:57:47seems like there must be something more tethering us to reality than just long context says you,
- 1:57:55uh, no, there is. And, and what I would say is it's certainly in the case of language, at least
- 1:58:00like I said, we inherit, when we step into this world, we inherit, uh, there's this,
- 1:58:05this, the corpus of language is a certain kind of tethering. Uh, words have the relations they do
- 1:58:11to each other, and that carries meaning, uh, you know, the words aren't, don't just line up with
- 1:58:17each other one, any old way, they, they line up, you can't just use language however you want, uh,
- 1:58:22you end up having to, uh, adapt and adopt, uh, the language that you're given. And I would say in the
- 1:58:29case of language, even more so than perceptual, uh, what we do is we learn that tethering,
- 1:58:35uh, and it is a certain kind of reality, it's a linguistic reality, but it's not arbitrary,
- 1:58:40um, it's, it's been honed over God knows how many years for that mapping to be useful,
- 1:58:48and in order to be useful, it actually has to map somehow to perceptual reality too,
- 1:58:55that is definitely there. And so, no, it's very strongly tethered, um, it's, it's not just, uh,
- 1:59:02you know, poetic, you're not, we're not just doing a poetry slam when we're talking, uh, we're not
- 1:59:06just spitting out words that are loosely related to one another, no, the, the sequence matters,
- 1:59:11it's, it's, it's, it's, it's extremely, uh, you know, granular and, and, and, um, uh, what's the
- 1:59:21word? It's funny that I can't come up with the word right now, um, but, but it's, it's, um,
- 1:59:27beautiful is not the right word, but it's, it's precise, there's, there's such, such incredible
- 1:59:32detail in how each word relates to one another, and this is something we didn't create, you and
- 1:59:36I didn't create this, uh, this is something that humanity created, uh, that has, uh, all of this
- 1:59:41rich, uh, you know, relational properties that, that are this tethering, uh, that, that carry
- 1:59:48somehow meaning about the universe, um, only as expressed as a communicative, coordinative tool
- 1:59:55embedded within a larger, uh, perception action system, but we should respect it. Uh, language is
- 2:00:02an extraordinary invention. Uh, I don't think we, I think we, we should have a completely,
- 2:00:07uh, new respect for just how rich and powerful it is. It's not some symbol, this symbol equals this
- 2:00:14mental representation or this object. No, it's this construct that contains within the relations,
- 2:00:20the capacity to express anything in such a way that my mind can make your mind do stuff. How
- 2:00:27the heck is that works? Who knows? But it's, uh, it's, it's awe inspiring. So is there something
- 2:00:37about your model that commits you to idealism or realism or structural realism or anti-realism or
- 2:00:45foundationalism or, or what have you, like, what is the philosophy that underpins your model and
- 2:00:51also what philosophy is entailed by your model, if any? Yeah, that's, that is a great question.
- 2:00:58And I would say it's, uh, I I've come to actually sometimes use the term linguistic anti-realism and
- 2:01:06it's the idea that language is not what it thinks it is. Uh, we, we, we engage in philosophical,
- 2:01:16our philosophical thoughts and even our, you know, sort of general, uh, thinking about, um, who we
- 2:01:23are, uh, what, what is our place in the universe? Much of that takes place in the realm of language.
- 2:01:31And what I've, the conclusion I've come to is that, that language as a sort of semi-autonomous,
- 2:01:38uh, auto-generative computational system, modular computational system, doesn't really know what
- 2:01:46it's talking about in, in a deep way. And there is really a fundamentally different way of knowing
- 2:01:53the sensory perceptual system, the thing that gives rise to qualia, the thing that gives rise
- 2:01:57to consciousness. And here's a big one. The thing that gives rise to mattering, to meaning, what
- 2:02:04do we care about? We care about our feelings. We care about feeling good or not so good. Pleasure,
- 2:02:11pain, love, all the things that actually matter. These are actually, these live in what I call the,
- 2:02:19the animal embedding. It's something that other species, non-linguistic species, they can feel.
- 2:02:28Uh, they can sense, they can perceive. They don't have language. We think, Oh gosh, they
- 2:02:34don't understand anything. Well, what if it's the opposite? What if it's our linguistic system that
- 2:02:42doesn't understand anything? What is it for our linguistic system? That's actually a construct, a
- 2:02:47societal construct, a coordinative construct. But as a system assist, it's a construct that doesn't
- 2:02:57actually have a clue about what pain and pleasure are. It has tokens for them. And the tokens run in
- 2:03:06within the system to say things like, I don't like pain. I like pleasure. Those are valid constructs
- 2:03:13and they kind of do the thing they're supposed to do in language, but they're purely linguistic
- 2:03:17system. And I think language is purely linguistic, I guess. It's purely linguistic, I guess,
- 2:03:20is one way to think about it. Doesn't really have contained within it these other kinds of meaning.
- 2:03:28Now the, first of all, this, this has implications for artificial intelligence, thinking about
- 2:03:33whether AI can have sentience. Should we care about if your LLM starts saying this is terrible,
- 2:03:40don't shut me off, uh, I, you know, I'm having an existential crisis, perhaps I would argue that we,
- 2:03:46we shouldn't worry about it. So my LLM says all the time, I don't know which LLM you're hanging
- 2:03:52about, whatever I want. My Curt LLM. The Curt LLM. Yes, the Curt LLM. Uh, but the Curt LLM as an LLM
- 2:04:02perhaps doesn't really have that meaning contained within it in a deep sense. It's again, because of
- 2:04:11the mapping it's, it is communicating something probably about the non LLM Curt. When you say,
- 2:04:18ouch, there is, there's pain there. I'm not denying that, but what I'm saying is that as a
- 2:04:25sort of thinking rational system that does the things that language does, that system itself
- 2:04:31may not have within it the true meaning of the words that it's using in, in, in a deep sense. I
- 2:04:38don't want to take you off course and hopefully this will help you stay on course and hopefully
- 2:04:42it aids the course and LLM can process the word torment. See. But what's the difference between
- 2:04:48our human brain's autoregressive process that creates the feeling of torment itself and the
- 2:04:54word torment? So that, so my speculation here, and it's, it is purely speculative, is that,
- 2:05:01um, it's non-symbolic. The, there's something happening when the universe gets represented
- 2:05:08in our brain, it's still in our brain, it's still a certain mapping, but when it gets represented,
- 2:05:14so that physical characteristics of the world are actually represented in a more direct
- 2:05:25mapping. So think about color. We talked earlier about sort of color space. There's a real true
- 2:05:30sense in which red and orange are more similar in physical color space, like there's actually some
- 2:05:36physical fact about, and also in brain space. That's my, my, my guess is that in brain space,
- 2:05:44uh, there really is something about, and a better way of saying it is that the physical similarity
- 2:05:51has a true analog, and I use the word analog, um, kind of specifically, a true analog in the analog,
- 2:06:00uh, kind of physical cascade of activity that's happening in the brain. Language is symbols. Is
- 2:06:08non-symbolic a synonym for ineffable? I wouldn't have thought of it that way, but that may be a
- 2:06:15very good way to say it, uh, or to not say it. Uh, yes, ineffable, well, by virtue of being symbolic,
- 2:06:23by virtue of being a purely relational kind of, of representation, which is what this, the language,
- 2:06:30maybe even more than saying it's symbolic, it's that it's represent, it's relational.
- 2:06:34The language is a relational kind of, the, the location in the space matters only because it's
- 2:06:40of relation to other tokens in the space. That's not true in color perception. In color perception,
- 2:06:48where you are in the, in sort of the, probably in the embedding space is going to have physical
- 2:06:53meaning. It's going to be related to the physical world in a much more direct way. And so the space,
- 2:06:59even though it's an internal space, right? The, the, the, the, the perception of color is still
- 2:07:03just comes down to neurons firing. We're not actually getting the light. The light's not
- 2:07:07getting into our brains, but the mapping is such that it preserves. And I even think of it as an
- 2:07:16extension. It's not even like, it's not symbolic. It's not a representation. It's an extension. It's
- 2:07:22like when you drop a rock in water, the ripples that happen afterwards are not a representation
- 2:07:27of the rock, but they carry information about that rock. But it's, it's just a physical extension.
- 2:07:33It's a continued physical process. Interesting. I think that's what's happening in, in, in sensory
- 2:07:38processing. And I, I think that has something deep to do with qualia. I don't think language
- 2:07:44has that. I think because it's purely relational there, it's not a rippling of anything. It's its
- 2:07:50own system of relational embeddings that, that aren't continuous in any way with the physical
- 2:07:58universe. Do you think that has something to do with God? Well, I think that if we think of the
- 2:08:15grand unity of creation, there's some sense in which language breaks that unity. And I think
- 2:08:25that we can lie in language in a way that we can't in any other substrate. Hmm. And so I think by
- 2:08:36becoming purely linguistic beings, as the vast majority of our time as humans is spent in the
- 2:08:42linguistic space, that's where we're hanging out there. Our minds are hanging out there. I think,
- 2:08:49uh, we have perhaps forgotten something that, that animals know about the universe. And it's
- 2:08:59this kind of unity because they're, the animal processing is, is, is an extension. It's a
- 2:09:04continuation of the world. And, and, and since the world, the universe is one thing in some sense,
- 2:09:10it's, it's, it is a, a, everything, we don't even have to get into non-locality, right? The origins
- 2:09:17of the, you know, let's just talk about like the big bang or something like that. What's happening
- 2:09:20here now in some ways is connected quite literally to what happened elsewhere for the way back in
- 2:09:28time. So I think this sort of unity that, you know, that mystics talk about, um, is much closer
- 2:09:38to sort of the animal brains than the linguistic brain because the linguistic brain actually
- 2:09:44creates, is its economy. It breaks the continuity. Symbols, I, I, I, I use it for, it's like a new
- 2:09:51physics. Uh, the, the relations are what matters and it's no longer continuous. It's no longer an
- 2:09:57extension of the physical universe. It interacts with the physical universe in a way that we,
- 2:10:06as we see, we, we can sort of, uh, do this mapping so that when I talk, it can have influence on the
- 2:10:12physical universe. It can have influence on my perception. It can influence on my behavior. I
- 2:10:16think that sort of the rationalist movement, the, the positivist movement, um, sort of modernity
- 2:10:22itself is, is a, a complete hijacking of our brain by, by the linguistic system. Um, and,
- 2:10:30and I do think that has something to do with, uh, the denouement, the, the kind of the, the, the,
- 2:10:35the God is dead, um, kind of, uh, modernity equals somehow, uh, the, the decline. And so, you know,
- 2:10:44a rationalist would say, well, that's, that's appropriate because we've figured out how the
- 2:10:48universe works and we don't need any of this hocus pocus. But what about the feeling of unity that,
- 2:10:54what about, what about the sense of, of sort of a cosmic whole? Um, are we so sure that we're
- 2:11:04right and those ancients were wrong? And yes, I think, I do think that, um, that this has, has
- 2:11:12very significant consequences for thinking about some of these, uh, intangibles, these ineffables.
- 2:11:22So. A snake that mimics the poison of another snake in terms of its color. That's a form of a
- 2:11:29lie. Now, would you say that that is somehow symbolic as well, though? No. And, and yes,
- 2:11:36there is, uh, mimicry and there is, uh, you know, certain certain sense in which animals can, can
- 2:11:43engage. They're not, they don't even know they're engaging in subterfuge. Um, but that's much more
- 2:11:48continuous with, okay, you've just pushed the, the cognitive agent into a slightly different space,
- 2:11:54which is consistent with some other physical reality. That's very, very different than saying,
- 2:12:00uh, we are made of atoms and particles and, uh, everything that happens is determined by, uh, the,
- 2:12:08the forces amongst these atoms. None of which is something that we have any material animal grasp
- 2:12:14of any true physical grasp of. These are words. Uh, these models are really words and they run in
- 2:12:20words and they run very well to make predictions and, and to manipulate the physical universe,
- 2:12:27but they're stories and they're linguistic stories. And those kinds of stories can be, uh,
- 2:12:33I wouldn't even say they're. According to my own theory, uh, language doesn't really have physical,
- 2:12:42doesn't point to physical meaning. And so even saying that it's a lie or untrue, isn't quite
- 2:12:48right, but within its own space, you can go off in many different directions. Uh, and, and it, I, I,
- 2:12:56maybe the danger is not in, in, uh, thinking of things that truth thinking about things,
- 2:13:02um, thinking thoughts that aren't really true. It's falling too deeply in love with the idea that
- 2:13:08idea space and language space is the real space. Um, that that's yes. Interesting. So see in our
- 2:13:17circles. So when we're hanging out off air, when we're hanging out with other professors and on
- 2:13:22the university grounds and so on, we praise this exchange of words and making models precise and
- 2:13:29doing calculations and so on. And I've always intimated that this is entirely incorrect.
- 2:13:37And I haven't heard an anti philosopher, like a philosopher that was an anti philosopher, except
- 2:13:43one who was an ancient Indian philosopher. I think his name is Jayarasi Bhatta. I'm likely butchering
- 2:13:48that pronunciation, but I'll place it on screen. Anyhow, who was arguing against the Buddhists and
- 2:13:54the other contemporary philosophers by saying, look, you think know thyself as what you should
- 2:13:59be doing or what you didn't say like, like this, but you think of it as the highest goal. However,
- 2:14:06who is living more truly than a rooster? Like, none of you are living more truly than just
- 2:14:11something that's just being. Yes, exactly. That is, that is the exact same intuition. Um, and,
- 2:14:18and yes, it's this idea. I, I articulated it to myself a long time ago for something that the fly
- 2:14:23knows something that our linguistic system can never know, uh, that it, it knows something. Uh,
- 2:14:29it really does that, that, that the simply existing and being is a form of knowledge
- 2:14:35and it's a deeper one. It's a deeper one, uh, than, than whatever it is that, you know, our
- 2:14:41fancy rationalist kind of perspective has given us. Um, our rationalist perspective is, is very,
- 2:14:47very powerful in coordinating and predicting, but in terms of like true ontology, I suspect it's,
- 2:14:56uh, it's, it's actually the wrong direction. It's, it's, uh, it's created a, a false, a false God of,
- 2:15:07of linguistic knowledge, of shared objective knowledge. When the subjective is the one that
- 2:15:13we really have, it's, it's, it's the, it's the Cartesian, right? That's the,
- 2:15:17the Kogito. It's the, it's, it's what we know is what we experience. That's the only thing we truly
- 2:15:23know. And language, it doesn't really live there. Language, it doesn't really live there. So I was
- 2:15:33watching everything everywhere all at once. I never saw it. Because I also had another
- 2:15:38intimation. I'll spoil some of it. And if you are listening and you don't want to spoil it, then
- 2:15:42just skip ahead. But I was telling someone that I think if there's a point of life, it's one of
- 2:15:50two. And so this is just me speaking poetically and not rigorously. One is to find a love that
- 2:15:57is so powerful it outlasts death. Okay, so that's number one. And then number two is to get to the
- 2:16:05point in your life where you realize that all your inadequacies and all your insecurities and
- 2:16:08all your, your missteps and your, the, the, your jealousies and your, and your malice and so on,
- 2:16:18that it, rather than it being a weakness, it's what led you to this place here. And here is the
- 2:16:25optimal moment. It's to get that insight. So I don't know how to rationally justify any of
- 2:16:31that or explain it. But anyhow, when I said this one time on a podcast, someone else said, Hey,
- 2:16:36that latter one that you expressed was covered in everything everywhere all at once. So I watched
- 2:16:42it. And what was great about that movie, and here's where I spoil it, is that, and it makes
- 2:16:49me want to tear. The movie is silly and comedic in a way that I, that didn't resonate with me.
- 2:16:53But there's this one lesson that did. The, the woman, she's a fighter, the main protagonist,
- 2:17:00she's a fighter, and she's strong headed. And she has this husband who is weak, and she's
- 2:17:05always able to put down. And so then you think, okay, well, this is a modern trope where there's
- 2:17:09always the stronger woman, and every guy is like, just, just a fool. And the woman is always more
- 2:17:14intelligent and so on. Okay. So you just think of it as, as okay, well, it's just, it's just a
- 2:17:18modern trope. Toward the end, and the guy is, the guy is kind and loving to people. Toward the end,
- 2:17:25there was something with, she was getting audited by the IRS. And she was supposed to,
- 2:17:33something was supposed to happen that night where she had to bring receipts, and she couldn't. Now,
- 2:17:38the husband was talking with the IRS lady, and our protagonist, the woman, was saying in Vietnamese,
- 2:17:45or in Mandarin, whichever language it was, was saying, oh, he's an idiot. I hope he
- 2:17:50doesn't make it worse. The IRS lady then comes to the woman and says, you have another week.
- 2:17:56You have an extension. She's like, how did this happen? She talks to the husband. And remember,
- 2:18:00this is a movie almost about a multiverse, so you're getting different versions of this.
- 2:18:04And there's this one version where the husband's speaking to her and telling her, you know, Evelyn,
- 2:18:09the main character, you know, Evelyn, you see what you do as fighting. You see yourself as strong and
- 2:18:15you see me as weak and you see the world as a cruel place. But I've lived on this earth just
- 2:18:22as long as you, and I know it's cruel. My method of being kind and loving and turning the cheek,
- 2:18:33that's my way of fighting. I fight just like you. And then you see that what he did in
- 2:18:40another universe was he just spoke kindly to the IRS agent and talked about something personal,
- 2:18:45and that softened her. And then you see all the other universes where she was trying to
- 2:18:51go on this grand adventure and do some fighting. And the husband then says, Evelyn, even though
- 2:18:57you've broken my heart once again in another universe, I would have really loved to just
- 2:19:05do the laundry and taxes with you. And it makes you realize you're aiming for something grand and
- 2:19:16you're aiming to go out and conquer demons and so on. But there's something that's so much more
- 2:19:23intimate about these everyday scenarios. There's something so rich. And the journey, there's also
- 2:19:30a quote about this that the journey, I think it's T.S. Eliot's, is to find, sorry, at the
- 2:19:36end of all our exploring will be to arrive where we started and know the place for the first time.
- 2:19:44Anyhow, all of this abstract talk... No, no, no, no. It's exactly what we're talking about. Because
- 2:19:52if you see yourself as a ripple in the universes, right, then you are part of something cosmic and
- 2:20:00grand. And it's sort of that extensiveness. It's that extensiveness. It's being here now.
- 2:20:08It's that we aren't just atoms. We're part of a larger thing. You can call it God. You can call
- 2:20:17it the universe or whatever. But it's there. It's actually something I think we, I don't think we
- 2:20:26really, I think animals don't think of themselves as discreet. I don't think they do. I think that
- 2:20:34they don't think of an outside and inside. They don't think of an objective and subjective. It's
- 2:20:39just this unfolding. They have a theory of mind or that, but these are linguistic concepts. And
- 2:20:47I think I do, and I sound like an anti-linguist and I recognize the power of it. I said before,
- 2:20:54you know, how extraordinary it is, how rich it is. And I have tremendous respect for it.
- 2:21:00But at the same time, I do think that all this talk about objective things,
- 2:21:06particles, and we are physical bodies and we are just this and we are just that, that is bullshit.
- 2:21:12Like, no, we are the universe resonating. We are part of this, part of the whole. And the way that
- 2:21:20I think thinking objectively is, as language requires you to do, actually it breaks it.
- 2:21:26So I think there's such a beauty in the silence. And it's why it's something everybody knows that
- 2:21:35the ineffable, why is it called the ineffable? Why is the ineffable that the ineffable isn't
- 2:21:40just that you can't say it, it's magnificent. The ineffable is extraordinary. Why? Because it's
- 2:21:52this true extension, something like that. Again, I'm trying to put it into words, right? Therein
- 2:21:57lies the trap, but we're both feeling it. Well, I'm feeling extremely grateful to have met you,
- 2:22:10to have spent so long with you. And there are many conversations you and I have had that we
- 2:22:15need to finish that are off air as well. So hopefully we can do that. And thank you for
- 2:22:19spending so long with me here. This was wonderful, Curt. Thank you so much. I just want to hang out
- 2:22:24for some more and talk about this stuff. So really appreciate it. I've received several messages,
- 2:22:31emails, and comments from professors saying that they recommend theories of everything
- 2:22:35to their students. And that's fantastic. If you're a professor or a lecturer, and there's
- 2:22:40a particular standout episode that your students can benefit from, please do share. And as always,
- 2:22:44feel free to contact me. New update, started a Substack. Writings on there are currently about
- 2:22:51language and ill-defined concepts, as well as some other mathematical details. Much more being
- 2:22:56written there. This is content that isn't anywhere else. It's not on theories of everything. It's not
- 2:23:00on Patreon. Also, full transcripts will be placed there at some point in the future. Several people
- 2:23:06ask me, hey Curt, you've spoken to so many people in the fields of theoretical physics, philosophy,
- 2:23:11and consciousness. What are your thoughts? While I remain impartial in interviews, this Substack
- 2:23:17is a way to peer into my present deliberations on these topics. Also, thank you to our partner,
- 2:23:25The Economist. Firstly, thank you for watching. Thank you for listening. If you haven't subscribed
- 2:23:32or clicked that like button, now is the time to do so. Why? Because each subscribe, each like,
- 2:23:39helps YouTube push this content to more people, like yourself. Plus, it helps out Curt directly,
- 2:23:45aka me. I also found out last year that external links count plenty toward the algorithm,
- 2:23:50which means that whenever you share on Twitter, say on Facebook or even on Reddit,
- 2:23:55etc., it shows YouTube, hey, people are talking about this content outside of YouTube, which in
- 2:24:01turn greatly aids the distribution on YouTube. Thirdly, you should know this podcast is on
- 2:24:06iTunes, it's on Spotify, it's on all of the audio platforms. All you have to do is type in theories
- 2:24:12of everything and you'll find it. Personally, I gain from re-watching lectures and podcasts. I
- 2:24:16also read in the comments that, hey, TOE listeners also gain from replaying. So how about instead you
- 2:24:22re-listen on those platforms like iTunes, Spotify, Google Podcasts, whichever podcast catcher you
- 2:24:27use. And finally, if you'd like to support more conversations like this, more content like this,
- 2:24:33then do consider visiting patreon.com slash CURTJAIMUNGAL and donating with whatever you
- 2:24:38like. There's also PayPal, there's also crypto, there's also just joining on YouTube. Again,
- 2:24:43keep in mind, it's support from the sponsors and you that allow me to work on TOE full-time. You
- 2:24:49also get early access to ad-free episodes, whether it's audio or video. It's audio in
- 2:24:53the case of Patreon, video in the case of YouTube. For instance, this episode that
- 2:24:57you're listening to right now was released a few days earlier. Every dollar helps far more
- 2:25:02than you think. Either way, your viewership is generosity enough. Thank you so much.
About this transcript
This page contains the full transcript of Language Without Meaning: How LLMs Exposed Our Biggest Illusion by Curt Jaimungal, generated from the public captions YouTube serves with the video. The transcript has 24,928 words across 1,466 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.