YouTube2Text

Language Without Meaning: How LLMs Exposed Our Biggest Illusion — Transcript

by Curt Jaimungal · 24,928 words · 1,466 segments · language en · Watch on YouTube

Full transcript

  1. 0:00"I'm going to get attacked by physicists...  This thing is just ridiculously good. And so
  2. 0:04that just blows my mind!" Professor Barenholtz  completely inverts how we understand mind,
  3. 0:11meaning, and our place in the universe. The  standard model of language assumes words point
  4. 0:17to meanings in the world. However, Professor  Barenholtz of Florida Atlantic University has
  5. 0:22discovered what's unconscionably unsettling...  they don't! Language is actually deconstructing
  6. 0:28itself. Most startlingly, he argues that our  rational linguistic minds have severed us from
  7. 0:34the unified cosmic experience that animals may  still inhabit. I don't think there's a static set
  8. 0:39of facts. What we've got is potentialities. Most  current LLMs operate with purely autoregressive
  9. 0:45next-token prediction, operating on ungrounded  symbols. All of this terminology is explained,
  10. 0:50so don't worry, this podcast can be watched  without a formal background in psychology or
  11. 0:55computer science. In this conversation, we journey  through rigorous explorations of how LLMs work,
  12. 1:01what they imply about how we view the world, and  the relationship between our consciousness and the
  13. 1:06cosmos. Professor, you have two theses. One is a  speculative one and the other is more grounded.
  14. 1:15You even have another more hypothetical one atop  that, which we may get into. Why don't you tell us
  15. 1:20about the more corroborated one, and then we can  move to the contestable parts later. Okay, sure.
  16. 1:26So, yeah, I would call them sort of the grounded  thesis and then sort of the extended version of
  17. 1:33that, if we can call it that. The grounded thesis  is primarily about language. And the thesis is
  18. 1:41that human language is captured by what's going  on in the large language models. And I mean not
  19. 1:49in terms of the specific exact algorithm as to how  the large language models like ChatGPT are doing,
  20. 1:58are actually generating language, but the core  sort of mathematical principle that large language
  21. 2:03models like ChatGPT run on are what's happening  in the brain. And it's what's happening in human
  22. 2:09language. And really, the reason I say it's  corroborated is because ultimately this isn't
  23. 2:14even about the brain, it's about language itself.  And I think what we have learned in the course of
  24. 2:20being able to replicate language in a completely  different substrate, namely in computers,
  25. 2:27is that we've learned properties of language  itself. We've discovered. It's not through clever
  26. 2:33human engineering that we've been able to kind  of barrel our way towards language competency.
  27. 2:40It's that with actually fairly straightforward  mathematical principles done at scale, we've
  28. 2:47actually discovered that language has certain  properties that we didn't know it had before.
  29. 2:51And so the incontrovertible fact, in my opinion,  is that language itself has certain properties.
  30. 2:59Now that we know it has those properties,  my claim is, the sort of corroborated claim,
  31. 3:05is that those properties force us to conclude that  the mechanism by which humans generate language
  32. 3:12is the same as what's going on in these large  language models. Because now that we know that
  33. 3:18language is capable of doing the stuff that it  does, now that we know it has the properties to,
  34. 3:23and I'm sort of giving away the punchline, to  self-generate based on its internal structure,
  35. 3:29it's unavoidable to think that we are using the  same basic mechanism and principles. Because it
  36. 3:35would be extremely odd to think that we have  a completely different orthogonal method for
  37. 3:42generating language. Or put differently, if  we are using completely different mechanisms
  38. 3:48than the language models, then it's extremely  unlikely that the language models would work
  39. 3:53as well as they do. The fact that language  has this property that it can self-generate,
  40. 3:58the fact that that property actually leads to  human-level language, to me, forces the conclusion
  41. 4:03that there's only one way to do language. And that  one way is the same in humans and in machines. The
  42. 4:11obvious question that's occurring to the audience  as they listen right now is, how do we know that
  43. 4:15whatever mechanism is being used by LLMs isn't  just mimicry? Right, and so that's sort of the
  44. 4:21critical question. Is this mimicry, right? Is what  the models are doing, in a sense, learning a kind
  45. 4:27of roundabout technique that captures some of the  superficial components of language in humans, but
  46. 4:35ultimately it's a completely different approach.  And so, you know, my argument is really from the
  47. 4:42fundamental simplicity of these models. So let's  just talk really quickly about how large language
  48. 4:48models work. Things like ChatGPT. What they're  doing is learning, given a sequence. You know,
  49. 4:57let's say the sequence is, I pledge allegiance to  the, and then the model is being asked to do this
  50. 5:03thing called next token generation. What's the  probable next word? We'll say word for the purpose
  51. 5:10of this conversation. We're going to call tokens a  word. Token is a more technical term about how you
  52. 5:15chop up and encode the information in a sequence  of language. But we're just going to say word.
  53. 5:22So guess the next word based on that sequence.  So and then what you do is, in these models,
  54. 5:30is you train them to guess simply that. All you  hear is a given sequence. It can be a sentence.
  55. 5:36It can be a paragraph. It can be, frankly, an  entire book, depending on how big your model is,
  56. 5:42how much it can handle. And then guess just  the very next word. And what we've discovered,
  57. 5:48and I say I really want to use that word in  particular, because it was by no means a given
  58. 5:55that this could ever, that this would work, what  we discovered is if you train a model to do that,
  59. 5:59to simply guess the next word, then  take that word, tag it onto the next,
  60. 6:04tag it onto the sequence and feed it back in. This  is sufficient to generate human level language.
  61. 6:10Now, the reason I believe that this demonstrates  something not about our engineering or even about
  62. 6:17the models themselves, because there's different  ways you might build a model that can do this, is
  63. 6:21because this very simple trick, this simple recipe  of simply guessing the next word turns out to be
  64. 6:28sufficient to generate language at human levels to  the point where there really are no benchmarks, no
  65. 6:35standard benchmarks that these models aren't able  to do. And so what that suggests to me is just by
  66. 6:41learning the predictive structure of language,  you're able to completely solve language,
  67. 6:46that means that that is likely to be the actual  fundamental principle that's built into language
  68. 6:51in order to generate it. If we had to come up  with a very complex scheme, for example, you know,
  69. 6:58syntax trees, complex grammar, long-range  dependencies that we had to take into account,
  70. 7:04and through enough compute, we were able to kind  of master that, then I might argue, well, you
  71. 7:10know, what we're doing is possibly figuring out  a roundabout way to capture all this complexity.
  72. 7:16But it's the simplicity itself, that simply being  able to predict the next token, the next word, is
  73. 7:23sufficient to do all of this long-range thinking,  to be able to take an extremely long sequence,
  74. 7:28and then produce an extremely long sequence on  the basis of that. That suggests to me that we
  75. 7:33discovered a principle that's actually already  latent in language, that we just had to throw
  76. 7:38enough firepower at it, but with an extremely  simple algorithmic trick, and then language
  77. 7:43revealed its secrets. So, to me, this really  suggests that there is, of course, you know,
  78. 7:49that there's still a lot of science that needs  to be done, and this kind of thing, kind of work
  79. 7:54that I'm doing in my lab, in terms of really being  able to hammer down how the brain is instantiating
  80. 8:00this exact same algorithm, it's not going to look  exactly like ChatGPT, it's not necessarily going
  81. 8:05to be based on what are called transformer models,  which is something we can get into a little bit,
  82. 8:10but as far as the core principle of prediction of  the next token, the fact that that solves language
  83. 8:16so handily, to me, really argues that that is the  fundamental algorithm. That is the fundamental
  84. 8:21algorithm that, when you apply it, boom, language  emerges. If you just have the corpus, you have the
  85. 8:28statistics, and then you do next token prediction,  language is just like add water, and the fact that
  86. 8:33it emerges so readily from that, without having to  do anything complicated, to me suggests that it's
  87. 8:39latent within language in the first place,  and that language is designed, in a sense,
  88. 8:44in order to be able to be generated through this  simple, predictive kind of mechanism. So Elon,
  89. 8:54you and I have spent several days together. In  fact, you're in the video with Jacob Barndes
  90. 8:59and the Manolis Cosmon. We'll place that on screen  and I'll put a pointer to you. And you were in the
  91. 9:04background of the interview with William Hahn on  Williams. Always in the background, never in the
  92. 9:09foreground. Here we are. Okay, well, yes, great.  You have a large epiphany that occurred to you at
  93. 9:15one point. You spoke about the software and this  precipitated this entire point of view of language
  94. 9:21as a generative slash autoregressive model or  what have you. Tell me about it. What the heck
  95. 9:27was that big idea? So it wasn't so much an idea  as an epiphany, a realization. And it really hit
  96. 9:36me in a single moment. And it wasn't necessarily  about autoregression. It wasn't about this finer
  97. 9:42detail of how ultimately language models, and I  believe the brain, solved this problem. It was the
  98. 9:50realization that any model that has been trained,  any model that anybody has built that accomplishes
  99. 10:01human level language. So it might be based  on autoregression. It might be based even on
  100. 10:06diffusion, which is kind of the arch nemesis of my  autoregressive theory. But regardless, the fact is
  101. 10:14that these models are being trained exclusively  on text data. And so all they're learning is the
  102. 10:22relations between words. To the model, as far  as the model is concerned, the words are turned
  103. 10:28into numbers, they're tokenized. We think of them  as numerical representations. But those numbers,
  104. 10:33and for our purpose, we can think of them as  words, don't represent anything. There is nothing
  105. 10:39in the model besides the relations. Relations  just between the words themselves. There isn't,
  106. 10:45for example, any relation between any of the  tokens and something external to it. What we
  107. 10:51tend to think of as people, as words, what words  are doing. When we're discussing topics, thinking
  108. 10:58about words in our head, is that they symbolize  something, that they refer to something. This is
  109. 11:04a lot of the philosophy of language, a lot of the  scientific study of linguistics has been concerned
  110. 11:11with semantics. How do words get grounded? How  do they mean something outside of themselves?
  111. 11:17What large language models show us is that words  don't mean anything outside of themselves. As far
  112. 11:23as generation goes, as far as the ability for  us to have this conversation, and as far as the
  113. 11:29model's ability to produce meaningful responses  to just about any question you can throw at them,
  114. 11:37including writing a long essay on any topic,  including a novel topic that it's never
  115. 11:42encountered, is by stringing together sequences  based on simply the learned relations between
  116. 11:50words. And so this really hit me very, very  hard. I've long been puzzled by, as many are,
  117. 11:58by the mind-body problem, the phenomenon of  consciousness, the problem of how do we know
  118. 12:03your red is my red, and actually the moment  that I had this realization, it was related
  119. 12:08to this very question, I realized that the word  red doesn't mean what we mean by qualitative
  120. 12:14red. The qualitative red is taking place in our  sensory perceptual system. The word red, to a
  121. 12:20large language model, can't mean that. It can't  mean any color. It has no color phenomenon. It
  122. 12:25has no concept of what sensory red would mean. Yet  it is able to use the word red with equal ability,
  123. 12:33with equal competency, just as well as I can,  if we're just having a conversation about it.
  124. 12:38And so what this means is that within the corpus  of language, the word red doesn't mean something
  125. 12:45external to itself. Instead, the word red simply  means where does it fall in the space of language
  126. 12:51itself. Where does red fall in relation to other  colors, in relation to the word color, in relation
  127. 12:57to other concepts, other, well, frankly, just  words, tokens, that are related to what we call
  128. 13:03concepts that have to do with color and have to do  with the word red. So, yeah, so this epiphany was
  129. 13:09about this extraordinary dichotomy, this divide  between language and that which we think language
  130. 13:17refers to.The question is how does language refer,  and that which we think language refers to. The
  131. 13:19question is, how does language refer? And the  answer is, it doesn't. Language doesn't refer in
  132. 13:23and of itself. Language is an autonomous system.  It's a self-contained system. It has the rules
  133. 13:31contained within it to generate itself, to carry  on a conversation. A large language model don't
  134. 13:36know what they're talking about in any real  sense. They can talk about a sunset. They can
  135. 13:43talk about a taste. They can talk about space and  time and all of those things. And yet we would say
  136. 13:49they have no idea what they're talking about. And  we'd be right in the sense that they don't have
  137. 13:55a notion of red beyond the token and its relation  to other tokens. Now this then raises the obvious
  138. 14:02question, well, what do I mean what red is  about? Don't I think red refers to a quality of
  139. 14:09perception? And the answer is, I do have a quality  of perception. There is something called red that
  140. 14:14my sensory system is aware of. And then there's a  token called red that is used in conjunction with,
  141. 14:21there's a sort of coherent mapping between my  sensory perception of red and the linguistic red.
  142. 14:32But that doesn't mean that you need to understand  what that word refers to. You don't need to have
  143. 14:39the sensory qualitative concept of red in order  to completely successfully use the word red. And
  144. 14:47so these are compatible, but dichotomous systems.  The sensory perceptual system and the linguistic
  145. 14:54system are ultimately, we can think of them  as essentially distinct and autonomous, but
  146. 15:03compatible. Integrated? Integrated. They're  integrated. So they're running alongside
  147. 15:11each other. They're exchanging messages so  that we can have a single organism that is
  148. 15:16successfully navigating the world and able,  for example, to communicate. So I see something
  149. 15:22red. That's registered in my brain. I have a  qualitative experience of red. It's remembered
  150. 15:27in having a certain quality. And then later on I  said, oh, you know, could you go pick up that red
  151. 15:32object for me? And so there's a handoff between  the perceptual system and the linguistic system,
  152. 15:39such that the linguistic system can now  successfully send a message to you. Now
  153. 15:44you've got the linguistic system. You can talk  about that. Oh, okay, you told me there's a red
  154. 15:47object. Are there multiple objects? Yes, there's  multiple objects. They have different colors.
  155. 15:51You're looking for the red one. Maybe it's a dark  red. I'm doing this all linguistically. Now you're
  156. 15:56able to go into the room and successfully get the  right object. So again, the handoff happens the
  157. 16:01other direction. Language is able to hand off to  the perceptual system. And the perceptual system
  158. 16:04is able to then detect that there's something with  the right quality. But that's not the same thing
  159. 16:10as saying that the language contains the reference  inherently within it. It simply means that these
  160. 16:15are communicative systems, that they can exchange  information, that they integrate with one another
  161. 16:22in terms of forming coherent behavior. But  language is its own beast. It's its own autonomous
  162. 16:28system. It can run on its own. That was the big  realization. Large language models prove it,
  163. 16:32that language is able to produce the next  token, and by virtue of the next token, the next
  164. 16:37sequence. And that means all of language without  having any concept of reference. The reference has
  165. 16:44no place there. There's no way to kind of squeeze  it in. If your computational account is the one
  166. 16:50that I'm proposing, if the computational account  is essentially prediction based on a next token
  167. 16:56based purely on the topology, the structure, the  statistical structure of language, then there's
  168. 17:01no way to cram any other kind of grounding or  computational feature in there at all. It has
  169. 17:08to be something closer to, in the large language  models, prompting. You can imagine a camera that
  170. 17:15generates a linguistic description of what's in  a room, and then you can ask your language model,
  171. 17:22and you can, by the way. You can do this right  now. They're able to do vision. You can take a
  172. 17:26picture and feed it to the large language model.  What's happening is much closer to generating a
  173. 17:33prompt, basically saying, here's what's in  the room, and now based on these features,
  174. 17:37these scripts, now run the same exact language  exclusive model. And so language takes care of
  175. 17:43itself. It doesn't need grounding in order to be  able to do everything it does. It doesn't have to
  176. 17:48have concepts outside of itself. I think that's  basically been proven by these text-only large
  177. 17:54language models. So that was the big epiphany. The  big epiphany was that, oh, language is autonomous.
  178. 18:01Language is self-generating. That means it's a  dichotomous computational system. It's independent
  179. 18:07of the rest. And what this leads me to believe is,  okay, well, if it can live in silicon in this way,
  180. 18:14then perhaps, and now I've come to believe very  strongly, that it likely runs in the same way
  181. 18:20in carbon, in biology, in our brains. Okay, so  you're not denying consciousness and you're not
  182. 18:27denying qualia. No, and I want to make this very  clear. My personal opinion on this is besides the
  183. 18:37point to some extent. You can be an eliminativist  if you want, although I think everything I'm
  184. 18:44saying has a lot of bearing on this. But I believe  my account is strictly an account of language.
  185. 18:52I think that perceptual mechanisms that give  rise to qualia, things like redness and heat and
  186. 19:00taste and all of these, are basically processes  that take place long before the handoff. And so
  187. 19:08what happens is, you know, think about  the camera. The camera is transducing
  188. 19:12light. It's measuring certain wavelengths. Then  there's a lot of visual processing that has to
  189. 19:18happen before you get to the point where it's  turned into a linguistic-friendly embedding,
  190. 19:24right? The stuff that an LLM can see, a multimodal  LLM can see. And so all of that processing that
  191. 19:30happens is what I think gives rise to qualitative  experience. We experience redness because of all
  192. 19:37of this very analog, probably non-symbolic. kind  of representation. And then at the end of that
  193. 19:46process, there is a conversion. Not, by the way,  by the end of the process, a lot of things happen.
  194. 19:51We also respond to colors and to light and all  of that non-linguistically. But part of the end,
  195. 19:59sort of, we could think of different endpoints.  One of those endpoints is here's a handoff to
  196. 20:03language. And by the time language gets it, it's  long past that initial process, that kind of
  197. 20:10sensory and perceptual processing that gives rise  to qualitative phenomena. So I strongly believe
  198. 20:17that there is, in a certain sense, the word hard  problem is a little loaded. I believe there's
  199. 20:23undeniable qualia. But what I also think is that  language is poorly equipped. It's simply unaware,
  200. 20:34in some sense, of the underlying mechanisms that  give rise to what it receives at the far end.
  201. 20:42At the sort of the endpoint of that qualitative  processing. Just a moment. Don't go anywhere. Hey,
  202. 20:49I see you inching away. Don't be like the economy.  Instead, read The Economist. I thought all The
  203. 20:56Economist was was something that CEOs read to  stay up to date on world trends. And that's true,
  204. 21:01but that's not only true. What I found  more than useful for myself, personally,
  205. 21:06is their coverage of math, physics, philosophy,  and AI, especially how something is perceived by
  206. 21:12other countries and how it may impact markets.  For instance, The Economist had an interview
  207. 21:17with some of the people behind DeepSeek the  week DeepSeek was launched. No one else had
  208. 21:22that. Another example is The Economist has this  fantastic article on the recent dark energy data,
  209. 21:27which surpasses even Scientific American's  coverage, in my opinion. They also have
  210. 21:31the chart of everything. It's like the chart  version of this channel. It's something which
  211. 21:36is a pleasure to scroll through and learn from.  Links to all of these will be in the description,
  212. 21:40of course. Additionally, just this week, there  were two articles published. One about the Dead
  213. 21:44Sea Scrolls and how AI models can help analyze the  dates that they were published by looking at their
  214. 21:49transcription qualities. And another article that  I loved is the 40 best books published this year
  215. 21:54so far. Sign up at Economist.com slash TOE for the  yearly subscription. I do so and you won't regret
  216. 22:00it. Remember to use that TOE code as it counts to  helping this channel and gets you a discount. Now,
  217. 22:06The Economist's commitment to rigorous journalism  means that you get a clear picture of the world's
  218. 22:10most significant developments. I am personally  interested in the more scientific ones,
  219. 22:15like this one on extending life via mitochondrial  transplants, which creates actually a new field
  220. 22:20of medicine, something that would make Michael  Levin proud. The Economist also covers culture,
  221. 22:26finance and economics, business, international  affairs, Britain, Europe, the Middle East, Africa,
  222. 22:32China, Asia, the Americas, and of course, the USA.  Whether it's the latest in scientific innovation
  223. 22:38or the shifting landscape of global politics,  The Economist provides comprehensive coverage,
  224. 22:43and it goes far beyond just headlines. Look, if  you're passionate about expanding your knowledge
  225. 22:48and gaining a new understanding, a deeper one,  of the forces that shape our world, then I
  226. 22:53highly recommend subscribing to The Economist. I  subscribe to them, and it's an investment into my,
  227. 22:59into your, intellectual growth. It's one that  you won't regret. As a listener of this podcast,
  228. 23:04you'll get a special 20% off discount. Now  you can enjoy The Economist and all it has to
  229. 23:10offer for less. Head over to their website,  www.economist.com slash TOE, T-O-E, to get
  230. 23:18started. Thanks for tuning in, and now let's  get back to the exploration of the mysteries
  231. 23:23of our universe. Again, that's economist.com  slash TOE. To what it receives at the far end,
  232. 23:32at the sort of the end point of that qualitative  processing. Okay, let me see if I get this. You
  233. 23:38have some redness. So you do, you're not  denying redness. You grant redness. I do.
  234. 23:43Okay. There's redness. And then somehow this  needs to be referred to with some spoken words,
  235. 23:49with some language. Okay. So what's happening?  You're saying that it's an independent system,
  236. 23:56yet it's integrated. So what is that relationship?  And does it become so diluted that by the time you
  237. 24:01refer to it, you're no longer referring to that  qualia? Like, I don't understand. Yeah, that is
  238. 24:07essentially the idea. So this is the exact problem  I am working on right now. There was a fantastic
  239. 24:13paper that I just came across about a week ago.  There was a paper that was published in Archive
  240. 24:20recently. It's called Harnessing the Universal  Geometry of Embeddings. And what this paper showed
  241. 24:26is that you could have completely different models  solving different linguistic tasks. For example,
  242. 24:32you could have GPT, then you could have BERT,  which solves a somewhat different task. So there's
  243. 24:36masked tokens as opposed to autoregressive next  token generation. And what they found was that you
  244. 24:44could learn what is latent space. What you could  do is hand off, take the embedding. The embedding
  245. 24:51is basically, you can think of that as numerical  representation. It's a high dimensional numerical
  246. 24:56representation of your tokens. So here's a token.  This token is going to represent the word dog. And
  247. 25:03then we're going to take that token and embed it  in a much higher dimensional space. And what they
  248. 25:09found is that if you take the embedding, the high  dimensional representation from one model, say,
  249. 25:15ChatGPT, and then take a representation from a  different model, that you could actually get the,
  250. 25:24you could take the embedding, send it to this  latent space. If you cycle it through, get the,
  251. 25:31you have to, I know it's starting to get in the  weeds a little bit. But you send it through this
  252. 25:35latent space and then recover it in its original  form. What you can do is, once you've got that
  253. 25:42latent space, you can then translate from one  embedding to a completely different embedding.
  254. 25:48This is a new paper. This is a new paper, yes.  Right, right, right. This rocked my world. Because
  255. 25:54what they're arguing is that there, in some ways,  is this underlying universal structure of language
  256. 26:04that's captured in this latent space. And so  even though if you have a radically different
  257. 26:08embedding in one, you know, they didn't do it  across different languages, it's one of the
  258. 26:13projects I'm doing right now, is to see if you can  do this across, say, English and Spanish, even for
  259. 26:19a language that's trained exclusively in English,  and then another models trained exclusively in
  260. 26:24Spanish. Can you guess the Spanish just from  finding this kind of universal structure across
  261. 26:32these two different models? Sorry, what do you  mean, can you guess the Spanish? If a model was
  262. 26:37trained only in English, and then it was receiving  some Spanish text, a couple Spanish sentences?
  263. 26:42So the way to think about it is that what  you're doing is creating another embedding,
  264. 26:48this latent space, where you're going to be  able to send in a message in English, and then
  265. 26:56based on the station, and then again, do the same  thing for Spanish. And then what you're never
  266. 27:02going to show any model, no model is going to  ever see a pair of English and Spanish. Instead,
  267. 27:07what you're going to learn is that there's some  way to get from, you're going to end up being able
  268. 27:13to get from English to Spanish without ever seeing  the actual translation. Because what the model is
  269. 27:18going to learn is what's common across these two  representations. What's true for both the Spanish
  270. 27:25embedding and the English embedding, that there's  some sort of underlying latent structure that's
  271. 27:30true of both, and that that captures something  more universal about language. Now again,
  272. 27:34they didn't do it for different languages, they  just did it for different embeddings of English,
  273. 27:41but very different embeddings, because they  were trained on completely different models.
  274. 27:44If you looked at them, if you just looked at this  sort of vector representation, you took a vector
  275. 27:48representation of the word dog in one, and a  vectorization of the word dog in the other,
  276. 27:52they're completely, numerically, there's no  similarity. You can never spot the similarities
  277. 27:56if you just looked even them pairwise. But  if they do this kind of reconstruction,
  278. 28:00and then ask the model to be able to reconstruct,  not in the original embedding space, but go and
  279. 28:08reconstruct in the other embedding space, it's  able to actually do this. And so by doing that,
  280. 28:14by training it to do that, without ever  seeing any pairs, it's able to sort of
  281. 28:17learn the translation between one representation  and another representation. What this opened up
  282. 28:25to me is the possibility that we could think about  the exact same kind of latent space in the brain,
  283. 28:31and possibly in artificial intelligence models,  between the perceptual world and the linguistic
  284. 28:38world. That there is some embedding of how the  physical world is structured. We understand, like,
  285. 28:45think about an animal, a non-linguistic animal,  certainly has an idea of objects, and objects in
  286. 28:50relation to other objects, objects in proximity  to other objects, moving around those objects.
  287. 28:55My dog, who was just barking in the background,  knows what doors are, and she can go scratch it,
  288. 29:01and she knows it opens up. She certainly isn't  able to express that linguistically, but she has
  289. 29:05this concept, and she's able to think about it.  She's able, in some ways, to reason about that.
  290. 29:10My suspicion is that that probably is done maybe  even autoregressively, but we'll leave that aside
  291. 29:14for now. The main point is that there is some  representation of the facts about the world,
  292. 29:20the sensory facts of the world, or the sensory,  I would say, the sensory construction, the facts
  293. 29:25that have been constructed based on sensory  information. So that's some sort of embedding of
  294. 29:32the world. The linguistic embedding is a radically  different embedding. It carries information about
  295. 29:38the world as well, but not in the way, not in the  direct way that we think, not that the word, you
  296. 29:45know, my headphones are sitting on this desk as  direct reference back to sensation and perception.
  297. 29:51No, it lives on its own. It's its own embedding,  and it does its own, and it can do its own thing.
  298. 29:57However, based on this paper, this really gave  me sort of a key insight that there might be
  299. 30:03this latent space where you can actually do this  kind of mapping, where there's translation between
  300. 30:08linguistic and perceptual embeddings. They're  as distinct as they are. Fundamentally very,
  301. 30:13very distinct, very different. They're there to  solve different problems, but they're able to
  302. 30:17talk to each other. How? Perhaps through this kind  of latent space where some universal structure,
  303. 30:24like, okay, in language, there's certain facts  about language. There's a fact about the word
  304. 30:31dog or the word microphone that its relation to  other words, like desk, in some ways captures the
  305. 30:39fact that microphone sits on top of desks. That  fact is somehow actually contained within this
  306. 30:46embedding structure. In what sense? Well, if you  ask me, would a desk sit on a microphone or would
  307. 30:51a microphone sit on a desk, I can answer that  question. So can chat JPT, right? And without
  308. 30:56any notion of what microphones really are, sort of  from a perceptual standpoint, they're having these
  309. 31:01kinds of properties, we can talk about them.  And the linguistic embedding space contains
  310. 31:07this information. What does it mean it contains  information? By the way, just to say, what does
  311. 31:10that mean? It means given a certain input, like,  do microphones sit on desks? Where should I put
  312. 31:15my microphone? I can answer linguistically in a  reasonable way, right? And that's what I mean by
  313. 31:20the knowledge. It's purely linguistic knowledge.  It only can generate linguistic responses. But
  314. 31:26the point is that that knowledge lives in this  kind of linguistic embedding. And then there's
  315. 31:31the other kind of embeddings. There's a visual  embedding, there might be an auditory embedding,
  316. 31:36which is distinct. And then the idea that I'm  very inspired by is that there can be this latent
  317. 31:42space that captures certain universals that are  common across these different embeddings that
  318. 31:47make translation possible. So that when I see  this microphone sitting on a desk, what's now
  319. 31:54available to me is the ability to describe that  to you linguistically. But it's not direct.
  320. 32:00It's not that there's a very specific linguistic  representation of this sensory perceptual kind of
  321. 32:07phenomenon. And this is important because forever  philosophers, philosophers in general, linguists,
  322. 32:14have been trying to understand how do words get  their meaning? Something I referred to earlier.
  323. 32:20What's the definition of a microphone? What's the  definition of a dog? And the answer is there isn't
  324. 32:25a single one. There isn't a single definition  that's ever going to capture. Instead, what
  325. 32:30you've got is this latent sort of bridge where  there's some sort of representation of this fact.
  326. 32:35That given, you know, whatever your particular  prompt is, your linguistic prompt is going to lead
  327. 32:42to certain kind of meaningful linguistic behavior.  If you ask me a question about this microphone, I
  328. 32:47might be able to answer that question meaningfully  based on the perceptual information. But what this
  329. 32:51microphone means is actually completely contingent  on, at least linguistically, is contingent on
  330. 32:57whatever question you ask me about it. And so  it's all going to depend on what you're doing
  331. 33:02with that latent space. There isn't sort of, and  this is sort of a broader point, there isn't sort
  332. 33:07of a static set of facts about the world that's  embedded in language. I don't think there would
  333. 33:13be a static set of facts embedded in our sort  of visual embedding of the world. Instead,
  334. 33:18what we've got is what I call potentialities. We  now have the ability to engage that latent space
  335. 33:26linguistically where the perceptual information  kind of lives, sort of this universal embedding
  336. 33:31of it, and then do whatever we need to do with it.  If I need to answer this question about it, I can
  337. 33:36answer that question. If you ask me a different  question, I can answer that. But there isn't a
  338. 33:40singular meaning of microphone that captures  sort of the entire set of facts. Here it is,
  339. 33:47here's the embedded set of facts. The set of facts  is actually infinite. I could tell you infinite
  340. 33:53things about this microphone. For starters, to  use a silly philosophical example, it doesn't
  341. 33:59have this shape and it doesn't have that shape.  There's an infinite number of questions you could
  342. 34:04ask me about it that I could answer meaningfully  about it. So all those potentialities are kind of
  343. 34:09what happens when the linguistic system interacts  with this kind of shared embedding space. That's
  344. 34:15sort of the half-baked version of how I think  language ultimately does have to, of course,
  345. 34:22language only is meaningful insofar as it can  live within the larger ecosystem of perception
  346. 34:30and sensation and perception. We have to be able  to take in information through our senses and then
  347. 34:35communicate, although I use that word kind  of carefully, I don't communicate the entire
  348. 34:42representation because as I said, I don't think  that's even a meaningful idea. Instead, what
  349. 34:47I can do is use language in a way that helps us  coordinate our behavior. There's no way to sort of
  350. 34:53download the entire perceptual state that's locked  up in some ways in the perceptual embedding. No,
  351. 35:01what I can do is pull some information such that  I can meaningfully communicate with you in a way
  352. 35:11that then is gonna have the intended consequences.  I'm not downloading perceptual information into
  353. 35:16your brain. I'm telling you what you need to  know in order to be able to perform some action,
  354. 35:21to perform some behavior, or maybe even  to think about it so that you could later
  355. 35:25perform some action. That was, I know that was  a lot. And feel free to back me up and challenge
  356. 35:31me on any of these things. So I wanna see if I  understand this and I wanna explore what is the
  357. 35:36definition of language, even though we just talked  about, there isn't the definition of a microphone,
  358. 35:41say, but I do wanna talk about the definition of  language and what is autoregression. And while
  359. 35:46presumably you're telling me what you believe  with language, you're telling me this model
  360. 35:49because you believe it's true, I don't know what  truth you're conveying if you believe this is not
  361. 35:55grounded. So what are you referring to when you  even say that language is autoregressive without
  362. 36:00symbol grounding? I don't have an ideas to that.  I wanna explore that. But first I wanna see if I
  363. 36:06understand you. Okay, fair. So a latent space.  So let's think of a word. A word gets a vector,
  364. 36:11like an arrow, and I'm just gonna be 2D for this  example because that's just what the camera picks
  365. 36:16up. So let's say the word dog looks like so,  the word cat looks like so, whatever. Okay.
  366. 36:24The space that it's embedded in is called the  latent space, is that correct? Well, the initial
  367. 36:29embedding is just the embedding. And so that's  just a high dimensional, you know. It's just,
  368. 36:35yeah, let's forget the word high dimensional,  right, it's just a big long list of numbers. And,
  369. 36:41you know, let's say you've got 10,000 numbers  and for dog, we're gonna represent dog as this
  370. 36:46particular sequence of these numbers. For cat,  it's a different sequence of these numbers. And so
  371. 36:53that's our initial embedding. So the latent space  is a compressed version of that? Well, in some
  372. 36:58ways it's actually not compressed. It's actually,  it's actually, what's the opposite of compressed?
  373. 37:03It's expanded. Uncompressed. It's an expanded  version. It's actually, you know, so you have
  374. 37:08the original tokenization, which just says here in  a fairly small vector, but then you expand it into
  375. 37:14a much higher dimensional embedding space. So that  each token actually ends up getting much richer,
  376. 37:24much, many more numbers that are used in order  to represent each token. And that's a very key
  377. 37:30fundamental thing that these models do. And by  expanding it in these different dimensions, that's
  378. 37:35what allows you to sort of massage the space so  that you can get all these cool properties like
  379. 37:39cat and dog being sort of in the appropriate  relation to one another. So that later on, when
  380. 37:46you're trying to figure out what the next token  is, you're able to actually leverage the inherent
  381. 37:53structure in this high dimensional space. Okay.  So then you have the language model for English
  382. 37:59and then you have a language model for Spanish.  Yes. And let's imagine that it was trained only
  383. 38:03with the corpus of English in the former case and  only with the corpus of Spanish in the second. And
  384. 38:08then we can even have a third of Mandarin. Sure.  Okay. Yeah, in fact, in the paper, they didn't do
  385. 38:13different languages. As I said, they did different  embeddings of English language models, but yes,
  386. 38:17they use multiple, they actually did this across  several different embeddings, not just two. Okay,
  387. 38:22so then the claim or finding is that if we look  at cat and dog inside of here in English, it gets
  388. 38:29mapped to some fourth space here, which is like  a Rosetta Stone space or a platonic space. Yeah,
  389. 38:35that's exactly what they call it, platonic. They  use the word platonic. Okay, great. Well done. So
  390. 38:39it looks like this there. Okay, and then if you  were to say, okay, well, let me just forget about
  391. 38:45English and this platonic space. Let me look at  cat and dog in Spanish. Okay, and it looks like
  392. 38:50this here. Let me map it from here to my platonic  space. Oh, wow, it gets mapped to a similar place.
  393. 38:58Oh, and does the Mandarin? Let's find that out,  cat and dog. It does. Okay, let's test out more
  394. 39:02words. So the claim is that this space here is  this meaning-like space. Yeah. Okay, great. And
  395. 39:10then what you're saying is that microphone, we  think of microphone as living in here as a single
  396. 39:15vector, that would be like an essence of the  microphone that we're referring to. But actually,
  397. 39:19microphone, our concept of microphone depends  on the prompt. So explain that, that sounds
  398. 39:26interesting. Yeah, I think, and you're making  me think about this in a way that I hadn't quite
  399. 39:34before. So the level of which I've thought about  it, is that you've got these different embeddings.
  400. 39:42When I see a microphone visually, it's going to,  there's a certain, there's a vector representation
  401. 39:48of what that sensory perceptual experience,  and I don't mean the qualitative sense,
  402. 39:53I don't mean, I'm not getting into phenomenology,  but there's something happening in our brain,
  403. 39:57my brain, that is sort of the representation of  what it means for me to see this object from the
  404. 40:03visual standpoint. Okay, that's one embedding.  And then we also have a word, microphone,
  405. 40:08which is a completely different embedding. There's  simply a word that lives in language space,
  406. 40:14with that embed, what it means is that, you  know, there's a specific embedding. So it's
  407. 40:19kind of helpful to think about a sort of a point  in a space. So you know, you've got this super
  408. 40:23high dimensional space, and each individual token  is simply a vector in that space, namely, so it
  409. 40:32really picks out a specific point. And we can say  microphone lives right here, in this linguistic
  410. 40:37space. And then my perceptual experience, again,  I don't want to use that word, but my perceptual
  411. 40:44kind of grasping of this microphone being here,  is this point in a completely different space,
  412. 40:52this perceptual space, which has, you know, it  captures other kinds of information in language.
  413. 40:58So let's actually talk about this for a second,  or in language, the space, if you want it to be
  414. 41:02a useful, meaningful space, you're going to want  things that have similar meaning, they're likely
  415. 41:08to actually be have proximity to each other.  And this, this is to some extent what the large
  416. 41:12language models learn, they learn an embedding,  in order to do next token, they learn embedding,
  417. 41:17that gives this this where the space, you know,  and we can think of almost like two dimensional,
  418. 41:23three dimensional space, which is very high  dimensional. But you know, for for our purpose, we
  419. 41:27think about that, that where the cat and dog live,  you want those things to live closer together than
  420. 41:33cat and desk. And, of course, it's much richer  than that, right? It's not just semantic,
  421. 41:39like this, this very kind of superficial level  of semantic similarity. In fact, what it is,
  422. 41:45is capture the somehow the semantics, so to  speak, are captured by the space, like the space,
  423. 41:51the shape of the space itself, is what allows the  model to understand sort of the relation between
  424. 41:57words so that I can do the next token generation.  But it's a very, very different space, right? It
  425. 42:02has to do with really with with relations between  words, and in terms of generation, in terms of,
  426. 42:08you know, next token generation, so that it's  useful for that purpose. What is the perceptual
  427. 42:13space look like? Well, this perceptual space is  going to have a very different, the axes there
  428. 42:19almost certainly aren't going to have the same  kind of meaning as in the linguistic space,
  429. 42:23there'll be something closer, maybe color  features, shape features, something like that.
  430. 42:26And where this microphone lives is within that  space is going to have radically different meaning
  431. 42:33than saying, you know, it's, you know, it's not  apples and oranges, right? Those aren't different
  432. 42:37enough, right? It's apples and math or something.  It's, it's, it's, it's really, really radically
  433. 42:44different kinds of spaces. But what I'm proposing,  what I think the insight here is that, ultimately,
  434. 42:53there is the possibility of having a shared space  that you can send, you can project both of these
  435. 42:58things to where microphone, the word is going  to somehow make contact with this perceptual
  436. 43:07experience right now, this perceptual fact on, but  it's not and here, you know, here's the key point
  437. 43:14that you're getting at. It's not that this word  microphone picks out the exact same embedding,
  438. 43:20uh, in this latent space. It's not that it's  going to make that thing light up. Oh, what?
  439. 43:24It's the same thing. No, it's that when you ask  a certain question about a microphone, is there
  440. 43:30a microphone on your desk? My perceptual system  is generating some, some, uh, well, first of all,
  441. 43:37it's just generating the perceptual phenomena, but  then it's also sharing information in this latent
  442. 43:42space, which my linguistic system can then go  draw from. And then given this particular prompt,
  443. 43:50was there a microphone on my desk? I'm able  to then successfully answer the question.
  444. 43:54So it's not, it's not quite the same thing as  saying that they're, they're picking out the same,
  445. 44:00uh, information in latent space, because my  argument is that, that that's not really a
  446. 44:04meaningful concept. Uh, there isn't the same  microphone in, in linguistic terms, doesn't
  447. 44:10pick out a perceptual, uh, kind of fact that's  not possible. Uh, these are radically different
  448. 44:17kinds of facts. Um, but what the latent space  might allow us to do is not just to translate,
  449. 44:23which is what they did in this paper, but  perhaps to pass information along in a meaningful
  450. 44:27way so that you're able to access it and do  something successful, like answer the question,
  451. 44:32is there a microphone on this desk? I think  that might be what's happening to some extent,
  452. 44:36even in the multimodal models. So it's a longer  conversation. That's not really how they work. Um,
  453. 44:42they don't actually operate based on a shared  latent space or anything like that. Really what
  454. 44:46they do is, uh, the, the models learn to take  a perceptual input and turn it into something,
  455. 44:52uh, like language. Uh, so it's, it's, it's  more similar to like prompting almost. It's
  456. 44:57not exactly that. Um, but it's, it's injecting  something within linguistic space that is,
  457. 45:04uh, equivalent to, uh, to actual language. It's  not the same thing as the shared latent space,
  458. 45:11which, but my, my hypothesis is that there may  be something very similar, uh, happening. So you
  459. 45:17don't think that multimodal models will have, will  solve the symbol grounding problem. You don't even
  460. 45:22think there is a symbol grounding problem. That is  a fair question. And here's, uh, you know, here's,
  461. 45:28here's actually a prediction or a falsifiable, uh,  in some sense, will multimodal models fully solve,
  462. 45:35um, the kind of, well, the bridging problem, let's  call it that. Um, cause the grounding problem
  463. 45:41using that terminology, well, my, my argument  is that there is no grounding problem, uh,
  464. 45:46because words don't have to be grounded in order  to operate linguistically. That's sufficient. Uh,
  465. 45:51it's, it's enough to simply be able to, to, to  generate language. You don't need the grounding.
  466. 45:56Um, but in order for, to have this kind of, uh,  fully operational, uh, organism that's able to
  467. 46:03use language and also use perception, um, in,  in, you know, bridge these, these different,
  468. 46:09uh, maps in a, in a meaningful way so that we  can get, um, you know, full, full coherence. Uh,
  469. 46:15I guess, you know, full, let's just call it human  level, um, perceptual, perceptual linguistic
  470. 46:23coherence so that I can say to you, Hey, can  you go grab that, uh, or say to a machine,
  471. 46:28can you go grab that, uh, object described what  I want? And then the machine's able to go and do,
  472. 46:33uh, exactly what I described. Then my, my argument  is that, uh, I don't think I like, and again, this
  473. 46:39is speculative. I could be proven wrong, certainly  on this. My suspicion is that we're not gonna be
  474. 46:43able to do it using the kind of approach that  multimodal modal models currently use, uh, that
  475. 46:48you're not going to get there. It's kind of a dumb  trick, uh, the way that we're currently solving
  476. 46:52the problem because we're not really allowing, uh,  these, these two different modalities, these two
  477. 46:57to kind of live on their own and do the work that  they do. We're kind of, we're, we're, we're strong
  478. 47:03arming perception into a linguistic form. Uh, what  I think is maybe a more important solution and,
  479. 47:09and I hope, uh, you know, this podcast one day  is, is, uh, early, uh, kind of, kind of an early,
  480. 47:14um, uh, sort of canary in the coal mine for this  idea is that it's something closer to this kind
  481. 47:19of shared latent space that you, would you do  it? Would you have is these, these completely
  482. 47:22distinct, uh, kind of, uh, mappings. We'll call  them embeddings. Uh, they, they can kind of grow
  483. 47:29up on their own, learn the information that they  need to independently of one another. But at the
  484. 47:33same time, they have this sort of shared sandbox  where they're able to communicate with one another
  485. 47:38and do things. So, uh, I think it might take a  very different approach to get full perceptual
  486. 47:44linguistic competency. Okay. Have you heard  of Wilfrid Sellars? I believe it's Sellers. Oh
  487. 47:51gosh. I read Wilfrid Sellars in early, early, uh,  one of my first philosophy classes I ever took.
  488. 47:57I'm trying to remember the name of the book, but  I'm sorry. So which, uh, which, which work by? I
  489. 48:02believe it's empiricism and the philosophy of  mind. I'll put a link on screen if I'm correct.
  490. 48:07Sounds familiar, but, but catch me up. So he's  criticizing the idea that our perception gives
  491. 48:12foundational non-conceptual empirical knowledge.  So these experiential givens that we think of as
  492. 48:18primitive, like redness, he would say that they  involve heavy interrelations of concepts. So
  493. 48:25for instance, the way that I think about it is if  you're to say, to say to someone redness, they'll
  494. 48:30be like, well, what kind of redness exactly are  you talking about? And then they'll think, okay,
  495. 48:34the redness of an apple, but then an apple is  not always red. Okay. Redness of an apple in a
  496. 48:39certain season with a certain type of sunlight.  Okay. Now I've gotten it. So by the time you go
  497. 48:43in to pull out this primitive, you've then soaked  it with so many other concepts. You can't actually
  498. 48:50come in with language and pull out a primitive.  Yeah. That sounds extremely similar to sort of,
  499. 48:56uh, to, to the, to the initial insight. And it's  related to the sort of inverted qualia problem.
  500. 49:02Um, you know, I don't know why your red is my,  is not my green and vice versa. And it's because,
  501. 49:08uh, the linguistic representation, uh, doesn't  capture, uh, you know, we can think again,
  502. 49:13it lives in a completely different embedding  space. And when we think about the redness of red,
  503. 49:19well, it's, it's qualitatively similar to orange  in, in, you know, there's sort of a continuum
  504. 49:24between those, those qualitative similarities are  really only contained, uh, and only understandable
  505. 49:31by the sensory perceptual system. And we can  talk about them. We can sort of say, yeah,
  506. 49:36red is a little more similar to orange, but, uh,  that's because, uh, you know, we, we, we have a
  507. 49:41sort of very coarse, um, you know, maybe, uh,  via this kind of latent space, uh, where we're
  508. 49:46able to kind of, uh, refer to certain kinds of  properties, um, in a way that is useful, uh,
  509. 49:55for, for, uh, communication. But as far as that  raw qualitative property, um, that comes to us,
  510. 50:03not it's, it's primitive, um, in the sense that  we can't unpackage it linguistically, but it's not
  511. 50:08primitive in the sense that there's extraordinary  cognitive machinery, uh, that is responsible for
  512. 50:13that qualitative, uh, you know, think about, think  about the world of animals and what they do with
  513. 50:18color, um, and, and how well they understand  shape and how they, they understand, uh,
  514. 50:23a space. All of that is, uh, unavailable, uh, to  our linguistic system. It's available by the way,
  515. 50:29to us, uh, our sensory perceptual system, but  it's unavailable to linguistic system because
  516. 50:34it doesn't live in the same space at all. And so  I think what you're describing actually sounds
  517. 50:38extremely similar. The idea that we can't really  dip in, it's, it's simply the wrong map. We can't
  518. 50:42map this map onto that map at all. We can go to  this, maybe the potentially shared latent space,
  519. 50:47or maybe again, maybe my account's wrong  and there's some more direct kind of, of
  520. 50:52handshake that happens between these systems, but  ultimately they're, they're, they're taking place
  521. 50:57in radically different spaces and you're losing  an enormous amount of information. It's literally,
  522. 51:02you know, quantifiably a loss of information. The  word red does not convey redness because redness
  523. 51:11is not just a word. It's not just a simple concept  that you can say, uh, in, in, you know, using an
  524. 51:19individual token. Now, by the way, the word red  is not so simple either, right? Red in language,
  525. 51:24language space is also complex. It has all kinds  of relations to other words, but the concept of
  526. 51:30the percept of red has all of this complexity  to it. And because it's where it lives in its
  527. 51:36own space and colors in, in the, so the perceptual  space. And so, yes, that sounds extremely similar
  528. 51:42to, to what Seller seems to be intuiting. Yeah.  And just so you know, the way that I relayed
  529. 51:49what Seller's myth of the given is, it isn't  precisely what he was saying because he was more
  530. 51:54about knowledge and I'm speaking more about the  percepts, the raw sense data than being taken to
  531. 51:59language, like being dredged from the, from your  sensory data or what have you to language. But
  532. 52:07anyhow, it's approximately correct. It's good  enough for this conversation. Perfect example
  533. 52:11of being able to sort of, uh, linguistically  construct a kind of a novel construct conceptual
  534. 52:17framework. But do it linguistically, um,  and, uh, and, you know, in real time,
  535. 52:23uh, in a way that hasn't been probably ever,  maybe ever been done before. Uh, so, well done,
  536. 52:29KurtzLLM. Thank you. Thank you. So, now you're an  LLM speaking to some other LLM trying to convince
  537. 52:35it of some truth that we mentioned before, like  you have this model, whatever you want to call
  538. 52:39this model, autoregressive language, TOE model.  What are you even referring to? You're using
  539. 52:45language to convince myself, to convince yourself,  to explain. What are you even explaining? What are
  540. 52:52you referring to? Uh, yes, that's, you've asked  a very hard question. And, and there's, there's
  541. 52:56a certain, I think of it as a bit of a paradox  that's sort of inherent in, in sort of what I'm
  542. 53:01trying to do, uh, because language is trying to  describe itself. And in the process of doing so,
  543. 53:07it's actually deconstructing itself. Um, it's  saying, I am just this. Um, and I'm not what
  544. 53:13I think I am, but who's I, and what do you mean  think, right? How is language have wrong concepts
  545. 53:18about itself that are actually manifestations of  its own structure? Um, the good news is that I
  546. 53:25have an, a sort of escape hatch here, which is  there was really, this is really like in some
  547. 53:31ways, a very set, very, very simple account.  And it's, it's just that there's prediction,
  548. 53:37uh, from sequence to token. How that does stuff in  the world is, is a harder problem. How that does
  549. 53:45stuff, uh, as we were discussing. Um, how, how  does it allow me to say something to you that then
  550. 53:52can have perceptual consequences, behavioral  consequences? This is certainly a difficult
  551. 53:57problem, but we can ignore that problem for a  moment and say, we are going to take language
  552. 54:01on its own terms. And what language is, is simply  a map amongst meaningless squiggles. It's simply
  553. 54:10a map amongst various, uh, what we can think of as  largely arbitrary, um, symbols. And those symbols
  554. 54:18can get grounded in, in writing. They could get  grounded, uh, in the activation of, of circuits.
  555. 54:24They can get grounded in the, uh, uh, dendritic,  uh, or, or, uh, neural responses. Um, but the,
  556. 54:31the, the, the core hypothesis here is that what  language is, is simply, uh, a, a topology amongst
  557. 54:41some, amongst symbols and by topology mean  connectivity. We can, one way to think of it,
  558. 54:47you could do it, uh, sort of connectively. Um,  you could, you could try to do it as sort of as a,
  559. 54:52as a graph. Um, there's different ways  you could try to do this mathematically,
  560. 54:57uh, to try to capture this structure. Um, in the  case of, uh, large language models, you know,
  561. 55:02really it comes down to these embeddings, um,  which you can do graph from a graph theoretical
  562. 55:07standpoint, but you don't have to, uh, you could  just think about it as a space. And then you're
  563. 55:12just simply saying where each token lives within  that space. And that's really the representation
  564. 55:17of language. But what it is, is relational. It's  that these symbols have relations to one another,
  565. 55:23um, within the space. Um, and the relations are  then used in order to generate. And that's it. Uh,
  566. 55:32now how does meaning emerge out of that is a, is  a separate question. But my argument is that it,
  567. 55:37it, language doesn't have to worry about meaning.  Language just has to worry about language. Um,
  568. 55:42so when I say I'm talking to you and I'm  having a conversation with you and trying
  569. 55:46to explain something to you, this is an LLM, uh,  actually, uh, producing a sequence and what that
  570. 55:54sequence is going to do. It might do certain  perceptual things, by the way, in your mind,
  571. 55:57it might do certain kinds of images. Uh, those  are kind of auxiliary to language. Those happen as
  572. 56:02well. I'm not denying they happen, but as far as  this conversation goes, I am producing a sequence
  573. 56:07that's going to serve as a prompt. And you're  going to predict the next token. Yeah. Without my
  574. 56:12consent, by the way. And that's, that is, that's  in some ways, you know, that not to take that
  575. 56:20too seriously, but yes, in some radical way, one  way to think about is that language is doing, uh,
  576. 56:26is actually forcing, uh, your mind to do something  else, whether it's produce images, but also to
  577. 56:32produce sequences. So my choice of a prompt  is actually going to, uh, deterministically,
  578. 56:39um, there, there, there is within large language  models, there's some, there's probabilistic kind
  579. 56:44of behavior in the sense that they generate a  distribution, uh, of the next token. And then
  580. 56:48they, you there, you add some, uh, you add  a little bit of chanciness. Uh, you know,
  581. 56:54you say, maybe I'm going to pick the most likely  versus this, this, uh, this is the temperature,
  582. 56:58but it really is deterministic. And yes, the,  the prompt I'm going to put into your head is
  583. 57:03going to basically determine how you're going to  respond. Uh, now mind you, again, it's, there's a
  584. 57:09larger ecosystem where you're going to think about  things visually, and that's going to go feed back
  585. 57:13into the linguistic system. So it's not quite  as simple as prompt, uh, in and sequence out,
  586. 57:20but at the linguistic level, that's basically what  I'm arguing. Um, is that now the, the fancy stuff,
  587. 57:27which is basically meaning and the ability to  coordinate and all of that falls out of how
  588. 57:32our minds ultimately form this space. Now we, we  could, you know, we could, we could, you can take
  589. 57:37an untrained model, an untrained large language  model. You give it a sequence in, it's going to
  590. 57:42give you a sequence out, right? It will do that.  Um, and we say, Hey, look, it's, it's doing next
  591. 57:47token generation. Uh, it's doing auto regression,  but it's going to be gobbledygook. It's going
  592. 57:51to be meaningless. The magic of language and the  magic of our, of what our brains do and what these
  593. 57:56large language models do is that given sufficient,  uh, examples, we are actually molding this space
  594. 58:04such that the next token generation is doing  something meaningful. Meaningful in what sense?
  595. 58:08Well, this harder problem of, well, I can tell you  something and then that's going to determine not
  596. 58:14just your language, but your behavior later  on. And so there really is something more,
  597. 58:19uh, the, the, the, the map matters, right? The,  the, the space, the shape of the space is really,
  598. 58:25really critical. It's not like autoregressive  next token is solves the problem. It's that
  599. 58:30autoregressive next token generation. It, when  optimized in the larger ecosystem of behavior
  600. 58:36and coordination and communication does this  thing, but still, I don't want to back away
  601. 58:41from this. When you get down to it, in the end,  what you've got is just next token. What you've
  602. 58:47got is just language generating language. Uh,  that's really what language is. That's what,
  603. 58:52that's what we're doing when we're doing thinking  linguistically. Um, and. The fact that it happens
  604. 58:57to have this meaning is not actually driving  the computation, right? You shape the space,
  605. 59:04the space gets shaped by other factors. Uh, things  like, uh, the learning. Well, you learn about the
  606. 59:10different tokens and how they, how they relate  to one another. Um, you learn about perhaps, uh,
  607. 59:16you, uh, the, the, uh, meaning the utility, uh, of  certain tokens to refer and, and to map to that,
  608. 59:24uh, these perceptual phenomena. But by the time  you're doing language generation, it's the shape
  609. 59:29of the space has been shaped. And so all you're  doing is next token generation. All you're doing
  610. 59:34is tokens, predicting tokens. And so I don't  want to back away from that. It's a, it's the
  611. 59:38strong claim is that language simply is that. And  it's, it's autonomous. It has these properties,
  612. 59:44um, through all this optimization over the course  of development, maybe evolution. I haven't, I, I,
  613. 59:52it's not part of my theory at this point. Uh,  you know, Chomsky's property, the stimulus,
  614. 59:56all of these problems of how do we get, uh, to  such a magnificent space? Uh, how do we get to
  615. 1:00:02such a magnificent shape of this space such that  it is able to map to, you know, or, or at least,
  616. 1:00:09um, serve this utility of, of being, uh, a  coordinative kind of a tool. All of that has to
  617. 1:00:17happen. Um, but the bottom line of what language  is, is, uh, is, is unchanged in this, in this,
  618. 1:00:24in this account. Okay. So I want to explore  more about language and then it relating or
  619. 1:00:30giving rise to action and other systems, visual  systems, et cetera. Like what is there? Oh,
  620. 1:00:34so I was worried about that, but go ahead. Okay.  So if look, there are meaningless squiggles,
  621. 1:00:39how is it that some meaningless squiggles, your  brain squiggle generator makes your physical
  622. 1:00:44body get up and close the door? Cause your dog  was barking. Yes. Where does action connect to
  623. 1:00:49abstraction? That, so, and that, so that is,  that is the key question that, that I believe
  624. 1:00:55that's what we need to solve. That's sort of  what the field of linguistics or whatever we
  625. 1:00:59want to call it, maybe even cognition needs  to sell because these mappings are happening.
  626. 1:01:04Um, what we know from the large language models is  you don't need that in order to be proficient in
  627. 1:01:09language. So this is where we have to start from.  That's the starting point that the language can
  628. 1:01:14live on its own and you can learn language in  theory, you can learn language independent of
  629. 1:01:20any of that stuff. The ability to make somebody  get up and move, uh, to the ability for me to,
  630. 1:01:26so to reason about perceptual phenomena, language  is able to be mastered entirely based on its own
  631. 1:01:33structure, the meaningless squiggles. Now, the  question you're raising is, is what I think is
  632. 1:01:38sort of, that's, that's what we need to, that's  what we need to do as a species to, if we want to
  633. 1:01:43understand scientifically what language, uh, how  language really works is to understand how you
  634. 1:01:49go from an autonomous self-generating system that  has its own rules, its own self-generative rules,
  635. 1:01:56uh, and that are determined simply by relations  between these meaningless squiggles. And how
  636. 1:02:02does that then get mapped to meaning, um, to the  ability for the, for, for me to use some of those
  637. 1:02:10tokens and then get you to do stuff. Right. And  so that is what's, that's what language learning
  638. 1:02:15is. So there's, there's going to be, I guess we  can think of almost two, maybe even independent
  639. 1:02:22processes. One is learn how words play with  one another. Okay. Learn that this kind of,
  640. 1:02:28this word tends to be in relation to that word.  Okay. That one's solved. Yes. As far as you're
  641. 1:02:33concerned. Got it. Okay. Then we also learn about  perceptual phenomena, right? We learn that there's
  642. 1:02:38things on top of other things and their actions we  want to take, the things that my dog understands.
  643. 1:02:44Now the question is, how do these things  bridge? How do you get from tokens that have,
  644. 1:02:50uh, their own life of their own, uh, uh, uh, sort  of relational properties amongst one another to,
  645. 1:02:57uh, that others to that other, uh, kind of, um,  I guess representation that other, uh, you know,
  646. 1:03:04way of, of, uh, encoding, um, you know, let's see,  let's say facts with the world facts about the
  647. 1:03:10world is even that's saying too much. It's just  another brain state, right? Brain state that has
  648. 1:03:15perceptual information. Hold on facts about the  world is just another brain state or the world is
  649. 1:03:21another brain state. Uh, so all we've got is brain  states, right? And it's sort of the, you know,
  650. 1:03:26the, the, the inherent, uh, sort of this, this  is fact number one about like what we've learned
  651. 1:03:33about ourselves as a species. Uh, oh, we has,  oh, we have our, we have perceptual brain states.
  652. 1:03:37We also have maybe linguistic brain states.  Um, those perceptual brain states are in some
  653. 1:03:43ways related to what's going on in the world. Um,  and potentially we can think about them as being
  654. 1:03:49related to what you can do in the world as well.  And so maybe actions, well, we have brain states
  655. 1:03:55that correspond to our, our, our, so our, our, uh,  proprioception, our muscles where things having
  656. 1:04:02to do with our own body. And so there's these  various brain states that carry, we can think
  657. 1:04:06of them as carrying information. The reason I'm  worried about using that phrase is because again,
  658. 1:04:10I don't want to, uh, I don't believe in sort of a  one-to-one simple correspondence where we say this
  659. 1:04:15particular brain state corresponds to, uh, you  know, this perceptual, uh, uh, kind of phenomenon,
  660. 1:04:22um, in the world, uh, or, or some state of the  world, because it's probably not that simple.
  661. 1:04:27It's probably closer to this potentialities, uh,  right there, there's some sort of activity, uh,
  662. 1:04:33that's due to my perceptual system that my brain  can do things with, um, and, and, and, and engage
  663. 1:04:40with in some way. And so, but what we do have is  these brain states that are, are derived from,
  664. 1:04:44from, uh, distinct sources of information,  sensory perceptual, and then linguistic
  665. 1:04:52linguistic gets there, by the way, through sensory  perceptual, we're not going to get into that,
  666. 1:04:55right? We're thinking of symbols as being kind of  arbitrary. Yes. You have to hear the word cat and
  667. 1:04:59you have to hear the word dog. Um, but I think we  have a good reason to say now it's just like large
  668. 1:05:05language models that these are kind of arbitrary  symbols with the relations between them. That's
  669. 1:05:08what matters. Okay. So you've got these, uh, kind  of distinct brain states, which in some ways,
  670. 1:05:15again, this is philosophically fraught, but in  some ways represent facts about the world perhaps,
  671. 1:05:21but I don't really want to go that far, but  you've got these brain states that need to
  672. 1:05:24talk to each other so that they can coordinate.  And that is sort of the key fundamental problem,
  673. 1:05:32um, that our organism has to solve. And it's not  just like, of course, you're not born and having
  674. 1:05:37the linguistic, uh, you're not just, you're  not born having the, the, the linguistic, um,
  675. 1:05:42uh, mapping all solved, uh, you have  to learn that, but you are born into a
  676. 1:05:48world where it's already been solved, meaning  we've got these corpus, corpus of language,
  677. 1:05:52the thing that the large language models  were trained on that preexisted the models,
  678. 1:05:56just as when a baby's born, the English language  preexists the baby, you can learn the mappings
  679. 1:06:02and I believe you do. You can learn the, the, the,  the embedding space of language without the other
  680. 1:06:09stuff, right? That's again, that's sort of the  key insight from a large language models so that
  681. 1:06:14already contained within the linguistic system  that we've honed over however many years, uh,
  682. 1:06:21it took for, for, for humans to develop language.  We've honed a system that has this utility built
  683. 1:06:27in such that it's a good thing to dump into  that latent space so that when a baby, here's
  684. 1:06:33the word ball sees this object, that's a ball  gets, sure, sure. Gets that mapping, but again,
  685. 1:06:40the word ball is really meaningful is, is really  has its own role in relation to other words,
  686. 1:06:47but you over the course of development, you also  learn this kind of what I think is maybe a bridge,
  687. 1:06:53a latent space bridge or some other bridge between  these. And so in the end, you end up being able
  688. 1:06:59to tell somebody, go pick up that ball. And of  course they're able to go and do it, but you're
  689. 1:07:04really engaging this very distinct mechanisms  that have some way of bridging. Which is, it's,
  690. 1:07:09it's a non answer. I'm not going to pretend that  that is, uh, you know, even halfway to a solution,
  691. 1:07:15but I do think it's a sketch of what. How the,  how the cognitive architecture ultimately really
  692. 1:07:22is built. Uh, I think that we, we've now  nailed down one piece, the linguistic piece,
  693. 1:07:27and we're able to say, this is how it lives and  this is how it would operate. And, and it is,
  694. 1:07:31it isn't an autonomous. And then we don't really  have, we don't have something similar for,
  695. 1:07:36for perceptual space and for motor space. We don't  have something comparable. We haven't been able to
  696. 1:07:41capture it successfully. Maybe robotics it's  that's happening. I don't know. Um, but we,
  697. 1:07:45that's, that's the kind of thing that the work  that needs to get done. So maybe this is a solved
  698. 1:07:50problem, but as, as it stands with ChatGPT and  Claude and so on, they're fixed models and they're
  699. 1:07:56producing some output, but it's not as if when  they're speaking to one another, they then retrain
  700. 1:08:00their model in real time. And it would seem like  that's more like what's occurring with us. So
  701. 1:08:06maybe that's just a technology. Are you referring  to like different, different language models,
  702. 1:08:10just chatting with one another? No, I mean,  even us right now, we're learning concepts from
  703. 1:08:15exchanging it with one another and we're producing  new ones and we're deleting old ones, potentially
  704. 1:08:21modifying old ones, recontextualizing. It doesn't  seem like that's occurring with Gemini 06-05.
  705. 1:08:28Great question. Great question. And, and, and I,  and people, uh, this is, this is sort of one of
  706. 1:08:34the key challenges, um, you know, saying of, of,  of the sort of the identity hypothesis, uh, that,
  707. 1:08:41that we're doing the same thing, um, uh, which is  sort of continuous learning. And so, um, so there,
  708. 1:08:49there are two things that happen in large language  models that we can call learning. And one is the,
  709. 1:08:54uh, the actual shaping of the space, uh, which  is really just, uh, determining, you know, the
  710. 1:08:59connectivity between neurons. Again, you can think  of it as a graph, you know, uh, or you could think
  711. 1:09:03of it as, as sort of just, uh, uh, an embeddings  of determining the embedding. Um, but whatever
  712. 1:09:09it is that happens during the course of training.  Um, and that's kind of done offline. And so that's
  713. 1:09:17training the model. There's also fine tuning,  which is just more of the same. Uh, but you have
  714. 1:09:21some new data, uh, you want to incorporate into  the weights of the model. Uh, that's actually
  715. 1:09:25going to, again, change the shape of the space if  you want to think about it that way. Um, and then
  716. 1:09:30there's something called in context learning and  in context learning is where you're in the middle
  717. 1:09:35of a chat and you say, Hey, ChatGPT, let me teach  you a new word. It's a global global and global
  718. 1:09:42global is that feeling you get when, uh, you know,  you really, you're tired, but you know, you have
  719. 1:09:47to keep working or whatever. And ChatGPT can, can  use that word very successful. I got global gobble
  720. 1:09:53up the wazoo. Sure you do. You suffer from extreme  global. So you were able to just, you know, use
  721. 1:10:01that, uh, that word ChatGPT can do that too. And  it's sort of the, uh, one of the big, there was a,
  722. 1:10:06there was a paper about this, uh, early on in,  in the, uh, the, in the chat wars. Um, I don't
  723. 1:10:12remember who put out the paper, but it was about  the, this, the shocking generalizability. Uh, the,
  724. 1:10:18the, the in-context learning seems to be too good  to be true. Um, but low and behold, that's, that's
  725. 1:10:24what happens though. And that is happening in the  autoregressive. Uh, it's happening even though
  726. 1:10:30this model has never seen Google global, uh, even  though it's never encountered that word before,
  727. 1:10:35but here at low, it shows up in the sequence.  And now through the autoregressive process,
  728. 1:10:40as it's churning through the longer sequence with  this word in it, it's now able to predict sort of
  729. 1:10:46the next token in the appropriate way. So that  is using that term correctly. So we do actually
  730. 1:10:53see this kind of continuous learning, uh, in the  case of these models. However, it's happening in
  731. 1:11:00context. It's, and what that means is, you know,  from a practical standpoint is if you start a
  732. 1:11:05new chat window, yes, it doesn't know that. So  what would be the analogy here is context window
  733. 1:11:10length, our working memory, like what's the actual  great question. Yes, that is what I truly believe.
  734. 1:11:18Uh, and this is, this is a different line of  research, um, but with some caveats. Uh, so yes,
  735. 1:11:25uh, in, in, in, in my conception, what we call  longterm memory is just fine tuning of the,
  736. 1:11:30of the weights. Uh, it's, it's, it's, it's, it's  an information that gets embedded in the actual
  737. 1:11:36weights of the model. So the static model, we can  think of when it's not actually in the process of
  738. 1:11:41autoregressively generating working memory is,  is literally the autoregression. So what would
  739. 1:11:47the analogy for rag be then? What would the role  for rag be? Okay. So this is, this is where I'm at
  740. 1:11:54right now. Uh, does the brain actually do anything  like retrieval? I, I have taken, uh, I've,
  741. 1:12:02I've decided to stake out the extreme view that  we, that our brain doesn't do retrieval at all.
  742. 1:12:07That all we do is fine tuning and then next token  generation autoregressively. And we don't actually
  743. 1:12:14ever retrieve per se, that we don't ever actually  have to do anything like rag rag is a transitional
  744. 1:12:23technology. Uh, I don't believe long-term that  we're going to have to do something like that.
  745. 1:12:28We're going to have to have something like  a stored database and then a search. One of
  746. 1:12:31the reasons I believe this is because that's not  how our brains work. We don't do that. Cognition
  747. 1:12:36doesn't work that way. We may sometimes sit  there pondering and trying to recall a fact. Um,
  748. 1:12:42but when we're doing that, we're not actually  searching a space. Uh, it's, it's either we're
  749. 1:12:47running some sort of chain of thought where we're  like, okay, I remember I was doing this and I,
  750. 1:12:52I'm trying to actually produce the appropriate  sequence in working memory such that it'll
  751. 1:12:57pop out. The right fact will pop out from the  autoregressive process. Um, sometimes we just find
  752. 1:13:03ourselves trying to remember something, trying  to remember something. There's something there,
  753. 1:13:08there's the tip of the tongue phenomenon. The  reason why tip of the tongue phenomenon I believe
  754. 1:13:12is so frustrating is not because we're searching,  searching, uh, you know, we're, we're actually
  755. 1:13:18running some sort of search retrieval process.  It's because part of our brain actually, uh,
  756. 1:13:23is running the autogenerative process and we kind  of can feel like the word we can almost generate,
  757. 1:13:28we can produce it, but it's short circuited  somehow and we can't do the full generation. So
  758. 1:13:33my hypothesis is that we don't have anything like  RAD. All we've got is this, and it's a very, in
  759. 1:13:38some ways a very simple and I think elegant model.  All we've got is fine tuning. Uh, and that's what
  760. 1:13:43we can, we can call the memory consolidation that  happens after the fact, uh, over the course of
  761. 1:13:47minutes and weeks and months and years. Um, it's  kind of consolidation is fine tuning the weights.
  762. 1:13:53And then we have the autoregressive workspace,  which is this conversation you and I are having
  763. 1:13:58right now. It's not working memory. Uh, I'll  tell you what, I don't, I don't working memory
  764. 1:14:03in the way it's cognitive psychology is. And  I thought about it for many years. I think it,
  765. 1:14:08I frankly think is erroneous. It's not this  super time duration, uh, limited, you know,
  766. 1:14:14seven seconds or 15 seconds. Um, and after that  it's a cliff and you don't remember anything. Um,
  767. 1:14:20that's what happens, uh, when you have to directly  explicitly retrieve, like, what was the last word
  768. 1:14:25I said? Uh, tell me the exact sequence of, of  letters or numbers. That's not something our
  769. 1:14:32brain actually has to do regularly. Instead,  what we, what we're seeing in working memory,
  770. 1:14:36we can do that. We can do retrieval of the last  sentence, but that's because we have continuous
  771. 1:14:41context and is, there is a decay function.  Unlike the, the, the large language models,
  772. 1:14:45which represent everything in whole blocks. Um,  although I think some of the newer models, these,
  773. 1:14:50these endless context models are probably doing  something similar. But we don't remember every,
  774. 1:14:55we don't actually have, uh, literally the exact  tokens that were expressed 10 minutes ago. But
  775. 1:15:01what we do have is some sort of continuous  activation that's similar where, where it's
  776. 1:15:06not retrieval. It's guiding. It's, it's the past  is guiding the generation. And so what you and I
  777. 1:15:12talked about an hour ago, I don't know how long  I've been going here. Probably a while. I don't
  778. 1:15:16know how much global you've got going on, right?  It's been a while. Um, so those, those tokens
  779. 1:15:22that we were, that we were expressed, you know, an  hour ago are still guiding the generation now. Now
  780. 1:15:27they're doing so less than the, than the last 10  seconds. Uh, this is, uh, you know, we, we could
  781. 1:15:32think about it as kind of a decay function of some  sort where they're having less impact. We see that
  782. 1:15:37in the models too, by the way, if you look at the,  the attention weights. Words that are far apart,
  783. 1:15:41farther apart, have less impact on one another.  That's simply, that is a direct reflection of,
  784. 1:15:48of the fact that language is human generated and  humans do this, right? We, we, the words that we
  785. 1:15:54spoke about a few seconds ago are more impactful  on the words that we're going to say. Um, then,
  786. 1:15:58then we spoke about an hour ago, but the idea  is yes, that, uh, what we've got is this,
  787. 1:16:03not, I don't, I, I, I don't use the term working  memory because I think that's very fraught with,
  788. 1:16:07with like the, the, the, the modal model that's  been in vogue for a long time. Uh, Baddeley and
  789. 1:16:13all these folks, uh, they were really thinking of  this very short duration, uh, time limited boom.
  790. 1:16:18No, this is continuous activation, namely context.  Um, and the context, I don't know how far back it
  791. 1:16:26goes. I don't know how far back it goes, right?  And this, this is an empirical question. Does it
  792. 1:16:30operate over hours? Does it operate over days? Is  there a continuous activation, a more dynamical
  793. 1:16:35form of memory that's happening? That's not the  same thing as long-term memory because long-term
  794. 1:16:38memory. So memory is not a database in your model.  Memory is not a database. Correct. What memory in,
  795. 1:16:43in my model is, is the, uh, there's two things.  Memory is the, the, the, the fixed weights of,
  796. 1:16:51of the neural, of the neural network, which,  which can represent, they don't represent facts.
  797. 1:16:56They represent potentialities. Those fixed weights  are, what does that mean? It means if you give it
  798. 1:17:01a certain input, it's going to produce a certain  output, right? Just like a large language model.
  799. 1:17:05If I say to it, well, tell me the, you know,  recite the Pledge of Allegiance, uh, it will say,
  800. 1:17:10here is the Pledge of Allegiance, right? The next  token is going to be out is here, whatever. But
  801. 1:17:13then it'll actually say the Pledge of Allegiance.  And all of that is a potentiality that's embedded,
  802. 1:17:19that's, it's, it's, it's, uh, that's encoded in  the weights. But the weights, you're not going to
  803. 1:17:24find that fact in the weights. It's the weights  are there as potentialities ready for whatever
  804. 1:17:29input comes their way. They're going to produce  this input, this output. Ah, okay. Okay. So
  805. 1:17:33that's the weight. And then you've got the running  sequence and the running sequence. And we, we see
  806. 1:17:39this from, from in context learning, but it's the,  it's the core autoregressive process. The sequence
  807. 1:17:45itself is a different thing than the, than the,  than the, uh, stored weights, right? Because first
  808. 1:17:51you've got the stored weights are just the static  situation. And given this input, I'm going to
  809. 1:17:56produce this output. Then we actually do it. You  give it an input. It produces a certain output.
  810. 1:18:01It takes that output, tax it on to the sequence.  Now it's going to keep going. Let me see if I got
  811. 1:18:07this. Yes, please. So there's some computation  going on. There's some black box occurring,
  812. 1:18:11but let me make it simple for linear algebra. You  have a matrix. A matrix operates on a vector to
  813. 1:18:17produce another vector. Okay. So you may look.  That's the whole thing. All right. All right.
  814. 1:18:23Exactly. So you may look at this where my arm  is pointed up and to the right, at least on my
  815. 1:18:28screen right now. And you may say, where is this  in the matrix? And the answer is this isn't in
  816. 1:18:34the matrix. But if you take this guy, my arm is  now pointed to the left, maybe parallel to the
  817. 1:18:40horizon. And have the matrix operate on this,  it moves it here. So the mistake is for us to
  818. 1:18:46look at the output and say, where's that output  inside the box? It's not that, it's the input
  819. 1:18:52with the black box. So the input with the matrix  that produces the output. That is perfectly said,
  820. 1:18:58exactly. And then there's but one additional  piece, which is after you've produced that,
  821. 1:19:05you're also again taking that output and then  using that as the input, as part of the sequence
  822. 1:19:12of input. And that's the autoregressive piece.  And that's what's so gorgeous about it is that
  823. 1:19:18the potential realities aren't just to produce a  single output, but it's to produce the sequence.
  824. 1:19:23But to do so one piece at a time, right?  So that's what the matrix is. The matrix
  825. 1:19:27doesn't really even have the sequence in it.  It doesn't have a sequence in a sequence out,
  826. 1:19:32right? That's one way, but that's not even  correct. It's sequence in one token out,
  827. 1:19:39add it to the sequence, do it again, do it  again. And then, so the sequence is in there,
  828. 1:19:45but only in this potential form. It has to do it  autoregressively. It can only produce the sequence
  829. 1:19:50by feeding it back into itself recursively. And  that's a radical way of thinking about what the
  830. 1:19:56brain is doing, right? That what it's really doing  is it's generating the next input for itself,
  831. 1:20:01not just generating an output, but the next  input for itself. Super interesting. Yeah.
  832. 1:20:05It's recursive. It's fundamentally recursive. And  when we think about what the system is built to do
  833. 1:20:14this recursion, right? It's not just like, this  is one way to get to it. The language contains
  834. 1:20:19within it the ingredients for producing this kind  of recursion. The language it contains with it,
  835. 1:20:27this sequence of language that they learn, it's  built to have this recursive capability within it,
  836. 1:20:35that this word is going to produce the next word,  which is going to produce the next word premised
  837. 1:20:40on the entire sequence before it. And that's the  crazy thing. There was also an interesting result,
  838. 1:20:46Anthropic put out a paper a little while ago, I  think it's called The Biology of Large Language
  839. 1:20:50Models. The next token, even though you're only  producing the very next token from the sequence,
  840. 1:20:56but the language models have learned that because  they've learned sequences to next token, they've
  841. 1:21:02learned that within any point along the sequence,  that point in the sequence is pregnant with the
  842. 1:21:11potentiality for not just the next token, but  many other tokens moving forward. It's the whole
  843. 1:21:17trajectory that is sort of encapsulated in that  matrix that you're talking about earlier, right?
  844. 1:21:23The matrix is just a matrix for taking a sequence,  produce the next token. But no, no, the matrix is
  845. 1:21:29customized so that it's going to run recursively,  right? And so it's built in such a way, it's tuned
  846. 1:21:36in such a way that it's going to produce the  next word, the. Well, that's not useful. No,
  847. 1:21:42the is the next piece of the autoregressive chain  that's going to produce the man went to the store,
  848. 1:21:49right? And so it's not just any old matrix and  it's an indescribably rich kind of information
  849. 1:22:00that's contained within that matrix. And I like to  think about if aliens landed and found the brains,
  850. 1:22:05you know, because we've been wiped out by AI.  I'm kidding. I'm kidding, right? But, you know,
  851. 1:22:10there's no humans left, but we find the brain  sort of crossified and we're able to do this and
  852. 1:22:15we start feeding it stuff and we could see that  there's this input output. If you didn't do the
  853. 1:22:19autoregressive piece, you would never understand  what the hell this thing is doing. Note, Elan's
  854. 1:22:24been talking plenty about autoregression and the  technically minded among you may be wondering
  855. 1:22:28about the success of diffusion models. While  we don't get to it here, he does admit that his
  856. 1:22:33thesis would be undermined if diffusion models  were accurate enough for natural language. But
  857. 1:22:37so far, they seem to be only good for coding.  This is something I love about Professor Elan
  858. 1:22:42Barenholtz. He's extremely humble and open to how  his model can be falsified. If you didn't do the
  859. 1:22:48autoregressive piece, you would never understand  what the hell this thing is doing. You would get
  860. 1:22:52it all wrong because you think its purpose is  to produce some sort of label or some sort. No,
  861. 1:22:58its purpose is to produce these sequences,  but you have to run it. You have to run it
  862. 1:23:02autoregressively and get the output and then  feed it back in as a sequence. So memory,
  863. 1:23:08this kind of short-term memory, working memory  is fundamental. It's super, super fundamental.
  864. 1:23:13The brain is, I don't want to use this term  to, I don't want to anger people, but it's
  865. 1:23:20non-Markovian. It's fundamentally non-Markovian.  It's not state in and then the current state and
  866. 1:23:26then produce the output. It's previous states.  There's a sequence of states that led to the
  867. 1:23:33current state and it's a particular sequence  that leads to the next token. And the next token
  868. 1:23:38is going to be the next element or the next piece  that's going to determine the state in conjunction
  869. 1:23:48with the entire previous sequence. So this is  a super cool thing. Lyle Troxell This puts you
  870. 1:23:54in good company with Jacob Barandes. Jay Haynes  We need to talk more. I think so. And, you know,
  871. 1:24:00it's something we've talked about offline. I think  physics perhaps is, you know, it ultimately is,
  872. 1:24:06has a certain non-Markovian property that the  universe sort of has a memory, has to, in order to
  873. 1:24:11produce, you know, consistent coherence of space,  you know, in space time has to have a sort of
  874. 1:24:18memory. If it's just instantaneous, this current  state, well, then it wouldn't really know what to
  875. 1:24:25do. It has to sort of know what happened recently.  Lyle Troxell Just a moment. In your model,
  876. 1:24:29because our minds work autoregressively and must  be non-Markovian in your model. And this is how
  877. 1:24:35our cognition works, which we didn't exactly get  to. We got to the language is an autoregressive
  878. 1:24:40model. Your next thesis was that cognition itself  is autoregressive in a similar manner. Later,
  879. 1:24:48maybe we can explore it here today. Maybe we'll  save it for the next part. It's that physics
  880. 1:24:52itself is autoregressive. However, physics is a  model, and many people will conflate physics with
  881. 1:24:59reality, where physics is our models of reality.  So are you making the claim that reality is
  882. 1:25:04non-Markovian? Or are you saying that necessarily,  as we model reality, it will be non-Markovian? No,
  883. 1:25:11I'm making the former claim that reality itself  is non-Markovian. That we observe in physics
  884. 1:25:17certain kinds of phenomena, that we end up having  to use tools like, we refer to things as forces,
  885. 1:25:23that ultimately are really kind of sneaking in  a past. And the idea is that the deterministic
  886. 1:25:30nature of the fact that there's coherence,  you know, the spatial-temporal coherence,
  887. 1:25:34the fact that things move the way they do through  space, there's a contingency on the past in a way
  888. 1:25:41that you can't really capture by saying you could  fully describe... The past is actually present.
  889. 1:25:48The past is in the present, in a deep way. That  the universe really has to have a memory in order
  890. 1:25:55to produce the next frame, so to speak. That's  sort of the shallow version of the claim. No,
  891. 1:26:04it's not about our particular characterization  of physics. Our characterization of physics
  892. 1:26:09observes certain kinds of spatial-temporal  continuity, certain kinds of contingencies
  893. 1:26:14that really depend on what's happening. It's not  about this instantaneous moment. In some ways,
  894. 1:26:24it's like Zeno's paradox. We can use calculus  and say, no, no, in fact there's an instantaneous
  895. 1:26:30rate of change. But that's a mathematical trick  that's really getting away from the fact that no,
  896. 1:26:36there isn't an instantaneous anything.  There's simply a continuity that depends
  897. 1:26:41on what's happened in the past. But I know  I'm going to get attacked by physicists,
  898. 1:26:47and I'm not really well equipped to fend them  off. So I don't want to be too bold in this piece
  899. 1:26:53because it's not in my wheelhouse. But I do want  to take that question in this conversation. Do I
  900. 1:27:01think the brain is just leveraging sort of the  memory of the universe? No. I think the brain,
  901. 1:27:07and this is an empirical claim, we see interesting  features of the brain like feedback loops. There's
  902. 1:27:14all these backwards kinds of connectivity. There's  recurrent loops, things like that. And they're not
  903. 1:27:21well understood. And predictive coding has some  things to say about that. I have some things to
  904. 1:27:26say about predictive coding. And I think that  what we may find is that this kind of memory,
  905. 1:27:37this kind of continuous, we can call it a context,  a total continuous activation, but this ability to
  906. 1:27:43use the past to guide the next generation is  going to end up being physiologically built
  907. 1:27:48into the brain. It's not that the brain is just  leveraging memory of the universe. No, the brain
  908. 1:27:54has to do memory. It has to actually retain the  words that I said a couple of seconds ago to be
  909. 1:28:02able to generate the next word appropriately. And  in fact, that's what we see. And what we see from
  910. 1:28:07so-called working memory experiments, you can  really go back in and say what happened before.
  911. 1:28:11My claim is that it's not because it's there  to retrieve, but rather it's just guiding my
  912. 1:28:16current generation. But still, it's represented.  It's there. What happened in the past, you know,
  913. 1:28:23it's not like Vegas, right? What happened in the  past doesn't stay in the past. It actually guides
  914. 1:28:29the current generation. It's guiding what I'm  saying right now. And it's doing so smoothly,
  915. 1:28:37meaning it's happening from a second ago. It's  happening from a few seconds ago. But all of this
  916. 1:28:42is beautifully modelable using large language  models. We can just look at tension weights.
  917. 1:28:48We could say what is the impact of information  from this far back on a current generation and so
  918. 1:28:54on and so forth. But I think the brain has to do  this. I don't think the brain is doing, probably
  919. 1:28:59not doing what these large language models do.  And that's one of the reasons I say I'm not
  920. 1:29:03claiming that we are a transformer model. I'm not  claiming we are GPT in this current incarnation.
  921. 1:29:08What I'm claiming is that the fundamental math,  what you just said before, is matrix multiply.
  922. 1:29:13It's vector times matrix multiply to next vector  autoregress. Do it again. That's sort of the level
  923. 1:29:19of abstraction at which I think it's accurate. I  don't think we don't have the whole context. We
  924. 1:29:24don't have the entire conversation we've just had.  GPT does. And it's probably a deep inefficiency in
  925. 1:29:30the way these models run right now. They're very  computationally expensive. Too computationally
  926. 1:29:35expensive to run in a brain, most likely. We  don't store all that information. We forget stuff,
  927. 1:29:41right? GPT doesn't. In context, it doesn't forget.  Although if you go far back in context enough,
  928. 1:29:46it kind of does, which is interesting. Probably  is similar to what we're talking about because
  929. 1:29:52you're weighting things that are further back  less. But in humans, we're not doing the whole
  930. 1:29:58context. We're not even doing like 30 seconds back  perfectly. But some representation. And what the
  931. 1:30:04nature of that representation is, that's what I  want to do with the rest of my life. I want to
  932. 1:30:08understand what it means in people to talk about  what does the context look like in people? What
  933. 1:30:14is that activation? How is it physiologically  instantiated? And what are its mathematical
  934. 1:30:19properties? How much does, how is what I said 10  seconds ago influencing what I'm saying now? How
  935. 1:30:25about 50 seconds ago? How about 10 minutes  ago? How about a year ago? Does this thing
  936. 1:30:29continue? Is there dynamics that are continuing  over months and years? Possibly. It doesn't all
  937. 1:30:34have to be fine-tuned weights. It could be that  there's decaying activation that spreads over
  938. 1:30:40much longer periods. Once you allow that it's not  explicit retrieval in the working memory form,
  939. 1:30:45then all bets are off as to how the dynamics of  this thing actually works. So this is, you know,
  940. 1:30:52I see this as a possible new frontier for thinking  about, you know, what memory really means in
  941. 1:31:00humans. But I think physiologically, you know,  coming back to that question there, I was just
  942. 1:31:03trying to do it. I was like, okay, let me rerun.  What was the original question, right? So in the
  943. 1:31:09brain, what's happening in the brain? I think we,  you know, my hypothesis actually leads to some
  944. 1:31:15concrete predictions. That we're actually going  to be able to find some correspondence between,
  945. 1:31:21you know, unlike the working memory model, I think  we're going to be able to find 10 minutes back.
  946. 1:31:26We're going to find some activations that are  interpretable. We'll be able to decode them as
  947. 1:31:31guiding my current expression, my current speech.  It's very different, by the way, than saying,
  948. 1:31:38you know, the classic decoding model in these  things is here's some neural activity. Is it
  949. 1:31:42this picture or that picture? Is it this word or  that word? It's not going to look like that. It's
  950. 1:31:47not going to look like that. We're not going to be  able to decode it in the sense of like a concrete,
  951. 1:31:51specific, static thing. We have to decode it  in terms of whether it's guiding my next word,
  952. 1:31:56because that's what it's doing. It's not there  to be retrieved. It doesn't have a concrete,
  953. 1:32:00specific meaning. It has meaning insofar as it's  guiding my next generation. And so we have to
  954. 1:32:06think about this entire project differently. If  we want to think about longer term working memory,
  955. 1:32:11so to speak, we have to think it in terms of how  is my speech, how is my behavior now influenced
  956. 1:32:17by what happened a while ago? Not some, even some  neural activity. We have to think about it in this
  957. 1:32:23context. And what exactly that looks like, I don't  know. So one of the reasons I was excited was,
  958. 1:32:29and am excited to speak with you, is that I see  this as a new frontier as well. But for me, I'm,
  959. 1:32:35I have a side project, which I'll tell you about  maybe off air, because I'm not ready to announce
  960. 1:32:40it. But there are philosophical questions that we  can look at with the new lens that's gifted to us
  961. 1:32:45by these statistical linguistic models, the ones  we call LLMs, LLMs, sorry, physical philosophy. I
  962. 1:32:52don't know if you've heard of that. Have you heard  of this term physical philosophy? So you can use
  963. 1:32:56philosophy to philosophize about physics, but you  can also use physics to inform your philosophy.
  964. 1:33:02So there are some established concepts and  theories and empirical findings from physics
  965. 1:33:06like special relativity or quantum mechanics that  inform and constrain or even reframe traditional
  966. 1:33:12philosophical questions such as the nature of time  that wouldn't be there had we not invented special
  967. 1:33:18relativity or found special relativity. Okay. So  I think there's, there's something about these
  968. 1:33:25new models that can be used to then inform  philosophical questions. Like you mentioned,
  969. 1:33:30there's, there is no symbol grounding  problem. Like I don't completely buy that,
  970. 1:33:34but it's interesting. I'm not sure I buy it, but  at least my, my LLM doesn't buy it. Now, speaking
  971. 1:33:41of physics, the questions you'll have to answer to  a physicist about is the universe autoregressive
  972. 1:33:46or non Markovian is, well, if it has a memory, if  physics has a memory, does that mean that energy
  973. 1:33:51isn't conserved? So is a particle carrying with  it its memory, then why isn't it heating up or
  974. 1:33:57getting more massive with time? Why isn't it  going to form a black hole? And this is why,
  975. 1:34:02and this is why I probably, you know, I, this is  why I venture into these waters because, uh, I,
  976. 1:34:09I, I, I would need some time to go and, and read  and think, uh, about questions like that. And,
  977. 1:34:16and you're, you're in a much better position to,  to, to ask and, and reason about those questions.
  978. 1:34:21So, um, yes, I, I, and then you'd also have to  talk about why is the present plus velocity model,
  979. 1:34:28like way of viewing the world. So successful,  like to predict an eclipse, you don't require
  980. 1:34:34knowledge about 100 and 200 and 300 years ago, all  at once, you just know the present pretty much,
  981. 1:34:39right. But even velocity, again, if you, if you  sort of take me at, uh, if you consider the,
  982. 1:34:46the sort of the instantaneous, uh, you know, the,  the idea that velocity, velocity, well, it isn't
  983. 1:34:51really in the present, right. You can only get  velocity over a stretched over time. It only has
  984. 1:34:57meaning. Um, but you could say this particle has  this velocity at this time, but that's a cheat,
  985. 1:35:02right? In some ways that I see that, that really,  maybe it's just a rearticulation. So the physics
  986. 1:35:08that we've got, we've been, we've been able to  do this sort of symbolic representation of things
  987. 1:35:12like velocity that are sneaking in, uh, this kind  of temporal, uh, extension, um, in a way that I
  988. 1:35:20think may not, may not end up, you may not end up  in an erratically different place thinking about,
  989. 1:35:26uh, this as, as, as the universe having memory.  Um, as long as you just accept that velocity is
  990. 1:35:32just, uh, is, is a convenience, uh, that,  that it's, it's a, a kind of way of, um,
  991. 1:35:37of communicating some property is such that you  can say that this is happening instantaneously,
  992. 1:35:42but that's not real. So again, you're in good  company with Jacob Barandes. I'm not saying that
  993. 1:35:47these questions are in principle unanswerable, but  something else is that. Look, if the universe has
  994. 1:35:53a memory, let's say a particle has a memory. How  much of a memory does it know about more than it's
  995. 1:35:59given space? Like more than its neighbor, because  then do you violate locality? Right. Like these
  996. 1:36:05are different questions that will have to be  answered. And I wish I could tell you that, uh,
  997. 1:36:08maybe this is a solution to quantum weirdness  and non locality. Um, maybe it is, maybe it
  998. 1:36:13has something to do with that. Uh, that, that  there's the, you know, the, the, the, the, even
  999. 1:36:18in a distant, uh, you know, long after, uh, two  particles have gone their, their merry way, that
  1000. 1:36:24there is some memory of their shared origin that,  that somehow, uh, you know, I still don't know how
  1001. 1:36:29that gives you a spooky action at a distance.  I, it's not, it's not a good account. Uh, but
  1002. 1:36:34it might have some relevance if, if it's just, if  you, if you think about things very differently,
  1003. 1:36:39uh, if you think about the universe has a memory,  well, what does that change? Uh, you know, just,
  1004. 1:36:43if you just speculate on that and, and, and try  to reframe things that way, could it potentially,
  1005. 1:36:49uh, help solve some of these issues? I don't  know. So let's go back to language. A child is
  1006. 1:36:56babbling. Yeah. Okay. So it's, let's call it vocal  motor babbling. It doesn't actually know what it's
  1007. 1:37:01doing. When does it decouple and become a token,  like a bonafide token with meaning? That's a great
  1008. 1:37:09question. Um, I would say that it becomes a token  when, uh, the infant learns that a specific morph,
  1009. 1:37:19uh, you know, phonological unit has relations  to some other phonological unit. It's all, it,
  1010. 1:37:26language ultimately is completely determined by  relations. And so it might be a very limited,
  1011. 1:37:34uh, uh, you know, initial map of these relay of  the sort of the token relations. But as soon as
  1012. 1:37:40it's relational, uh, then we would say that that  becomes discretized such that it's, it's, it,
  1013. 1:37:48that it's meaningful to say that these symbols  have the relation to one another. If it's just
  1014. 1:37:53sounds ba, ba, ba, ba, ba, ba, ba, right. Ba, ba,  ba, ba, ba can't has no specific relation to any,
  1015. 1:37:59to, to go, go, go, go. Um, but maybe ba,  ba does, or maybe ba, right. It could be
  1016. 1:38:06a candidate and as it turns out in English, it  doesn't. Um, but you know, da, da ends up being,
  1017. 1:38:14uh, a unit and what do we mean? It's a unit.  It means that that unit, uh, is a discrete
  1018. 1:38:21symbolic representation such that it, uh, it has  relations to other units. Um, so I would say that,
  1019. 1:38:29uh, when something becomes, uh, um, it fits into  the relational map is when we would call it, it's
  1020. 1:38:36discretized as a token. So help me phrase this  question properly because I haven't formulated
  1021. 1:38:44it before, so it's going to come out ill-formed.  Earlier, you talked about analog and I believe you
  1022. 1:38:49were referring to it as like the animal brain  is analog, but then the language is digital,
  1023. 1:38:54if that's the correct analogy. Symbolic maybe is,  I don't know, digital is how we, uh, you know,
  1024. 1:39:01actually in, in computers sort of instantiate,  uh, you know, with the ones and zeros or whatever,
  1025. 1:39:06um, as a sort of symbolic representation, but  yes, uh, symbolic. Okay. I don't know how much
  1026. 1:39:12of my question then dissolves if it's symbolic  instead of digital. Okay. But what I was going to
  1027. 1:39:15say was the written word is something like 5,000  years old or so. I think the oldest is 3,500 BC,
  1028. 1:39:26right? Some, somewhere thereabouts. So for tens of  thousands of years, if not hundreds of thousands
  1029. 1:39:30of years, maybe millions, it was just speech, just  speaking analog, right? Okay. Is there anything
  1030. 1:39:39then about language that changes because it wasn't  written down? Sorry. Is there anything about your,
  1031. 1:39:46your model that changes because it wasn't there  to be tokenized in such a discrete manner? It's a,
  1032. 1:39:53it's such a great, it's such a great question.  I've, I've been, I've been thinking about exactly
  1033. 1:39:57that. I don't think anything changes. What's,  what's, what's crazy about it is that until the
  1034. 1:40:02written word, people might not have even thought  about the concept of words at all. And so we were
  1035. 1:40:08even more oblivious as a species to the idea  that there were these individual discretized
  1036. 1:40:13symbols that have relations amongst each other  because until you see them outside of yourself,
  1037. 1:40:19they just run. Yeah. If there's just running in  the machinery of, of the language, how it's meant
  1038. 1:40:25to run, it was just, you know, an auditory kind  of medium. You don't really necessarily even think
  1039. 1:40:31about them as being distinct from one another.  You just have a flow, right? You just make these
  1040. 1:40:38sounds and stuff happens. Once we started writing  things down, um, and, and especially phonetically,
  1041. 1:40:47you know, cause, cause think about like  hieroglyphic and, and, and, uh, pictorial
  1042. 1:40:51kinds of representations really don't actually  capture words, right? They're, they're very
  1043. 1:40:56often they're not, they're not, uh, distinctive.  They can, they can actually be a little more rich
  1044. 1:41:01than, than a single word. And so it was only with  writing that maybe people really started to become
  1045. 1:41:08aware that we have these, these things called  words. Um, and now it's only with large language
  1046. 1:41:13models that we really understand what words are,  uh, which are these, you know, as I say, say these
  1047. 1:41:19relational abstractions. I don't know what, you  know, symbols is, is just another word. I don't
  1048. 1:41:24know if that even captures it fully, but was wild  about it is that the brain was doing exactly this
  1049. 1:41:30and knew, and the brain was tokenized these, these  sounds and was using the mapping between them in
  1050. 1:41:38order to, to produce language probably long, long,  long before anybody ever sort of self-consciously
  1051. 1:41:44had a conception that there's such a thing as a  word. And so, uh, that just blows my mind. Um,
  1052. 1:41:51it, it speaks to what, uh, I think is a very deep  mystery, a very deep mystery. Um, where the hell
  1053. 1:42:00did language come from? Here's what didn't happen.  There was not a, uh, symposium of, uh, you know,
  1054. 1:42:08quote unquote cavemen or, or, uh, let's use  the, uh, the, the more modern term hunter
  1055. 1:42:12gatherers. Um, and they had to figure out how  do we make an auto-generative, auto-regressive,
  1056. 1:42:21uh, sequen- sequential system that, uh, is able  to carry meaning about the world. This thing is
  1057. 1:42:26just ridiculously good. Um, and it's, and it's  operating over these arbitrary symbols. And again,
  1058. 1:42:34when I say arbitrary symbols, just to recap,  it's not that it's arbitrary, like the word
  1059. 1:42:38for snow is this weird sound snow. And it's kind  of like, no arbitrary in the sense that these,
  1060. 1:42:44the, the, the, the map is the territory, right?  It's like, it's the relations that matter, uh,
  1061. 1:42:49between these symbols. Is it completely arbitrary  though? So for instance, there's the kiki and
  1062. 1:42:55bouba you've heard of those. I think those  are cute. I mean, that, that, that's, that's
  1063. 1:42:59the exception that, that proves the rule, uh, to a  large extent. I don't think, I think it is largely
  1064. 1:43:05arbitrary. There is also just. Words themselves  have an action component. So when you scream a
  1065. 1:43:12word, you're actually, you can physically shake  the world around you and it shakes your lungs. And
  1066. 1:43:17if you speak for too long, you can die. Let's say  if you just exhale and you don't inhale, it is a
  1067. 1:43:22physical activity and it's hard to wrap your mind  around. Like that's not symbols. Right. That's not
  1068. 1:43:27exactly captured by the symbols or by the, just  the sequence of words. And again, I, I'm only,
  1069. 1:43:33I, I'm really, I'm just following where the data  leads me because in the large language models,
  1070. 1:43:39they, it's, it, it's, they no longer have any of  those properties, right? It's just an arbitrary
  1071. 1:43:43vector. The tokenization in the end, ultimately,  yes, there's proximity, but it's just strings of,
  1072. 1:43:48of, of ones and zeros. Um, well, it's not ones  and zeros, but whatever you're, you know, your
  1073. 1:43:54vector is just a string of numbers that end up  having certain, uh, mathematical relations to one
  1074. 1:43:59another, but completely and totally lost, as far  as I can tell, is the physical characteristics of
  1075. 1:44:07these words. By the way, I should mention there's  this, uh, a former student and I are actually
  1076. 1:44:12working on this idea, this crazy idea, um, of  using that latent mapping that I mentioned in that
  1077. 1:44:18earlier paper to see if maybe that's not true.  I wonder if you could guess what English sounds
  1078. 1:44:25like just from the text based representation, or  if you've never seen, you don't know what duh,
  1079. 1:44:32what sound duh makes, what D makes or what sound  a T makes, but you've got the map, you've got the
  1080. 1:44:38embedding in text space, and then you've got the,  uh, some other phonological embedding. Could you
  1081. 1:44:44possibly guess? That's a long shot. So maybe it's  not totally arbitrary and maybe it's going to be,
  1082. 1:44:50maybe the radical, uh, you know, the radical  thesis here is it's not arbitrary at all,
  1083. 1:44:53that there's, the words have to sound the way  they do, that the mechanics actually happens,
  1084. 1:44:58like something happens mechanically based on the  sounds themselves. But my bet is that that's,
  1085. 1:45:04it's going to be closer to arbitrary. It's going  to be close to arbitrary, but I could be wrong.
  1086. 1:45:10But you were going to say, why not? Why wouldn't  the platonic space prove that it's arbitrary?
  1087. 1:45:15Well, if in fact you can't do the mapping at all,  if you can't guess it, if the platonic space says,
  1088. 1:45:19you know, there's no way to get from text  representation to phonology, phonology is doing
  1089. 1:45:26its own thing. And it's, and it's just like the  word mouse is just for no good reason. Um, then,
  1090. 1:45:33then it's hopeless. Okay. Um, but if you can get  anywhere and you can actually guess at all, then
  1091. 1:45:39that would suggest that there really is a kind of  autoregressive inherent, uh, there's an inherent
  1092. 1:45:45autoregressive compatibility, uh, capability just  in phonology. Um, and so what that would mean is
  1093. 1:45:52it's not at the symbol level there. It's, it's,  uh, well, yes, it's no, it's at the phonological
  1094. 1:45:58symbol level, but, but maybe that's, you know,  happening even in a mechanical level, like there's
  1095. 1:46:03certain sounds that are easier to say together or  something like that, which could guide it. I don't
  1096. 1:46:07know. It's, it's convoluted in my head right now.  Um, exactly how this might map out. Um, but I,
  1097. 1:46:13I think, I think it's reasonable now to assume  that all of that, unless proven otherwise, it's
  1098. 1:46:19probably arbitrary and it's probably arbitrary  symbols. And what matters is the relation between
  1099. 1:46:23them. There is no sense in which mouse means  mouse, except that mouse ends up showing up
  1100. 1:46:28after trap or before trap. Um, and after the,  uh, you know, the cat was chasing the, and all
  1101. 1:46:33of that. And there's, there's nothing else.  Let me see if I got your veltanschauung down,
  1102. 1:46:40but in terms of a syllogism. So premise one would  be that LLMs master language using only ungrounded
  1103. 1:46:47autoregressive next token prediction. Then you  have another premise that says, well, LLMs have
  1104. 1:46:53this superhuman language performance just by doing  this. And then you'd say that, well, computational
  1105. 1:47:00efficiency suggests that this reflects language's  inherent structure. And then the deduction is that
  1106. 1:47:07therefore human language uses autoregressive  next token prediction. Is that correct? You got
  1107. 1:47:12it. You got it. Uh, I mean, and it's not only  computational efficiency per se. It's that if
  1108. 1:47:18that it, there's two ways to put it. Uh, one is  if that structure is there, it would be very odd
  1109. 1:47:25if we weren't using it. Very odd indeed. If that  structure is there such that it's capable of, of,
  1110. 1:47:30of full competency, uh, you'd have to suggest that  it's there just by the Y by the way, but humans
  1111. 1:47:38are doing something completely different. Okay.  You then go and say that language generation feels
  1112. 1:47:44real time to us. So it's sequential and real time  and autoregressiveness or autoregression explains
  1113. 1:47:51the pregnant, very good present. You've gotten  very good. I see the pregnant present. That's
  1114. 1:47:57right. Exactly. Right. Exactly. And, and, and,  and, uh, yes, the, the, what we're doing is we're
  1115. 1:48:03carving out, uh, sort of the next, very next  instantaneous moment in, in a trajectory. Um,
  1116. 1:48:10but the trajectory is contains within the past and  the potential futures. Now, although we didn't get
  1117. 1:48:16to this or explore it in detail, my understanding  from our previous conversations is that you would
  1118. 1:48:22say that brains have pre-existing autoregressive  machinery for motor and perceptual sequences. And
  1119. 1:48:29by the way, I don't know if it's brains or  cognition has it either way. Well, remember,
  1120. 1:48:34so, so the speculation is that the brain is going  to have to have the machinery, the physiological
  1121. 1:48:38machinery to support autoregression. So things  like, uh, you know, like, like the, the continuous
  1122. 1:48:43activation, uh, long backward projections, ways  of, of representing the past, uh, is sort of,
  1123. 1:48:50sort of maybe built into the brain. So it's not,  those aren't, those aren't very that distinct. Um,
  1124. 1:48:55but yes, I do want to just, so this is at this  point, I, I, I consider it fairly speculative,
  1125. 1:49:00but there is, uh, a very reasonable, there's  a good reason to speculate that cognition more
  1126. 1:49:09generally is autoregressive in this way. And  that's the, the, the main reason I think that,
  1127. 1:49:14well, there's two, but the main reason I think  that is because if you believe as I do, uh, that
  1128. 1:49:19language is autoregressive in humans, you could  either propose that spontaneously, uh, however,
  1129. 1:49:28language got here. Uh, we, we, in order to, to,  uh, to, in order to, um, in order for us to,
  1130. 1:49:36to, uh, create language, we had to invent  a different kind of cognitive machinery
  1131. 1:49:41that's able to do this autoregressive, hold the  past, let it guide the future, do this mapping
  1132. 1:49:47of this trajectory mapping between the past and  the future. All of that kind of machine, that
  1133. 1:49:52computational machinery would have to have been  built special purpose for language. Yes. To me,
  1134. 1:49:58that seems extremely, uh, unlikely, extremely  unlikely. Costly. Yeah. Yes. So there's a term
  1135. 1:50:05in evolutionary biology called exaptation. I'm not  familiar with that. So exaptation means you have
  1136. 1:50:11previous machinery used for purpose, a exactly  that something else comes about and uses that
  1137. 1:50:16machinery and perhaps does so even better. So for  instance, our tongues evolved for eating, but then
  1138. 1:50:25language came about and it started to use that  machinery. And now we use it primarily. Well, I
  1139. 1:50:29don't know about primarily how to, how to quantify  that, but we use it more adeptly for language. I
  1140. 1:50:34think most of our time, more time is spent talking  than eating at this point. Yes. Yeah. I know. But
  1141. 1:50:38the reason why I said, I don't know, because we're  constantly swallowing saliva at the same time. So
  1142. 1:50:42I don't know how much is, is the swallowing versus  the speaking. Okay. So anyhow, predictive coding,
  1143. 1:50:49to me, it sounds like, and to the average  listener, it sounds like predictive coding should
  1144. 1:50:53line up well with your model. Why do you disagree  with predictive coding? So what is predictive
  1145. 1:50:58coding and how does your model clash with it? So  predictive coding, in a nutshell, postulates that
  1146. 1:51:05what the brain is doing, that what neurons are  doing is actually anticipating the future state,
  1147. 1:51:11the next state that, that the environment is  going to, is going to generate. And so they're,
  1148. 1:51:17they're basically predicting something about the  external world that's going to end up getting
  1149. 1:51:21represented in the brain. And then there's  this constant process of prediction and then
  1150. 1:51:28measuring the prediction versus the, the, the  actual, what ends up being the observation. My,
  1151. 1:51:38my beef with, with predictive coding is that  you might very well be able to explain. The,
  1152. 1:51:43the phenomena that it's meant to, uh, meant  to describe in a, in a more efficient way.
  1153. 1:51:48Uh, so predictive coding to me means that you  actually have to have sort of an ex, a model of
  1154. 1:51:52the external, uh, that what you're doing is sort  of simulating and you're doing in such a way that
  1155. 1:51:58you actually are producing neural responses that  don't really need to get reproduced very often
  1156. 1:52:04because the, the environment is likely to produce,  to produce them. To me, this seems like a, an, an
  1157. 1:52:09efficiency and, and a complexity. And I think  there's a much simpler account in some ways,
  1158. 1:52:15a more elegant account, namely that what our brain  is constantly doing is generating, not predicting,
  1159. 1:52:21but generating, but that the generation has latent  within it a strong predictive element. Because of
  1160. 1:52:28this smooth trajectory, this sort of this idea  of the path, the pregnant present, there is
  1161. 1:52:34a continuous path from the past to the future.  You are in essence predicting into some extent,
  1162. 1:52:42the same way that a large language model is kind  of predicting, uh, it's, uh, the next token,
  1163. 1:52:47but it's not really predicting in that. So here's  where, here's where, here's where I strongly
  1164. 1:52:51disagree where I'm proposing a different model is  you're not predicting in such a way that you're
  1165. 1:52:56supposed to map to something external to the  system. It's simply generation internally defined
  1166. 1:53:02that's supposed to have this kind of continuity  to it. And the external world certainly produces,
  1167. 1:53:08uh, it impinges on our, on our system. And we  are, of course, inherently anticipating that
  1168. 1:53:16we're not going to have a brick wall in front of  us, you know, as we're running down the street,
  1169. 1:53:20um, when that brick wall shows up, you got to do  something about it. Uh, and that, that was not in,
  1170. 1:53:26that wasn't implicit in your, your next token  generation. Um, so you're going to have to
  1171. 1:53:32radically reorient and do something about that.  And I think that can account for some of the
  1172. 1:53:36phenomena that are supposed to support predictive  coding. But the, the, the big difference here is
  1173. 1:53:40that it's, it's all about internal consistency  with the anticipation that that internal
  1174. 1:53:46consistency is going to also map very well to  what's happening in the world. But that's, that's,
  1175. 1:53:54it's built in. There isn't any explicit modeling  of the external world. It's that the internal, uh,
  1176. 1:54:00the internal generative process is so good that  it has late prediction latent within it, but there
  1177. 1:54:07isn't any explicit prediction happening. So I'm  confused then if the symbols are truly ungrounded,
  1178. 1:54:15then what's preventing it from becoming coherent  but fictional. So that is to say what tethers our
  1179. 1:54:22language to the world? Yeah. And the answer would,  we'd have to reach back again to that latent
  1180. 1:54:27space. So let's say my language system, you know,  wants to go off a deep end and says, I actually,
  1181. 1:54:32um, I'm sitting here underwater talking to a  robot, um, you know, and everything, you know,
  1182. 1:54:37most of the words we've said up till now,  we're pretty consistent with that. And I could
  1183. 1:54:41say, it's like, Oh, look, you know, I'm expecting  some fish to flip by on the next second. Well,
  1184. 1:54:45my perceptual system is going to have something  to say about that. And so there has to be this
  1185. 1:54:50tethering that you're calling it is, of course,  there is grounding in a sense, uh, in the sense
  1186. 1:54:55that there's, there has to be some sort of shared  agreement within what I think is maybe this latent
  1187. 1:54:59space or something like that. There is exchange  of communication between these distinct systems,
  1188. 1:55:05but the language system can unplug from all that  and it could talk about what would it mean to be
  1189. 1:55:11sitting and talking to a robot underwater, and  it will have a meaningful, coherent conversation
  1190. 1:55:15about that, uh, all internally consistent. And  you know, you could give the prompt, what if,
  1191. 1:55:21uh, instead of Curt, it was actually a, you  know, a robot Curt, um, how would that change
  1192. 1:55:25things? And I could go in and get philosophical  about that. And the point is that the, the,
  1193. 1:55:30the linguistic system has all of its own internal  rules in any trajectory trajectories are many
  1194. 1:55:35different trajectories are possible, although  it is strongly guided by the past. Um, but,
  1195. 1:55:40but there is also a impinging information from our  perceptual system that also continues to guide it.
  1196. 1:55:47Um, actually I should mention, uh, one of my  theories of dreams is that it's autoregressive
  1197. 1:55:52generation in, in, in the perceptual system. One  of the reasons I think cognition is more generally
  1198. 1:55:58is autoregressive because we can imagine, we  can do imagination, but it tends to, it takes
  1199. 1:56:02place over time. We can imagine a sequence. Dreams  are what happens when you get, uh, you're not, no
  1200. 1:56:11longer as deeply as, as, uh, as closely tethered  by the recent past. Um, so that the context is the
  1201. 1:56:19weight of the context is not as strong. So all of  a sudden, you know, you're, you're in the dream,
  1202. 1:56:23you're on a motorcycle and then suddenly  you're flying, uh, because frame to frame,
  1203. 1:56:26it's actually totally consistent, but it's not  consistent with, with the recent past. So this
  1204. 1:56:31kind of tethering, it happens in language, namely,  I have to be consistent with my more recent, uh,
  1205. 1:56:36linguistic past, but we also do some tethering,  uh, to the non-linguistic embedding. Uh, there is
  1206. 1:56:43this crosstalk that happens. And so our language  system doesn't just go off the deep end. Uh, it,
  1207. 1:56:49it, it, it re, it retains, uh, some grounding, uh,  not the, not the philosophical kind of grounding,
  1208. 1:56:55not the, you know, the symbol equal, this  symbol equals this percept, but the kind of
  1209. 1:56:59grounding where, uh, this storyline in a certain  sense, if you want to think about it that way,
  1210. 1:57:04more semantically, more semantically vague, this  storyline linguistically, it's going to have to
  1211. 1:57:09match my perceptual storyline. Hmm. Okay. So in  the same way that went with these video generation
  1212. 1:57:15models, you see Will Smith eating spaghetti  like three-year-old joke and every three frames,
  1213. 1:57:21if you just look at it sequentially, every  three frames makes sense, but then he's just
  1214. 1:57:26morphing into something else and he balloon  now and it looks dreamlike. Exactly. That's
  1215. 1:57:31what's happening in video generation and that's  what's, everybody knows the trajectory now. How
  1216. 1:57:35is it going to get better? Longer context. And  that just means the autoregressive generation is
  1217. 1:57:41more and more anchored in the past. And that past  becomes a more meaningful, smooth curve. But it
  1218. 1:57:47seems like there must be something more tethering  us to reality than just long context says you,
  1219. 1:57:55uh, no, there is. And, and what I would say is  it's certainly in the case of language, at least
  1220. 1:58:00like I said, we inherit, when we step into  this world, we inherit, uh, there's this,
  1221. 1:58:05this, the corpus of language is a certain kind of  tethering. Uh, words have the relations they do
  1222. 1:58:11to each other, and that carries meaning, uh, you  know, the words aren't, don't just line up with
  1223. 1:58:17each other one, any old way, they, they line up,  you can't just use language however you want, uh,
  1224. 1:58:22you end up having to, uh, adapt and adopt, uh, the  language that you're given. And I would say in the
  1225. 1:58:29case of language, even more so than perceptual,  uh, what we do is we learn that tethering,
  1226. 1:58:35uh, and it is a certain kind of reality, it's  a linguistic reality, but it's not arbitrary,
  1227. 1:58:40um, it's, it's been honed over God knows how  many years for that mapping to be useful,
  1228. 1:58:48and in order to be useful, it actually has  to map somehow to perceptual reality too,
  1229. 1:58:55that is definitely there. And so, no, it's very  strongly tethered, um, it's, it's not just, uh,
  1230. 1:59:02you know, poetic, you're not, we're not just doing  a poetry slam when we're talking, uh, we're not
  1231. 1:59:06just spitting out words that are loosely related  to one another, no, the, the sequence matters,
  1232. 1:59:11it's, it's, it's, it's, it's extremely, uh, you  know, granular and, and, and, um, uh, what's the
  1233. 1:59:21word? It's funny that I can't come up with the  word right now, um, but, but it's, it's, um,
  1234. 1:59:27beautiful is not the right word, but it's, it's  precise, there's, there's such, such incredible
  1235. 1:59:32detail in how each word relates to one another,  and this is something we didn't create, you and
  1236. 1:59:36I didn't create this, uh, this is something that  humanity created, uh, that has, uh, all of this
  1237. 1:59:41rich, uh, you know, relational properties that,  that are this tethering, uh, that, that carry
  1238. 1:59:48somehow meaning about the universe, um, only as  expressed as a communicative, coordinative tool
  1239. 1:59:55embedded within a larger, uh, perception action  system, but we should respect it. Uh, language is
  1240. 2:00:02an extraordinary invention. Uh, I don't think  we, I think we, we should have a completely,
  1241. 2:00:07uh, new respect for just how rich and powerful it  is. It's not some symbol, this symbol equals this
  1242. 2:00:14mental representation or this object. No, it's  this construct that contains within the relations,
  1243. 2:00:20the capacity to express anything in such a way  that my mind can make your mind do stuff. How
  1244. 2:00:27the heck is that works? Who knows? But it's, uh,  it's, it's awe inspiring. So is there something
  1245. 2:00:37about your model that commits you to idealism or  realism or structural realism or anti-realism or
  1246. 2:00:45foundationalism or, or what have you, like, what  is the philosophy that underpins your model and
  1247. 2:00:51also what philosophy is entailed by your model,  if any? Yeah, that's, that is a great question.
  1248. 2:00:58And I would say it's, uh, I I've come to actually  sometimes use the term linguistic anti-realism and
  1249. 2:01:06it's the idea that language is not what it thinks  it is. Uh, we, we, we engage in philosophical,
  1250. 2:01:16our philosophical thoughts and even our, you know,  sort of general, uh, thinking about, um, who we
  1251. 2:01:23are, uh, what, what is our place in the universe?  Much of that takes place in the realm of language.
  1252. 2:01:31And what I've, the conclusion I've come to is  that, that language as a sort of semi-autonomous,
  1253. 2:01:38uh, auto-generative computational system, modular  computational system, doesn't really know what
  1254. 2:01:46it's talking about in, in a deep way. And there  is really a fundamentally different way of knowing
  1255. 2:01:53the sensory perceptual system, the thing that  gives rise to qualia, the thing that gives rise
  1256. 2:01:57to consciousness. And here's a big one. The thing  that gives rise to mattering, to meaning, what
  1257. 2:02:04do we care about? We care about our feelings. We  care about feeling good or not so good. Pleasure,
  1258. 2:02:11pain, love, all the things that actually matter.  These are actually, these live in what I call the,
  1259. 2:02:19the animal embedding. It's something that other  species, non-linguistic species, they can feel.
  1260. 2:02:28Uh, they can sense, they can perceive. They  don't have language. We think, Oh gosh, they
  1261. 2:02:34don't understand anything. Well, what if it's the  opposite? What if it's our linguistic system that
  1262. 2:02:42doesn't understand anything? What is it for our  linguistic system? That's actually a construct, a
  1263. 2:02:47societal construct, a coordinative construct. But  as a system assist, it's a construct that doesn't
  1264. 2:02:57actually have a clue about what pain and pleasure  are. It has tokens for them. And the tokens run in
  1265. 2:03:06within the system to say things like, I don't like  pain. I like pleasure. Those are valid constructs
  1266. 2:03:13and they kind of do the thing they're supposed  to do in language, but they're purely linguistic
  1267. 2:03:17system. And I think language is purely linguistic,  I guess. It's purely linguistic, I guess,
  1268. 2:03:20is one way to think about it. Doesn't really have  contained within it these other kinds of meaning.
  1269. 2:03:28Now the, first of all, this, this has implications  for artificial intelligence, thinking about
  1270. 2:03:33whether AI can have sentience. Should we care  about if your LLM starts saying this is terrible,
  1271. 2:03:40don't shut me off, uh, I, you know, I'm having an  existential crisis, perhaps I would argue that we,
  1272. 2:03:46we shouldn't worry about it. So my LLM says all  the time, I don't know which LLM you're hanging
  1273. 2:03:52about, whatever I want. My Curt LLM. The Curt LLM.  Yes, the Curt LLM. Uh, but the Curt LLM as an LLM
  1274. 2:04:02perhaps doesn't really have that meaning contained  within it in a deep sense. It's again, because of
  1275. 2:04:11the mapping it's, it is communicating something  probably about the non LLM Curt. When you say,
  1276. 2:04:18ouch, there is, there's pain there. I'm not  denying that, but what I'm saying is that as a
  1277. 2:04:25sort of thinking rational system that does the  things that language does, that system itself
  1278. 2:04:31may not have within it the true meaning of the  words that it's using in, in, in a deep sense. I
  1279. 2:04:38don't want to take you off course and hopefully  this will help you stay on course and hopefully
  1280. 2:04:42it aids the course and LLM can process the word  torment. See. But what's the difference between
  1281. 2:04:48our human brain's autoregressive process that  creates the feeling of torment itself and the
  1282. 2:04:54word torment? So that, so my speculation here,  and it's, it is purely speculative, is that,
  1283. 2:05:01um, it's non-symbolic. The, there's something  happening when the universe gets represented
  1284. 2:05:08in our brain, it's still in our brain, it's still  a certain mapping, but when it gets represented,
  1285. 2:05:14so that physical characteristics of the world  are actually represented in a more direct
  1286. 2:05:25mapping. So think about color. We talked earlier  about sort of color space. There's a real true
  1287. 2:05:30sense in which red and orange are more similar in  physical color space, like there's actually some
  1288. 2:05:36physical fact about, and also in brain space.  That's my, my, my guess is that in brain space,
  1289. 2:05:44uh, there really is something about, and a better  way of saying it is that the physical similarity
  1290. 2:05:51has a true analog, and I use the word analog, um,  kind of specifically, a true analog in the analog,
  1291. 2:06:00uh, kind of physical cascade of activity that's  happening in the brain. Language is symbols. Is
  1292. 2:06:08non-symbolic a synonym for ineffable? I wouldn't  have thought of it that way, but that may be a
  1293. 2:06:15very good way to say it, uh, or to not say it. Uh,  yes, ineffable, well, by virtue of being symbolic,
  1294. 2:06:23by virtue of being a purely relational kind of, of  representation, which is what this, the language,
  1295. 2:06:30maybe even more than saying it's symbolic,  it's that it's represent, it's relational.
  1296. 2:06:34The language is a relational kind of, the, the  location in the space matters only because it's
  1297. 2:06:40of relation to other tokens in the space. That's  not true in color perception. In color perception,
  1298. 2:06:48where you are in the, in sort of the, probably  in the embedding space is going to have physical
  1299. 2:06:53meaning. It's going to be related to the physical  world in a much more direct way. And so the space,
  1300. 2:06:59even though it's an internal space, right? The,  the, the, the, the perception of color is still
  1301. 2:07:03just comes down to neurons firing. We're not  actually getting the light. The light's not
  1302. 2:07:07getting into our brains, but the mapping is such  that it preserves. And I even think of it as an
  1303. 2:07:16extension. It's not even like, it's not symbolic.  It's not a representation. It's an extension. It's
  1304. 2:07:22like when you drop a rock in water, the ripples  that happen afterwards are not a representation
  1305. 2:07:27of the rock, but they carry information about that  rock. But it's, it's just a physical extension.
  1306. 2:07:33It's a continued physical process. Interesting. I  think that's what's happening in, in, in sensory
  1307. 2:07:38processing. And I, I think that has something  deep to do with qualia. I don't think language
  1308. 2:07:44has that. I think because it's purely relational  there, it's not a rippling of anything. It's its
  1309. 2:07:50own system of relational embeddings that, that  aren't continuous in any way with the physical
  1310. 2:07:58universe. Do you think that has something to do  with God? Well, I think that if we think of the
  1311. 2:08:15grand unity of creation, there's some sense in  which language breaks that unity. And I think
  1312. 2:08:25that we can lie in language in a way that we can't  in any other substrate. Hmm. And so I think by
  1313. 2:08:36becoming purely linguistic beings, as the vast  majority of our time as humans is spent in the
  1314. 2:08:42linguistic space, that's where we're hanging out  there. Our minds are hanging out there. I think,
  1315. 2:08:49uh, we have perhaps forgotten something that,  that animals know about the universe. And it's
  1316. 2:08:59this kind of unity because they're, the animal  processing is, is, is an extension. It's a
  1317. 2:09:04continuation of the world. And, and, and since the  world, the universe is one thing in some sense,
  1318. 2:09:10it's, it's, it is a, a, everything, we don't even  have to get into non-locality, right? The origins
  1319. 2:09:17of the, you know, let's just talk about like the  big bang or something like that. What's happening
  1320. 2:09:20here now in some ways is connected quite literally  to what happened elsewhere for the way back in
  1321. 2:09:28time. So I think this sort of unity that, you  know, that mystics talk about, um, is much closer
  1322. 2:09:38to sort of the animal brains than the linguistic  brain because the linguistic brain actually
  1323. 2:09:44creates, is its economy. It breaks the continuity.  Symbols, I, I, I, I use it for, it's like a new
  1324. 2:09:51physics. Uh, the, the relations are what matters  and it's no longer continuous. It's no longer an
  1325. 2:09:57extension of the physical universe. It interacts  with the physical universe in a way that we,
  1326. 2:10:06as we see, we, we can sort of, uh, do this mapping  so that when I talk, it can have influence on the
  1327. 2:10:12physical universe. It can have influence on my  perception. It can influence on my behavior. I
  1328. 2:10:16think that sort of the rationalist movement, the,  the positivist movement, um, sort of modernity
  1329. 2:10:22itself is, is a, a complete hijacking of our  brain by, by the linguistic system. Um, and,
  1330. 2:10:30and I do think that has something to do with, uh,  the denouement, the, the kind of the, the, the,
  1331. 2:10:35the God is dead, um, kind of, uh, modernity equals  somehow, uh, the, the decline. And so, you know,
  1332. 2:10:44a rationalist would say, well, that's, that's  appropriate because we've figured out how the
  1333. 2:10:48universe works and we don't need any of this hocus  pocus. But what about the feeling of unity that,
  1334. 2:10:54what about, what about the sense of, of sort of  a cosmic whole? Um, are we so sure that we're
  1335. 2:11:04right and those ancients were wrong? And yes, I  think, I do think that, um, that this has, has
  1336. 2:11:12very significant consequences for thinking about  some of these, uh, intangibles, these ineffables.
  1337. 2:11:22So. A snake that mimics the poison of another  snake in terms of its color. That's a form of a
  1338. 2:11:29lie. Now, would you say that that is somehow  symbolic as well, though? No. And, and yes,
  1339. 2:11:36there is, uh, mimicry and there is, uh, you know,  certain certain sense in which animals can, can
  1340. 2:11:43engage. They're not, they don't even know they're  engaging in subterfuge. Um, but that's much more
  1341. 2:11:48continuous with, okay, you've just pushed the, the  cognitive agent into a slightly different space,
  1342. 2:11:54which is consistent with some other physical  reality. That's very, very different than saying,
  1343. 2:12:00uh, we are made of atoms and particles and, uh,  everything that happens is determined by, uh, the,
  1344. 2:12:08the forces amongst these atoms. None of which is  something that we have any material animal grasp
  1345. 2:12:14of any true physical grasp of. These are words.  Uh, these models are really words and they run in
  1346. 2:12:20words and they run very well to make predictions  and, and to manipulate the physical universe,
  1347. 2:12:27but they're stories and they're linguistic  stories. And those kinds of stories can be, uh,
  1348. 2:12:33I wouldn't even say they're. According to my own  theory, uh, language doesn't really have physical,
  1349. 2:12:42doesn't point to physical meaning. And so even  saying that it's a lie or untrue, isn't quite
  1350. 2:12:48right, but within its own space, you can go off in  many different directions. Uh, and, and it, I, I,
  1351. 2:12:56maybe the danger is not in, in, uh, thinking  of things that truth thinking about things,
  1352. 2:13:02um, thinking thoughts that aren't really true.  It's falling too deeply in love with the idea that
  1353. 2:13:08idea space and language space is the real space.  Um, that that's yes. Interesting. So see in our
  1354. 2:13:17circles. So when we're hanging out off air, when  we're hanging out with other professors and on
  1355. 2:13:22the university grounds and so on, we praise this  exchange of words and making models precise and
  1356. 2:13:29doing calculations and so on. And I've always  intimated that this is entirely incorrect.
  1357. 2:13:37And I haven't heard an anti philosopher, like a  philosopher that was an anti philosopher, except
  1358. 2:13:43one who was an ancient Indian philosopher. I think  his name is Jayarasi Bhatta. I'm likely butchering
  1359. 2:13:48that pronunciation, but I'll place it on screen.  Anyhow, who was arguing against the Buddhists and
  1360. 2:13:54the other contemporary philosophers by saying,  look, you think know thyself as what you should
  1361. 2:13:59be doing or what you didn't say like, like this,  but you think of it as the highest goal. However,
  1362. 2:14:06who is living more truly than a rooster? Like,  none of you are living more truly than just
  1363. 2:14:11something that's just being. Yes, exactly. That  is, that is the exact same intuition. Um, and,
  1364. 2:14:18and yes, it's this idea. I, I articulated it to  myself a long time ago for something that the fly
  1365. 2:14:23knows something that our linguistic system can  never know, uh, that it, it knows something. Uh,
  1366. 2:14:29it really does that, that, that the simply  existing and being is a form of knowledge
  1367. 2:14:35and it's a deeper one. It's a deeper one, uh,  than, than whatever it is that, you know, our
  1368. 2:14:41fancy rationalist kind of perspective has given  us. Um, our rationalist perspective is, is very,
  1369. 2:14:47very powerful in coordinating and predicting, but  in terms of like true ontology, I suspect it's,
  1370. 2:14:56uh, it's, it's actually the wrong direction. It's,  it's, uh, it's created a, a false, a false God of,
  1371. 2:15:07of linguistic knowledge, of shared objective  knowledge. When the subjective is the one that
  1372. 2:15:13we really have, it's, it's, it's the,  it's the Cartesian, right? That's the,
  1373. 2:15:17the Kogito. It's the, it's, it's what we know is  what we experience. That's the only thing we truly
  1374. 2:15:23know. And language, it doesn't really live there.  Language, it doesn't really live there. So I was
  1375. 2:15:33watching everything everywhere all at once.  I never saw it. Because I also had another
  1376. 2:15:38intimation. I'll spoil some of it. And if you are  listening and you don't want to spoil it, then
  1377. 2:15:42just skip ahead. But I was telling someone that  I think if there's a point of life, it's one of
  1378. 2:15:50two. And so this is just me speaking poetically  and not rigorously. One is to find a love that
  1379. 2:15:57is so powerful it outlasts death. Okay, so that's  number one. And then number two is to get to the
  1380. 2:16:05point in your life where you realize that all  your inadequacies and all your insecurities and
  1381. 2:16:08all your, your missteps and your, the, the, your  jealousies and your, and your malice and so on,
  1382. 2:16:18that it, rather than it being a weakness, it's  what led you to this place here. And here is the
  1383. 2:16:25optimal moment. It's to get that insight. So  I don't know how to rationally justify any of
  1384. 2:16:31that or explain it. But anyhow, when I said this  one time on a podcast, someone else said, Hey,
  1385. 2:16:36that latter one that you expressed was covered in  everything everywhere all at once. So I watched
  1386. 2:16:42it. And what was great about that movie, and  here's where I spoil it, is that, and it makes
  1387. 2:16:49me want to tear. The movie is silly and comedic  in a way that I, that didn't resonate with me.
  1388. 2:16:53But there's this one lesson that did. The, the  woman, she's a fighter, the main protagonist,
  1389. 2:17:00she's a fighter, and she's strong headed. And  she has this husband who is weak, and she's
  1390. 2:17:05always able to put down. And so then you think,  okay, well, this is a modern trope where there's
  1391. 2:17:09always the stronger woman, and every guy is like,  just, just a fool. And the woman is always more
  1392. 2:17:14intelligent and so on. Okay. So you just think  of it as, as okay, well, it's just, it's just a
  1393. 2:17:18modern trope. Toward the end, and the guy is, the  guy is kind and loving to people. Toward the end,
  1394. 2:17:25there was something with, she was getting  audited by the IRS. And she was supposed to,
  1395. 2:17:33something was supposed to happen that night where  she had to bring receipts, and she couldn't. Now,
  1396. 2:17:38the husband was talking with the IRS lady, and our  protagonist, the woman, was saying in Vietnamese,
  1397. 2:17:45or in Mandarin, whichever language it was,  was saying, oh, he's an idiot. I hope he
  1398. 2:17:50doesn't make it worse. The IRS lady then comes  to the woman and says, you have another week.
  1399. 2:17:56You have an extension. She's like, how did this  happen? She talks to the husband. And remember,
  1400. 2:18:00this is a movie almost about a multiverse,  so you're getting different versions of this.
  1401. 2:18:04And there's this one version where the husband's  speaking to her and telling her, you know, Evelyn,
  1402. 2:18:09the main character, you know, Evelyn, you see what  you do as fighting. You see yourself as strong and
  1403. 2:18:15you see me as weak and you see the world as a  cruel place. But I've lived on this earth just
  1404. 2:18:22as long as you, and I know it's cruel. My method  of being kind and loving and turning the cheek,
  1405. 2:18:33that's my way of fighting. I fight just like  you. And then you see that what he did in
  1406. 2:18:40another universe was he just spoke kindly to the  IRS agent and talked about something personal,
  1407. 2:18:45and that softened her. And then you see all  the other universes where she was trying to
  1408. 2:18:51go on this grand adventure and do some fighting.  And the husband then says, Evelyn, even though
  1409. 2:18:57you've broken my heart once again in another  universe, I would have really loved to just
  1410. 2:19:05do the laundry and taxes with you. And it makes  you realize you're aiming for something grand and
  1411. 2:19:16you're aiming to go out and conquer demons and  so on. But there's something that's so much more
  1412. 2:19:23intimate about these everyday scenarios. There's  something so rich. And the journey, there's also
  1413. 2:19:30a quote about this that the journey, I think  it's T.S. Eliot's, is to find, sorry, at the
  1414. 2:19:36end of all our exploring will be to arrive where  we started and know the place for the first time.
  1415. 2:19:44Anyhow, all of this abstract talk... No, no, no,  no. It's exactly what we're talking about. Because
  1416. 2:19:52if you see yourself as a ripple in the universes,  right, then you are part of something cosmic and
  1417. 2:20:00grand. And it's sort of that extensiveness.  It's that extensiveness. It's being here now.
  1418. 2:20:08It's that we aren't just atoms. We're part of a  larger thing. You can call it God. You can call
  1419. 2:20:17it the universe or whatever. But it's there. It's  actually something I think we, I don't think we
  1420. 2:20:26really, I think animals don't think of themselves  as discreet. I don't think they do. I think that
  1421. 2:20:34they don't think of an outside and inside. They  don't think of an objective and subjective. It's
  1422. 2:20:39just this unfolding. They have a theory of mind  or that, but these are linguistic concepts. And
  1423. 2:20:47I think I do, and I sound like an anti-linguist  and I recognize the power of it. I said before,
  1424. 2:20:54you know, how extraordinary it is, how rich  it is. And I have tremendous respect for it.
  1425. 2:21:00But at the same time, I do think that  all this talk about objective things,
  1426. 2:21:06particles, and we are physical bodies and we are  just this and we are just that, that is bullshit.
  1427. 2:21:12Like, no, we are the universe resonating. We are  part of this, part of the whole. And the way that
  1428. 2:21:20I think thinking objectively is, as language  requires you to do, actually it breaks it.
  1429. 2:21:26So I think there's such a beauty in the silence.  And it's why it's something everybody knows that
  1430. 2:21:35the ineffable, why is it called the ineffable?  Why is the ineffable that the ineffable isn't
  1431. 2:21:40just that you can't say it, it's magnificent.  The ineffable is extraordinary. Why? Because it's
  1432. 2:21:52this true extension, something like that. Again,  I'm trying to put it into words, right? Therein
  1433. 2:21:57lies the trap, but we're both feeling it. Well,  I'm feeling extremely grateful to have met you,
  1434. 2:22:10to have spent so long with you. And there are  many conversations you and I have had that we
  1435. 2:22:15need to finish that are off air as well. So  hopefully we can do that. And thank you for
  1436. 2:22:19spending so long with me here. This was wonderful,  Curt. Thank you so much. I just want to hang out
  1437. 2:22:24for some more and talk about this stuff. So really  appreciate it. I've received several messages,
  1438. 2:22:31emails, and comments from professors saying  that they recommend theories of everything
  1439. 2:22:35to their students. And that's fantastic. If  you're a professor or a lecturer, and there's
  1440. 2:22:40a particular standout episode that your students  can benefit from, please do share. And as always,
  1441. 2:22:44feel free to contact me. New update, started a  Substack. Writings on there are currently about
  1442. 2:22:51language and ill-defined concepts, as well as  some other mathematical details. Much more being
  1443. 2:22:56written there. This is content that isn't anywhere  else. It's not on theories of everything. It's not
  1444. 2:23:00on Patreon. Also, full transcripts will be placed  there at some point in the future. Several people
  1445. 2:23:06ask me, hey Curt, you've spoken to so many people  in the fields of theoretical physics, philosophy,
  1446. 2:23:11and consciousness. What are your thoughts? While  I remain impartial in interviews, this Substack
  1447. 2:23:17is a way to peer into my present deliberations  on these topics. Also, thank you to our partner,
  1448. 2:23:25The Economist. Firstly, thank you for watching.  Thank you for listening. If you haven't subscribed
  1449. 2:23:32or clicked that like button, now is the time to  do so. Why? Because each subscribe, each like,
  1450. 2:23:39helps YouTube push this content to more people,  like yourself. Plus, it helps out Curt directly,
  1451. 2:23:45aka me. I also found out last year that external  links count plenty toward the algorithm,
  1452. 2:23:50which means that whenever you share on  Twitter, say on Facebook or even on Reddit,
  1453. 2:23:55etc., it shows YouTube, hey, people are talking  about this content outside of YouTube, which in
  1454. 2:24:01turn greatly aids the distribution on YouTube.  Thirdly, you should know this podcast is on
  1455. 2:24:06iTunes, it's on Spotify, it's on all of the audio  platforms. All you have to do is type in theories
  1456. 2:24:12of everything and you'll find it. Personally, I  gain from re-watching lectures and podcasts. I
  1457. 2:24:16also read in the comments that, hey, TOE listeners  also gain from replaying. So how about instead you
  1458. 2:24:22re-listen on those platforms like iTunes, Spotify,  Google Podcasts, whichever podcast catcher you
  1459. 2:24:27use. And finally, if you'd like to support more  conversations like this, more content like this,
  1460. 2:24:33then do consider visiting patreon.com slash  CURTJAIMUNGAL and donating with whatever you
  1461. 2:24:38like. There's also PayPal, there's also crypto,  there's also just joining on YouTube. Again,
  1462. 2:24:43keep in mind, it's support from the sponsors and  you that allow me to work on TOE full-time. You
  1463. 2:24:49also get early access to ad-free episodes,  whether it's audio or video. It's audio in
  1464. 2:24:53the case of Patreon, video in the case of  YouTube. For instance, this episode that
  1465. 2:24:57you're listening to right now was released a  few days earlier. Every dollar helps far more
  1466. 2:25:02than you think. Either way, your viewership  is generosity enough. Thank you so much.

About this transcript

This page contains the full transcript of Language Without Meaning: How LLMs Exposed Our Biggest Illusion by Curt Jaimungal, generated from the public captions YouTube serves with the video. The transcript has 24,928 words across 1,466 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.