YouTube2Text

Do AI Tokenomics Matter More Than Model Benchmarks? — Transcript

by The TWIML AI Podcast with Sam Charrington · 10,758 words · 1,711 segments · language en · Watch on YouTube

Full transcript

  1. 0:00We started to enter a phase of AI that's
  2. 0:02not just about making models smarter.
  3. 0:04It's also about making them economically
  4. 0:07sustainable.
  5. 0:08As reasoning models consume more tokens,
  6. 0:10context windows continue to grow, and
  7. 0:13agents become embedded in more products
  8. 0:15and workflows, the economics of these
  9. 0:17systems are becoming impossible to
  10. 0:19ignore. That's given rise to a new
  11. 0:21conversation around tokenomics, how we
  12. 0:24think about the costs, incentives, and
  13. 0:25tradeoffs shaping the next generation of
  14. 0:27AI.
  15. 0:29One person who's been thinking deeply
  16. 0:30about this is Stanford Professor and Big
  17. 0:33Spin co-founder Chris Potts. His recent
  18. 0:36work argues that measuring AI progress
  19. 0:38requires looking beyond model benchmarks
  20. 0:40to ask a different question. What are
  21. 0:42our tokens actually buying us?
  22. 0:45Here's Chris explaining how he thinks
  23. 0:47about tokenomics.
  24. 0:48>> Another interesting moment to be in as
  25. 0:51we're all being made aware of the true
  26. 0:53costs of all this AI usage. The analogy
  27. 0:55here is like it used to cost me $20 to
  28. 0:58take a ride share to the airport for
  29. 1:00Uber or Lyft, and now it costs 90. But
  30. 1:02it's more like 20 to like 500 or
  31. 1:05something, right? And I think what's
  32. 1:06happening is that the big providers are
  33. 1:08testing the waters on charging us the
  34. 1:11true costs plus whatever profit they
  35. 1:13need to make as they all try to gear up
  36. 1:15for IPOs and so forth. And in turn, that
  37. 1:18is very quickly leading people to ask
  38. 1:20questions like what is the return on
  39. 1:22investment for all these tokens that we
  40. 1:24have purchased. And it's a very tricky
  41. 1:26area to be in because what does it mean
  42. 1:28to think about value in this context?
  43. 1:31>> I'm Sam Charrington, and this is the
  44. 1:33Twilio AI podcast. For over a decade,
  45. 1:36I've been exploring the ideas and
  46. 1:37innovation shaping the future of AI
  47. 1:39through conversations like this one that
  48. 1:41help you understand what's real, what's
  49. 1:44next, and what matters. Let's jump in.
  50. 1:55I want to say thanks for coming on. I've
  51. 1:57been looking forward to this
  52. 1:59conversation and I think where I'd love
  53. 2:03to start us off is to really dig into
  54. 2:07your background and how it kind of got
  55. 2:10you, you know, to to where you are now.
  56. 2:13>> Yeah, my background is in linguistics.
  57. 2:16Um, linguistics proper, not even natural
  58. 2:18language processing. I did my PhD on,
  59. 2:21among many other things, swears.
  60. 2:24What swears are like, why we swear, what
  61. 2:26information they encode, what kind of
  62. 2:27taboos exist around them, and so forth.
  63. 2:30And that was actually the trigger that
  64. 2:32got me into NLP because I wanted a lot
  65. 2:35of data of people swearing. I wanted to
  66. 2:37know what the context was like, what
  67. 2:39their intentions were. So, I turned to
  68. 2:41corpora.
  69. 2:42And from there you start using NLP
  70. 2:45toolkits to add structure to those
  71. 2:47corpora. And then after a few years,
  72. 2:49maybe you're writing your own tools for
  73. 2:51doing that work.
  74. 2:52And then when you look back after 18
  75. 2:54years or whatever it's been, you're just
  76. 2:56an AI person or an NLP person. But that
  77. 2:59is the true story and I I feel like if I
  78. 3:01had to, I could trace the lineage of
  79. 3:03every one of my current projects back to
  80. 3:05my fascination with why we care when
  81. 3:07someone drops an F-bomb.
  82. 3:10>> [laughter]
  83. 3:12>> So, are you a an F-bomb dropper or did
  84. 3:15you come at it from the perspective of
  85. 3:17trying to understand these others?
  86. 3:19>> [laughter]
  87. 3:20>> I think very infrequently in my life.
  88. 3:23Um,
  89. 3:24on for my linguistics class, semantics
  90. 3:26and pragmatics, which is about
  91. 3:27linguistic meaning, on the final day we
  92. 3:29always do a class on swearing. And I
  93. 3:32review the history and we kind of tie
  94. 3:34all the course themes together. My
  95. 3:36handouts for that are full of swears.
  96. 3:39But I only swear once in the lecture. I
  97. 3:42I I present the result that people
  98. 3:44remember things better if the utterance
  99. 3:46contains a swear because it has a kind
  100. 3:48of emotional resonance, very primitive
  101. 3:50reaction. And so in that moment I pick
  102. 3:53some fact from the course, some trivial
  103. 3:55thing,
  104. 3:56and I restate it with a swear, and then
  105. 3:58I say, "All of you will remember this
  106. 3:59for eternity."
  107. 4:01But other than that, I'm very shy about
  108. 4:03it in the class.
  109. 4:04>> That's funny. I'm sure there is loads of
  110. 4:08research on this, but
  111. 4:11uh
  112. 4:11I grew up in New York City, and as a New
  113. 4:14Yorker, I think that swearing is just
  114. 4:16kind of part of my natural language and
  115. 4:19way of communicating.
  116. 4:21And I married a a Midwestern girl, and
  117. 4:24she doesn't tolerate it at all. She
  118. 4:26doesn't do it, she doesn't tolerate it,
  119. 4:28she won't tolerate it for me. And it
  120. 4:31made for We've been married for 30
  121. 4:33years, so I adapt quickly, apparently.
  122. 4:35But uh
  123. 4:37it, you know, for a long time, it took a
  124. 4:39lot of restraint to like change that way
  125. 4:44of communicating, particularly when I'm
  126. 4:45communicating about something that I'm
  127. 4:47excited about or emotional about or, you
  128. 4:49know, want to convey the importance of
  129. 4:53um It's a really interesting topic, and
  130. 4:56uh
  131. 4:58um Well, we're not going to turn the
  132. 5:00podcast into a podcast about swearing,
  133. 5:03but I imagine there's enough research
  134. 5:05there that we could if we wanted to.
  135. 5:06>> It's a fascinating area, yeah, because
  136. 5:08it gets right to the heart of the
  137. 5:10culture that we've constructed and how
  138. 5:12it relates to our usage and everything
  139. 5:14else about us. Yeah, it's fascinating
  140. 5:16that we have swears. When the old swears
  141. 5:18lose their power, we invent new ones. We
  142. 5:20pretend like nobody should use them, but
  143. 5:22as you say, people use them all the
  144. 5:23time, and it feels like an important
  145. 5:25part of being a language user that we've
  146. 5:27got them available to us. Yes, endless
  147. 5:30string of questions.
  148. 5:32>> I'd love to hear your take on
  149. 5:35kind of a linguist in the age of modern
  150. 5:39AI, you know, transformers, statistical
  151. 5:42models, you know, this is a
  152. 5:44uh NLP used to be kind of coming from a
  153. 5:47linguistic perspective and now the
  154. 5:49entire field is shifted to
  155. 5:51statistical perspective and I'd love to
  156. 5:53hear your reflections on
  157. 5:56being on the other side of that
  158. 5:57transition
  159. 5:59as well as maybe more importantly ways
  160. 6:01that you think that kind of the
  161. 6:02traditional foundational linguistics is
  162. 6:04still important to
  163. 6:07the the way we think about AI today.
  164. 6:09>> These questions are on my mind all the
  165. 6:11time yeah because I operate at the
  166. 6:13intersection of all these different
  167. 6:14fields and I will say it's it's useful
  168. 6:16to distinguish in this context
  169. 6:18linguistics you know and people in my
  170. 6:20department at Stanford study language
  171. 6:22and social identity historical
  172. 6:24linguistics the structure of language
  173. 6:27and they're just doing scientific
  174. 6:29investigation of language as a human
  175. 6:31phenomenon and they are not
  176. 6:32technologists and they're not trying to
  177. 6:34inform technology.
  178. 6:35So their project is
  179. 6:37interestingly impacted by technological
  180. 6:40developments. For NLP people who are of
  181. 6:42course participating directly in the
  182. 6:44engineering project they're affected in
  183. 6:47a very different way by the rise of gen
  184. 6:48AI and the kind of homogeneous nature of
  185. 6:51the solutions that people now adopt in
  186. 6:53that space.
  187. 6:54So for the linguists I feel like this is
  188. 6:57the most exciting moment that anyone
  189. 6:58could have dreamed of. I feel incredibly
  190. 7:00privileged to be alive in this moment
  191. 7:04where humans encounter for the very
  192. 7:06first time non-human creatures that use
  193. 7:09our language very fluently.
  194. 7:11I think it's weirding us all out but
  195. 7:13from the point of view of understanding
  196. 7:15the human capacity for language what a
  197. 7:18gift because you can ask about the
  198. 7:20mechanisms
  199. 7:22which are different from humans but
  200. 7:23obviously sufficient for achieving a
  201. 7:25certain kind of behavioral performance.
  202. 7:28Um
  203. 7:29we can think about them as investigative
  204. 7:31tools. I mean we train them on the
  205. 7:32internet they're basically incredibly
  206. 7:35powerful distributional learners and we
  207. 7:36can all learn a lot from them about the
  208. 7:38true structure of language by just
  209. 7:40looking at the kinds of things that they
  210. 7:42learn. And it really gets at the heart
  211. 7:44of core questions in linguistics about
  212. 7:47how much of language learning is innate
  213. 7:49and the nature of our capacity and
  214. 7:51whether it's statistical or symbolic.
  215. 7:53All those things come flooding in in a
  216. 7:55completely fresh way. And so whatever
  217. 7:57your reaction to language models
  218. 8:00is, it should be a significant one,
  219. 8:02right? This should be causing you to
  220. 8:04rethink key questions. And that's all
  221. 8:07you could hope for as a scientist, that
  222. 8:08you have new angles, new perspectives,
  223. 8:11new questions reopen. That's been
  224. 8:13incredible.
  225. 8:15For NLP, I think it's a more uncertain
  226. 8:18prospect because
  227. 8:20pre
  228. 8:22the the arrival of like pre-trained
  229. 8:25models, which for me would be like the
  230. 8:27Elmo model back in 2017, 2018. Before
  231. 8:31that, there was still a lot of
  232. 8:33statistical work, of course, and we were
  233. 8:34in the deep learning era.
  234. 8:36But you could still, for example, do a
  235. 8:38PhD that was entirely about some
  236. 8:40specific phenomenon and maybe some very
  237. 8:42specific tweak to a model. So you could
  238. 8:44say, "I'm going to work on summarization
  239. 8:45and I've got a new idea about how to do
  240. 8:47that well using deep learning models."
  241. 8:50And that could be your PhD. And what we
  242. 8:52started to see 2018, 2019, 2020,
  243. 8:55especially with the arrival of GPT-3,
  244. 8:57that that was a very uncertain prospect
  245. 8:59because you might wake up one morning to
  246. 9:01find that you had been completely
  247. 9:02scooped. That with essentially no
  248. 9:04effort, one of these large pre-training
  249. 9:06runs had done better than you at the
  250. 9:08thing that you'd worked so hard on.
  251. 9:10And that caused an interesting, probably
  252. 9:12overall productive, but interesting and
  253. 9:14challenging crisis for people,
  254. 9:16especially students who were trying to
  255. 9:18figure out what to do next with their
  256. 9:20PhD research. But I think all of us felt
  257. 9:23a kind of real uncertainty in that
  258. 9:24moment.
  259. 9:25>> Yeah, I remember the anxiety of that
  260. 9:28time and
  261. 9:31I always felt it was kind of expressed
  262. 9:32as
  263. 9:34you know, is research in NLP
  264. 9:36fundamentally like scale limited or do
  265. 9:39you need a certain degree of scale that
  266. 9:41only a handful of organizations have to
  267. 9:43do foundational research and is everyone
  268. 9:47else going to be relegated to like
  269. 9:50poking the poking the beast and seeing
  270. 9:53what it does?
  271. 9:54And I'm curious do you feel
  272. 9:57like that was an anxiety that's passed
  273. 9:59or is it still very present? Has it
  274. 10:02panned out quite like that? How How do
  275. 10:04you you know what how's it been resolved
  276. 10:06for you?
  277. 10:07>> Also fascinating. Not resolved. It's
  278. 10:09something I discuss a lot with my
  279. 10:11collaborators and with my students.
  280. 10:13We're all trying to figure this out in
  281. 10:14this moment. I will say one concrete
  282. 10:17thing we did was orient a lot of our
  283. 10:19research toward interpretability.
  284. 10:22Uh just the project of understanding how
  285. 10:24these models end up being so good at
  286. 10:26such hard tasks.
  287. 10:28And the reason we did that is it's
  288. 10:29relatively inexpensive and it's also an
  289. 10:31area where clearly you would be
  290. 10:33explicitly hoping that models would get
  291. 10:35better because then there would be more
  292. 10:37to explain.
  293. 10:39Versus if you were doing that
  294. 10:41summarization project, you might quietly
  295. 10:43be hoping that there wasn't going to be
  296. 10:44so much progress so that you could make
  297. 10:46the progress. Like let's hope the next
  298. 10:48model isn't good at summarization. I
  299. 10:50want to be the star of that show. That's
  300. 10:52as I said very uncertain but if you're
  301. 10:53doing mech interp, you're like let's get
  302. 10:55the new model released because now we're
  303. 10:56going to have even more structure to
  304. 10:58find, even more to explain. And that
  305. 11:00felt like a very productive choice. I
  306. 11:02don't want to leave out the fact that
  307. 11:03it's also cheaper to do this research
  308. 11:05and that is significant.
  309. 11:07And then I would say that right now a
  310. 11:08lot of us are in a moment of thinking we
  311. 11:11should do stuff that is weird and
  312. 11:13creative and out of the mainstream. We
  313. 11:16should be thinking about trying to
  314. 11:18achieve the next big thing because
  315. 11:20competing with these massively
  316. 11:22resourced, incredibly creative and
  317. 11:24talented teams is just not a winning
  318. 11:26game. So let's play a different game and
  319. 11:29hope that that's as they say where the
  320. 11:31puck is going, not where it is.
  321. 11:33>> And what are some examples of that kind
  322. 11:35of thinking?
  323. 11:36>> We've been thinking a lot about
  324. 11:37architectures cuz I have a lot of
  325. 11:38complaints about current architectures.
  326. 11:41And I would say the other main theme
  327. 11:42right now for us in my group is thinking
  328. 11:45about data.
  329. 11:46You know, data have strange and wondrous
  330. 11:49properties. I think we don't understand
  331. 11:50how data affect models.
  332. 11:52And that has all sorts of implications
  333. 11:54for security and safety and also the
  334. 11:56nature of the learning that these models
  335. 11:58do. It really data is fundamental. It's
  336. 12:00all data-driven learning. And so telling
  337. 12:02the full causal story from data to final
  338. 12:04model state
  339. 12:06feels like it will just be significant
  340. 12:08for lots of questions.
  341. 12:11But I wouldn't want to leave out the
  342. 12:12architecture one because
  343. 12:14I feel like the architecture everyone
  344. 12:16has arrived at, these stacked
  345. 12:18transformers that we make very deep and
  346. 12:20very large, are tremendously
  347. 12:22inefficient.
  348. 12:24You would hope they were using all that
  349. 12:26depth and all that representational
  350. 12:27power to learn modular recursive
  351. 12:32functions for things and all sorts of
  352. 12:33exciting stuff. It is not what we find
  353. 12:35and that seems like a real opportunity
  354. 12:37to just level up and do better and maybe
  355. 12:40we could get massively more capable
  356. 12:42models with half the depth and a quarter
  357. 12:45of the representational width. And that
  358. 12:47would be transformative for the
  359. 12:48economics of AI in addition to leading
  360. 12:50to all sorts of exciting things for
  361. 12:52capabilities.
  362. 12:53>> Yeah, it's funny and maybe a bit
  363. 12:55validating for me to hear you say that
  364. 12:57because whenever I architect arch
  365. 13:00whenever I articulate a thought in that
  366. 13:02direction
  367. 13:03with
  368. 13:04uh particularly with folks that are you
  369. 13:06know, coming from the frontier labs or
  370. 13:09you know, the the
  371. 13:11um
  372. 13:12essentially the frontier labs
  373. 13:14I get back this kind of feeling that
  374. 13:16yeah, you're just not bitter lesson
  375. 13:17piled enough. Like structures, that's
  376. 13:21old school thinking, you know, you're
  377. 13:23just trying to like train some features.
  378. 13:26Just collect a lot of data, throw it at
  379. 13:28the model, and that's all you need.
  380. 13:30>> Okay, but here's my response to them.
  381. 13:33Let's say rewind to 2017. We've got the
  382. 13:35transformer.
  383. 13:36It's got absolute positional encodings,
  384. 13:39and it's got a particular structure for
  385. 13:42its uh MLP layer, which is pretty narrow
  386. 13:44and pretty dense,
  387. 13:46and a certain structure to its
  388. 13:48activations and its uh layer norms.
  389. 13:50That's 2017.
  390. 13:53The bitter lesson pill thing to do would
  391. 13:54be to scale that up.
  392. 13:56But just consider, for example, how much
  393. 13:59it would cost to use the N squared
  394. 14:02attention and the absolute positional
  395. 14:03encodings but have a context window of 1
  396. 14:05million. This is the bitter lesson pill
  397. 14:08thing, right? Just keep scaling. But it
  398. 14:09would be absurd. It would cost trillions
  399. 14:11of dollars to produce models that we all
  400. 14:13interact with right now. What did people
  401. 14:15do instead? They thought hard about
  402. 14:16locality, and they thought about how
  403. 14:19like positional encodings should be
  404. 14:21favoring local relationships. They
  405. 14:23completely rethought the MLPs so that
  406. 14:25it's now wide and sparse. Everyone did
  407. 14:27careful work on the activation functions
  408. 14:30to make sure there weren't weird
  409. 14:31outliers so that they could quantize in
  410. 14:33a good way. And so forth and so on. All
  411. 14:36of this analysis work built on
  412. 14:37intuitions about data and learning led
  413. 14:40to the model that we have now, which is
  414. 14:41like a ship of Theseus compared to the
  415. 14:432017 transformer. The only thing that
  416. 14:44survives is attention and the feed
  417. 14:47forward layer.
  418. 14:48And I claim for you that none of that
  419. 14:50stuff is bitter lesson pill. That was
  420. 14:52all analysis work that was meant to save
  421. 14:54based on priors in the data and priors
  422. 14:56about how they knew learning would
  423. 14:57happen. So I go back at them. You're not
  424. 15:00bitter lesson pill enough, apparently.
  425. 15:02Although, this is a reductio, I think.
  426. 15:06>> Oh, I love this. That's such a great
  427. 15:07response.
  428. 15:08I think it also really calls out the
  429. 15:13relationship between data, mech and
  430. 15:16terp, and efficiency, like core themes
  431. 15:19that you've been focused on and how
  432. 15:21they,
  433. 15:22you know, interrelate and support one
  434. 15:24another.
  435. 15:25>> Yeah, absolutely. And this relates to
  436. 15:26one of my hot takes, you know, it's very
  437. 15:27fashionable, especially among inter
  438. 15:29researchers, but I think in general for
  439. 15:30people to say, "We don't understand how
  440. 15:33these models work. It is also very
  441. 15:34mysterious to us." But the truth is that
  442. 15:37people in the field have very deep
  443. 15:39intuitions about how these models work,
  444. 15:41and that is the causal factor in us
  445. 15:44making so much progress, because they
  446. 15:46could think analytically, "What would
  447. 15:48the structure of positional encodings
  448. 15:50and attention be I could do this at
  449. 15:52million context scale." You can only
  450. 15:54achieve that kind of thing based on deep
  451. 15:56analysis and insight, not by just
  452. 15:58guessing.
  453. 15:59And so when people say, "Oh, we don't
  454. 16:01know understand." I say, "I think you
  455. 16:02understand much better than you're
  456. 16:04letting on. I think you understand at
  457. 16:06least as well as my car mechanic
  458. 16:08understands how my car works." There are
  459. 16:10mysteries,
  460. 16:11but you can take a lot of action and be
  461. 16:13very effective improving things.
  462. 16:15>> Why do you think they say
  463. 16:17that? Why do you think they say that
  464. 16:18they don't understand the models?
  465. 16:20There's just got to be some payback
  466. 16:21there.
  467. 16:22>> It's probably a paradox of expertise,
  468. 16:24right? So the more you do know, the more
  469. 16:26you feel like there are also mysteries,
  470. 16:28and it's hard to step back from that and
  471. 16:30be objective and say, "Yeah, well, we
  472. 16:31did make a phenomenal amount of
  473. 16:33progress, and that can't be just because
  474. 16:35of happenstance. That was because we
  475. 16:37know a lot." But all you see as an
  476. 16:39expert is all the things that are still
  477. 16:41to be explained.
  478. 16:43Um partly also it's just a narrative in
  479. 16:45the field, and it does stretch back to
  480. 16:46days when I think we had very little
  481. 16:48understanding of how these models work.
  482. 16:50Possibly because a lot of them weren't
  483. 16:51that good. There was very little to
  484. 16:53explain. And so that's just been slow to
  485. 16:55catch up with how much progress we have
  486. 16:57made in understanding the kind of
  487. 16:59intuitive human-level mechanisms that
  488. 17:01these models are operating with.
  489. 17:03>> I also wanted to ask you about DSPY. I
  490. 17:07forgot about this as we were talking
  491. 17:09earlier, but you were involved in DSPY,
  492. 17:13which um
  493. 17:14well, I'll let you talk about it, but
  494. 17:16I'm curious
  495. 17:18uh how it connects into your research
  496. 17:21and like
  497. 17:22uh you know, some of these pillars that
  498. 17:24we've we've talked about.
  499. 17:26>> Oh, there's lots of wonderful strands.
  500. 17:28And what a meta strand I could offer you
  501. 17:30cuz we were talking about being
  502. 17:31strategic with research, this does stem
  503. 17:34from Omar Khattab, my student. He's the
  504. 17:36the the visionary behind DSPy and still
  505. 17:38its lead.
  506. 17:39And he just had the intuition early on
  507. 17:41that we should rethink what it means to
  508. 17:43make a scientific contribution.
  509. 17:45Previously, we thought in terms of
  510. 17:46papers as the beginning and the end of
  511. 17:48all of this kind of thing that you would
  512. 17:50contribute.
  513. 17:52We should instead, he said, think about
  514. 17:54projects and about empowering people.
  515. 17:56And so for him, the paper is one part of
  516. 17:59a broader contribution that might
  517. 18:01actually be centered on
  518. 18:03an open-source or open-weights release
  519. 18:06that would allow people to do big
  520. 18:08things. And that's where you find impact
  521. 18:10and that's the nature of a contribution
  522. 18:12going forward.
  523. 18:13And DSPy is a kind of embodiment of
  524. 18:15that. Although, he made a similar
  525. 18:16investment with the ColBERT retrieval
  526. 18:18model.
  527. 18:19And then people built on what he did,
  528. 18:22and then you really saw it take off
  529. 18:23where open-source contributions made it
  530. 18:25easier and easier to use that
  531. 18:26technology, leading to more and more
  532. 18:28impact. And of course, DSPy is another
  533. 18:31wonderful example because in investing
  534. 18:34in this community
  535. 18:35and in the open-source resource itself,
  536. 18:38he built a huge following.
  537. 18:40There are lots of startups, mine
  538. 18:41included, where the core tech stack for
  539. 18:43the LLMs is built on DSPy, and that has
  540. 18:46made life so much easier. And then of
  541. 18:48course, it was a platform for him and
  542. 18:50for us to really think in an innovative
  543. 18:52way about prompt optimization and
  544. 18:55agentic workflows and all of those
  545. 18:57things.
  546. 18:58>> Yeah, I was thinking not too long ago
  547. 19:00the degree to which model strength
  548. 19:06as a correlate to model size, I suppose,
  549. 19:08and capability
  550. 19:11has
  551. 19:12kind of overcome the need for an
  552. 19:15explicit framework like DSPy DS Pi.
  553. 19:19>> Yeah, there are kind of two levels to
  554. 19:20that. The one would be just the
  555. 19:21engineering side where
  556. 19:24uh DS Pi is great four years ago because
  557. 19:27it's kind of hard to construct the code
  558. 19:29around one of these systems in a way
  559. 19:30that's modular and reproducible and so
  560. 19:32forth because packing out something
  561. 19:34where you've got a prompt string in the
  562. 19:36middle of your code with some slots in
  563. 19:38it, it's very error-prone and it leads
  564. 19:40to bad system designs. And DS Pi solved
  565. 19:42that. And you could think that the need
  566. 19:44for that is diminishing somewhat because
  567. 19:47now we all specify these systems in
  568. 19:49English and have the coding agents do
  569. 19:51them.
  570. 19:52>> And even before that there were, you
  571. 19:54know, another hundred frameworks that
  572. 19:56solved that particular part of the
  573. 19:57puzzle.
  574. 19:58>> Oh, yeah, there's always competition and
  575. 19:59I think at that level of just thinking
  576. 20:01about programming interfaces and APIs,
  577. 20:04they can all learn from each other. And
  578. 20:05so like, you know, DS Pi learned a lot
  579. 20:07from PyTorch in terms of layer-wise
  580. 20:09design and the kind of modularity that
  581. 20:11introduced. And then, of course, you
  582. 20:13would hope that everyone kind of slurps
  583. 20:15up all these interesting innovations and
  584. 20:17it leads to everyone being better.
  585. 20:18There's lots of evidence of that at the
  586. 20:20level of interfaces. I would maintain
  587. 20:22for you that even if we have agents
  588. 20:23actually writing the code for these
  589. 20:24systems, it's great for us and for them
  590. 20:27if they write it in something that
  591. 20:28actually expresses these systems as
  592. 20:30modular components so that we can audit
  593. 20:32them, so that they can change them. It
  594. 20:34just feels like good engineering
  595. 20:35practices for any agent to think in a
  596. 20:37modular way. And that's what DS Pi
  597. 20:40encodes. The other side is like the
  598. 20:42prompt optimization side and a belief
  599. 20:45people have that the need to be careful
  600. 20:47with your prompts is diminishing over
  601. 20:49time.
  602. 20:50I understand that narrative, but people
  603. 20:52should also, for example, just run like
  604. 20:55a simple annotation study where they use
  605. 20:57a few different models or the same model
  606. 20:59a few times on slightly different data.
  607. 21:01They will be blown away by the amount of
  608. 21:04variation that still exists.
  609. 21:06To be charitable, let's say that these
  610. 21:08LLMs disagree about fundamental facts
  611. 21:10about how to label certain texts or what
  612. 21:13kind of response to give.
  613. 21:14We all kind of slip past this because we
  614. 21:16feel like, "Hey, they're smart and
  615. 21:17they're good and they're getting
  616. 21:18better." But if you quantify it, it's
  617. 21:20pretty disturbing. And the next step
  618. 21:22from that is to think about having all
  619. 21:24those agents optimize a prompt so that
  620. 21:26their behavior is at least consistent.
  621. 21:28And then you're right back at that DESP
  622. 21:30vision.
  623. 21:31>> It's interesting that you say that
  624. 21:33because
  625. 21:35I don't I don't feel like that
  626. 21:36necessarily aligns with my recent
  627. 21:39experience. And in particular,
  628. 21:42one thing that I've noticed that's been
  629. 21:44surprising is
  630. 21:47how
  631. 21:49well aligned, I guess. Maybe that's not
  632. 21:52the right word, but how similar the
  633. 21:53responses I get to uh uh
  634. 21:56query across different models. So, for
  635. 21:58example,
  636. 21:59you know, these are, you know, often
  637. 22:02kind of what I would call like a casual
  638. 22:04prompt, a casual query, something that I
  639. 22:06might, you know, uh type into Google.
  640. 22:09And it will now generate uh an LLM
  641. 22:12response for me in this kind of AI mode.
  642. 22:15Um
  643. 22:17and I'll take the same thing and put it
  644. 22:19into ChatGPT and maybe Claude. And it
  645. 22:23surprises me that the
  646. 22:26the responses are often very, very
  647. 22:28similar. Like, you know, very similar
  648. 22:30structure, very similar facts, very
  649. 22:33similar citations.
  650. 22:36And
  651. 22:37you know, I
  652. 22:39stepping back, like, there are lots of
  653. 22:40ways that they could answer or approach
  654. 22:42these different questions. But it seems
  655. 22:44like, you know, the models or the
  656. 22:45training or the system prompts or
  657. 22:48something is all kind of converged on
  658. 22:50something that makes the models express
  659. 22:52themselves, you know, very similarly,
  660. 22:55which, you know, seems to be at odds
  661. 22:57with uh
  662. 22:58you know, the the last thing you said
  663. 22:59about the need to optimize prompts or
  664. 23:02the impact of the individual prompt.
  665. 23:04>> I'm open-minded, but for example, like
  666. 23:06we just did a we did a we did a thing
  667. 23:07recently, we were writing a grant and we
  668. 23:09needed a title and you want to be
  669. 23:10strategic with these titles. So, we come
  670. 23:12up with a whole bunch of them ourselves
  671. 23:13and then we all disagree on what would
  672. 23:15be the best. So, let's find out what the
  673. 23:17agents think. So, ask a few Anthropic
  674. 23:19models and a few um
  675. 23:21GPT models, which of these five titles,
  676. 23:24which is the best?
  677. 23:25So, you get a different answer from all
  678. 23:27of them along with a detailed rationale
  679. 23:29about why obviously, of course, the
  680. 23:31choice that the model has made in that
  681. 23:32moment is the best one.
  682. 23:34This is great because then we can think
  683. 23:36about which one of these arguments is
  684. 23:37most persuasive.
  685. 23:39But if you were hoping for consistency
  686. 23:41at an subjective labeling task, which
  687. 23:43this is one,
  688. 23:44uh you can see right there that you're
  689. 23:45going to have a real problem unless you
  690. 23:47give very specific criteria
  691. 23:50and then you're kind of also
  692. 23:51constructing a prompt for them and you
  693. 23:53might want to manage them differently.
  694. 23:55There is a real I don't have evidence
  695. 23:57for this yet, but we have an intuition
  696. 23:59at Big Spin in the research we've done
  697. 24:01that you get a kind of paradox that the
  698. 24:03more requirements you add actually the
  699. 24:05more variation you'll see because the
  700. 24:07different models will key into different
  701. 24:08subparts of the requirements and since
  702. 24:11they do it very concertedly,
  703. 24:13you can actually get systematically
  704. 24:15biased behavior from something that you
  705. 24:17thought was a very good specification.
  706. 24:20And that again calls for this idea that
  707. 24:22what you need to do is figure out what
  708. 24:23the labels ought to look like and then
  709. 24:25have some automatic optimization process
  710. 24:28get the model there.
  711. 24:30And that's what things like Jeppa and
  712. 24:32MeetPro are for.
  713. 24:34>> A topic that I really wanted to A topic
  714. 24:37that I would really like to dig in to
  715. 24:39with you based on our our previous
  716. 24:42conversation was the idea of tokenomics.
  717. 24:45Uh it's something that people are
  718. 24:47talking about a lot recently.
  719. 24:50Um I think, you know, folks that use
  720. 24:52Claude code, for example, have like a a
  721. 24:55visceral experience with Anthropic
  722. 24:57changing the terms around usage, but
  723. 24:59it's happening under the covers with all
  724. 25:01of these large providers.
  725. 25:04And so, I think
  726. 25:06way more now than
  727. 25:09you know, 6 months ago, like we're all
  728. 25:13a little antsy with the relationship we
  729. 25:14have with these, you know, big model
  730. 25:16providers and the the value that we get.
  731. 25:19And you recently
  732. 25:20wrote an article about this. You know,
  733. 25:23talk a little bit about
  734. 25:24a how to how
  735. 25:27how it ties into kind of your broader
  736. 25:29research, uh but also some of the things
  737. 25:32that you uh found when you started to
  738. 25:34dig into this area.
  739. 25:36>> Yeah, another interesting moment to be
  740. 25:39in
  741. 25:39as we're all being made aware of the
  742. 25:42true costs of all this AI usage. I saw a
  743. 25:44tweet from Ed Zitron
  744. 25:47just a screenshot from someone who was
  745. 25:49noticing that Copilot was telling them
  746. 25:52that their month their bill last month
  747. 25:53was $500. And if they keep up the way
  748. 25:56they are with Copilot's pilots new
  749. 25:57billing, it will be $11,000
  750. 26:00in the next month.
  751. 26:01>> Wow. Wow.
  752. 26:02>> Which is real sticker shock. And you
  753. 26:04know, the analogy here is like I used to
  754. 26:06cost me $20 to take a ride share to the
  755. 26:09airport, Uber or Lyft, and now it costs
  756. 26:1290.
  757. 26:12>> I use that analogy as well.
  758. 26:14>> But it's more [laughter] like 20 to like
  759. 26:16500 or something, right?
  760. 26:19>> Right.
  761. 26:19>> Um
  762. 26:20>> If only the slope will be as shallow as
  763. 26:21Uber, right? [laughter]
  764. 26:22>> That's right. We start to wish for those
  765. 26:24easier stories.
  766. 26:26Yes, and so what will happen? I mean, I
  767. 26:28think what's happening is that the big
  768. 26:30providers are testing the waters on
  769. 26:32charging us
  770. 26:33the true costs plus whatever profit they
  771. 26:35need to make as they all try to gear up
  772. 26:37for IPOs and so forth.
  773. 26:39And in turn, that is very quickly
  774. 26:41leading people to ask questions like
  775. 26:43what is the return on investment for all
  776. 26:45these tokens that we have purchased?
  777. 26:49And it's a very tricky area to be in
  778. 26:51because what does it mean to think about
  779. 26:53value in this context? Even if we focus
  780. 26:55in on people who are doing just coding
  781. 26:58with coding agents,
  782. 27:00can we agree on what it means to add
  783. 27:02value? And maybe we have a few measures
  784. 27:03in mind like making a pull request or
  785. 27:07uh committed lines of code that last in
  786. 27:09the repo for a while or
  787. 27:11documentation touched or skill files
  788. 27:14created.
  789. 27:15But we might also worry that that's not
  790. 27:17capturing the value for many kinds of
  791. 27:19sessions we have which are more
  792. 27:21open-ended and about discovery.
  793. 27:24So that's the first question is just
  794. 27:25solving this value issue, right? We just
  795. 27:27agree on what it would mean to add value
  796. 27:28for a for a coding agent.
  797. 27:31>> Now I I like the this line of inquiry
  798. 27:35because
  799. 27:37to me it's it's the response to this
  800. 27:39thing that drives me crazy which is
  801. 27:42oh you know
  802. 27:44big
  803. 27:45uh you know tech company CEO this year
  804. 27:4895% of our code will be generated by you
  805. 27:51know AI.
  806. 27:53It's like
  807. 27:55yeah, hey what does that really mean at
  808. 27:58that level? Like what's what's what are
  809. 28:00the details beneath there but uh is that
  810. 28:02a good thing or a bad [laughter] thing?
  811. 28:05>> That's a oh another dimension, right?
  812. 28:07Which is is that code a liability or an
  813. 28:09asset?
  814. 28:09>> Right. Right. Right. [laughter] And this
  815. 28:12idea of like uh you articulate it as
  816. 28:14kind of code longevity in the code base.
  817. 28:16That's an interesting way to think about
  818. 28:17it. There's probably a lot of
  819. 28:18interesting ways to think about it that
  820. 28:20very few are thinking about right now.
  821. 28:22>> And all of these fall victim to the
  822. 28:24standard thing that once you make it a
  823. 28:25metric, it's no longer useful to you.
  824. 28:27Like if we said oh let's get it's
  825. 28:29completion of projects, right? Well then
  826. 28:31everyone would just have many projects
  827. 28:32that they completed but they could all
  828. 28:34be liabilities and add very little
  829. 28:36value. So but one framework we could
  830. 28:38offer that we did in the research you
  831. 28:40alluded to is let's think about this
  832. 28:42like economist might. So we might have
  833. 28:43like a consumer price index and the
  834. 28:46first step will be what's the basket of
  835. 28:48goods that we're going to consider, You
  836. 28:49you know, in that standard land it would
  837. 28:51be like the the price of eggs and the
  838. 28:54cost of rent and other kinds of tangible
  839. 28:56goods. What are engineering goods that
  840. 28:58we might track?
  841. 28:59>> And eggs might be a summary or
  842. 29:02uh pull request, a bug fix or something
  843. 29:04like that.
  844. 29:05>> Or we could think broadly cuz we both
  845. 29:07use these coding agents,
  846. 29:09um
  847. 29:09requirement discovery, right? Knowledge
  848. 29:12accumulation. These are things that we
  849. 29:14don't currently track, of course, even
  850. 29:16as engineers, but might be behind our
  851. 29:19intuition that these coding agents are
  852. 29:20making us productive even if it's not
  853. 29:22reflected in the PR counts or whatever,
  854. 29:24right? I mean, in a sophisticated
  855. 29:26approach you might say, "I don't want
  856. 29:27more PRs because this is just a certain
  857. 29:30kind of um
  858. 29:31busy work that doesn't relate to the
  859. 29:33actual goals I have." What are the
  860. 29:34actual goals? It's completing valuable
  861. 29:36projects and so forth. If I could do it
  862. 29:37with fewer PRs,
  863. 29:39um but I had, you know, really robust
  864. 29:41code, I'd be possibly happy with that.
  865. 29:44So, we got to figure out what the basket
  866. 29:46of goods is, but then we could start to
  867. 29:47track it relative to token usage, and
  868. 29:49that would be the consumer price index.
  869. 29:51So, for any time period we could just
  870. 29:52say I've got my tokens spent and I've
  871. 29:54got my goods produced. Tokens divided by
  872. 29:57goods produced is a pretty rough measure
  873. 30:00of um
  874. 30:01the purchasing power of the tokens in
  875. 30:03those time periods.
  876. 30:05Then you would do the standard compute
  877. 30:06consumer price index thing of making
  878. 30:08what they call a hedonic adjustment. So,
  879. 30:10you could just say maybe quality is
  880. 30:11improving over time. So, you'd pick some
  881. 30:14measure for that and make an adjustment
  882. 30:15to the line.
  883. 30:17And when we did that study, we did code
  884. 30:19survival. So, um the number of lines of
  885. 30:21code that survives more than 4 days in
  886. 30:23the repository. We made an adjustment
  887. 30:25upward because that rate is going up.
  888. 30:27>> That's surprisingly short.
  889. 30:30>> 4 days?
  890. 30:30>> 4 days survival?
  891. 30:32>> You could make So, again, all this is
  892. 30:33around measurement and I'm happy to just
  893. 30:35be starting this dis this discourse
  894. 30:37because we can see it's important to the
  895. 30:38economics of AI and it seems like the
  896. 30:40work isn't being done at a high enough
  897. 30:42rate for us to get a clear picture. So,
  898. 30:44we could make it longer and maybe the
  899. 30:45adjustment would be different. I think
  900. 30:47currently for the data we have, which is
  901. 30:49this SweetChat benchmark, which was
  902. 30:51released by researchers at Stanford.
  903. 30:53It's about 6,000 real coding sessions,
  904. 30:56all the metadata, everything you'd want.
  905. 30:59What we see with Opus 4.6 usage in the
  906. 31:02time period we have, which is February
  907. 31:04to mid-April of this year,
  908. 31:06a decline in the purchasing power of
  909. 31:08tokens. That CPI is going down.
  910. 31:12And
  911. 31:13again, I just want to open the question,
  912. 31:15is it because we have the wrong basket
  913. 31:16of goods, or is it because we're
  914. 31:19actually getting less value from these
  915. 31:20tokens? The The The one thing I can say
  916. 31:24that's kind of definitely a causal
  917. 31:25factor here
  918. 31:27is that in February of this year, most
  919. 31:29of the tokens went to producing code,
  920. 31:32which relates to the outcomes we just
  921. 31:33talked about. By mid-April, it was quite
  922. 31:37split between code generation, thinking,
  923. 31:40and also explanation to the user.
  924. 31:43And so, that split now is going to have
  925. 31:44an effect on the things we're measuring,
  926. 31:46and that might be cause for reflection.
  927. 31:48There's value in those explanations
  928. 31:50that's not reflected in PRs, but might
  929. 31:52be reflected in something like knowledge
  930. 31:53discovery.
  931. 31:54>> And I see that coming up within the same
  932. 31:58time frame, it's become very common to
  933. 32:01now talk about the token efficiency of a
  934. 32:04new model that's been released. Uh
  935. 32:07with the implication being
  936. 32:09uh tokens of, you know, internal use
  937. 32:12tokens, thinking tokens versus, you
  938. 32:14know, per token of output, I guess, is
  939. 32:16maybe a way to think about it.
  940. 32:18>> Another fascinating dimension, and this
  941. 32:19actually relates all the way back to the
  942. 32:21theme of efficiency for these
  943. 32:22architectures. So, here's a claim I'll
  944. 32:24make for you.
  945. 32:26Based on my read of the literature on
  946. 32:28inference time scaling,
  947. 32:30what's sometimes called test time
  948. 32:32scaling, which is just having the models
  949. 32:34generate lots of tokens
  950. 32:36at the moment that you ask them a
  951. 32:37question. So, those scaling trends,
  952. 32:39everything we're seeing now is
  953. 32:40completely in line with those
  954. 32:42predictions, which is
  955. 32:44you get pretty good gains for a while
  956. 32:46with the more tokens you spend on a log
  957. 32:49scale. So, this is jumping up quite a
  958. 32:50lot, but you do see it reflected in
  959. 32:52performance improvements, but it
  960. 32:54flattens out over time.
  961. 32:56And it's not like this curve skyrockets.
  962. 33:00It's sobering. You got to spend a lot of
  963. 33:02tokens for small gains in performance.
  964. 33:05We all knew this.
  965. 33:07We all knew this, and we're just seeing
  966. 33:09it now play out. And when people talk
  967. 33:10about token efficiency and worry about
  968. 33:12this, I think what they're seeing is
  969. 33:13just the real lesson of what we already
  970. 33:15projected from inference time scaling.
  971. 33:17>> And this is independent of the approach
  972. 33:20to inference time scaling you're taking,
  973. 33:22whether it's
  974. 33:23you know, multiple parallel
  975. 33:26you know, multiple parallel inferences
  976. 33:28or some kind of oracle or you know, any
  977. 33:32number of other schemes. It's just
  978. 33:34fundamental to inference time scaling.
  979. 33:36>> It's a great question, right? I think we
  980. 33:39know that it's independent of some of
  981. 33:41those things, like the parallel work
  982. 33:43versus having it do lots of long chains.
  983. 33:45But some of the other factors you
  984. 33:47mentioned, I think we just don't know,
  985. 33:48and that's why I said it relates back to
  986. 33:50the question of efficiency for these
  987. 33:51architectures.
  988. 33:53If we made a fundamental change to how
  989. 33:55the models work,
  990. 33:56maybe these tradeoffs would be very
  991. 33:58different. I mean, after all, so all of
  992. 34:00this stuff is a kind of patch job on the
  993. 34:03fact that there's no recursion in the
  994. 34:05depth. It's a fixed depth. And so, the
  995. 34:08only recursion we can get, the
  996. 34:09open-ended notion of computation is by
  997. 34:11generation. But if we had models that
  998. 34:13could be recursive, maybe fewer tokens
  999. 34:16for larger gains. I think we don't know.
  1000. 34:19Yeah, I mean, in the end, we're going to
  1001. 34:20spend the cost on compute or tokens.
  1002. 34:24So, this might not affect our bills in
  1003. 34:25the end, but it is a fascinating
  1004. 34:27question. What are the true scaling laws
  1005. 34:29and what's possible in this space? And
  1006. 34:31you're right to push back. We talk about
  1007. 34:33these things like they were like
  1008. 34:34platonic ideals of laws.
  1009. 34:37Scaling law invokes that.
  1010. 34:39You're right, but even for the scaling
  1011. 34:40laws for pre-training, you know, there's
  1012. 34:42lots to discover there. And many of the
  1013. 34:45stories of progress are actually like
  1014. 34:47transcending the scaling law. And we see
  1015. 34:49like better improvements than those laws
  1016. 34:51predicted because everyone worked so
  1017. 34:53hard behind the scenes to do very
  1018. 34:54innovative things, which maybe relates
  1019. 34:56to our bitter lesson discussion.
  1020. 34:58>> Any particular example come to mind of
  1021. 35:00that?
  1022. 35:01>> Data usage and the nature of the data
  1023. 35:02really matters, and overtraining the
  1024. 35:04models really matters, which is kind of
  1025. 35:05pushing up against the standard scaling
  1026. 35:07law presentation. And now I'm just going
  1027. 35:09to speculate. I should check on this,
  1028. 35:11but things like mixture of experts might
  1029. 35:13have really flipped the script on what
  1030. 35:15it means to count parameters and in turn
  1031. 35:17how these laws relate. And then I think
  1032. 35:19maybe even also stuff like the context
  1033. 35:20window and so forth. This is another
  1034. 35:22thing to check, but I just speculate
  1035. 35:24that we've seen larger gains from
  1036. 35:26pre-training than you would have
  1037. 35:28predicted by those early scaling laws
  1038. 35:30papers, suggesting that there is some
  1039. 35:32innovative thing that was happening on
  1040. 35:34top of pure scaling.
  1041. 35:35>> Thinking about the concept of a a market
  1042. 35:39basket, one kind of pushback that
  1043. 35:42comes up for me is
  1044. 35:44in the you know, the real economy, you
  1045. 35:47know, eggs is different than milk is
  1046. 35:50different than
  1047. 35:52you know, beef, etc., etc. And they're
  1048. 35:56all
  1049. 35:58influenced by different factors.
  1050. 36:01You know, production, for example.
  1051. 36:04Whereas
  1052. 36:06what you've done with the this kind of
  1053. 36:08CPI basket with tokens is kind of like
  1054. 36:13more like analogies. Like here's the
  1055. 36:15typical bundle of work and you know,
  1056. 36:17what it requires from a consumptive
  1057. 36:19perspective, but
  1058. 36:21the tokens aren't fundamentally
  1059. 36:22different. Like they're the same tokens.
  1060. 36:24It's just like how much it takes to do
  1061. 36:25this versus how much it takes to do that
  1062. 36:27versus how much it takes to do that.
  1063. 36:29Um
  1064. 36:31you know, tell me what I'm what I'm
  1065. 36:32missing there and
  1066. 36:34you know, what does kind of
  1067. 36:36characterizing these products, you know,
  1068. 36:39give you in your analysis?
  1069. 36:41>> Yeah, fascinating to think about. One
  1070. 36:43thing I could insert there is the tokens
  1071. 36:45are different at the level of being used
  1072. 36:47for code generation or skill file
  1073. 36:50writing or explanation or thinking,
  1074. 36:52right? Those are different kinds of
  1075. 36:54tokens that probably do feel tangibly
  1076. 36:56different to us. So, is that an element
  1077. 36:58in your thinking?
  1078. 36:59>> I think I was thinking from our
  1079. 37:02conversation that you had 10 different
  1080. 37:04almost like tasks, like 10 different
  1081. 37:06types of tasks from the domain of code
  1082. 37:09generation, which you know, if they were
  1083. 37:12all kind of largely code generation, you
  1084. 37:14know, that is the part that had some
  1085. 37:16dissonance for me. But, if you're
  1086. 37:18talking about like if your your basket
  1087. 37:20is like creative writing versus, you
  1088. 37:23know, a few code generation things that
  1089. 37:25are kind of in different uh versus, you
  1090. 37:29know,
  1091. 37:30summarization versus editorial
  1092. 37:33commenting feedback, those, you know,
  1093. 37:35may be more fundamental.
  1094. 37:37>> Yeah, so I think this is very
  1095. 37:39significant. And we have done some
  1096. 37:41research on this as well at the level of
  1097. 37:43what kinds of session types exist and in
  1098. 37:45turn what kinds of users are there. So,
  1099. 37:47you might notice of your own behavior. I
  1100. 37:48guess this is reflected in your comment
  1101. 37:50that sometimes you want a quick check-in
  1102. 37:52on a question. Sometimes you want a
  1103. 37:54quick um PR to get fired off. Sometimes
  1104. 37:57you want to be in a mode of deep
  1105. 37:58collaboration.
  1106. 38:00Sometimes you're partnering with the AI,
  1107. 38:02sometimes you're delegating the work and
  1108. 38:03so forth and so on.
  1109. 38:05And the outcome measures that we choose
  1110. 38:07should be sensitive to this. We
  1111. 38:09shouldn't penalize the agent if your
  1112. 38:11chat interaction with it about some
  1113. 38:13scientific question didn't lead to a PR.
  1114. 38:15It was never on the table in the first
  1115. 38:17place.
  1116. 38:18Whereas, if you're trying to get some
  1117. 38:20work delegated that's actually a coding
  1118. 38:22task and all it does is chat with you,
  1119. 38:24that would feel quite unproductive.
  1120. 38:27So, we need to bring that in and that
  1121. 38:28would be a higher level discovery
  1122. 38:29process of what people are trying to do
  1123. 38:32and so forth.
  1124. 38:33>> In thinking about the the notion of
  1125. 38:36value, is this
  1126. 38:38something that you're anticipating
  1127. 38:42Like it it strikes me that that's an
  1128. 38:43entire, you know, research thread that,
  1129. 38:46you know, one could go into. I don't
  1130. 38:47know if that's a linguistics or a
  1131. 38:49linguist or an economist or a computer
  1132. 38:52scientist. You know, probably
  1133. 38:53interdisciplinary.
  1134. 38:56Like most interesting questions. Um, but
  1135. 38:58is that, you know, is that something
  1136. 39:00that you're working on or
  1137. 39:03um, was it something that you put out
  1138. 39:05there for someone to take up and run
  1139. 39:07with?
  1140. 39:08>> I am not sure. I can tell you the
  1141. 39:10the lineage of this idea is that we
  1142. 39:13founded the startup Big Spin because we
  1143. 39:15would like to see more people benefit
  1144. 39:18from AI.
  1145. 39:19Whether you love it or hate it, it's
  1146. 39:20here
  1147. 39:22and I would like the benefits to be more
  1148. 39:24evenly distributed. And I can tell that
  1149. 39:26that will mean
  1150. 39:28bringing on board many more people than
  1151. 39:31currently benefit from AI. Right now, I
  1152. 39:33would say that it's mostly experts
  1153. 39:35deriving real value
  1154. 39:37and a lot of the world is currently even
  1155. 39:39trying to figure out what this is all
  1156. 39:40about as a tool or an entity in their
  1157. 39:42lives.
  1158. 39:44So, we would like to have more access
  1159. 39:46and more productivity
  1160. 39:48and that implies making the user
  1161. 39:49experiences much better.
  1162. 39:52Um, figuring out what interactional
  1163. 39:54patterns lead to success for people,
  1164. 39:56meeting them where they are in a kind of
  1165. 39:58adaptive way. The whole list of things
  1166. 40:00that you might worry about if you were a
  1167. 40:01product manager who had some deployed AI
  1168. 40:03product.
  1169. 40:04And I think by that route and from that
  1170. 40:07perspective, we just ended up worrying
  1171. 40:09about our own token usage increasing and
  1172. 40:12wondering whether there's real value
  1173. 40:14there and it just happened to collide
  1174. 40:16actually just like 3 weeks ago
  1175. 40:18with this emerging narrative on the back
  1176. 40:20I think of all these rumors about IPOs
  1177. 40:23about what the return on investment was
  1178. 40:25and then all these CEOs came out and
  1179. 40:26said oh our spend was enormous and we
  1180. 40:28want to scale back and we're walking
  1181. 40:30back our claims from a few months ago
  1182. 40:31and that is just a fascinating thing to
  1183. 40:33witness in the
  1184. 40:34in the narrative here.
  1185. 40:36>> I'm wondering are you also does this
  1186. 40:38research also attempt to project
  1187. 40:40forward? In theory you could
  1188. 40:44um you know create a model for you know
  1189. 40:48Anthropic's cost and spend you know
  1190. 40:51based on you know publicly available
  1191. 40:53data and some presumptions and
  1192. 40:56give us a sense for
  1193. 40:59you know how close we are to paying full
  1194. 41:01freight for our tokens versus you know
  1195. 41:03if we're only paying 10% for our tokens
  1196. 41:06you know you could then project you know
  1197. 41:08what that cost might look like uh
  1198. 41:11you know over time as we're paying more
  1199. 41:13and more of the the full cost.
  1200. 41:15>> Yeah I don't have again fascinating
  1201. 41:16questions I don't have resolving
  1202. 41:18answers. I am glad I am not tasked in
  1203. 41:20some organization with projecting spend
  1204. 41:22on all of this stuff because I think it
  1205. 41:23would be basically impossible. For the
  1206. 41:25time period that I was describing for
  1207. 41:27our little um CPI experiment Anthropic
  1208. 41:30changed the default reasoning on the
  1209. 41:32model at least two times. So we see like
  1210. 41:35it start they launched it with default
  1211. 41:37reasoning high. We have a mysterious
  1212. 41:39sudden rise in the token usage which we
  1213. 41:41cannot explain. And there's a new
  1214. 41:43baseline they lowered it to medium as
  1215. 41:45the default they patched a bunch of bugs
  1216. 41:47that were related to context management
  1217. 41:49and then turned it back up to high.
  1218. 41:51And all of these things have an effect
  1219. 41:53on the total token output as you can
  1220. 41:54imagine they also changed the default
  1221. 41:55context window which meant people could
  1222. 41:57swallow up much more
  1223. 41:59uh stuff at any given moment.
  1224. 42:02So imagine trying to predict what token
  1225. 42:04spend is going to be like when you have
  1226. 42:05all these exogenous events in addition
  1227. 42:09to
  1228. 42:10changes that we don't even know about
  1229. 42:12and questions about where the value
  1230. 42:13actually lies.
  1231. 42:15Very difficult. And then, you know, the
  1232. 42:17true cost of a token, the estimates vary
  1233. 42:19wildly for every dollar we spend, it
  1234. 42:21could be as low as two and as high as
  1235. 42:2320. And I think this is just because
  1236. 42:26it's hard to factor in things like R&D
  1237. 42:28and future build-out and depreciation
  1238. 42:30and all of that stuff. I think at the
  1239. 42:32current moment we just don't know, but
  1240. 42:34there couldn't be a more significant
  1241. 42:35question for the global economy,
  1242. 42:37basically, than
  1243. 42:40where the value is and who's going to
  1244. 42:42pay and how much.
  1245. 42:44>> Yeah, in your article you coined the
  1246. 42:46term tokenflation to
  1247. 42:49describe, at least the recent behavior
  1248. 42:51of token economics. I imagine you see
  1249. 42:55that continuing.
  1250. 42:56>> Seems to be continuing. Yeah. Yeah,
  1251. 42:58that's certainly the picture that we get
  1252. 43:00from the CPI, a picture of tokenflation.
  1253. 43:02Yes, your token is not buying you what
  1254. 43:04it once did, according to everything we
  1255. 43:06can think to measure here.
  1256. 43:08And even adjusting for models getting
  1257. 43:10better, right? That's critical there.
  1258. 43:12Cuz if it was just a story of models
  1259. 43:13thinking more and being more robust, and
  1260. 43:16we were all getting exponentially better
  1261. 43:17outcomes from this, then the spend would
  1262. 43:19look completely rational.
  1263. 43:21But that's not the picture that we see,
  1264. 43:23and so we have to do some hard thinking
  1265. 43:24about what's going to happen and how to
  1266. 43:27improve the situation.
  1267. 43:28>> Let's dig into that a little bit more.
  1268. 43:30Your How would you articulate what
  1269. 43:32you're seeing? The models are
  1270. 43:35getting, quote unquote, better.
  1271. 43:37Um you know, there's a set of open
  1272. 43:39questions about uh are the reported ways
  1273. 43:43that models are better actually
  1274. 43:47reflective of some intrinsic betterness,
  1275. 43:49like in that question brings uh
  1276. 43:52it is often about like benchmarking and
  1277. 43:56uh learning the benchmarks, overfitting,
  1278. 43:58that kind of thing.
  1279. 43:59Uh and then there's
  1280. 44:02the kind of question of chattiness and
  1281. 44:05and the
  1282. 44:07the volume of thought that it
  1283. 44:10requires a given model generation to
  1284. 44:12produce an answer. What are other
  1285. 44:14factors that you see?
  1286. 44:16>> Yeah, we could pick that apart as well.
  1287. 44:17So, and this relates to
  1288. 44:21uh a line I've had consistently, which
  1289. 44:23is that we should think in terms of
  1290. 44:24systems, not in terms of models. So, in
  1291. 44:27the data that we've got, Sonnet or Opus
  1292. 44:30and Sonnet 4.5 versus 4.6, those two
  1293. 44:33generation changes,
  1294. 44:35those are real model changes, I assume.
  1295. 44:36I think they did something very
  1296. 44:37substantive at the level of the weights.
  1297. 44:40Um and everybody immediately saw that
  1298. 44:42that led to like a 5x increase in token
  1299. 44:44usage. And this was related to the
  1300. 44:47introduction of adaptive thinking.
  1301. 44:50Now, fix that. That's the level shift
  1302. 44:52that we already took, and maybe we're
  1303. 44:54seeing improvements there that are
  1304. 44:56worthwhile. It gets hard to say, but
  1305. 44:58let's assume there was a level up in
  1306. 44:59improvement.
  1307. 45:01Then for the period that we did our CPI
  1308. 45:03experiment for, that's a fixed model,
  1309. 45:05Opus 4.6.
  1310. 45:07So, all the code improvements that we
  1311. 45:09saw in the data relate to the product.
  1312. 45:12This has to relate to things like them
  1313. 45:14turning the knobs on the adaptive
  1314. 45:15thinking,
  1315. 45:16changing things about the system prompt,
  1316. 45:19changing things at the level of the
  1317. 45:20product, and that's where the
  1318. 45:21improvements were.
  1319. 45:23And so, that shows you that even for a
  1320. 45:24fixed model, we can get very different
  1321. 45:26outcomes for these things because they
  1322. 45:27really are sophisticated engineered
  1323. 45:29systems at this point.
  1324. 45:31>> Yeah.
  1325. 45:31And so, was the product in this case
  1326. 45:34specifically Claude Code or
  1327. 45:37>> Oh, yeah. And so, we don't There's tons
  1328. 45:38of stuff there.
  1329. 45:40Yes, I believe we know that these are
  1330. 45:41all Claude Code sessions that we kept in
  1331. 45:43our data. Sweet Chat is broader than
  1332. 45:44that and involves a couple of other
  1333. 45:46coding agents, but I think I can say
  1334. 45:48that all our data are Claude Code
  1335. 45:49sessions using Opus 4.6.
  1336. 45:52>> Have you seen any evidence that
  1337. 45:55changes
  1338. 45:56via API usage experience
  1339. 45:59uh similarly dramatic uh variation in
  1340. 46:04performance.
  1341. 46:06>> Oh, fascinating. To kind of control for
  1342. 46:08a lot of that product level stuff, all
  1343. 46:10the prompts that are hidden from us, all
  1344. 46:12of those affordances. Yeah, I don't
  1345. 46:13know, but that's a nice thing to think
  1346. 46:15about because it gives us a more things
  1347. 46:18that we can control for and more things
  1348. 46:19that are knowable.
  1349. 46:20>> So, kind of in parallel to the model
  1350. 46:24evolution, there's also evolution of the
  1351. 46:28user. You've alluded to this a little
  1352. 46:30bit about kind of your concept is that
  1353. 46:32most AI users now are experts.
  1354. 46:37Uh talk a little bit about the role of
  1355. 46:40expertise. I think this is also kind of
  1356. 46:43echoing back to our conversation about
  1357. 46:45DSPy and like prompt optimization. You
  1358. 46:48know, you've done some research into how
  1359. 46:50folks are using these models and the
  1360. 46:51role of, you know, AI fluency. Tell us
  1361. 46:54about that research.
  1362. 46:55>> Oh, yeah. First, I should say, so the um
  1363. 46:58distribution of users across expertise
  1364. 47:00levels.
  1365. 47:02I'd I So, I guess the nuanced picture
  1366. 47:04I'd offer is that the people deriving a
  1367. 47:06lot of value from AI in the current
  1368. 47:08moment tend to be experts. It must be
  1369. 47:10the case that most users of AI are are
  1370. 47:13beginners, just because the numbers are
  1371. 47:15so large and expertise can't be that
  1372. 47:17widely distributed yet. And that's a
  1373. 47:19very interesting thing because I think
  1374. 47:21probably most things are getting
  1375. 47:22designed for those experts implicitly or
  1376. 47:24explicitly. But, for the whole economic
  1377. 47:27picture to work out, many more people
  1378. 47:30need to derive value from this via one
  1379. 47:32avenue or another.
  1380. 47:34And so, that does shine a light on this
  1381. 47:36expertise thing as a real factor.
  1382. 47:38And the headline result there actually
  1383. 47:40builds on something that Anthropic did.
  1384. 47:42They have this AI fluency index, and
  1385. 47:43their core observation in that work is
  1386. 47:46that experts display an augmentative
  1387. 47:49style.
  1388. 47:50They iterate with the AI. They push
  1389. 47:53back. They complain. They change their
  1390. 47:55requirements.
  1391. 47:56It's a really collaborative mode.
  1392. 47:58Whereas novices, low-fluency users,
  1393. 48:02delegate. So, they trust in the AI. They
  1394. 48:05let it do its thing. They accept the
  1395. 48:06responses uncritically.
  1396. 48:08And our contribution is to just show
  1397. 48:10that this is a causal factor in success
  1398. 48:14with these products right now. Experts
  1399. 48:16can do harder things more reliably as a
  1400. 48:18result of all that friction they
  1401. 48:21introduce, all that pushback.
  1402. 48:24Whereas novice users, they accept, but
  1403. 48:27they end up accepting the wrong thing,
  1404. 48:28and they're not able to level up from
  1405. 48:30the basic tasks that they think to start
  1406. 48:33with.
  1407. 48:34And that's obviously significant, and it
  1408. 48:35feels so tantalizing because pushing
  1409. 48:38back is a natural human behavior. I feel
  1410. 48:42like we could encourage everyone in the
  1411. 48:43world to do this. We probably need to
  1412. 48:45get them out of the mode of thinking
  1413. 48:47it's a superintelligence. You should
  1414. 48:48just trust it. That has been the
  1415. 48:50narrative for a while, but we're seeing
  1416. 48:52in the current moment, and possibly for
  1417. 48:54the foreseeable future, is that you got
  1418. 48:56to complain, collaborate, introduce
  1419. 48:58yourself, push back, all that stuff that
  1420. 49:00I think we do, that we take that for
  1421. 49:02granted, right?
  1422. 49:03>> Yeah. Yeah. And so, from a methodology
  1423. 49:07perspective, how did you approach
  1424. 49:09exploring this?
  1425. 49:10>> Hey, we built on the work that Anthropic
  1426. 49:12did,
  1427. 49:13um, which what they set up a nice
  1428. 49:15framework with some independent research
  1429. 49:17who were doing this kind of usability
  1430. 49:19stuff.
  1431. 49:20And we just have an annotation protocol.
  1432. 49:23Like we can talk in detail if you want
  1433. 49:24about this, but at Big Science, we have
  1434. 49:26lots of these best practices around
  1435. 49:28having language models essentially
  1436. 49:29collaborate on annotation projects to
  1437. 49:32kind of triangulate on the truth and
  1438. 49:34factor out their individual biases.
  1439. 49:36So, we do that stuff, and we apply all
  1440. 49:38these fluency markers, and then
  1441. 49:39separately we do a thing of estimating
  1442. 49:41task complexity, and looking for signs
  1443. 49:44of visible and invisible failures.
  1444. 49:47And so, it's the connection between the
  1445. 49:49fluency markers and the task complexity
  1446. 49:52success metrics. That was our
  1447. 49:55contribution there and that's where you
  1448. 49:57can see high fluency users are the ones
  1449. 49:59doing harder tasks. Paradoxically,
  1450. 50:01there's more signs of failure for them
  1451. 50:03uh because they complain, they push
  1452. 50:05back, they're trying harder things.
  1453. 50:07But as part of all that friction,
  1454. 50:09they're successful with harder things as
  1455. 50:10well.
  1456. 50:11>> And if you were to try to apply this
  1457. 50:14insight
  1458. 50:15from the perspective of someone in an
  1459. 50:17organization that's trying to you know
  1460. 50:19help or guide their organization to
  1461. 50:22be more successful with AI, like what do
  1462. 50:25you think are the key lessons of this
  1463. 50:27fluency work?
  1464. 50:28>> If it's an org that's just starting out
  1465. 50:30and wants people to figure out how this
  1466. 50:32could be part of the organization's
  1467. 50:33mission, it would just be that push back
  1468. 50:35message. And you could do an experiment
  1469. 50:37where you interact with it about
  1470. 50:39something where you're a world expert.
  1471. 50:40We're all an expert in something.
  1472. 50:42Engage in a discourse with one of the
  1473. 50:44best models about something you're an
  1474. 50:45expert in and see how often you feel you
  1475. 50:47have to push back and this could be a
  1476. 50:49kind of a lesson you say.
  1477. 50:50>> Aha.
  1478. 50:51>> For other spheres where I don't know the
  1479. 50:53answer, it might be just as errorful.
  1480. 50:55That could be a good visceral thing. If
  1481. 50:57the org is very far along, I think the
  1482. 50:59main thing to do right now is to have a
  1483. 51:01team of these LLMs interacting to
  1484. 51:03improve things. For example, at Big
  1485. 51:05Spin, I didn't set this up. Our founding
  1486. 51:07engineer is very future forward on
  1487. 51:10agents and he's incredible at this. And
  1488. 51:13when we do PRs now, the first round of
  1489. 51:15review is the agents all interacting,
  1490. 51:18collaborating, disagreeing. They do the
  1491. 51:20first round of comments. They do the
  1492. 51:22first round of code updates. Only after
  1493. 51:24they've resolved things do we look at a
  1494. 51:26PR. So the final human stage should be
  1495. 51:29very high value and the agents did all
  1496. 51:31that work. But when you have one agent
  1497. 51:33do it, they often just reinforce
  1498. 51:35themselves and you don't get good
  1499. 51:36outcomes. It's that team of rivals thing
  1500. 51:39that is transformative.
  1501. 51:40>> You know, I think it's interesting
  1502. 51:41because you know, on the one hand like
  1503. 51:44of of course that makes sense.
  1504. 51:46But on other hand it it there's
  1505. 51:48something
  1506. 51:50you know, it also implies that you
  1507. 51:53shouldn't be using these things in areas
  1508. 51:56where you don't have enough expertise to
  1509. 51:58evaluate the answer. Yet
  1510. 52:01that's where you most need the
  1511. 52:04assistance, the support. Um so
  1512. 52:07>> Well, and again and this is a little bit
  1513. 52:09worrisome about the overall narrative
  1514. 52:11around AI.
  1515. 52:13The place where we can get around this
  1516. 52:15is with software development because
  1517. 52:17let's say that I'm trying to accomplish
  1518. 52:18something in a language that I don't
  1519. 52:20know how to code in.
  1520. 52:22I can have the agent do work for me
  1521. 52:24because probably in the end I can run
  1522. 52:26the program and look at the results.
  1523. 52:28And that's what mattered to me is that I
  1524. 52:30run the results and I see and if I don't
  1525. 52:33see what I want then I can complain and
  1526. 52:35we can iterate.
  1527. 52:36That verification step that doesn't
  1528. 52:38imply I have comprehensive knowledge, it
  1529. 52:40just implies that I know what I want to
  1530. 52:41see in the end is so critical and I
  1531. 52:44think this is a causal factor in models
  1532. 52:46being so good at coding because it's
  1533. 52:48like the ultimate verifiable domain for
  1534. 52:50them.
  1535. 52:51But as soon as we leave that and go even
  1536. 52:53into something like the legal realm
  1537. 52:55where the requirements are strict but
  1538. 52:57they're not codified in code and they
  1539. 52:59have ambiguity about them. This whole
  1540. 53:02picture falls apart.
  1541. 53:05And you then are back at what you just
  1542. 53:06said which is this awful kind of paradox
  1543. 53:08is like, yeah, use AI but in the end
  1544. 53:11unless you're expert enough to evaluate
  1545. 53:13every single one of its responses, you
  1546. 53:14might be in real trouble.
  1547. 53:17I don't know how to get out of this
  1548. 53:18because the the verification step is
  1549. 53:20like we go to trial but this is very
  1550. 53:22consequential.
  1551. 53:24>> Yeah, [laughter] that's expensive.
  1552. 53:26Yeah, that's funny. I mean it does make
  1553. 53:28me think a little bit about
  1554. 53:30you know, some of the types of errors
  1555. 53:32that we're trying to avoid are
  1556. 53:34factuality and
  1557. 53:37you know, there is
  1558. 53:38a temptation to say, well, let's just
  1559. 53:40throw more tokens at it. Like I'll have
  1560. 53:42a a critic model that uh uh evaluates
  1561. 53:45everything that the
  1562. 53:48you know, is generated by the primary
  1563. 53:49model. But, then you go back to my
  1564. 53:51observation that these models tend
  1565. 53:54to correlate uh in their responses as
  1566. 53:57well.
  1567. 53:58Um
  1568. 53:59yeah, it's it's super interesting.
  1569. 54:01>> That's a good point. Yeah, from my
  1570. 54:03picture, we want real diversity of
  1571. 54:04perspectives. This is just like, you
  1572. 54:06know, red teaming for humans. This is
  1573. 54:09most successful when you have a really
  1574. 54:11diverse team of people who think
  1575. 54:12creatively and differently. And if every
  1576. 54:14one of the members of that team is
  1577. 54:16thinking in a homogeneous way, they miss
  1578. 54:17all of the crucial things.
  1579. 54:19Same exact issue. If all of the code
  1580. 54:21review agents are biased in the same
  1581. 54:23way, they will miss exactly the same
  1582. 54:24class of bugs, and then we're all sunk.
  1583. 54:27Yeah, I don't know how you'd encourage
  1584. 54:28this diversity in the ecosystem. We're
  1585. 54:30probably, as you say, converging towards
  1586. 54:31some kind of one model. Um but, I think
  1587. 54:34for my picture, we need diversity. Yeah,
  1588. 54:36we got to keep those open weights models
  1589. 54:38going or something, cuz they're the
  1590. 54:39weird players in the space.
  1591. 54:41>> For sure, for sure. So, we've talked
  1592. 54:43about uh efficiency, interpretability,
  1593. 54:48tokenomics,
  1594. 54:50uh uh uh
  1595. 54:51fluency.
  1596. 54:53Yeah,
  1597. 54:54you're involved in a a lot of different
  1598. 54:56research directions. Excellent,
  1599. 54:57excellent. Where are What's next for
  1600. 54:59you? Where do you see either Where do
  1601. 55:01you see this all going kind of
  1602. 55:02externally, but also like where is your
  1603. 55:04research going?
  1604. 55:05>> Yeah, this is great. Um
  1605. 55:08and I'm I as I said before, we're trying
  1606. 55:09to think in weird and creative ways
  1607. 55:11about what the future could hold.
  1608. 55:13And I encourage my students to do this.
  1609. 55:15And they're smart, so they say, "All
  1610. 55:16right, Chris, I'll think along those
  1611. 55:17lines, but what's your answer to this
  1612. 55:19question?" So, I do have an answer.
  1613. 55:20[laughter] And it's really shooting for
  1614. 55:22the moon here, which would be what about
  1615. 55:24the architectural innovation that would
  1616. 55:26upend the whole story around the stack
  1617. 55:29transformer and the way we need to do
  1618. 55:31data center build out to even get
  1619. 55:33incremental gains in performance. That
  1620. 55:35could be upended, and it would come from
  1621. 55:37some very innovative thing around maybe
  1622. 55:39recursive use of you the building blocks
  1623. 55:42that we've got.
  1624. 55:43So, architectures, we should think, and
  1625. 55:45when people say, "Oh, no, we don't need
  1626. 55:46more architectures. The transformer is
  1627. 55:48good enough." That's where we should
  1628. 55:50push back as academics doing something
  1629. 55:52more clever and more scrappy that could
  1630. 55:54change the world. And the other one is
  1631. 55:57thinking in the inter space much more
  1632. 55:59about data. And that's just because I
  1633. 56:02want to tell the true story of how we go
  1634. 56:04from data to model capabilities, but it
  1635. 56:06also
  1636. 56:07checks a box for me on connecting
  1637. 56:10interpretability to safety. It has been
  1638. 56:13hard for me to connect those two things.
  1639. 56:15We have found some ways to do it, but
  1640. 56:16it's not a slam dunk as a narrative,
  1641. 56:18even though it's the dominant narrative.
  1642. 56:20But, I will say that when we get into
  1643. 56:22things like data poisoning from
  1644. 56:24innocuous examples, this is probably a
  1645. 56:26growing societal concern. There is
  1646. 56:29evidence that with very few examples
  1647. 56:31planted in a pre-training data set, you
  1648. 56:34can have a significant influence on the
  1649. 56:36outlook and preferences and quirks of
  1650. 56:38the final model.
  1651. 56:40So, can we detect those examples? What's
  1652. 56:43the nature of those attacks? How well
  1653. 56:45hidden could they be? What's the
  1654. 56:46smallest number of examples? And why
  1655. 56:48does it happen? These are all going to
  1656. 56:50be very pressing questions. And so,
  1657. 56:53again, it's just a data-oriented
  1658. 56:54question that's very alive for me in the
  1659. 56:56current moment.
  1660. 56:58>> On the architecture front, are there
  1661. 57:03is there research that you're seeing or
  1662. 57:06doing that is, you know, as yet under
  1663. 57:09the radar that you think is, you know,
  1664. 57:12promising and or underappreciated?
  1665. 57:14>> I think you had my student Julie Calleja
  1666. 57:16on,
  1667. 57:18and she is an advocate for byte-level
  1668. 57:19models, essentially tokenizer-free
  1669. 57:21models. I think that's a big part of the
  1670. 57:23future. It's a critical thing if you
  1671. 57:25want to have truly multilingual models
  1672. 57:27that are also equitable in terms of how
  1673. 57:28many tokens they charge us for, getting
  1674. 57:30back to that earlier theme, but also
  1675. 57:33Julie's perspective is that this is
  1676. 57:34speculative, but I think there's
  1677. 57:35something to this that the it's a kind
  1678. 57:38of inference time scaling because you do
  1679. 57:40more compute at test time
  1680. 57:42um because you have more tokens and
  1681. 57:44therefore more opportunities to
  1682. 57:46build on interesting things.
  1683. 57:49So, that could be a big part of the
  1684. 57:50future and the other one would be
  1685. 57:52recursive architectures as I said. But
  1686. 57:55if you want to go all the way out, you
  1687. 57:56could think, "Why do we always assume
  1688. 57:58we're going to do gradient-based
  1689. 57:59learning?" There are lots of
  1690. 58:00alternatives to that and nobody is
  1691. 58:03exploring them because everyone takes it
  1692. 58:05as a truism. We're all in our very
  1693. 58:07narrow row here without even really
  1694. 58:08realizing it.
  1695. 58:10Who knows what's outside in this garden?
  1696. 58:13It's very risky as a research bet
  1697. 58:15because
  1698. 58:16only one in a thousand of these ideas
  1699. 58:18will pay off.
  1700. 58:20But what's the point of being an
  1701. 58:21academic researcher if you're not going
  1702. 58:23to take that kind of risks? That's what
  1703. 58:25That's what we're positioned to do.
  1704. 58:26>> Well, Chris, thanks so much for jumping
  1705. 58:28on and sharing a bit about what you're
  1706. 58:30working on. It's uh very cool stuff.
  1707. 58:33>> Thank you. What a wonderful
  1708. 58:34conversation. It gave me lots of new
  1709. 58:35things to think about.
  1710. 58:37>> Awesome. Awesome. [music] Thanks so
  1711. 58:38much.

About this transcript

This page contains the full transcript of Do AI Tokenomics Matter More Than Model Benchmarks? by The TWIML AI Podcast with Sam Charrington, generated from the public captions YouTube serves with the video. The transcript has 10,758 words across 1,711 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.