Do AI Tokenomics Matter More Than Model Benchmarks? — Transcript
Full transcript
- 0:00We started to enter a phase of AI that's
- 0:02not just about making models smarter.
- 0:04It's also about making them economically
- 0:07sustainable.
- 0:08As reasoning models consume more tokens,
- 0:10context windows continue to grow, and
- 0:13agents become embedded in more products
- 0:15and workflows, the economics of these
- 0:17systems are becoming impossible to
- 0:19ignore. That's given rise to a new
- 0:21conversation around tokenomics, how we
- 0:24think about the costs, incentives, and
- 0:25tradeoffs shaping the next generation of
- 0:27AI.
- 0:29One person who's been thinking deeply
- 0:30about this is Stanford Professor and Big
- 0:33Spin co-founder Chris Potts. His recent
- 0:36work argues that measuring AI progress
- 0:38requires looking beyond model benchmarks
- 0:40to ask a different question. What are
- 0:42our tokens actually buying us?
- 0:45Here's Chris explaining how he thinks
- 0:47about tokenomics.
- 0:48>> Another interesting moment to be in as
- 0:51we're all being made aware of the true
- 0:53costs of all this AI usage. The analogy
- 0:55here is like it used to cost me $20 to
- 0:58take a ride share to the airport for
- 1:00Uber or Lyft, and now it costs 90. But
- 1:02it's more like 20 to like 500 or
- 1:05something, right? And I think what's
- 1:06happening is that the big providers are
- 1:08testing the waters on charging us the
- 1:11true costs plus whatever profit they
- 1:13need to make as they all try to gear up
- 1:15for IPOs and so forth. And in turn, that
- 1:18is very quickly leading people to ask
- 1:20questions like what is the return on
- 1:22investment for all these tokens that we
- 1:24have purchased. And it's a very tricky
- 1:26area to be in because what does it mean
- 1:28to think about value in this context?
- 1:31>> I'm Sam Charrington, and this is the
- 1:33Twilio AI podcast. For over a decade,
- 1:36I've been exploring the ideas and
- 1:37innovation shaping the future of AI
- 1:39through conversations like this one that
- 1:41help you understand what's real, what's
- 1:44next, and what matters. Let's jump in.
- 1:55I want to say thanks for coming on. I've
- 1:57been looking forward to this
- 1:59conversation and I think where I'd love
- 2:03to start us off is to really dig into
- 2:07your background and how it kind of got
- 2:10you, you know, to to where you are now.
- 2:13>> Yeah, my background is in linguistics.
- 2:16Um, linguistics proper, not even natural
- 2:18language processing. I did my PhD on,
- 2:21among many other things, swears.
- 2:24What swears are like, why we swear, what
- 2:26information they encode, what kind of
- 2:27taboos exist around them, and so forth.
- 2:30And that was actually the trigger that
- 2:32got me into NLP because I wanted a lot
- 2:35of data of people swearing. I wanted to
- 2:37know what the context was like, what
- 2:39their intentions were. So, I turned to
- 2:41corpora.
- 2:42And from there you start using NLP
- 2:45toolkits to add structure to those
- 2:47corpora. And then after a few years,
- 2:49maybe you're writing your own tools for
- 2:51doing that work.
- 2:52And then when you look back after 18
- 2:54years or whatever it's been, you're just
- 2:56an AI person or an NLP person. But that
- 2:59is the true story and I I feel like if I
- 3:01had to, I could trace the lineage of
- 3:03every one of my current projects back to
- 3:05my fascination with why we care when
- 3:07someone drops an F-bomb.
- 3:10>> [laughter]
- 3:12>> So, are you a an F-bomb dropper or did
- 3:15you come at it from the perspective of
- 3:17trying to understand these others?
- 3:19>> [laughter]
- 3:20>> I think very infrequently in my life.
- 3:23Um,
- 3:24on for my linguistics class, semantics
- 3:26and pragmatics, which is about
- 3:27linguistic meaning, on the final day we
- 3:29always do a class on swearing. And I
- 3:32review the history and we kind of tie
- 3:34all the course themes together. My
- 3:36handouts for that are full of swears.
- 3:39But I only swear once in the lecture. I
- 3:42I I present the result that people
- 3:44remember things better if the utterance
- 3:46contains a swear because it has a kind
- 3:48of emotional resonance, very primitive
- 3:50reaction. And so in that moment I pick
- 3:53some fact from the course, some trivial
- 3:55thing,
- 3:56and I restate it with a swear, and then
- 3:58I say, "All of you will remember this
- 3:59for eternity."
- 4:01But other than that, I'm very shy about
- 4:03it in the class.
- 4:04>> That's funny. I'm sure there is loads of
- 4:08research on this, but
- 4:11uh
- 4:11I grew up in New York City, and as a New
- 4:14Yorker, I think that swearing is just
- 4:16kind of part of my natural language and
- 4:19way of communicating.
- 4:21And I married a a Midwestern girl, and
- 4:24she doesn't tolerate it at all. She
- 4:26doesn't do it, she doesn't tolerate it,
- 4:28she won't tolerate it for me. And it
- 4:31made for We've been married for 30
- 4:33years, so I adapt quickly, apparently.
- 4:35But uh
- 4:37it, you know, for a long time, it took a
- 4:39lot of restraint to like change that way
- 4:44of communicating, particularly when I'm
- 4:45communicating about something that I'm
- 4:47excited about or emotional about or, you
- 4:49know, want to convey the importance of
- 4:53um It's a really interesting topic, and
- 4:56uh
- 4:58um Well, we're not going to turn the
- 5:00podcast into a podcast about swearing,
- 5:03but I imagine there's enough research
- 5:05there that we could if we wanted to.
- 5:06>> It's a fascinating area, yeah, because
- 5:08it gets right to the heart of the
- 5:10culture that we've constructed and how
- 5:12it relates to our usage and everything
- 5:14else about us. Yeah, it's fascinating
- 5:16that we have swears. When the old swears
- 5:18lose their power, we invent new ones. We
- 5:20pretend like nobody should use them, but
- 5:22as you say, people use them all the
- 5:23time, and it feels like an important
- 5:25part of being a language user that we've
- 5:27got them available to us. Yes, endless
- 5:30string of questions.
- 5:32>> I'd love to hear your take on
- 5:35kind of a linguist in the age of modern
- 5:39AI, you know, transformers, statistical
- 5:42models, you know, this is a
- 5:44uh NLP used to be kind of coming from a
- 5:47linguistic perspective and now the
- 5:49entire field is shifted to
- 5:51statistical perspective and I'd love to
- 5:53hear your reflections on
- 5:56being on the other side of that
- 5:57transition
- 5:59as well as maybe more importantly ways
- 6:01that you think that kind of the
- 6:02traditional foundational linguistics is
- 6:04still important to
- 6:07the the way we think about AI today.
- 6:09>> These questions are on my mind all the
- 6:11time yeah because I operate at the
- 6:13intersection of all these different
- 6:14fields and I will say it's it's useful
- 6:16to distinguish in this context
- 6:18linguistics you know and people in my
- 6:20department at Stanford study language
- 6:22and social identity historical
- 6:24linguistics the structure of language
- 6:27and they're just doing scientific
- 6:29investigation of language as a human
- 6:31phenomenon and they are not
- 6:32technologists and they're not trying to
- 6:34inform technology.
- 6:35So their project is
- 6:37interestingly impacted by technological
- 6:40developments. For NLP people who are of
- 6:42course participating directly in the
- 6:44engineering project they're affected in
- 6:47a very different way by the rise of gen
- 6:48AI and the kind of homogeneous nature of
- 6:51the solutions that people now adopt in
- 6:53that space.
- 6:54So for the linguists I feel like this is
- 6:57the most exciting moment that anyone
- 6:58could have dreamed of. I feel incredibly
- 7:00privileged to be alive in this moment
- 7:04where humans encounter for the very
- 7:06first time non-human creatures that use
- 7:09our language very fluently.
- 7:11I think it's weirding us all out but
- 7:13from the point of view of understanding
- 7:15the human capacity for language what a
- 7:18gift because you can ask about the
- 7:20mechanisms
- 7:22which are different from humans but
- 7:23obviously sufficient for achieving a
- 7:25certain kind of behavioral performance.
- 7:28Um
- 7:29we can think about them as investigative
- 7:31tools. I mean we train them on the
- 7:32internet they're basically incredibly
- 7:35powerful distributional learners and we
- 7:36can all learn a lot from them about the
- 7:38true structure of language by just
- 7:40looking at the kinds of things that they
- 7:42learn. And it really gets at the heart
- 7:44of core questions in linguistics about
- 7:47how much of language learning is innate
- 7:49and the nature of our capacity and
- 7:51whether it's statistical or symbolic.
- 7:53All those things come flooding in in a
- 7:55completely fresh way. And so whatever
- 7:57your reaction to language models
- 8:00is, it should be a significant one,
- 8:02right? This should be causing you to
- 8:04rethink key questions. And that's all
- 8:07you could hope for as a scientist, that
- 8:08you have new angles, new perspectives,
- 8:11new questions reopen. That's been
- 8:13incredible.
- 8:15For NLP, I think it's a more uncertain
- 8:18prospect because
- 8:20pre
- 8:22the the arrival of like pre-trained
- 8:25models, which for me would be like the
- 8:27Elmo model back in 2017, 2018. Before
- 8:31that, there was still a lot of
- 8:33statistical work, of course, and we were
- 8:34in the deep learning era.
- 8:36But you could still, for example, do a
- 8:38PhD that was entirely about some
- 8:40specific phenomenon and maybe some very
- 8:42specific tweak to a model. So you could
- 8:44say, "I'm going to work on summarization
- 8:45and I've got a new idea about how to do
- 8:47that well using deep learning models."
- 8:50And that could be your PhD. And what we
- 8:52started to see 2018, 2019, 2020,
- 8:55especially with the arrival of GPT-3,
- 8:57that that was a very uncertain prospect
- 8:59because you might wake up one morning to
- 9:01find that you had been completely
- 9:02scooped. That with essentially no
- 9:04effort, one of these large pre-training
- 9:06runs had done better than you at the
- 9:08thing that you'd worked so hard on.
- 9:10And that caused an interesting, probably
- 9:12overall productive, but interesting and
- 9:14challenging crisis for people,
- 9:16especially students who were trying to
- 9:18figure out what to do next with their
- 9:20PhD research. But I think all of us felt
- 9:23a kind of real uncertainty in that
- 9:24moment.
- 9:25>> Yeah, I remember the anxiety of that
- 9:28time and
- 9:31I always felt it was kind of expressed
- 9:32as
- 9:34you know, is research in NLP
- 9:36fundamentally like scale limited or do
- 9:39you need a certain degree of scale that
- 9:41only a handful of organizations have to
- 9:43do foundational research and is everyone
- 9:47else going to be relegated to like
- 9:50poking the poking the beast and seeing
- 9:53what it does?
- 9:54And I'm curious do you feel
- 9:57like that was an anxiety that's passed
- 9:59or is it still very present? Has it
- 10:02panned out quite like that? How How do
- 10:04you you know what how's it been resolved
- 10:06for you?
- 10:07>> Also fascinating. Not resolved. It's
- 10:09something I discuss a lot with my
- 10:11collaborators and with my students.
- 10:13We're all trying to figure this out in
- 10:14this moment. I will say one concrete
- 10:17thing we did was orient a lot of our
- 10:19research toward interpretability.
- 10:22Uh just the project of understanding how
- 10:24these models end up being so good at
- 10:26such hard tasks.
- 10:28And the reason we did that is it's
- 10:29relatively inexpensive and it's also an
- 10:31area where clearly you would be
- 10:33explicitly hoping that models would get
- 10:35better because then there would be more
- 10:37to explain.
- 10:39Versus if you were doing that
- 10:41summarization project, you might quietly
- 10:43be hoping that there wasn't going to be
- 10:44so much progress so that you could make
- 10:46the progress. Like let's hope the next
- 10:48model isn't good at summarization. I
- 10:50want to be the star of that show. That's
- 10:52as I said very uncertain but if you're
- 10:53doing mech interp, you're like let's get
- 10:55the new model released because now we're
- 10:56going to have even more structure to
- 10:58find, even more to explain. And that
- 11:00felt like a very productive choice. I
- 11:02don't want to leave out the fact that
- 11:03it's also cheaper to do this research
- 11:05and that is significant.
- 11:07And then I would say that right now a
- 11:08lot of us are in a moment of thinking we
- 11:11should do stuff that is weird and
- 11:13creative and out of the mainstream. We
- 11:16should be thinking about trying to
- 11:18achieve the next big thing because
- 11:20competing with these massively
- 11:22resourced, incredibly creative and
- 11:24talented teams is just not a winning
- 11:26game. So let's play a different game and
- 11:29hope that that's as they say where the
- 11:31puck is going, not where it is.
- 11:33>> And what are some examples of that kind
- 11:35of thinking?
- 11:36>> We've been thinking a lot about
- 11:37architectures cuz I have a lot of
- 11:38complaints about current architectures.
- 11:41And I would say the other main theme
- 11:42right now for us in my group is thinking
- 11:45about data.
- 11:46You know, data have strange and wondrous
- 11:49properties. I think we don't understand
- 11:50how data affect models.
- 11:52And that has all sorts of implications
- 11:54for security and safety and also the
- 11:56nature of the learning that these models
- 11:58do. It really data is fundamental. It's
- 12:00all data-driven learning. And so telling
- 12:02the full causal story from data to final
- 12:04model state
- 12:06feels like it will just be significant
- 12:08for lots of questions.
- 12:11But I wouldn't want to leave out the
- 12:12architecture one because
- 12:14I feel like the architecture everyone
- 12:16has arrived at, these stacked
- 12:18transformers that we make very deep and
- 12:20very large, are tremendously
- 12:22inefficient.
- 12:24You would hope they were using all that
- 12:26depth and all that representational
- 12:27power to learn modular recursive
- 12:32functions for things and all sorts of
- 12:33exciting stuff. It is not what we find
- 12:35and that seems like a real opportunity
- 12:37to just level up and do better and maybe
- 12:40we could get massively more capable
- 12:42models with half the depth and a quarter
- 12:45of the representational width. And that
- 12:47would be transformative for the
- 12:48economics of AI in addition to leading
- 12:50to all sorts of exciting things for
- 12:52capabilities.
- 12:53>> Yeah, it's funny and maybe a bit
- 12:55validating for me to hear you say that
- 12:57because whenever I architect arch
- 13:00whenever I articulate a thought in that
- 13:02direction
- 13:03with
- 13:04uh particularly with folks that are you
- 13:06know, coming from the frontier labs or
- 13:09you know, the the
- 13:11um
- 13:12essentially the frontier labs
- 13:14I get back this kind of feeling that
- 13:16yeah, you're just not bitter lesson
- 13:17piled enough. Like structures, that's
- 13:21old school thinking, you know, you're
- 13:23just trying to like train some features.
- 13:26Just collect a lot of data, throw it at
- 13:28the model, and that's all you need.
- 13:30>> Okay, but here's my response to them.
- 13:33Let's say rewind to 2017. We've got the
- 13:35transformer.
- 13:36It's got absolute positional encodings,
- 13:39and it's got a particular structure for
- 13:42its uh MLP layer, which is pretty narrow
- 13:44and pretty dense,
- 13:46and a certain structure to its
- 13:48activations and its uh layer norms.
- 13:50That's 2017.
- 13:53The bitter lesson pill thing to do would
- 13:54be to scale that up.
- 13:56But just consider, for example, how much
- 13:59it would cost to use the N squared
- 14:02attention and the absolute positional
- 14:03encodings but have a context window of 1
- 14:05million. This is the bitter lesson pill
- 14:08thing, right? Just keep scaling. But it
- 14:09would be absurd. It would cost trillions
- 14:11of dollars to produce models that we all
- 14:13interact with right now. What did people
- 14:15do instead? They thought hard about
- 14:16locality, and they thought about how
- 14:19like positional encodings should be
- 14:21favoring local relationships. They
- 14:23completely rethought the MLPs so that
- 14:25it's now wide and sparse. Everyone did
- 14:27careful work on the activation functions
- 14:30to make sure there weren't weird
- 14:31outliers so that they could quantize in
- 14:33a good way. And so forth and so on. All
- 14:36of this analysis work built on
- 14:37intuitions about data and learning led
- 14:40to the model that we have now, which is
- 14:41like a ship of Theseus compared to the
- 14:432017 transformer. The only thing that
- 14:44survives is attention and the feed
- 14:47forward layer.
- 14:48And I claim for you that none of that
- 14:50stuff is bitter lesson pill. That was
- 14:52all analysis work that was meant to save
- 14:54based on priors in the data and priors
- 14:56about how they knew learning would
- 14:57happen. So I go back at them. You're not
- 15:00bitter lesson pill enough, apparently.
- 15:02Although, this is a reductio, I think.
- 15:06>> Oh, I love this. That's such a great
- 15:07response.
- 15:08I think it also really calls out the
- 15:13relationship between data, mech and
- 15:16terp, and efficiency, like core themes
- 15:19that you've been focused on and how
- 15:21they,
- 15:22you know, interrelate and support one
- 15:24another.
- 15:25>> Yeah, absolutely. And this relates to
- 15:26one of my hot takes, you know, it's very
- 15:27fashionable, especially among inter
- 15:29researchers, but I think in general for
- 15:30people to say, "We don't understand how
- 15:33these models work. It is also very
- 15:34mysterious to us." But the truth is that
- 15:37people in the field have very deep
- 15:39intuitions about how these models work,
- 15:41and that is the causal factor in us
- 15:44making so much progress, because they
- 15:46could think analytically, "What would
- 15:48the structure of positional encodings
- 15:50and attention be I could do this at
- 15:52million context scale." You can only
- 15:54achieve that kind of thing based on deep
- 15:56analysis and insight, not by just
- 15:58guessing.
- 15:59And so when people say, "Oh, we don't
- 16:01know understand." I say, "I think you
- 16:02understand much better than you're
- 16:04letting on. I think you understand at
- 16:06least as well as my car mechanic
- 16:08understands how my car works." There are
- 16:10mysteries,
- 16:11but you can take a lot of action and be
- 16:13very effective improving things.
- 16:15>> Why do you think they say
- 16:17that? Why do you think they say that
- 16:18they don't understand the models?
- 16:20There's just got to be some payback
- 16:21there.
- 16:22>> It's probably a paradox of expertise,
- 16:24right? So the more you do know, the more
- 16:26you feel like there are also mysteries,
- 16:28and it's hard to step back from that and
- 16:30be objective and say, "Yeah, well, we
- 16:31did make a phenomenal amount of
- 16:33progress, and that can't be just because
- 16:35of happenstance. That was because we
- 16:37know a lot." But all you see as an
- 16:39expert is all the things that are still
- 16:41to be explained.
- 16:43Um partly also it's just a narrative in
- 16:45the field, and it does stretch back to
- 16:46days when I think we had very little
- 16:48understanding of how these models work.
- 16:50Possibly because a lot of them weren't
- 16:51that good. There was very little to
- 16:53explain. And so that's just been slow to
- 16:55catch up with how much progress we have
- 16:57made in understanding the kind of
- 16:59intuitive human-level mechanisms that
- 17:01these models are operating with.
- 17:03>> I also wanted to ask you about DSPY. I
- 17:07forgot about this as we were talking
- 17:09earlier, but you were involved in DSPY,
- 17:13which um
- 17:14well, I'll let you talk about it, but
- 17:16I'm curious
- 17:18uh how it connects into your research
- 17:21and like
- 17:22uh you know, some of these pillars that
- 17:24we've we've talked about.
- 17:26>> Oh, there's lots of wonderful strands.
- 17:28And what a meta strand I could offer you
- 17:30cuz we were talking about being
- 17:31strategic with research, this does stem
- 17:34from Omar Khattab, my student. He's the
- 17:36the the visionary behind DSPy and still
- 17:38its lead.
- 17:39And he just had the intuition early on
- 17:41that we should rethink what it means to
- 17:43make a scientific contribution.
- 17:45Previously, we thought in terms of
- 17:46papers as the beginning and the end of
- 17:48all of this kind of thing that you would
- 17:50contribute.
- 17:52We should instead, he said, think about
- 17:54projects and about empowering people.
- 17:56And so for him, the paper is one part of
- 17:59a broader contribution that might
- 18:01actually be centered on
- 18:03an open-source or open-weights release
- 18:06that would allow people to do big
- 18:08things. And that's where you find impact
- 18:10and that's the nature of a contribution
- 18:12going forward.
- 18:13And DSPy is a kind of embodiment of
- 18:15that. Although, he made a similar
- 18:16investment with the ColBERT retrieval
- 18:18model.
- 18:19And then people built on what he did,
- 18:22and then you really saw it take off
- 18:23where open-source contributions made it
- 18:25easier and easier to use that
- 18:26technology, leading to more and more
- 18:28impact. And of course, DSPy is another
- 18:31wonderful example because in investing
- 18:34in this community
- 18:35and in the open-source resource itself,
- 18:38he built a huge following.
- 18:40There are lots of startups, mine
- 18:41included, where the core tech stack for
- 18:43the LLMs is built on DSPy, and that has
- 18:46made life so much easier. And then of
- 18:48course, it was a platform for him and
- 18:50for us to really think in an innovative
- 18:52way about prompt optimization and
- 18:55agentic workflows and all of those
- 18:57things.
- 18:58>> Yeah, I was thinking not too long ago
- 19:00the degree to which model strength
- 19:06as a correlate to model size, I suppose,
- 19:08and capability
- 19:11has
- 19:12kind of overcome the need for an
- 19:15explicit framework like DSPy DS Pi.
- 19:19>> Yeah, there are kind of two levels to
- 19:20that. The one would be just the
- 19:21engineering side where
- 19:24uh DS Pi is great four years ago because
- 19:27it's kind of hard to construct the code
- 19:29around one of these systems in a way
- 19:30that's modular and reproducible and so
- 19:32forth because packing out something
- 19:34where you've got a prompt string in the
- 19:36middle of your code with some slots in
- 19:38it, it's very error-prone and it leads
- 19:40to bad system designs. And DS Pi solved
- 19:42that. And you could think that the need
- 19:44for that is diminishing somewhat because
- 19:47now we all specify these systems in
- 19:49English and have the coding agents do
- 19:51them.
- 19:52>> And even before that there were, you
- 19:54know, another hundred frameworks that
- 19:56solved that particular part of the
- 19:57puzzle.
- 19:58>> Oh, yeah, there's always competition and
- 19:59I think at that level of just thinking
- 20:01about programming interfaces and APIs,
- 20:04they can all learn from each other. And
- 20:05so like, you know, DS Pi learned a lot
- 20:07from PyTorch in terms of layer-wise
- 20:09design and the kind of modularity that
- 20:11introduced. And then, of course, you
- 20:13would hope that everyone kind of slurps
- 20:15up all these interesting innovations and
- 20:17it leads to everyone being better.
- 20:18There's lots of evidence of that at the
- 20:20level of interfaces. I would maintain
- 20:22for you that even if we have agents
- 20:23actually writing the code for these
- 20:24systems, it's great for us and for them
- 20:27if they write it in something that
- 20:28actually expresses these systems as
- 20:30modular components so that we can audit
- 20:32them, so that they can change them. It
- 20:34just feels like good engineering
- 20:35practices for any agent to think in a
- 20:37modular way. And that's what DS Pi
- 20:40encodes. The other side is like the
- 20:42prompt optimization side and a belief
- 20:45people have that the need to be careful
- 20:47with your prompts is diminishing over
- 20:49time.
- 20:50I understand that narrative, but people
- 20:52should also, for example, just run like
- 20:55a simple annotation study where they use
- 20:57a few different models or the same model
- 20:59a few times on slightly different data.
- 21:01They will be blown away by the amount of
- 21:04variation that still exists.
- 21:06To be charitable, let's say that these
- 21:08LLMs disagree about fundamental facts
- 21:10about how to label certain texts or what
- 21:13kind of response to give.
- 21:14We all kind of slip past this because we
- 21:16feel like, "Hey, they're smart and
- 21:17they're good and they're getting
- 21:18better." But if you quantify it, it's
- 21:20pretty disturbing. And the next step
- 21:22from that is to think about having all
- 21:24those agents optimize a prompt so that
- 21:26their behavior is at least consistent.
- 21:28And then you're right back at that DESP
- 21:30vision.
- 21:31>> It's interesting that you say that
- 21:33because
- 21:35I don't I don't feel like that
- 21:36necessarily aligns with my recent
- 21:39experience. And in particular,
- 21:42one thing that I've noticed that's been
- 21:44surprising is
- 21:47how
- 21:49well aligned, I guess. Maybe that's not
- 21:52the right word, but how similar the
- 21:53responses I get to uh uh
- 21:56query across different models. So, for
- 21:58example,
- 21:59you know, these are, you know, often
- 22:02kind of what I would call like a casual
- 22:04prompt, a casual query, something that I
- 22:06might, you know, uh type into Google.
- 22:09And it will now generate uh an LLM
- 22:12response for me in this kind of AI mode.
- 22:15Um
- 22:17and I'll take the same thing and put it
- 22:19into ChatGPT and maybe Claude. And it
- 22:23surprises me that the
- 22:26the responses are often very, very
- 22:28similar. Like, you know, very similar
- 22:30structure, very similar facts, very
- 22:33similar citations.
- 22:36And
- 22:37you know, I
- 22:39stepping back, like, there are lots of
- 22:40ways that they could answer or approach
- 22:42these different questions. But it seems
- 22:44like, you know, the models or the
- 22:45training or the system prompts or
- 22:48something is all kind of converged on
- 22:50something that makes the models express
- 22:52themselves, you know, very similarly,
- 22:55which, you know, seems to be at odds
- 22:57with uh
- 22:58you know, the the last thing you said
- 22:59about the need to optimize prompts or
- 23:02the impact of the individual prompt.
- 23:04>> I'm open-minded, but for example, like
- 23:06we just did a we did a we did a thing
- 23:07recently, we were writing a grant and we
- 23:09needed a title and you want to be
- 23:10strategic with these titles. So, we come
- 23:12up with a whole bunch of them ourselves
- 23:13and then we all disagree on what would
- 23:15be the best. So, let's find out what the
- 23:17agents think. So, ask a few Anthropic
- 23:19models and a few um
- 23:21GPT models, which of these five titles,
- 23:24which is the best?
- 23:25So, you get a different answer from all
- 23:27of them along with a detailed rationale
- 23:29about why obviously, of course, the
- 23:31choice that the model has made in that
- 23:32moment is the best one.
- 23:34This is great because then we can think
- 23:36about which one of these arguments is
- 23:37most persuasive.
- 23:39But if you were hoping for consistency
- 23:41at an subjective labeling task, which
- 23:43this is one,
- 23:44uh you can see right there that you're
- 23:45going to have a real problem unless you
- 23:47give very specific criteria
- 23:50and then you're kind of also
- 23:51constructing a prompt for them and you
- 23:53might want to manage them differently.
- 23:55There is a real I don't have evidence
- 23:57for this yet, but we have an intuition
- 23:59at Big Spin in the research we've done
- 24:01that you get a kind of paradox that the
- 24:03more requirements you add actually the
- 24:05more variation you'll see because the
- 24:07different models will key into different
- 24:08subparts of the requirements and since
- 24:11they do it very concertedly,
- 24:13you can actually get systematically
- 24:15biased behavior from something that you
- 24:17thought was a very good specification.
- 24:20And that again calls for this idea that
- 24:22what you need to do is figure out what
- 24:23the labels ought to look like and then
- 24:25have some automatic optimization process
- 24:28get the model there.
- 24:30And that's what things like Jeppa and
- 24:32MeetPro are for.
- 24:34>> A topic that I really wanted to A topic
- 24:37that I would really like to dig in to
- 24:39with you based on our our previous
- 24:42conversation was the idea of tokenomics.
- 24:45Uh it's something that people are
- 24:47talking about a lot recently.
- 24:50Um I think, you know, folks that use
- 24:52Claude code, for example, have like a a
- 24:55visceral experience with Anthropic
- 24:57changing the terms around usage, but
- 24:59it's happening under the covers with all
- 25:01of these large providers.
- 25:04And so, I think
- 25:06way more now than
- 25:09you know, 6 months ago, like we're all
- 25:13a little antsy with the relationship we
- 25:14have with these, you know, big model
- 25:16providers and the the value that we get.
- 25:19And you recently
- 25:20wrote an article about this. You know,
- 25:23talk a little bit about
- 25:24a how to how
- 25:27how it ties into kind of your broader
- 25:29research, uh but also some of the things
- 25:32that you uh found when you started to
- 25:34dig into this area.
- 25:36>> Yeah, another interesting moment to be
- 25:39in
- 25:39as we're all being made aware of the
- 25:42true costs of all this AI usage. I saw a
- 25:44tweet from Ed Zitron
- 25:47just a screenshot from someone who was
- 25:49noticing that Copilot was telling them
- 25:52that their month their bill last month
- 25:53was $500. And if they keep up the way
- 25:56they are with Copilot's pilots new
- 25:57billing, it will be $11,000
- 26:00in the next month.
- 26:01>> Wow. Wow.
- 26:02>> Which is real sticker shock. And you
- 26:04know, the analogy here is like I used to
- 26:06cost me $20 to take a ride share to the
- 26:09airport, Uber or Lyft, and now it costs
- 26:1290.
- 26:12>> I use that analogy as well.
- 26:14>> But it's more [laughter] like 20 to like
- 26:16500 or something, right?
- 26:19>> Right.
- 26:19>> Um
- 26:20>> If only the slope will be as shallow as
- 26:21Uber, right? [laughter]
- 26:22>> That's right. We start to wish for those
- 26:24easier stories.
- 26:26Yes, and so what will happen? I mean, I
- 26:28think what's happening is that the big
- 26:30providers are testing the waters on
- 26:32charging us
- 26:33the true costs plus whatever profit they
- 26:35need to make as they all try to gear up
- 26:37for IPOs and so forth.
- 26:39And in turn, that is very quickly
- 26:41leading people to ask questions like
- 26:43what is the return on investment for all
- 26:45these tokens that we have purchased?
- 26:49And it's a very tricky area to be in
- 26:51because what does it mean to think about
- 26:53value in this context? Even if we focus
- 26:55in on people who are doing just coding
- 26:58with coding agents,
- 27:00can we agree on what it means to add
- 27:02value? And maybe we have a few measures
- 27:03in mind like making a pull request or
- 27:07uh committed lines of code that last in
- 27:09the repo for a while or
- 27:11documentation touched or skill files
- 27:14created.
- 27:15But we might also worry that that's not
- 27:17capturing the value for many kinds of
- 27:19sessions we have which are more
- 27:21open-ended and about discovery.
- 27:24So that's the first question is just
- 27:25solving this value issue, right? We just
- 27:27agree on what it would mean to add value
- 27:28for a for a coding agent.
- 27:31>> Now I I like the this line of inquiry
- 27:35because
- 27:37to me it's it's the response to this
- 27:39thing that drives me crazy which is
- 27:42oh you know
- 27:44big
- 27:45uh you know tech company CEO this year
- 27:4895% of our code will be generated by you
- 27:51know AI.
- 27:53It's like
- 27:55yeah, hey what does that really mean at
- 27:58that level? Like what's what's what are
- 28:00the details beneath there but uh is that
- 28:02a good thing or a bad [laughter] thing?
- 28:05>> That's a oh another dimension, right?
- 28:07Which is is that code a liability or an
- 28:09asset?
- 28:09>> Right. Right. Right. [laughter] And this
- 28:12idea of like uh you articulate it as
- 28:14kind of code longevity in the code base.
- 28:16That's an interesting way to think about
- 28:17it. There's probably a lot of
- 28:18interesting ways to think about it that
- 28:20very few are thinking about right now.
- 28:22>> And all of these fall victim to the
- 28:24standard thing that once you make it a
- 28:25metric, it's no longer useful to you.
- 28:27Like if we said oh let's get it's
- 28:29completion of projects, right? Well then
- 28:31everyone would just have many projects
- 28:32that they completed but they could all
- 28:34be liabilities and add very little
- 28:36value. So but one framework we could
- 28:38offer that we did in the research you
- 28:40alluded to is let's think about this
- 28:42like economist might. So we might have
- 28:43like a consumer price index and the
- 28:46first step will be what's the basket of
- 28:48goods that we're going to consider, You
- 28:49you know, in that standard land it would
- 28:51be like the the price of eggs and the
- 28:54cost of rent and other kinds of tangible
- 28:56goods. What are engineering goods that
- 28:58we might track?
- 28:59>> And eggs might be a summary or
- 29:02uh pull request, a bug fix or something
- 29:04like that.
- 29:05>> Or we could think broadly cuz we both
- 29:07use these coding agents,
- 29:09um
- 29:09requirement discovery, right? Knowledge
- 29:12accumulation. These are things that we
- 29:14don't currently track, of course, even
- 29:16as engineers, but might be behind our
- 29:19intuition that these coding agents are
- 29:20making us productive even if it's not
- 29:22reflected in the PR counts or whatever,
- 29:24right? I mean, in a sophisticated
- 29:26approach you might say, "I don't want
- 29:27more PRs because this is just a certain
- 29:30kind of um
- 29:31busy work that doesn't relate to the
- 29:33actual goals I have." What are the
- 29:34actual goals? It's completing valuable
- 29:36projects and so forth. If I could do it
- 29:37with fewer PRs,
- 29:39um but I had, you know, really robust
- 29:41code, I'd be possibly happy with that.
- 29:44So, we got to figure out what the basket
- 29:46of goods is, but then we could start to
- 29:47track it relative to token usage, and
- 29:49that would be the consumer price index.
- 29:51So, for any time period we could just
- 29:52say I've got my tokens spent and I've
- 29:54got my goods produced. Tokens divided by
- 29:57goods produced is a pretty rough measure
- 30:00of um
- 30:01the purchasing power of the tokens in
- 30:03those time periods.
- 30:05Then you would do the standard compute
- 30:06consumer price index thing of making
- 30:08what they call a hedonic adjustment. So,
- 30:10you could just say maybe quality is
- 30:11improving over time. So, you'd pick some
- 30:14measure for that and make an adjustment
- 30:15to the line.
- 30:17And when we did that study, we did code
- 30:19survival. So, um the number of lines of
- 30:21code that survives more than 4 days in
- 30:23the repository. We made an adjustment
- 30:25upward because that rate is going up.
- 30:27>> That's surprisingly short.
- 30:30>> 4 days?
- 30:30>> 4 days survival?
- 30:32>> You could make So, again, all this is
- 30:33around measurement and I'm happy to just
- 30:35be starting this dis this discourse
- 30:37because we can see it's important to the
- 30:38economics of AI and it seems like the
- 30:40work isn't being done at a high enough
- 30:42rate for us to get a clear picture. So,
- 30:44we could make it longer and maybe the
- 30:45adjustment would be different. I think
- 30:47currently for the data we have, which is
- 30:49this SweetChat benchmark, which was
- 30:51released by researchers at Stanford.
- 30:53It's about 6,000 real coding sessions,
- 30:56all the metadata, everything you'd want.
- 30:59What we see with Opus 4.6 usage in the
- 31:02time period we have, which is February
- 31:04to mid-April of this year,
- 31:06a decline in the purchasing power of
- 31:08tokens. That CPI is going down.
- 31:12And
- 31:13again, I just want to open the question,
- 31:15is it because we have the wrong basket
- 31:16of goods, or is it because we're
- 31:19actually getting less value from these
- 31:20tokens? The The The one thing I can say
- 31:24that's kind of definitely a causal
- 31:25factor here
- 31:27is that in February of this year, most
- 31:29of the tokens went to producing code,
- 31:32which relates to the outcomes we just
- 31:33talked about. By mid-April, it was quite
- 31:37split between code generation, thinking,
- 31:40and also explanation to the user.
- 31:43And so, that split now is going to have
- 31:44an effect on the things we're measuring,
- 31:46and that might be cause for reflection.
- 31:48There's value in those explanations
- 31:50that's not reflected in PRs, but might
- 31:52be reflected in something like knowledge
- 31:53discovery.
- 31:54>> And I see that coming up within the same
- 31:58time frame, it's become very common to
- 32:01now talk about the token efficiency of a
- 32:04new model that's been released. Uh
- 32:07with the implication being
- 32:09uh tokens of, you know, internal use
- 32:12tokens, thinking tokens versus, you
- 32:14know, per token of output, I guess, is
- 32:16maybe a way to think about it.
- 32:18>> Another fascinating dimension, and this
- 32:19actually relates all the way back to the
- 32:21theme of efficiency for these
- 32:22architectures. So, here's a claim I'll
- 32:24make for you.
- 32:26Based on my read of the literature on
- 32:28inference time scaling,
- 32:30what's sometimes called test time
- 32:32scaling, which is just having the models
- 32:34generate lots of tokens
- 32:36at the moment that you ask them a
- 32:37question. So, those scaling trends,
- 32:39everything we're seeing now is
- 32:40completely in line with those
- 32:42predictions, which is
- 32:44you get pretty good gains for a while
- 32:46with the more tokens you spend on a log
- 32:49scale. So, this is jumping up quite a
- 32:50lot, but you do see it reflected in
- 32:52performance improvements, but it
- 32:54flattens out over time.
- 32:56And it's not like this curve skyrockets.
- 33:00It's sobering. You got to spend a lot of
- 33:02tokens for small gains in performance.
- 33:05We all knew this.
- 33:07We all knew this, and we're just seeing
- 33:09it now play out. And when people talk
- 33:10about token efficiency and worry about
- 33:12this, I think what they're seeing is
- 33:13just the real lesson of what we already
- 33:15projected from inference time scaling.
- 33:17>> And this is independent of the approach
- 33:20to inference time scaling you're taking,
- 33:22whether it's
- 33:23you know, multiple parallel
- 33:26you know, multiple parallel inferences
- 33:28or some kind of oracle or you know, any
- 33:32number of other schemes. It's just
- 33:34fundamental to inference time scaling.
- 33:36>> It's a great question, right? I think we
- 33:39know that it's independent of some of
- 33:41those things, like the parallel work
- 33:43versus having it do lots of long chains.
- 33:45But some of the other factors you
- 33:47mentioned, I think we just don't know,
- 33:48and that's why I said it relates back to
- 33:50the question of efficiency for these
- 33:51architectures.
- 33:53If we made a fundamental change to how
- 33:55the models work,
- 33:56maybe these tradeoffs would be very
- 33:58different. I mean, after all, so all of
- 34:00this stuff is a kind of patch job on the
- 34:03fact that there's no recursion in the
- 34:05depth. It's a fixed depth. And so, the
- 34:08only recursion we can get, the
- 34:09open-ended notion of computation is by
- 34:11generation. But if we had models that
- 34:13could be recursive, maybe fewer tokens
- 34:16for larger gains. I think we don't know.
- 34:19Yeah, I mean, in the end, we're going to
- 34:20spend the cost on compute or tokens.
- 34:24So, this might not affect our bills in
- 34:25the end, but it is a fascinating
- 34:27question. What are the true scaling laws
- 34:29and what's possible in this space? And
- 34:31you're right to push back. We talk about
- 34:33these things like they were like
- 34:34platonic ideals of laws.
- 34:37Scaling law invokes that.
- 34:39You're right, but even for the scaling
- 34:40laws for pre-training, you know, there's
- 34:42lots to discover there. And many of the
- 34:45stories of progress are actually like
- 34:47transcending the scaling law. And we see
- 34:49like better improvements than those laws
- 34:51predicted because everyone worked so
- 34:53hard behind the scenes to do very
- 34:54innovative things, which maybe relates
- 34:56to our bitter lesson discussion.
- 34:58>> Any particular example come to mind of
- 35:00that?
- 35:01>> Data usage and the nature of the data
- 35:02really matters, and overtraining the
- 35:04models really matters, which is kind of
- 35:05pushing up against the standard scaling
- 35:07law presentation. And now I'm just going
- 35:09to speculate. I should check on this,
- 35:11but things like mixture of experts might
- 35:13have really flipped the script on what
- 35:15it means to count parameters and in turn
- 35:17how these laws relate. And then I think
- 35:19maybe even also stuff like the context
- 35:20window and so forth. This is another
- 35:22thing to check, but I just speculate
- 35:24that we've seen larger gains from
- 35:26pre-training than you would have
- 35:28predicted by those early scaling laws
- 35:30papers, suggesting that there is some
- 35:32innovative thing that was happening on
- 35:34top of pure scaling.
- 35:35>> Thinking about the concept of a a market
- 35:39basket, one kind of pushback that
- 35:42comes up for me is
- 35:44in the you know, the real economy, you
- 35:47know, eggs is different than milk is
- 35:50different than
- 35:52you know, beef, etc., etc. And they're
- 35:56all
- 35:58influenced by different factors.
- 36:01You know, production, for example.
- 36:04Whereas
- 36:06what you've done with the this kind of
- 36:08CPI basket with tokens is kind of like
- 36:13more like analogies. Like here's the
- 36:15typical bundle of work and you know,
- 36:17what it requires from a consumptive
- 36:19perspective, but
- 36:21the tokens aren't fundamentally
- 36:22different. Like they're the same tokens.
- 36:24It's just like how much it takes to do
- 36:25this versus how much it takes to do that
- 36:27versus how much it takes to do that.
- 36:29Um
- 36:31you know, tell me what I'm what I'm
- 36:32missing there and
- 36:34you know, what does kind of
- 36:36characterizing these products, you know,
- 36:39give you in your analysis?
- 36:41>> Yeah, fascinating to think about. One
- 36:43thing I could insert there is the tokens
- 36:45are different at the level of being used
- 36:47for code generation or skill file
- 36:50writing or explanation or thinking,
- 36:52right? Those are different kinds of
- 36:54tokens that probably do feel tangibly
- 36:56different to us. So, is that an element
- 36:58in your thinking?
- 36:59>> I think I was thinking from our
- 37:02conversation that you had 10 different
- 37:04almost like tasks, like 10 different
- 37:06types of tasks from the domain of code
- 37:09generation, which you know, if they were
- 37:12all kind of largely code generation, you
- 37:14know, that is the part that had some
- 37:16dissonance for me. But, if you're
- 37:18talking about like if your your basket
- 37:20is like creative writing versus, you
- 37:23know, a few code generation things that
- 37:25are kind of in different uh versus, you
- 37:29know,
- 37:30summarization versus editorial
- 37:33commenting feedback, those, you know,
- 37:35may be more fundamental.
- 37:37>> Yeah, so I think this is very
- 37:39significant. And we have done some
- 37:41research on this as well at the level of
- 37:43what kinds of session types exist and in
- 37:45turn what kinds of users are there. So,
- 37:47you might notice of your own behavior. I
- 37:48guess this is reflected in your comment
- 37:50that sometimes you want a quick check-in
- 37:52on a question. Sometimes you want a
- 37:54quick um PR to get fired off. Sometimes
- 37:57you want to be in a mode of deep
- 37:58collaboration.
- 38:00Sometimes you're partnering with the AI,
- 38:02sometimes you're delegating the work and
- 38:03so forth and so on.
- 38:05And the outcome measures that we choose
- 38:07should be sensitive to this. We
- 38:09shouldn't penalize the agent if your
- 38:11chat interaction with it about some
- 38:13scientific question didn't lead to a PR.
- 38:15It was never on the table in the first
- 38:17place.
- 38:18Whereas, if you're trying to get some
- 38:20work delegated that's actually a coding
- 38:22task and all it does is chat with you,
- 38:24that would feel quite unproductive.
- 38:27So, we need to bring that in and that
- 38:28would be a higher level discovery
- 38:29process of what people are trying to do
- 38:32and so forth.
- 38:33>> In thinking about the the notion of
- 38:36value, is this
- 38:38something that you're anticipating
- 38:42Like it it strikes me that that's an
- 38:43entire, you know, research thread that,
- 38:46you know, one could go into. I don't
- 38:47know if that's a linguistics or a
- 38:49linguist or an economist or a computer
- 38:52scientist. You know, probably
- 38:53interdisciplinary.
- 38:56Like most interesting questions. Um, but
- 38:58is that, you know, is that something
- 39:00that you're working on or
- 39:03um, was it something that you put out
- 39:05there for someone to take up and run
- 39:07with?
- 39:08>> I am not sure. I can tell you the
- 39:10the lineage of this idea is that we
- 39:13founded the startup Big Spin because we
- 39:15would like to see more people benefit
- 39:18from AI.
- 39:19Whether you love it or hate it, it's
- 39:20here
- 39:22and I would like the benefits to be more
- 39:24evenly distributed. And I can tell that
- 39:26that will mean
- 39:28bringing on board many more people than
- 39:31currently benefit from AI. Right now, I
- 39:33would say that it's mostly experts
- 39:35deriving real value
- 39:37and a lot of the world is currently even
- 39:39trying to figure out what this is all
- 39:40about as a tool or an entity in their
- 39:42lives.
- 39:44So, we would like to have more access
- 39:46and more productivity
- 39:48and that implies making the user
- 39:49experiences much better.
- 39:52Um, figuring out what interactional
- 39:54patterns lead to success for people,
- 39:56meeting them where they are in a kind of
- 39:58adaptive way. The whole list of things
- 40:00that you might worry about if you were a
- 40:01product manager who had some deployed AI
- 40:03product.
- 40:04And I think by that route and from that
- 40:07perspective, we just ended up worrying
- 40:09about our own token usage increasing and
- 40:12wondering whether there's real value
- 40:14there and it just happened to collide
- 40:16actually just like 3 weeks ago
- 40:18with this emerging narrative on the back
- 40:20I think of all these rumors about IPOs
- 40:23about what the return on investment was
- 40:25and then all these CEOs came out and
- 40:26said oh our spend was enormous and we
- 40:28want to scale back and we're walking
- 40:30back our claims from a few months ago
- 40:31and that is just a fascinating thing to
- 40:33witness in the
- 40:34in the narrative here.
- 40:36>> I'm wondering are you also does this
- 40:38research also attempt to project
- 40:40forward? In theory you could
- 40:44um you know create a model for you know
- 40:48Anthropic's cost and spend you know
- 40:51based on you know publicly available
- 40:53data and some presumptions and
- 40:56give us a sense for
- 40:59you know how close we are to paying full
- 41:01freight for our tokens versus you know
- 41:03if we're only paying 10% for our tokens
- 41:06you know you could then project you know
- 41:08what that cost might look like uh
- 41:11you know over time as we're paying more
- 41:13and more of the the full cost.
- 41:15>> Yeah I don't have again fascinating
- 41:16questions I don't have resolving
- 41:18answers. I am glad I am not tasked in
- 41:20some organization with projecting spend
- 41:22on all of this stuff because I think it
- 41:23would be basically impossible. For the
- 41:25time period that I was describing for
- 41:27our little um CPI experiment Anthropic
- 41:30changed the default reasoning on the
- 41:32model at least two times. So we see like
- 41:35it start they launched it with default
- 41:37reasoning high. We have a mysterious
- 41:39sudden rise in the token usage which we
- 41:41cannot explain. And there's a new
- 41:43baseline they lowered it to medium as
- 41:45the default they patched a bunch of bugs
- 41:47that were related to context management
- 41:49and then turned it back up to high.
- 41:51And all of these things have an effect
- 41:53on the total token output as you can
- 41:54imagine they also changed the default
- 41:55context window which meant people could
- 41:57swallow up much more
- 41:59uh stuff at any given moment.
- 42:02So imagine trying to predict what token
- 42:04spend is going to be like when you have
- 42:05all these exogenous events in addition
- 42:09to
- 42:10changes that we don't even know about
- 42:12and questions about where the value
- 42:13actually lies.
- 42:15Very difficult. And then, you know, the
- 42:17true cost of a token, the estimates vary
- 42:19wildly for every dollar we spend, it
- 42:21could be as low as two and as high as
- 42:2320. And I think this is just because
- 42:26it's hard to factor in things like R&D
- 42:28and future build-out and depreciation
- 42:30and all of that stuff. I think at the
- 42:32current moment we just don't know, but
- 42:34there couldn't be a more significant
- 42:35question for the global economy,
- 42:37basically, than
- 42:40where the value is and who's going to
- 42:42pay and how much.
- 42:44>> Yeah, in your article you coined the
- 42:46term tokenflation to
- 42:49describe, at least the recent behavior
- 42:51of token economics. I imagine you see
- 42:55that continuing.
- 42:56>> Seems to be continuing. Yeah. Yeah,
- 42:58that's certainly the picture that we get
- 43:00from the CPI, a picture of tokenflation.
- 43:02Yes, your token is not buying you what
- 43:04it once did, according to everything we
- 43:06can think to measure here.
- 43:08And even adjusting for models getting
- 43:10better, right? That's critical there.
- 43:12Cuz if it was just a story of models
- 43:13thinking more and being more robust, and
- 43:16we were all getting exponentially better
- 43:17outcomes from this, then the spend would
- 43:19look completely rational.
- 43:21But that's not the picture that we see,
- 43:23and so we have to do some hard thinking
- 43:24about what's going to happen and how to
- 43:27improve the situation.
- 43:28>> Let's dig into that a little bit more.
- 43:30Your How would you articulate what
- 43:32you're seeing? The models are
- 43:35getting, quote unquote, better.
- 43:37Um you know, there's a set of open
- 43:39questions about uh are the reported ways
- 43:43that models are better actually
- 43:47reflective of some intrinsic betterness,
- 43:49like in that question brings uh
- 43:52it is often about like benchmarking and
- 43:56uh learning the benchmarks, overfitting,
- 43:58that kind of thing.
- 43:59Uh and then there's
- 44:02the kind of question of chattiness and
- 44:05and the
- 44:07the volume of thought that it
- 44:10requires a given model generation to
- 44:12produce an answer. What are other
- 44:14factors that you see?
- 44:16>> Yeah, we could pick that apart as well.
- 44:17So, and this relates to
- 44:21uh a line I've had consistently, which
- 44:23is that we should think in terms of
- 44:24systems, not in terms of models. So, in
- 44:27the data that we've got, Sonnet or Opus
- 44:30and Sonnet 4.5 versus 4.6, those two
- 44:33generation changes,
- 44:35those are real model changes, I assume.
- 44:36I think they did something very
- 44:37substantive at the level of the weights.
- 44:40Um and everybody immediately saw that
- 44:42that led to like a 5x increase in token
- 44:44usage. And this was related to the
- 44:47introduction of adaptive thinking.
- 44:50Now, fix that. That's the level shift
- 44:52that we already took, and maybe we're
- 44:54seeing improvements there that are
- 44:56worthwhile. It gets hard to say, but
- 44:58let's assume there was a level up in
- 44:59improvement.
- 45:01Then for the period that we did our CPI
- 45:03experiment for, that's a fixed model,
- 45:05Opus 4.6.
- 45:07So, all the code improvements that we
- 45:09saw in the data relate to the product.
- 45:12This has to relate to things like them
- 45:14turning the knobs on the adaptive
- 45:15thinking,
- 45:16changing things about the system prompt,
- 45:19changing things at the level of the
- 45:20product, and that's where the
- 45:21improvements were.
- 45:23And so, that shows you that even for a
- 45:24fixed model, we can get very different
- 45:26outcomes for these things because they
- 45:27really are sophisticated engineered
- 45:29systems at this point.
- 45:31>> Yeah.
- 45:31And so, was the product in this case
- 45:34specifically Claude Code or
- 45:37>> Oh, yeah. And so, we don't There's tons
- 45:38of stuff there.
- 45:40Yes, I believe we know that these are
- 45:41all Claude Code sessions that we kept in
- 45:43our data. Sweet Chat is broader than
- 45:44that and involves a couple of other
- 45:46coding agents, but I think I can say
- 45:48that all our data are Claude Code
- 45:49sessions using Opus 4.6.
- 45:52>> Have you seen any evidence that
- 45:55changes
- 45:56via API usage experience
- 45:59uh similarly dramatic uh variation in
- 46:04performance.
- 46:06>> Oh, fascinating. To kind of control for
- 46:08a lot of that product level stuff, all
- 46:10the prompts that are hidden from us, all
- 46:12of those affordances. Yeah, I don't
- 46:13know, but that's a nice thing to think
- 46:15about because it gives us a more things
- 46:18that we can control for and more things
- 46:19that are knowable.
- 46:20>> So, kind of in parallel to the model
- 46:24evolution, there's also evolution of the
- 46:28user. You've alluded to this a little
- 46:30bit about kind of your concept is that
- 46:32most AI users now are experts.
- 46:37Uh talk a little bit about the role of
- 46:40expertise. I think this is also kind of
- 46:43echoing back to our conversation about
- 46:45DSPy and like prompt optimization. You
- 46:48know, you've done some research into how
- 46:50folks are using these models and the
- 46:51role of, you know, AI fluency. Tell us
- 46:54about that research.
- 46:55>> Oh, yeah. First, I should say, so the um
- 46:58distribution of users across expertise
- 47:00levels.
- 47:02I'd I So, I guess the nuanced picture
- 47:04I'd offer is that the people deriving a
- 47:06lot of value from AI in the current
- 47:08moment tend to be experts. It must be
- 47:10the case that most users of AI are are
- 47:13beginners, just because the numbers are
- 47:15so large and expertise can't be that
- 47:17widely distributed yet. And that's a
- 47:19very interesting thing because I think
- 47:21probably most things are getting
- 47:22designed for those experts implicitly or
- 47:24explicitly. But, for the whole economic
- 47:27picture to work out, many more people
- 47:30need to derive value from this via one
- 47:32avenue or another.
- 47:34And so, that does shine a light on this
- 47:36expertise thing as a real factor.
- 47:38And the headline result there actually
- 47:40builds on something that Anthropic did.
- 47:42They have this AI fluency index, and
- 47:43their core observation in that work is
- 47:46that experts display an augmentative
- 47:49style.
- 47:50They iterate with the AI. They push
- 47:53back. They complain. They change their
- 47:55requirements.
- 47:56It's a really collaborative mode.
- 47:58Whereas novices, low-fluency users,
- 48:02delegate. So, they trust in the AI. They
- 48:05let it do its thing. They accept the
- 48:06responses uncritically.
- 48:08And our contribution is to just show
- 48:10that this is a causal factor in success
- 48:14with these products right now. Experts
- 48:16can do harder things more reliably as a
- 48:18result of all that friction they
- 48:21introduce, all that pushback.
- 48:24Whereas novice users, they accept, but
- 48:27they end up accepting the wrong thing,
- 48:28and they're not able to level up from
- 48:30the basic tasks that they think to start
- 48:33with.
- 48:34And that's obviously significant, and it
- 48:35feels so tantalizing because pushing
- 48:38back is a natural human behavior. I feel
- 48:42like we could encourage everyone in the
- 48:43world to do this. We probably need to
- 48:45get them out of the mode of thinking
- 48:47it's a superintelligence. You should
- 48:48just trust it. That has been the
- 48:50narrative for a while, but we're seeing
- 48:52in the current moment, and possibly for
- 48:54the foreseeable future, is that you got
- 48:56to complain, collaborate, introduce
- 48:58yourself, push back, all that stuff that
- 49:00I think we do, that we take that for
- 49:02granted, right?
- 49:03>> Yeah. Yeah. And so, from a methodology
- 49:07perspective, how did you approach
- 49:09exploring this?
- 49:10>> Hey, we built on the work that Anthropic
- 49:12did,
- 49:13um, which what they set up a nice
- 49:15framework with some independent research
- 49:17who were doing this kind of usability
- 49:19stuff.
- 49:20And we just have an annotation protocol.
- 49:23Like we can talk in detail if you want
- 49:24about this, but at Big Science, we have
- 49:26lots of these best practices around
- 49:28having language models essentially
- 49:29collaborate on annotation projects to
- 49:32kind of triangulate on the truth and
- 49:34factor out their individual biases.
- 49:36So, we do that stuff, and we apply all
- 49:38these fluency markers, and then
- 49:39separately we do a thing of estimating
- 49:41task complexity, and looking for signs
- 49:44of visible and invisible failures.
- 49:47And so, it's the connection between the
- 49:49fluency markers and the task complexity
- 49:52success metrics. That was our
- 49:55contribution there and that's where you
- 49:57can see high fluency users are the ones
- 49:59doing harder tasks. Paradoxically,
- 50:01there's more signs of failure for them
- 50:03uh because they complain, they push
- 50:05back, they're trying harder things.
- 50:07But as part of all that friction,
- 50:09they're successful with harder things as
- 50:10well.
- 50:11>> And if you were to try to apply this
- 50:14insight
- 50:15from the perspective of someone in an
- 50:17organization that's trying to you know
- 50:19help or guide their organization to
- 50:22be more successful with AI, like what do
- 50:25you think are the key lessons of this
- 50:27fluency work?
- 50:28>> If it's an org that's just starting out
- 50:30and wants people to figure out how this
- 50:32could be part of the organization's
- 50:33mission, it would just be that push back
- 50:35message. And you could do an experiment
- 50:37where you interact with it about
- 50:39something where you're a world expert.
- 50:40We're all an expert in something.
- 50:42Engage in a discourse with one of the
- 50:44best models about something you're an
- 50:45expert in and see how often you feel you
- 50:47have to push back and this could be a
- 50:49kind of a lesson you say.
- 50:50>> Aha.
- 50:51>> For other spheres where I don't know the
- 50:53answer, it might be just as errorful.
- 50:55That could be a good visceral thing. If
- 50:57the org is very far along, I think the
- 50:59main thing to do right now is to have a
- 51:01team of these LLMs interacting to
- 51:03improve things. For example, at Big
- 51:05Spin, I didn't set this up. Our founding
- 51:07engineer is very future forward on
- 51:10agents and he's incredible at this. And
- 51:13when we do PRs now, the first round of
- 51:15review is the agents all interacting,
- 51:18collaborating, disagreeing. They do the
- 51:20first round of comments. They do the
- 51:22first round of code updates. Only after
- 51:24they've resolved things do we look at a
- 51:26PR. So the final human stage should be
- 51:29very high value and the agents did all
- 51:31that work. But when you have one agent
- 51:33do it, they often just reinforce
- 51:35themselves and you don't get good
- 51:36outcomes. It's that team of rivals thing
- 51:39that is transformative.
- 51:40>> You know, I think it's interesting
- 51:41because you know, on the one hand like
- 51:44of of course that makes sense.
- 51:46But on other hand it it there's
- 51:48something
- 51:50you know, it also implies that you
- 51:53shouldn't be using these things in areas
- 51:56where you don't have enough expertise to
- 51:58evaluate the answer. Yet
- 52:01that's where you most need the
- 52:04assistance, the support. Um so
- 52:07>> Well, and again and this is a little bit
- 52:09worrisome about the overall narrative
- 52:11around AI.
- 52:13The place where we can get around this
- 52:15is with software development because
- 52:17let's say that I'm trying to accomplish
- 52:18something in a language that I don't
- 52:20know how to code in.
- 52:22I can have the agent do work for me
- 52:24because probably in the end I can run
- 52:26the program and look at the results.
- 52:28And that's what mattered to me is that I
- 52:30run the results and I see and if I don't
- 52:33see what I want then I can complain and
- 52:35we can iterate.
- 52:36That verification step that doesn't
- 52:38imply I have comprehensive knowledge, it
- 52:40just implies that I know what I want to
- 52:41see in the end is so critical and I
- 52:44think this is a causal factor in models
- 52:46being so good at coding because it's
- 52:48like the ultimate verifiable domain for
- 52:50them.
- 52:51But as soon as we leave that and go even
- 52:53into something like the legal realm
- 52:55where the requirements are strict but
- 52:57they're not codified in code and they
- 52:59have ambiguity about them. This whole
- 53:02picture falls apart.
- 53:05And you then are back at what you just
- 53:06said which is this awful kind of paradox
- 53:08is like, yeah, use AI but in the end
- 53:11unless you're expert enough to evaluate
- 53:13every single one of its responses, you
- 53:14might be in real trouble.
- 53:17I don't know how to get out of this
- 53:18because the the verification step is
- 53:20like we go to trial but this is very
- 53:22consequential.
- 53:24>> Yeah, [laughter] that's expensive.
- 53:26Yeah, that's funny. I mean it does make
- 53:28me think a little bit about
- 53:30you know, some of the types of errors
- 53:32that we're trying to avoid are
- 53:34factuality and
- 53:37you know, there is
- 53:38a temptation to say, well, let's just
- 53:40throw more tokens at it. Like I'll have
- 53:42a a critic model that uh uh evaluates
- 53:45everything that the
- 53:48you know, is generated by the primary
- 53:49model. But, then you go back to my
- 53:51observation that these models tend
- 53:54to correlate uh in their responses as
- 53:57well.
- 53:58Um
- 53:59yeah, it's it's super interesting.
- 54:01>> That's a good point. Yeah, from my
- 54:03picture, we want real diversity of
- 54:04perspectives. This is just like, you
- 54:06know, red teaming for humans. This is
- 54:09most successful when you have a really
- 54:11diverse team of people who think
- 54:12creatively and differently. And if every
- 54:14one of the members of that team is
- 54:16thinking in a homogeneous way, they miss
- 54:17all of the crucial things.
- 54:19Same exact issue. If all of the code
- 54:21review agents are biased in the same
- 54:23way, they will miss exactly the same
- 54:24class of bugs, and then we're all sunk.
- 54:27Yeah, I don't know how you'd encourage
- 54:28this diversity in the ecosystem. We're
- 54:30probably, as you say, converging towards
- 54:31some kind of one model. Um but, I think
- 54:34for my picture, we need diversity. Yeah,
- 54:36we got to keep those open weights models
- 54:38going or something, cuz they're the
- 54:39weird players in the space.
- 54:41>> For sure, for sure. So, we've talked
- 54:43about uh efficiency, interpretability,
- 54:48tokenomics,
- 54:50uh uh uh
- 54:51fluency.
- 54:53Yeah,
- 54:54you're involved in a a lot of different
- 54:56research directions. Excellent,
- 54:57excellent. Where are What's next for
- 54:59you? Where do you see either Where do
- 55:01you see this all going kind of
- 55:02externally, but also like where is your
- 55:04research going?
- 55:05>> Yeah, this is great. Um
- 55:08and I'm I as I said before, we're trying
- 55:09to think in weird and creative ways
- 55:11about what the future could hold.
- 55:13And I encourage my students to do this.
- 55:15And they're smart, so they say, "All
- 55:16right, Chris, I'll think along those
- 55:17lines, but what's your answer to this
- 55:19question?" So, I do have an answer.
- 55:20[laughter] And it's really shooting for
- 55:22the moon here, which would be what about
- 55:24the architectural innovation that would
- 55:26upend the whole story around the stack
- 55:29transformer and the way we need to do
- 55:31data center build out to even get
- 55:33incremental gains in performance. That
- 55:35could be upended, and it would come from
- 55:37some very innovative thing around maybe
- 55:39recursive use of you the building blocks
- 55:42that we've got.
- 55:43So, architectures, we should think, and
- 55:45when people say, "Oh, no, we don't need
- 55:46more architectures. The transformer is
- 55:48good enough." That's where we should
- 55:50push back as academics doing something
- 55:52more clever and more scrappy that could
- 55:54change the world. And the other one is
- 55:57thinking in the inter space much more
- 55:59about data. And that's just because I
- 56:02want to tell the true story of how we go
- 56:04from data to model capabilities, but it
- 56:06also
- 56:07checks a box for me on connecting
- 56:10interpretability to safety. It has been
- 56:13hard for me to connect those two things.
- 56:15We have found some ways to do it, but
- 56:16it's not a slam dunk as a narrative,
- 56:18even though it's the dominant narrative.
- 56:20But, I will say that when we get into
- 56:22things like data poisoning from
- 56:24innocuous examples, this is probably a
- 56:26growing societal concern. There is
- 56:29evidence that with very few examples
- 56:31planted in a pre-training data set, you
- 56:34can have a significant influence on the
- 56:36outlook and preferences and quirks of
- 56:38the final model.
- 56:40So, can we detect those examples? What's
- 56:43the nature of those attacks? How well
- 56:45hidden could they be? What's the
- 56:46smallest number of examples? And why
- 56:48does it happen? These are all going to
- 56:50be very pressing questions. And so,
- 56:53again, it's just a data-oriented
- 56:54question that's very alive for me in the
- 56:56current moment.
- 56:58>> On the architecture front, are there
- 57:03is there research that you're seeing or
- 57:06doing that is, you know, as yet under
- 57:09the radar that you think is, you know,
- 57:12promising and or underappreciated?
- 57:14>> I think you had my student Julie Calleja
- 57:16on,
- 57:18and she is an advocate for byte-level
- 57:19models, essentially tokenizer-free
- 57:21models. I think that's a big part of the
- 57:23future. It's a critical thing if you
- 57:25want to have truly multilingual models
- 57:27that are also equitable in terms of how
- 57:28many tokens they charge us for, getting
- 57:30back to that earlier theme, but also
- 57:33Julie's perspective is that this is
- 57:34speculative, but I think there's
- 57:35something to this that the it's a kind
- 57:38of inference time scaling because you do
- 57:40more compute at test time
- 57:42um because you have more tokens and
- 57:44therefore more opportunities to
- 57:46build on interesting things.
- 57:49So, that could be a big part of the
- 57:50future and the other one would be
- 57:52recursive architectures as I said. But
- 57:55if you want to go all the way out, you
- 57:56could think, "Why do we always assume
- 57:58we're going to do gradient-based
- 57:59learning?" There are lots of
- 58:00alternatives to that and nobody is
- 58:03exploring them because everyone takes it
- 58:05as a truism. We're all in our very
- 58:07narrow row here without even really
- 58:08realizing it.
- 58:10Who knows what's outside in this garden?
- 58:13It's very risky as a research bet
- 58:15because
- 58:16only one in a thousand of these ideas
- 58:18will pay off.
- 58:20But what's the point of being an
- 58:21academic researcher if you're not going
- 58:23to take that kind of risks? That's what
- 58:25That's what we're positioned to do.
- 58:26>> Well, Chris, thanks so much for jumping
- 58:28on and sharing a bit about what you're
- 58:30working on. It's uh very cool stuff.
- 58:33>> Thank you. What a wonderful
- 58:34conversation. It gave me lots of new
- 58:35things to think about.
- 58:37>> Awesome. Awesome. [music] Thanks so
- 58:38much.
About this transcript
This page contains the full transcript of Do AI Tokenomics Matter More Than Model Benchmarks? by The TWIML AI Podcast with Sam Charrington, generated from the public captions YouTube serves with the video. The transcript has 10,758 words across 1,711 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.