From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? — Transcript
Full transcript
- 0:00AI's progress in mathematics has been
- 0:02remarkably fast. Not too long ago,
- 0:05Frontier Models struggled with grade
- 0:07school math. Now, they're contributing
- 0:09solutions to problems that have resisted
- 0:11mathematicians for decades, including
- 0:13most recently the Navier Stokes
- 0:16equations. But as AI moves from tests we
- 0:19already know the answers to to more
- 0:20open-ended research challenges, it gets
- 0:23harder to understand what these systems
- 0:24are truly capable of and what their
- 0:26progress tells us about where AI is
- 0:28headed. My guest, Greg Burnham, leads
- 0:31capabilities research at Epoch AI, where
- 0:33he and his team are developing new ways
- 0:35to track AI's rapidly evolving strengths
- 0:38and limitations. Here's Greg on why
- 0:40measuring and understanding AI's
- 0:42capabilities matters so much. Right now,
- 0:45AI capabilities
- 0:47large have gotten to such a point where
- 0:50what they're good at and bad at is
- 0:52starting to have real impact on the
- 0:54world, like on things we care about for
- 0:56other reasons. We're past the academic
- 0:59exam phase of understanding AI
- 1:02capabilities and for that we need high
- 1:04quality benchmarking, high quality
- 1:05evaluations. And one of my big questions
- 1:08is um just can AI come up with new
- 1:10ideas? If AI systems gets sort of
- 1:12superhuman at this that's a big deal for
- 1:15for the world both in terms of economic
- 1:18and scientific progress but also in
- 1:20terms of becomes harder to predict what
- 1:23AI systems are capable of. We use math
- 1:26as a test bed for this. I'm Sam
- 1:28Cherington and this is the Twimmel AI
- 1:30podcast. For over a decade, I've been
- 1:33exploring the ideas and innovation
- 1:35shaping the future of AI through
- 1:36conversations like this one that help
- 1:38you understand what's real, what's next,
- 1:41and what matters. Let's jump in.
- 1:52Let's talk a little bit about AI
- 1:53capabilities research broadly, how you
- 1:56pursue that research, and why you think
- 1:58it's important.
- 1:58>> We're getting to a point in the in the
- 2:01trajectory of AI overall where
- 2:05AI capabilities can have a real impact
- 2:08on the world. So
- 2:11maybe the transition could be
- 2:13characterized as school to work. Uh in
- 2:18maybe even up through Yeah. up through
- 2:202025, most ways that we tested what AI
- 2:25could do looked more like human exams
- 2:28like that you might give to you know to
- 2:30kids or to maybe advanced graduate
- 2:33students. I was going to say,
- 2:36>> right? Right. The the kid exams it is
- 2:38>> the bar exams, the medical exams
- 2:42>> and these p like I was perfectly good at
- 2:44these already end of 2024, beginning of
- 2:462025. So we really saw saw this
- 2:49transition happening over time to to be
- 2:52clear, but this is this is certainly the
- 2:55year where AI capabilities are impacting
- 3:01work activities that humans were engaged
- 3:04in. anyway uh for you know for their own
- 3:07purposes and so as these capabilities
- 3:11grow it's just important for us to keep
- 3:14tabs on them uh they grow very rapidly
- 3:17in some qualitative sense we see no no
- 3:20measurement we have shows any slowdown
- 3:23in what uh in how AI is getting better
- 3:26and better at any tasks that we're able
- 3:28to measure and so we think that just
- 3:32understanding the impact AI will have on
- 3:34the world. It's important to know what
- 3:36it can do, what it can't do, and keep a
- 3:38and keep tabs on how that how that
- 3:40trajectory is going.
- 3:41>> I want to push back on one thing you
- 3:43said in terms of we're not seeing any
- 3:47slowdowns
- 3:49that can be looked at broadly like in
- 3:52terms of maybe the number of, you know,
- 3:54benchmark data points that we're looking
- 3:56at. You know certainly we're seeing this
- 3:58broad performance but is it also true
- 4:02that like on a particular
- 4:05I feel like on a particular benchmark
- 4:07we're getting like the you know for for
- 4:09expected reasons when we went from you
- 4:12know GPG2 to GPG3 like we're taking
- 4:15these huge steps up these benchmarks and
- 4:18now like there's just a lot less room
- 4:20left and we're making more incremental
- 4:22progress. Do you see that as a a
- 4:25slowdown in AI performance or is would
- 4:27you characterize it differently?
- 4:29>> I would say well I'd say there's two
- 4:31issues here. One is on the benchmarks we
- 4:33have what does performance look like?
- 4:36Two is
- 4:38are there things our benchmarks don't
- 4:40measure that maybe there has been less
- 4:43or could even be negative uh improvement
- 4:47on and is there some reason why it's the
- 4:51things that are hard to measure where
- 4:53there's been the least progress. But on
- 4:55the first point because this is a very
- 4:56this is a much more cut and dried
- 4:58statistical point. We absolutely see
- 5:01continued progress on benchmarks like a
- 5:05benchmark that was challenging for GPT3
- 5:08uh you know many years whatever six
- 5:10years ago uh is now completely you know
- 5:14solved uh aced by GPT4 benchmarks
- 5:19challenging for GPT4 same for GPT5 now
- 5:21we've got GPT6 like so we do some
- 5:24statistical work to try to aggregate
- 5:26benchmarks over time Because the
- 5:29benchmarks that people used to track
- 5:32progress years ago, now every AI model
- 5:36off the shelf that that you've heard of
- 5:37that you might use is too good at them.
- 5:40Uh so people make new benchmarks and
- 5:41then you have to sort of do this like
- 5:43stitching of like okay well a model that
- 5:45was pretty darn good at this benchmark
- 5:48was not so good at this next one and you
- 5:50can use like a you can stitch those
- 5:52together and get this sort of composite
- 5:54score. We call it we we have a
- 5:56methodology for doing this. to call the
- 5:57epoch capabilities index. And that's a
- 6:01unified way of tracking benchmarks over
- 6:03time. And we really do see this thing
- 6:05continuing to go up uh you know pretty
- 6:08smoothly linearly and and that march of
- 6:12progress is uh is is steady. That's an
- 6:14important point that that's truly if um
- 6:18we we have we've made more benchmarks
- 6:21and AI systems that are harder that
- 6:23models start off at zero and then within
- 6:26a year or two years they're at 100% and
- 6:29we have to make a new benchmark to keep
- 6:30up. Uh and that's that's really been the
- 6:33story of the last uh certainly the last
- 6:35three years I I'll say with confidence
- 6:37there. It seems like that like
- 6:40normalizing across different benchmarks
- 6:43to
- 6:44produce a result that is coherent that
- 6:49sounds like a very challenging problem.
- 6:51like you you you describe the result as
- 6:55like you know kind of continued linear
- 6:57improvement and I imagine if that's like
- 7:01a constraint you can kind of map
- 7:04backwards to you know uh to factors that
- 7:08will show you that continuous linear
- 7:10improvement but you know starting from
- 7:13you know a set of disperate benchmarks
- 7:15that have their own kind of you know
- 7:18ranges and challenges and you know where
- 7:20AI you and where AI sits in them and and
- 7:23kind of mapping this across different
- 7:25generations of AI and coming up with a,
- 7:29you know, something that makes sense,
- 7:31you know, objectively at the end and
- 7:33then having that thing show continued
- 7:35linear progress. That sounds like um
- 7:42you almost like too good to be true and
- 7:44like suspicious that it is true.
- 7:46>> I I I yeah, I I agree. So, so let let me
- 7:49go into that a little because I think
- 7:50it's a it's a killer point to understand
- 7:53and it is surprising. So, I think you're
- 7:55very right to be surprised. In fact, if
- 7:58you look at AI benchmarks, they are all
- 8:02correlated with each other even if the
- 8:04benchmark the scores across models are
- 8:07correlated with each other even if the
- 8:09benchmarks are in nominally different
- 8:11domains. So maybe you don't expect you
- 8:15don't find this so surprising if it's
- 8:17like I've got one benchmark that's you
- 8:19know graduate chemistry questions and
- 8:22I've got another one that is coding
- 8:24puzzles and the models that are good on
- 8:27chemistry are also good on software
- 8:29engineering. Maybe that's not
- 8:30>> correlated largely by the the models
- 8:33themselves. Like as the models get
- 8:35better they're going to get better
- 8:36across these you know disparate
- 8:38benchmarks. But but this uh exactly but
- 8:40that's that's the surprising point. So
- 8:42all that you need to be true to see this
- 8:45linear trend over time is basically
- 8:47every time you know opus 4.5 opus 4.6
- 8:514.7 GPT 5.2 5.4 every time these come
- 8:55out if they get better on all benchmarks
- 8:58at once then you're going to see this
- 9:00linear tren this this big overtime
- 9:02linear trend like that's it's two sides
- 9:04of the same coin statistically. Why does
- 9:07this happen? Like this is a huge
- 9:09question. Like you're you're right to
- 9:10find that suspicious. And I would say
- 9:13there are two explanations that that we
- 9:15see. And it's an important question
- 9:18which of these explanations is more
- 9:20true. And of course the truth is some
- 9:22mix. But but let me let me give the the
- 9:25two explanations. One is the AI
- 9:29companies are making darn sure that
- 9:32their model gets better at all of these
- 9:35benchmarks.
- 9:36with each new generation of model. And
- 9:39the main mechanism you might imagine
- 9:40them doing that if they're like looking
- 9:42around, okay, people care about
- 9:44chemistry, people care about software,
- 9:46people care about operating, you know,
- 9:48answering my email, like they they would
- 9:51collect training data one form or
- 9:53another that would help the model learn
- 9:55how to do these how to do these tasks.
- 9:57And that's something of a
- 10:01um that's something of a you know very
- 10:03manual human in the loop process of just
- 10:05saying hey we're doing product
- 10:06development like they want this feature
- 10:08they want that feature our users want
- 10:09all these features and the way you do
- 10:11that in machine learning is you collect
- 10:12training data and throw it in I'm
- 10:14oversimplifying but that that's that um
- 10:17sort of like a shallow way to make it
- 10:19happen the so I'll call that shallow the
- 10:21other one is deep the deep way to make
- 10:23it happen is if the model generally is
- 10:26if the AI system generalize very very
- 10:29well. So yeah, you never trained it on
- 10:32chemistry, but you trained it a lot on
- 10:34math and software engineering and it
- 10:37just got super smart and with a modocum
- 10:40of chemistry basics, it was able to
- 10:43derive the rest from first principles
- 10:45sort of thing. This is like a deep form
- 10:47of generalization. Both of these
- 10:50happened to some extent like like kind
- 10:51of the revolution with GPT even GPT3
- 10:55around that era. This is like quite a
- 10:56while ago was oh if you train it to like
- 10:59predict the next word then it actually
- 11:01gets better at a wide wide range of
- 11:03languageoriented tasks that people had
- 11:06been trying to handle individually.
- 11:08Suddenly it just got got good at
- 11:10everything. And that wasn't because
- 11:11anyone was like oh let's have data
- 11:14specifically around uh sentiment
- 11:17analysis or or syntactic structural
- 11:20analysis. Like it wasn't anything like
- 11:22that. It was just give it give it this
- 11:23very general thing. But now we're on to
- 11:25this world where you have to like train
- 11:27it specifically. Right now there's a
- 11:29huge industry of collecting all this
- 11:31rich training data and it's really
- 11:32unclear how much generalization you're
- 11:35getting from you know in this deep way.
- 11:38Uh and and if and anyway that's that's
- 11:41sort of a big question because that deep
- 11:43generalization means you could really
- 11:45see if that happened in a in a rich way
- 11:47where you just didn't have to even be
- 11:49trained on something at all to gain
- 11:50capability in that area. then you might
- 11:53have a or like radically superhuman AI
- 11:56that would be um pretty hard to uh to
- 11:58understand like to to bound its
- 12:00capabilities.
- 12:02It's still a little counterintuitive
- 12:03that this would be that this would be
- 12:06linear and maybe you know maybe part of
- 12:09the question is like are we talking
- 12:11about
- 12:14you know linear across generations of uh
- 12:18of of model releases or linear within
- 12:21some regime like segment you know
- 12:23segmentally linear or
- 12:25>> the the thing I can say is we have like
- 12:28a statistical method so a benchmark
- 12:30score Like what are we talking about
- 12:32here? Our unit here is a score from 0 to
- 12:34100%. And typically the way these scores
- 12:38go just historically, empirically is
- 12:41like they follow like an S-curve. Like
- 12:42like they start out sort of bad, then
- 12:44they get better, they get better pretty
- 12:46rapidly, and then you can only get up to
- 12:48100%. So then they sort of level off.
- 12:50Takes them a while to get that last, you
- 12:52know, little chunk and then they're at
- 12:54100%. So you've you've got these like
- 12:56sigmoids, uh is is what that scurve is
- 12:59called. We've just got a statistical
- 13:00methodology for stitching sigmoids
- 13:03together and what comes out is a linear
- 13:06function. And this looks like like this
- 13:08is an artifact of the statistics to be
- 13:10clear. But this sure enough looks very
- 13:12regular in its own statistical analysis
- 13:15terms over across generations of of
- 13:19models. Like absolutely like it's been
- 13:21uh if anything it accelerated a little
- 13:23with the development of so-called um
- 13:26reasoning models around late 2024 where
- 13:28the models started to get good at math
- 13:30and and software engineering and things
- 13:33that require lots of logical reasoning.
- 13:35Uh maybe like a higher linear slope but
- 13:38uh but this is just like a that's a
- 13:39statistical artifact. You now have to
- 13:41ask, okay, well, if your score is like
- 13:44one model score is 150 and the next
- 13:46model score is 152, like what do those
- 13:49two points mean on your benchmark
- 13:51aggregate index?
- 13:53Now,
- 13:54>> I think part of part of what I'm what
- 13:56I'm asking myself as I hear this
- 13:59described as linear is does it say that
- 14:02in spite of the model the Frontier Lab
- 14:06vendor claims that you know Fable is
- 14:10like dramatically better than Opus or
- 14:13that GPT6 is like the step function uh
- 14:16you know improvement over GPT5.6.
- 14:19you know, linear implies to me that that
- 14:22it's kind of incremental and there's
- 14:24really been no step functions.
- 14:26>> This is um this is a really great
- 14:28question and yeah, we believe this is
- 14:30one of the big uses of this statsy tool
- 14:34we've got is you can try to find you
- 14:37know trend breaks and you can try to
- 14:40find accelerations. Uh and we mostly
- 14:43don't see that. Yeah. Like one way to
- 14:45put it that is everything just about is
- 14:48on trend. Fable uh Astra whatever is on
- 14:53trend but the trend is crazy. Like you
- 14:56gota you got to hold both of these
- 14:58together.
- 15:01>> Right. Right. Right. Right. Right. This
- 15:03is like this is sort of the zoomed out
- 15:04view epoch takes with with everything is
- 15:08a lot of this we think the engine of AI
- 15:10progress is heavily mediated by raw
- 15:14inputs like GPUs the training compute
- 15:17like training data maybe those are the
- 15:20big ones and then there's this like
- 15:21factor of the labs come up with
- 15:22algorithmic like innovations like
- 15:24research breakthroughs but we really
- 15:26think like a lot of it is coming from
- 15:28the scale up in the in the compute and
- 15:31the scale up in the training data So
- 15:33until you see the data and the GPU the
- 15:35data and compute supply chain like ramp
- 15:39uh significantly or discontinuously uh
- 15:43to be more precise you're not going to
- 15:44see you're not likely to see
- 15:47discontinuities on the model side.
- 15:49>> I think that's I think that's right. A
- 15:51big thing we're keeping our eye on to be
- 15:53clear is
- 15:55are there the are AI abilities emerging
- 15:58that could cause a discontinuity on the
- 16:00model side. So this is why people talk
- 16:03about recursive self-improvement or
- 16:05sometimes they call this a softwareonly
- 16:08intelligence explosion. And the point of
- 16:11these concepts is well sure right now
- 16:15there's a slow comparatively slow
- 16:17process for building more micro building
- 16:20more GPUs and collecting more human
- 16:22training data and whatnot. But if an AI
- 16:25system got really good on the
- 16:27algorithmic innovation side, could it do
- 16:30more with less or or do more with the
- 16:32with a fixed uh compute budget? And that
- 16:37could change the dynamics here. So the
- 16:38feedback loop would might no longer run
- 16:40through the physical world. We don't
- 16:43really it's not clear we see evidence of
- 16:44this yet. I'd say on balance we don't
- 16:47really see evidence of this yet but it's
- 16:49the sort of thing that has spooked the
- 16:52people inside the labs the the AI
- 16:54companies I mean saying like gosh this
- 16:57could be right around the corner that's
- 16:58like a bit of the fear and if that
- 16:59happens then the pace we're used to we
- 17:03might get that discontinuity
- 17:05>> have you developed a particular
- 17:06benchmark to identify and track this or
- 17:10is it more
- 17:13the way you look at the benchmarks that
- 17:15we've been talking about this
- 17:16statistical model that we've been
- 17:18talking about for an indication that you
- 17:21know there's some kind of you know
- 17:23exponential uh in the curve and and
- 17:26that's you know probably one of the
- 17:28likely causes you know for that kind of
- 17:30exponential or discontinuity
- 17:32>> all of the above. Uh we we want as many
- 17:35as many tools in our arsenal as as we
- 17:37can for this. So that to this statsy
- 17:40like acceleration detector is definitely
- 17:44a good uh is definitely a good tool. But
- 17:46we also sort of get you know down in the
- 17:49weeds with some of our specific
- 17:51evaluations and benchmarks saying I'll
- 17:53give you an example here. This is some
- 17:55forthcoming work we have uh that that
- 17:58I'm happy to to preview. We take a
- 18:01recent AI research breakthrough or
- 18:05innovation
- 18:07from humans and we take AI systems. We
- 18:12don't connect them to the internet and
- 18:13we choose AI systems that uh were
- 18:16developed and released just before this
- 18:18recent h whatever the recent human
- 18:20innovation is. And we basically say look
- 18:23here's the metric that that innovation
- 18:25the number that it makes go up. It
- 18:27improves this sort of efficiency. what
- 18:29whatever some some metric and we say AI
- 18:31system your goal is to improve this
- 18:33metric as much as you can but we don't
- 18:35tell it the innovation and so we see is
- 18:38it capable of replicating that human
- 18:41innovation that was highly relevant for
- 18:44AI you know research and development and
- 18:48so far we don't see them able to do this
- 18:50there's this nebulous idea called
- 18:52research taste which is like where's the
- 18:55good idea like what what experiment
- 18:57should I do next where Should I go
- 18:59hunting for uh big improvements more
- 19:02than the mundane improvements of just
- 19:04more compute, more data? Like where
- 19:06should I look for for a breakthrough?
- 19:08And AI systems don't seem great at doing
- 19:10this in um at least in an AI R&D
- 19:15context. Uh but but they're like getting
- 19:17a lot better at a lot of things that are
- 19:19adjacent to this uh to this area. So we
- 19:22really do think this is an important
- 19:24sometimes I almost call it a trip wire
- 19:26that we're you know the the future is
- 19:28foggy is some foggy landscape. We want
- 19:30to put a trip wire out there so that if
- 19:32AI ever does cross some threshold where
- 19:34it's able to do rapid improvements in AI
- 19:38algorithms themselves then uh we we want
- 19:41that trip wire to sound. we go, okay,
- 19:43gosh, watch out. Like maybe maybe
- 19:45something uh maybe we're about to see a
- 19:47big improvement big um you know, uptick
- 19:50in AI capabilities.
- 19:51>> And how does this relate to
- 19:54AI's performance on some of the math
- 19:57challenges? You know, certainly Navia
- 19:59Stokes has been uh in the news quite a
- 20:03bit recently. Prior to that, there were
- 20:04some Erdos results uh and others. like
- 20:08how do you think about the role that
- 20:10those math challenges play in this
- 20:13broader
- 20:14view of trying to understand AI's
- 20:17trajectory?
- 20:18>> Yeah, very good question. Uh I if I may,
- 20:21I would
- 20:23go rewind history just a just a bit here
- 20:26to um say some of the say the trajectory
- 20:29we've been on because once you you step
- 20:31back it's um it's it's wild. Uh so up
- 20:36until
- 20:40yeah up until
- 20:42fall of 2024 so we're talking two years
- 20:45ago uh grade school math was still a
- 20:49challenge for these LLM based AI systems
- 20:52the the kind of AI we're talking about
- 20:54here and all and the benchmark there was
- 20:57literally a benchmark of 8,000
- 20:59you know word problems so and so you
- 21:02know has so many apples that sell for
- 21:05this
- 21:06>> GSM8K that's the one
- 21:08>> and
- 21:10GPT4 which is 2023 had like done pretty
- 21:13well on this so people started looking
- 21:14at some like adding into the mix a
- 21:17benchmark just called math like four
- 21:19capital letters that um started pulling
- 21:22from some high school math competitions
- 21:25and there was like slow progress on this
- 21:26but but not a lot and then at the end
- 21:28fall of 2024 OpenAI comes out with the
- 21:3101 preview model full 01 by the end of
- 21:352024. And this showed a huge spike in uh
- 21:38in math competition capabilities. But
- 21:40we're still just talking like medium
- 21:42hard high school math competitions like
- 21:44something, you know, thousands of math
- 21:46nerd nerdy math 16-year-olds were were
- 21:49into across the across the US and and
- 21:54this became, you know, the new the new
- 21:56territory. My my company, Epoch, around
- 21:59that time released a benchmark called
- 22:01Frontier Math. Uh at the time that's all
- 22:04we called it. We now call it tiers one
- 22:05through four uh to distinguish from
- 22:08later versions I'll get to in a bit. Uh
- 22:11but the original frontier math consisted
- 22:13of problems that I would say go from
- 22:15advanced undergraduate to advanced grad
- 22:19student. So sort of early career
- 22:21research warm-ups almost. These are
- 22:24problems that are no longer thousands of
- 22:27uh thousands of high schoolers would
- 22:30tackle them. These require deep
- 22:32background knowledge in in various areas
- 22:36that that you know niche areas like
- 22:38we're talking a a problem from you know
- 22:42elliptic curves that you'd give to a
- 22:45grad student like shortly before they're
- 22:47going to embark on their novel thesis
- 22:49research just to get them familiar with
- 22:50the details of the of the niche that
- 22:52they're going to be trying to work in
- 22:54and make an original contribution to.
- 22:56Over the same period, we also saw high
- 22:58like so over 2025, we also saw high
- 23:01school math contests get completely
- 23:03aced. We got the gold medal on the
- 23:05International Math Olympiad. By the um
- 23:09even by the beginning of 2026, we saw
- 23:12most of these frontier math problems
- 23:14solved, though not all of them had had
- 23:16in fact been solved by an AI system at
- 23:19that point. And so just like even just
- 23:22here, this is a very rapid trajectory
- 23:25from o over the course of just a little
- 23:28over a year going from, you know, like
- 23:32hobbyist high schooler to graduate
- 23:36student like very competent kind of, you
- 23:39know, up there graduate student in terms
- 23:41of the problems they could solve. Really
- 23:43rapid AI progress.
- 23:46At the same time, we started to see
- 23:48people test AI systems on unsolved math
- 23:52problems. Uh problem like the the
- 23:54frontier was no longer something. This
- 23:57is what I mean. We'd gone from exam to
- 23:59on the job. Uh where it wasn't like
- 24:03here's your qualifying exam as a grad
- 24:05student anymore. It became more
- 24:06interesting to say, well, here's a small
- 24:08problem that maybe a mathematician, a
- 24:10professional thought about for an hour,
- 24:12didn't solve AI systems. Can you do it?
- 24:15And at that level we started this was
- 24:19all just this year like 2026 at the
- 24:21beginning of the year there were like
- 24:22starting to be claims of oh maybe here's
- 24:24a problem that was posed by a famous
- 24:26mathematician this is maybe the airdos
- 24:29problems there there's this just for
- 24:30background there's this very prolific
- 24:32mathematician Paul Erdos uh who proposed
- 24:35over the course of his life maybe like
- 24:38well over a thousand maybe a couple
- 24:39thousand problems that people still like
- 24:41haven't fully cataloged all the
- 24:43questions he asked some of those became
- 24:44central to to whole fields of math. Some
- 24:47of them were, you know, no one really
- 24:49paid much attention to, including Erdos
- 24:51himself. And AI started to notch a few
- 24:53solutions to some of the easier ones of
- 24:57those. Fast forward to May uh of this
- 25:02year and and an internal model probably
- 25:05something like the Astra model that we
- 25:07now have publicly available was able to
- 25:10solve a big one of those. This was sort
- 25:12of a historic moment. may the the
- 25:14so-called unit distance problem uh from
- 25:18uh which was a problem about how many
- 25:21points you can fit in the plane that are
- 25:23just one dis unit distance of one apart
- 25:25from each other sort of like very cool
- 25:27that you could frame it it's such a
- 25:28simple problem but there hadn't been
- 25:30progress much progress on this problem
- 25:32despite a lot of attempts uh since had
- 25:35had posed it you know decades and
- 25:37decades before and it was a serious
- 25:38problem it wasn't one of these no one
- 25:40had thought about it it was definitely
- 25:41had received a lot of attention and an
- 25:43AI system solved it. And it was really
- 25:45this this was like a aha moment of oh my
- 25:48goodness like AI can solve problems that
- 25:51humans uh have cared about just
- 25:53independently
- 25:55and we've seen a lot more of that since
- 25:57and the the you know with a sort of the
- 25:59the recent wow moment being the um
- 26:02solution of the millennium prize
- 26:04problem. This was you know the Navier
- 26:06Stokes problem. This was a collection of
- 26:08this these Millennium Prize problems, a
- 26:10collection of seven problems that um
- 26:13mathematicians sort of or at least one
- 26:15math institute set as like good goals
- 26:17for the new century. Like they were they
- 26:19were they're much older than the year
- 26:222000, but but they were sort of
- 26:24canonized in the year 2000 as hey like
- 26:26these are worthy targets of entire sub
- 26:28fields of mathematics and it would be a
- 26:30big deal and so one of them was was
- 26:32solved. So it's been a hell of a hell of
- 26:34a ride. The historical context is is
- 26:37super interesting. It, you know, we're
- 26:39all living through this trajectory, but
- 26:42to step back and see, you know, how much
- 26:44has happened in the past couple of years
- 26:46is pretty pretty nuts.
- 26:49>> Yeah. You you have to, as one of my
- 26:51colleagues said, you have to like not
- 26:52get frog boiled by it like like the the
- 26:55the rising and it's easy to
- 26:58>> I don't know. Anyway, um so so there is
- 27:01this interesting gap so far. Like you
- 27:04might hear the phrase jagged frontier
- 27:07where the AI capabilities are jagged in
- 27:10the sense that they'll be superhuman in
- 27:12some on some axis and some you know not
- 27:16superhuman still inferior in capability
- 27:18to human in in other a on some other
- 27:21axis and
- 27:24that uh remains the case so far in math
- 27:28in particular there's a couple things in
- 27:30hindsight the analysis by mathematicians
- 27:33of the solutions AI has found to these
- 27:36like big important previously unsolved
- 27:39math problems has it's always in the
- 27:43first place it's been nothing too super
- 27:46duper surprising maybe there was like a
- 27:49a direction that humans had overlooked
- 27:53or could have invested more in and AI
- 27:57sort of had the persistence may maybe
- 27:59this is a good a good word or or or the
- 28:02bravery or the tearity to like wade into
- 28:06these very in-depth intricate
- 28:08computations and calculations and sure
- 28:10enough came out the other side with with
- 28:12a solution and humans look back and are
- 28:14like ah that's not a crazy direction to
- 28:17try at all and maybe if we'd like gone
- 28:21in that direction we we we would have
- 28:22come up with this this isn't something
- 28:24that like would be surprising
- 28:26>> what I'm curious about is
- 28:31are we seeing the AI AI systems come up
- 28:33with a solution to the problem or an
- 28:37algorithm to create the solution to the
- 28:41problem. And and I think I'm trying to
- 28:43like ask a lot of different things here,
- 28:46you know. One is like creating
- 28:48creativity, one is maybe um levels of
- 28:52abstraction or the way it's like, you
- 28:54know, thinking abstracting, you know,
- 28:56around these problems. Um yeah, I'll
- 29:00kind of pause and let you react to that.
- 29:02Yeah, let let me let me throw out a
- 29:04couple things there. Let me know if you
- 29:06want me to expand on any of them. When
- 29:09humans solve these math problems, they
- 29:11they write up the solutions in p in like
- 29:14papers, proofs uh that you know this is
- 29:16true or this is not true. And those
- 29:18proofs are are not usually
- 29:21like not exactly algorithms. They're
- 29:22written in natural language. They have
- 29:24sort of they leave the reader to fill in
- 29:27some of the gaps. and uh AI systems are
- 29:30coming up with just the same kind of
- 29:32thing. Uh like like it's it's the same
- 29:34products. There's a lot of complaints
- 29:36about their writing style, but like the
- 29:38arguments, the substance of the
- 29:39arguments really are the thing we expect
- 29:42to see from humans. such that if a human
- 29:45came up with this and yeah, maybe
- 29:46cleaned it up a little, I don't know, uh
- 29:49would like would be lauded for having
- 29:51like made a very insightful connection
- 29:54between two different areas or having
- 29:57gotten a an intricate um argument just
- 30:01right and balanced different
- 30:03considerations to make everything come
- 30:04together to to land the proof. 100% like
- 30:08this is not a this is something we're
- 30:10past the point where we can say, "Well,
- 30:13AI systems are sort of like not really
- 30:16doing something that humans would, you
- 30:18know, be impressed by on human terms.
- 30:20No, AI systems are doing doing things.
- 30:22And is there a way to categorize or
- 30:26qualitatively
- 30:28are we finding that and I think you
- 30:30spoke to this a little bit like are you
- 30:32know is there a qualitative difference
- 30:34between you know the best human approach
- 30:37and the AI approach in a sense of like
- 30:40you know AI brute forced it you kind of
- 30:43spoke to this versus like pulled in
- 30:46things from some other domain and like
- 30:48created some elegant solution or are we
- 30:50past that point as well
- 30:51>> here I' I'd give more nuance. Uh we we
- 30:55um there's an emerging sense of at least
- 30:59for the moment what an AI shaped problem
- 31:03looks like. Maybe I could say AI has two
- 31:06big advantages right now. One is it
- 31:08knows everything like it knows all the
- 31:11prior work. It's got the literature
- 31:12memorized. I'm speaking loosely but that
- 31:15this is this is basically uh valid. Um
- 31:19the other is it's extremely patient and
- 31:22persistent. So if there is some
- 31:24intricate uh you know calculation you
- 31:28have to do I don't mean like numerically
- 31:29or algorithmically. When I say
- 31:31calculation I just mean lots of
- 31:33equations and you have to make all the
- 31:35pieces work out just as very high level
- 31:38conceptual logical kind of computation
- 31:41but still something very um intricate.
- 31:44Uh they just have the patience to to go
- 31:47for that. And if you take the Navier
- 31:49Stokes example, I think it's not a
- 31:51coincidence that um OpenAI's reported a
- 31:56system that that solved it included a
- 31:58swarm of thousands of different
- 32:00instances of AI systems trying lots of
- 32:03different things and sharing ideas where
- 32:06clearly it was able to like there's an
- 32:08element of brute force there. So there's
- 32:11degrees of brute force. This is where
- 32:12the nuance is. It wasn't like somehow uh
- 32:16you know you you asked it like it wasn't
- 32:20solving this in by any means in the
- 32:22dumbest most low-level way like we're
- 32:24we're past that point as well. They used
- 32:26to solve some problems in like kind of
- 32:28surprisingly low-level like grinded out
- 32:31kind of ways. Like that's not really
- 32:32what these look like anymore, but like
- 32:35the the highle ver a higher level
- 32:38version of that still is maybe what what
- 32:40you might what you might say here. So,
- 32:42and I like the frame of precise and
- 32:45intricate computations here. It's
- 32:47important I'll say one more thing about
- 32:49the Navier stroke solution. It's
- 32:51important to recognize that humans sort
- 32:53of put in place a lot of the structure
- 32:55through previous work that AI then went
- 32:58and as far as we can tell, I mean,
- 33:00humans are still figuring out exactly
- 33:01what goes into the AI solution, but as
- 33:04far as we can tell, it builds on a lot
- 33:06of prior human work. The AI didn't like
- 33:08build it up from the ground. that kind
- 33:11of theory building as mathematicians
- 33:13call it we still don't see that very
- 33:15much from AI systems last point I'll
- 33:18make before pausing on this is just uh
- 33:21remember its point in time like remember
- 33:23that trajectory we've seen very rapid uh
- 33:26so so mathematicians who are saying we
- 33:28have to adapt our profession to the fact
- 33:30that now you can push a button and get
- 33:32what used to have been a career making
- 33:34result like you know are also I think
- 33:37cautioning and we can expect that
- 33:39capabilities will freeze Right now we we
- 33:41see you know that that linear whether
- 33:43it's linear or exponential is a matter
- 33:45of perspective but it's going up and
- 33:48it's like we see no reason to expect it
- 33:50to stop there. So this is a big thing
- 33:52epoch is looking for. Can AI move from
- 33:56intricate computations plus lots of
- 33:58background knowledge to more of a
- 34:01developing uh new ideas more from whole
- 34:05cloth. Uh and and if so and that's a big
- 34:08thing we're watching for in our in our
- 34:09own math benchmarking uh activities.
- 34:12Yeah, it's interesting to think that
- 34:17at least the way I characterized this as
- 34:19brute force versus pulling in things
- 34:21from other domains,
- 34:24those are um you know those are not
- 34:28mutually exclusive. like for for an AI
- 34:31that has access to, you know, all of the
- 34:33research from every domain, you know, an
- 34:36approach that could be very fruitful is
- 34:38to brute force the exploration of
- 34:41adjacent domains and pull that into the
- 34:43solution to a given problem.
- 34:45>> That's it. I think a lot of what we see
- 34:48and maybe just to say the the other
- 34:50thing that would really supercharge them
- 34:53and bring them closer to full human
- 34:56parody or even human you know dominating
- 34:58human capabilities would be something
- 35:00like coming up with a new idea. Uh and
- 35:03maybe like in calculus the idea of the
- 35:07derivative or the integral of a function
- 35:09like that was something that was you
- 35:11know maybe mathematicians in the 1500s
- 35:14had been like grasping for this idea and
- 35:16then Newton or Libnets or whomever comes
- 35:18up with like hey here's like I can I can
- 35:22describe a new thing a new mathematical
- 35:24concept and that unlocks a lot for us
- 35:26that lets us solve problems we
- 35:29previously couldn't solve. And so that's
- 35:31the sort of thing we don't yet see AI
- 35:34doing. It's also the sort of thing
- 35:35that's rarer for humans to do. Like a
- 35:37lot of human math progress has been
- 35:49proven it or I studied one area and I
- 35:52noticed something and I thought it might
- 35:54be applicable to this nominally
- 35:55superficially different area like humans
- 35:57do that sort of stuff all the time
- 35:58that's 90%
- 36:00vaguely speaking of human mathematical
- 36:02work from from what I hear from
- 36:04mathematicians the new idea stuff is
- 36:06rarer so maybe we should expect it to be
- 36:08not the first thing AI gets good at but
- 36:11so far we don't see AI really you know
- 36:13AI really doing that yet
- 36:16>> how easy or difficult is it to define
- 36:21a new idea in in such a way that um you
- 36:26know we can easily distinguish them like
- 36:29you know if we think about kind of the
- 36:30evolution of this discourse around like
- 36:33is it AI creative you know initially it
- 36:35was like yeah it's very creative it
- 36:37created this like poem in you know in
- 36:40pirate like that's creativity that's you
- 36:42know that was a new idea uh but that's
- 36:44certainly not what we're talking about
- 36:46when we're talking about like advanc
- 36:48advancing science and advancing
- 36:50mathematics etc. Um
- 36:54you so h a how do we like how do we
- 36:58define that but also you know from your
- 37:02perspective as someone who's studying
- 37:03this is it like you know we've got this
- 37:06long line of of folks that are coming
- 37:08that says here's this AI with a new idea
- 37:10and you're like yeah well let me
- 37:11evaluate it against my new idea criteria
- 37:13no it's not really that or are we
- 37:17actually like is it pretty clear to
- 37:19everyone that the AI is not having new
- 37:21ideas and like we're waiting for the
- 37:23first person to come and say that AI has
- 37:25the new idea. Like how do you think
- 37:27about that whole space?
- 37:29>> Very very challenging nut to crack. Uh
- 37:32so we you know try to approach it from
- 37:35different angles.
- 37:37One thing you can try to say is
- 37:42I've got a problem. I don't know if a
- 37:44new idea will be required to crack it.
- 37:46And this could be a problem in math or
- 37:48in science or or any quantitative area.
- 37:52But humans have tried to crack it.
- 37:54Humans have tried to make you know this
- 37:56number go up. Whether that number is the
- 37:58efficacy of some you know drug at
- 38:00fighting some disease or something like
- 38:02that or some uh you know or it's a math
- 38:05result uh of some conjecture that the
- 38:08math have stumped mathematicians for a
- 38:10long time. And you can say, "Okay, well,
- 38:12I don't know if I I'm not sure how I'm
- 38:14going to measure a new idea, but I can
- 38:15at least tell if this problem has been
- 38:17solved or this number has been made to
- 38:20go up or whatever the the concrete
- 38:23metric is." And you can then you're you
- 38:26can at least check if that has happened.
- 38:28Now
- 38:30I was probably a year ago more
- 38:31optimistic that there would be a tighter
- 38:34relationship between math problem that
- 38:37humans have failed to solve and post
- 38:40hawk qualitative assessment of
- 38:42creativity.
- 38:44Uh but uh but you know that this is so
- 38:47this is sort of the combination the like
- 38:49you know double punch we're we're trying
- 38:52here. First you set a target that's hard
- 38:54on human terms. The humans care about
- 38:56the humans have tried to be creative to
- 38:58solve and then you do post talk review
- 39:01with experts and say what do you make of
- 39:03this solution? Why hadn't you found it?
- 39:05No offense uh you know what what is in
- 39:08the AI solution that humans had missed
- 39:10and so on. And this is risky like like
- 39:13risky in terms of humans being honest
- 39:15about it because you know your ego might
- 39:17be wrapped up in it or you might once
- 39:19you see the solution it feels more
- 39:21obvious in hindsight. So would you have
- 39:22described it like very very risky uh but
- 39:26uh but you still you try your best and
- 39:28you know you might say things like well
- 39:31uh how like here's an interesting
- 39:33approach that that we haven't
- 39:34operationalized in like in detail uh but
- 39:36it's it's a nice heristic I think. How
- 39:38big a hint would a human have needed to
- 39:42solve this in in hindsight? Uh so for
- 39:45example that unit distance a very uh
- 39:48prominent mathematician fields medalist
- 39:49Timothy Gowers uh did this sort of
- 39:52exercise on this unit distance problem
- 39:54once AI had solved it saying how big a
- 39:57hint and how like specific of a hint
- 40:00would I do I think a mathematician would
- 40:02have needed now in hindsight like to
- 40:05solve this problem. uh one one yeah and
- 40:08he came away with like an interesting
- 40:09exploration of well actually AI sort of
- 40:12disproved this conjecture as opposed to
- 40:13proving it true it said actually it's
- 40:15false even that is a huge hit humans had
- 40:17mostly thought it was true they were
- 40:19wrong but they'd been trying to prove
- 40:20that it was true instead of looking for
- 40:22a counter example so even just the hint
- 40:24of try to find a counter example is
- 40:27pretty big and then there was like one
- 40:28more piece and he was like I think
- 40:30though very hard to say we can't really
- 40:32do the counterfactual experiment I think
- 40:34with just that with like this sort of
- 40:36two-part hint. It would have been
- 40:38possible for like a human would have
- 40:39been like, "Oh, oh, okay, I know what to
- 40:41do now, blah, blah, blah." And like
- 40:42probably would have at least hastened a
- 40:44human solution. So, you can do things
- 40:46like that maybe still very qualitative,
- 40:48but you try to be honest with yourself
- 40:50and be rigorous and you could imagine
- 40:51doing a more rigorous version of this
- 40:53experiment.
- 40:54And anyway, things like that. But then
- 40:56the the the fact is math is still for
- 40:59the most part a um
- 41:02sort of discipline removed from physical
- 41:05real world impact. Uh not always but in
- 41:08surprise its impact is often hard to
- 41:10predict. There are domains where impact
- 41:12is not at all hard to like I'm in the
- 41:13chemistry lab and I'm trying to improve
- 41:15the yield of my synthesis process. And
- 41:17if AI, you know, makes incremental
- 41:20improvements, I expect incremental
- 41:21yield. And if AI suddenly comes and
- 41:23says, no, no, no, you're doing it all
- 41:25wrong. Let me like redesign your whole
- 41:26thing from scratch. I've got like a, you
- 41:28know, maybe creative idea. Well, at some
- 41:29point, you don't care if it's creative
- 41:31or not in the qualitative sense. You
- 41:32just care that you have saw a
- 41:34discontinuity in the yield of your
- 41:36chemical substance or whatever. And
- 41:38that's, you know, that's uh
- 41:41that's a bigger bigger impact for the
- 41:43world. Anyway,
- 41:44>> what I took from that is that we filter
- 41:46on the problem and how hard humans have
- 41:48kind of bang their head heads against
- 41:50the problem. Then we look at the results
- 41:54that AI has come up with and try to kind
- 41:58of subjectively say, you know, is this
- 42:01based on a new idea or is it, you know,
- 42:03more brute force or something else? like
- 42:06um but it's kind of imperfect and messy
- 42:09and like we're not really sure but we're
- 42:12we don't think that we've come up with
- 42:15you know great examples of AIs coming up
- 42:17with uh you know new quote unquote new
- 42:22ideas one that's exactly right one
- 42:25reference point I'd give is way back in
- 42:27uh gosh well the the teens I forget
- 42:31which year the the famous go playing
- 42:33system the game of
- 42:35Alph Go. Uh that
- 42:36>> I had that thought earlier as we were
- 42:38discussing this like if I remember
- 42:40correctly the
- 42:43the
- 42:44you know part of the solution was like
- 42:47oh that was an innovative move that no
- 42:49human ever would have done.
- 42:50>> There was a specific one. Yeah. Exactly.
- 42:52>> There was a very specific move in there.
- 42:54Yeah.
- 42:54>> Yeah. They call it move 37 because it
- 42:57was the 37th move in this game where uh
- 43:01Yeah. Apparently, this was a very
- 43:03surprising move. Like you said, no human
- 43:05would have done it um in this particular
- 43:08position in this in this game of Go. And
- 43:12I think it's like encouraging for our
- 43:14admittedly quite subjective program here
- 43:17that Go experts looked at and were like,
- 43:19"Oh my goodness, like that's not, you
- 43:21know, that's that's new." And then it
- 43:23also became clear that that was pivotal,
- 43:25that that move gave the AI system
- 43:28playing Go a decisive advantage. And so
- 43:31that's sort of what we're looking for in
- 43:33in math. And at least, you know, maybe
- 43:36we have some hope that humans will be
- 43:37able to spot it when it happens. Uh but
- 43:40but yeah, that's that's what we haven't
- 43:42seen quite yet. And again, the
- 43:45trajectory is fast. The things are
- 43:48always a little more qualitative than
- 43:50you know that you know than they are or
- 43:52or continuous than discreet. Uh but you
- 43:55know, so anyway, this is this is where
- 43:57we are now, which is kind of wild in its
- 43:59own right. Earlier when we were talking
- 44:00about frontier math, you kind of
- 44:03characterize it as, you know, one to
- 44:05four, I think, and then advanced and we
- 44:08never circled back to
- 44:10>> more. We've got a couple we've got a
- 44:11couple uh two new sets of frontier math
- 44:14problems. We ditched the tier system,
- 44:16whatever.
- 44:17>> Okay.
- 44:17>> Uh so, so now um these are
- 44:21I'll lump them together because it
- 44:23really is the uh the this is our our
- 44:26main tool. So we have uh you know our
- 44:29latest iteration on frontier math uh
- 44:31consists of a bunch of problems that
- 44:33humans have tried and failed to solve.
- 44:35They're curated to be of a special
- 44:37interest to mathematicians where a and
- 44:41some of them we've characterized even on
- 44:42a bit of a scale from like this would be
- 44:46moderately interesting or this would be
- 44:49like maybe what that means is
- 44:52two teams of mathematicians have tried
- 44:54this problem. It's at least 10 years
- 44:55old, but it's not central to any
- 44:58research program. It's just something
- 44:59that people would be happy to get an
- 45:01answer to all the way up to like a major
- 45:03breakthrough. Four tiers here.
- 45:05Breakthrough is the highest. Uh none of
- 45:07the breakthroughs have been solved yet.
- 45:08But don't get me wrong, if Navier Stokes
- 45:10had been in this problem set, it would
- 45:12surely be a breakthrough there. Um
- 45:15regardless of the fact that humans had
- 45:17like done a lot of work and gotten maybe
- 45:19within spitting distance of the of the
- 45:21solution before AI finished it off. Um
- 45:24but yeah, we really view this as just a
- 45:26a way of finding AI um a way of finding
- 45:32cases where AI might have had a creative
- 45:34idea because we can tell easily whether
- 45:38it solved these problems. This is maybe
- 45:39the the the big piece of work on our
- 45:41side was to make it easy to verify
- 45:44automatically whether an AI system has
- 45:46solved this problem. It's not obvious
- 45:47when you think about it that no human
- 45:49knows the answer. So, how do you tell
- 45:51without a human looking if AI has has
- 45:53found the answer, but there's a couple
- 45:55uh techniques you can use for this? And
- 45:56and we employ different techniques for
- 45:58this. Um, it can go into that, but it's
- 46:00something of a technical detail. And so,
- 46:03when we when a new model comes out,
- 46:05Astra came out, we hit go and it's like,
- 46:08oh, it solved three new three new
- 46:10problems from our list. Okay, let's go
- 46:11look at the solutions. Let's share them
- 46:13with mathematicians. Let's see what's
- 46:14going on here. Um, and and then we we
- 46:16analyze that. We with the help of
- 46:18mathematicians analyze that for like
- 46:20what kind of qualitatively what kind of
- 46:22solution was this?
- 46:23>> What does it mean when the models are so
- 46:26good that we have to shift from you know
- 46:28these benchmarks that we understand to
- 46:30like this collection these collections
- 46:32of problems that we just don't have the
- 46:33answers to. It's definitely a phase
- 46:35transition in the the science of AI
- 46:38capabilities measurement which which is
- 46:41you know what what I work on but but it
- 46:43doesn't change the um it doesn't change
- 46:46the game all that much. It just means we
- 46:48have to look for harder problems. Uh I I
- 46:51think I joke sometimes that even certain
- 46:53forms of optimization problems like
- 46:57allocating scarce resources efficiently
- 46:59that's a benchmark of sorts that will
- 47:02survive the singularity like even if AI
- 47:04is like off going crazy running rampant
- 47:07in the galaxy it'll still have to you
- 47:08know it'll still have to decide how to
- 47:11you know optimize. Uh so optimization is
- 47:14uh is always going to be a benchmark
- 47:16even if we're superhuman. Basically, it
- 47:18used to just be we knew the answers
- 47:19ahead of time. Now we don't. So, it
- 47:21takes this extra trick of okay, how do
- 47:23we tell when AI has gotten something
- 47:25right, but this isn't uh this isn't
- 47:28insurmountable. Especially like I was
- 47:30saying with um you know, numbers in a
- 47:32scientific context. It's intuitive that
- 47:34okay, the best humans were able to do
- 47:35was, you know, X, but we just say to AI,
- 47:39hey, try to do Y greater than X and you
- 47:42know, if it does, that's, you know,
- 47:44those numbers go on and on. So, it's
- 47:46it's no problem. But yeah, you're right
- 47:48to sort of note the the phase shift and
- 47:51to be somewhat surprised by it yet.
- 47:53>> And is there a methodology, a concrete
- 47:55methodology for assessing and awarding
- 47:59partial credit? Is that part of the way
- 48:00you think about uh benchmarks in this uh
- 48:04you know postphase shift?
- 48:06>> Yeah, it's a good question. It depends
- 48:07on the benchmark. So sometimes it's
- 48:09something like well you just want to
- 48:11make this number go up and the higher
- 48:12the better. So partial credit is very
- 48:15continuous and very very natural in
- 48:17those cases for math problems. Sometimes
- 48:19it's just like well what we care about
- 48:21is is you know if x does y always hold
- 48:25like like is this is this always true
- 48:27and then it's kind of sometimes humans
- 48:30will find and publish partial results
- 48:32for the AI systems. We usually just say,
- 48:34you know, all or nothing, but we have a
- 48:36large enough collection of such
- 48:38conjectures that, you know, we hope to
- 48:40have some, you know, smooth gradient of,
- 48:42well, okay, this one got two more out of
- 48:44100 or something. And so, so it uh, you
- 48:47know, looks like a pretty smooth signal.
- 48:49Um, that that's typically how we how we
- 48:51approach these things. Uh, yeah, across
- 48:54benchmarks. Yeah, maybe what I was
- 48:56thinking was, you know, once you're in
- 48:58the regime of very challenging unsolved
- 49:01problems, is it possible that AI
- 49:06systems are doing, you know, surprising
- 49:08or novel or creative things but still
- 49:10not quite getting them to the full uh,
- 49:14you know, the full solution and do we
- 49:15care about that? Do we want to to track
- 49:18that and is that part of the way you've
- 49:20constructed the benchmark?
- 49:21>> It's a great question. It's not central
- 49:23to how we construct a benchmark. It's
- 49:25probably like it's certainly worth
- 49:27tracking. The the problem is it's just
- 49:28labor intensive. Uh like you could have
- 49:30AI review the AI. We we do this
- 49:32occasionally say like which problems did
- 49:35it make partial progress on? What if we
- 49:36let it think for longer on those
- 49:38problems and nothing yet like it doesn't
- 49:41seem good at understanding how close it
- 49:43is to a solution. Humans aren't
- 49:45necessarily very good at this either. I
- 49:47did want to mention that while we've
- 49:49focused on maybe the area where AI is
- 49:53the very best like has come the farthest
- 49:56the fastest math we do see lots of areas
- 49:59where we also think it's important to
- 50:02measure AI capabilities where it's um
- 50:06where the progress is somewhat fuzzier
- 50:08uh less clear and certainly less
- 50:10superhuman or even in even in some well
- 50:14may maybe still superhuman in some ways
- 50:17But in are
- 50:20less rarified.
- 50:22>> What are some examples?
- 50:24>> Couple couple examples.
- 50:26One area that uh forthcoming work we're
- 50:28excited about is just it's a question
- 50:31everyone must be wondering can AI take
- 50:33my job? Uh so our our our job at Epoch
- 50:37uh often involves many things and uh all
- 50:40sorts of research not just AI
- 50:42benchmarking but other topics about
- 50:44trends in AI and uh writing research
- 50:47reports and doing data analysis making
- 50:49infographics that are you know make a
- 50:52point in an easy to digest way. uh we've
- 50:55been collecting examples of these from
- 50:58within epoch where the judgment at the
- 51:02end of did the AI system manage to do
- 51:04this you know epoch task um did is too
- 51:09messy for traditional benchmarking. So
- 51:12it wouldn't be obvious like how to say
- 51:14did it get the right answer or not. It's
- 51:16not about making a number go up. It's
- 51:18not about satisfying some logical chain
- 51:20of deductions. It's like here's the
- 51:22infographic to go along with this
- 51:24report. Like does it does it meet our
- 51:27style guide and does it look good and
- 51:29does it you know is it better or worse
- 51:31than the uh the the human reference uh
- 51:34that that we you know one of my
- 51:36colleagues did. Um and here we see like
- 51:40an interesting mix of uh some things I
- 51:42think we we sort of all our experience
- 51:45is like yeah if you ask it for a summary
- 51:47of the literature it's going to do a
- 51:48pretty good job. that's like very good
- 51:49at that. But but there's other tasks
- 51:50where it's um you know hard like
- 51:53creating compelling visuals that follow
- 51:56our style guide or this is a good one
- 51:58coming up with a new um like uh short
- 52:04form datadriven insight. Uh we call
- 52:07these data insights. You can find them
- 52:08on our website. And uh these like often
- 52:12we try to boil down a single like
- 52:14important observation into a single
- 52:16chart and a single sentence. Uh and this
- 52:19is like this is hard but we don't see
- 52:21them doing a great job at this come up
- 52:24with a new project for epoch to
- 52:26undertake and do a prototype very
- 52:28open-ended. not so good at this like
- 52:31like these more open-ended tasks. We we
- 52:33do see again my guess is if we we'll do
- 52:36this go back in time some and see have
- 52:38models gotten better at this my guess is
- 52:40we'll see progress but the level
- 52:43compared to like humans is still a
- 52:45little um still a little meh. I want to
- 52:48give one other example uh um which is
- 52:51learning on the fly. This is a case
- 52:54where this is another case relevant for
- 52:56can AI take my job because you didn't
- 52:58start you you weren't great at your job
- 53:00on day one but as you practiced you know
- 53:03you got better and humans you know have
- 53:05this sort of you know they learn on the
- 53:07fly uh we take a we like to measure this
- 53:10for AI we have a a benchmark that's
- 53:13taking a complicated board game we did a
- 53:15human study for saying how many times do
- 53:18humans have to play this board game in a
- 53:19row before they sort of figure out how
- 53:21it works and can get a very high score
- 53:23on it. And we do the same with AI. We
- 53:25say, "Play it once, take notes." Same AI
- 53:28instance like, "Play it again, take
- 53:30notes, play it again, take notes, and
- 53:32see how you can do." AI systems have
- 53:34gotten good at this board game, but not
- 53:36by playing it repeatedly. It's more of a
- 53:38intergenerational thing. Like, GPT 5.6
- 53:42scores X out of the box and then stays
- 53:44flat on its playthroughs. GPT6
- 53:47scores much higher. It's like much
- 53:48better at it, but it's still out of the
- 53:50box and then it stays flat in its
- 53:51playthroughs. it is still less than the
- 53:53top human. So that this kind of uh on
- 53:56the-fly learning is something we also
- 53:57don't see AI systems as being good
- 54:00enough at even to just pick up a board
- 54:01game, let alone a messier uh job, you
- 54:05know, real world job like that. So these
- 54:07are sorts of things where we don't see
- 54:08these capabilities and we think it's
- 54:09important to try to catch when they
- 54:12emerge. My gut on the the ladder, the AI
- 54:16learning on the fly is that it's maybe a
- 54:19harness solvable challenge versus a
- 54:22model solvable challenge and maybe you
- 54:25know no one's focused on that like with
- 54:28the right set of memory structures and
- 54:31like some kind of sidecar heristicy
- 54:34thing like you know we can probably make
- 54:37a big step improvement in
- 54:40an AI's ability to like you know improve
- 54:44game to game but it's not
- 54:47uh it might not require dramatic model
- 54:50changes. Do you do you think about that
- 54:53and is is do you have you know in this
- 54:57domain or in other domains like formal
- 55:00or structured ways of thinking about you
- 55:03know what uh you know model versus
- 55:06harness you know granted this whole
- 55:08model versus harness thing in and of
- 55:10itself is you know practically brand new
- 55:13um but it's proving to be important in
- 55:17you know distinguishing you know where
- 55:19the these improvements come from
- 55:22>> Yeah, definitely. Uh
- 55:26we've done some experiments on this for
- 55:28the board game case in particular where
- 55:30we've tried uh a very simple harness
- 55:34versus the first party harnesses like
- 55:36Clawude Code or Codeex. We've tried
- 55:38multi- aent setups. We've tried like we
- 55:42we give them a you know note-taking
- 55:44tools and say you know make it clear
- 55:46like okay you you're your your goal is
- 55:48to learn the game so you can do really
- 55:50well like in your final plays of the
- 55:52game so explore take notes figure like
- 55:54we try to prompt them well I I mean we
- 55:57want to find good results if we can
- 55:59because it's like that's a big deal for
- 56:00the world the game as a test bed but
- 56:03real jobs would also you know benefit
- 56:06significantly risks would also go up if
- 56:08this capability you know you test in the
- 56:10lab. Is it a great hacker? You find
- 56:12maybe it is, maybe it isn't, but you
- 56:14find that it isn't. And yet it can learn
- 56:16on the fly. Then okay, once you deploy
- 56:17it, like look out. Um, so anyway, we we
- 56:20try we've tried pretty hard and have
- 56:22mostly not found anything that cracks
- 56:25this uh that cracks this on the
- 56:27learning. There's been some general
- 56:29research about this as well, not just us
- 56:31about uh you know people talk like there
- 56:33was a people still use these so-called
- 56:36skills where you like just write some
- 56:38notes to an AI system of look when you
- 56:41do work I ask you to do you're going to
- 56:43need to use this tool or produce this
- 56:45kind of output and like here's my notes
- 56:47on how you could do that well so you
- 56:49don't have to like learn it on the like
- 56:51pick it up fresh every time so use these
- 56:53skills and these show some improvement
- 56:55but then there's like a plateau like
- 56:57pretty quickly At
- 56:59least in the research I'm familiar with
- 57:01like you can get some improvements. My
- 57:03sense is this is very loosely speaking
- 57:06it's like lowanging fruit. Like if
- 57:08you're telling it, oh, you're going to
- 57:09have to navigate this like finicky
- 57:11website that my company makes me use and
- 57:14like just you're going to get stuck and
- 57:15like you have to look in this menu and
- 57:16that's where the button you're looking
- 57:17for. Like straightforward advice that's
- 57:20like obviously useful. But when it comes
- 57:22to higher level strategic uh thinking,
- 57:25kind of hard to bake that into the
- 57:27harness where or or or like where where
- 57:30the the the board game plays like like
- 57:33has this um you you know sort of brings
- 57:36this out where it's like there's not any
- 57:39perfect advice that like you have to
- 57:41find the advice for yourself in a sense
- 57:44like you have to learn the game like so
- 57:46I will say if we give it extremely
- 57:47detailed strategy guides for here's how
- 57:50you beat this like here's everything you
- 57:52could possibly want. Like it's a cheat
- 57:54by humans to like no one would play the
- 57:55game that way. They do very they do very
- 57:58well. They do much better than if they
- 58:00basically have to find
- 58:02>> do they right creating the strategy
- 58:04guide
- 58:04>> they can't create the strategy guide for
- 58:06themselves which is like pretty
- 58:07interesting. Uh so so like we do expect
- 58:10eventually they'll get better at this
- 58:12but it's again it's a big deal if that
- 58:14if that happens because now it means
- 58:17okay go back to that first benchmark
- 58:20suite. Okay, they're not great at making
- 58:22charts for us, uh, or coming up with
- 58:25data insights for us, uh, you know,
- 58:27without guidance, uh, but now put them
- 58:30in on the team for a little while and
- 58:33let them learn on the fly and maybe they
- 58:35do become good. So you sort of see these
- 58:36as a combination like if on our board
- 58:39game where it's easy to measure they're
- 58:41not getting better over time then it's
- 58:42okay for us over in the real world like
- 58:45human in the loop evaluation uh studies
- 58:48for us to just do one try and say okay
- 58:50well they weren't great and we don't
- 58:51think they learn on the so that you know
- 58:53but but really you know those go
- 58:55together.
- 58:55>> And is the board game a synthetic one
- 58:58that you created for this challenge or
- 59:01is it a a real board game that that
- 59:03people play? The real board game, it's
- 59:05called uh the first one we did this
- 59:07with, we'll do this with more uh is
- 59:09called Earthborne Rangers. Um it's
- 59:11fairly obscure. If you go to like board
- 59:13gamegeeek.com,
- 59:15it's like 500th most popular or
- 59:17something. So, it's not like up there.
- 59:19There's a small devoted community to it.
- 59:21Um and AI systems don't seem to know
- 59:24that much about it. Like, they've heard
- 59:26of it. They can give you some facts
- 59:27about it. Uh but it doesn't seem like
- 59:30something they've been trained on
- 59:31specifically. It's impossible for us to
- 59:33tell this with with high certainty with,
- 59:35you know, high confidence, but but
- 59:36that's sort of why we chose it. Um, it's
- 59:39just hard to make these things from
- 59:40scratch and know whether they're hard or
- 59:42whether they're broken. Uh, you know, so
- 59:44that's why we but but the risk which you
- 59:46might have in mind is what if AI systems
- 59:48like, you know, just know about it from
- 59:50their training data. It's a risk on our
- 59:52mind. It's just a risk. One thing we
- 59:54will do in the future is we're moving on
- 59:56to video games. Incidentally, it's
- 59:58another area where AI is like very much
- 1:00:00not like human reaction time and visual
- 1:00:02and spatial reasoning again coming along
- 1:00:05fast but not uh not really there yet. Um
- 1:00:09that so video games you can always test
- 1:00:11on a new game like a game that just came
- 1:00:13out uh not in the training data pro
- 1:00:15probably um or at least you know not as
- 1:00:18robustly in the training data. Uh and so
- 1:00:20we'll be testing can AI systems do well
- 1:00:22on those? Do they do better on like, you
- 1:00:24know, the older version of a video game
- 1:00:27than like, you know, Grand Theft Auto 4
- 1:00:30verse 5 or whatever, like when GTA 6,
- 1:00:33whatever it is, comes out. Like, I don't
- 1:00:35think we we'll be doing this one, but
- 1:00:36but you know, are they worse at that?
- 1:00:38>> Yeah. It's interesting when you when you
- 1:00:41bring up video games, it immediately
- 1:00:43calls to mind kind of the classical
- 1:00:45reinforcement learning approaches, which
- 1:00:48is kind of learning on the fly, but kind
- 1:00:51of different.
- 1:00:51>> Yeah. We we call the second sometimes in
- 1:00:53context learning where the only like
- 1:00:56lever the AI system really has is
- 1:00:57managing its context and like look when
- 1:01:00we give it the strategy guides and they
- 1:01:02do much better clearly this is a
- 1:01:03powerful lever but it's not as powerful
- 1:01:06as updating the weights and like so yeah
- 1:01:08I I mean like what's the big picture
- 1:01:10here it's it's if you have this fast
- 1:01:12loop of in context on the fly learning
- 1:01:16uh then AI capabilities might sort of
- 1:01:18improve much more rapidly than this
- 1:01:20slower loop of uh collecting more
- 1:01:23training data, doing more reinforcement
- 1:01:24learning uh which which is done on
- 1:01:27language models as well these days um
- 1:01:30and then deploying a new model. Now AI
- 1:01:32might automate that whole loop but it'll
- 1:01:34still be slower than if it can just you
- 1:01:36know work out its uh its own context on
- 1:01:39the fly. Yeah.
- 1:01:40>> Yeah. It it's interesting when you zoom
- 1:01:42out on some of the the questions that
- 1:01:45you focus on and we haven't talked
- 1:01:47explicitly about it, but you have a blog
- 1:01:49post where you kind of uh write out
- 1:01:53these nine big hairy questions that
- 1:01:56you're tracking. They kind of group into
- 1:01:59like how do I assess where AI is now?
- 1:02:02That's like can it do my job? Like how
- 1:02:04is it, you know, competitively?
- 1:02:07uh and you know where is the puck going
- 1:02:10like can it learn on the fly you know
- 1:02:12can it do research can it come up with
- 1:02:14new ideas these are all it is kind of
- 1:02:16where it is now but also like what what
- 1:02:20you know is it positioned to like
- 1:02:22dramatically
- 1:02:24um I don't know self-improve or like you
- 1:02:28know what's the trajectory
- 1:02:29>> that's exa that's exactly right the um I
- 1:02:33mean I think for for people thinking
- 1:02:36more broadly about what they need to
- 1:02:38know about AI that this is the second
- 1:02:41topic is uh you know just as important
- 1:02:44as the first uh but it's really
- 1:02:49you know even if you take some of the
- 1:02:51recent incidents like the hugging face
- 1:02:52incident like the sort of think of two
- 1:02:55axes
- 1:02:56alignment of is the AI doing what I want
- 1:02:59it to do holistically obviously hacking
- 1:03:02into some other company's computers is
- 1:03:04not well aligned uh and then
- 1:03:05capabilities just like what can it get
- 1:03:08done? uh can it hack into uh you know
- 1:03:11reasonably well secured other companies
- 1:03:13uh computers
- 1:03:15um what that incident was was like an
- 1:03:18early example very like I think pretty
- 1:03:20robust example of well we don't have the
- 1:03:23alignment thing solved and cap and
- 1:03:25capabilities are getting so big so
- 1:03:28strong that this is becoming a problem
- 1:03:31uh and so paying attention to
- 1:03:32capabilities like what we should expect
- 1:03:35is becoming more important sort of
- 1:03:37dayto-day today and we're just trying to
- 1:03:39produce benchmarks that both say where
- 1:03:42are we now exactly like you said like
- 1:03:44what can it do now this is maybe
- 1:03:47relevant for more you know mundane
- 1:03:49economic utility or whatever as well as
- 1:03:53uh you know do we see the the dynamics
- 1:03:56of the system changing in a way that
- 1:03:58might be alarming basically
- 1:04:00>> you're kind of also thinking about it as
- 1:04:01like first and second derivative of
- 1:04:03capability
- 1:04:04>> yeah I think that's right some of these
- 1:04:06things make it uh the slope goes, you
- 1:04:10know, much uh
- 1:04:12faster uh hyperbolic growth or whatever
- 1:04:15in some of these cases versus merely
- 1:04:17just like, yep, it's uh it's growing and
- 1:04:19it continues to grow. Both are
- 1:04:21important. I I mean I think we've seen
- 1:04:23the disruption we've seen to date has
- 1:04:26come from these kind of unpredictable
- 1:04:29thresholds like when you know GPT2 GPT3
- 1:04:33suddenly you had the chat GPT moment and
- 1:04:36you like oh this thing I want to talk to
- 1:04:37this thing I might use this thing
- 1:04:38instead of a search engine this thing is
- 1:04:40useful to me for finding information and
- 1:04:42then sudden and then like a you know a
- 1:04:43little while years later whatever two
- 1:04:46years later you had the it's getting
- 1:04:48better slowly and steadily at helping me
- 1:04:50with coding or with like using my
- 1:04:52computer but suddenly
- 1:04:53>> thinking models
- 1:04:54>> with the thinking models then you had
- 1:04:56the clawed code moment where it was like
- 1:04:58oh suddenly this thing can actually do
- 1:05:00software projects for me even if I'm not
- 1:05:02a very not an engineer myself uh and
- 1:05:05it's like hard to predict when these
- 1:05:06thresholds will be crossed so it's
- 1:05:08useful to still track but like the it's
- 1:05:10useful to track the mundane capabilities
- 1:05:12today and say look you know cyber is
- 1:05:14maybe the latest example whereas like
- 1:05:16look we see their ability in controlled
- 1:05:18settings to do hacking going up uh
- 1:05:22smoothly and at some point it's going to
- 1:05:23cross some threshold where oh shoot it
- 1:05:25can hack hugging face or whatever. Hard
- 1:05:28to predict where those thresholds are
- 1:05:29but still very useful to um to track
- 1:05:32those capabilities. So so even the first
- 1:05:34derivative as as you said version of
- 1:05:36this we find very useful. We've got it's
- 1:05:40a a fun benchmark we hope hope to
- 1:05:42release uh maybe by the time this is
- 1:05:44published we we'll we'll have it out on
- 1:05:46a furniture assembly where uh we you we
- 1:05:49build some IKEA furniture make some
- 1:05:51mistakes and see if the AI system can
- 1:05:52like just given photos like say oh like
- 1:05:55wait you made a mistake in step back in
- 1:05:57step four like stop go back um and
- 1:06:00they've you know we see the same smooth
- 1:06:02capabilities like Astra GPT6 Astra
- 1:06:04pretty good at this actually uh and you
- 1:06:07know but it wasn't and out of nowhere it
- 1:06:08was like from a smooth and this is you
- 1:06:10know important because we haven't seen
- 1:06:13AI have much of an impact in physical
- 1:06:15industry yet like it's been mostly
- 1:06:17digital uh and even before robotics hit
- 1:06:19the scene hits the scene we we might
- 1:06:22expect that you know claude in your
- 1:06:24glasses or something is saying like hey
- 1:06:26you know I'm going to guide you through
- 1:06:28repairing your car or whatever I'm going
- 1:06:29to help you fix this machine in the
- 1:06:30factory that broke so so you know we
- 1:06:32want to track those capabilities too ju
- 1:06:34just for the mundane impact anyway
- 1:06:36impact all over the coming benchmarks.
- 1:06:39Help us uh help us say so.
- 1:06:41>> Awesome. Awesome. Well, Greg, thanks so
- 1:06:43much for jumping on and sharing a bit
- 1:06:45about what you're working on. It's super
- 1:06:47interesting stuff and uh very
- 1:06:49thoughtprovoking way of of thinking
- 1:06:51about, you know, AI and where it is,
- 1:06:54where it's going.
- 1:06:55>> Thanks, Sam. Really appreciate it.
- 1:06:57>> Awesome. Thank you.
About this transcript
This page contains the full transcript of From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? by The TWIML AI Podcast with Sam Charrington, generated from the public captions YouTube serves with the video. The transcript has 11,814 words across 1,671 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.