YouTube2Text

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? — Transcript

by The TWIML AI Podcast with Sam Charrington · 11,814 words · 1,671 segments · language en · Watch on YouTube

Full transcript

  1. 0:00AI's progress in mathematics has been
  2. 0:02remarkably fast. Not too long ago,
  3. 0:05Frontier Models struggled with grade
  4. 0:07school math. Now, they're contributing
  5. 0:09solutions to problems that have resisted
  6. 0:11mathematicians for decades, including
  7. 0:13most recently the Navier Stokes
  8. 0:16equations. But as AI moves from tests we
  9. 0:19already know the answers to to more
  10. 0:20open-ended research challenges, it gets
  11. 0:23harder to understand what these systems
  12. 0:24are truly capable of and what their
  13. 0:26progress tells us about where AI is
  14. 0:28headed. My guest, Greg Burnham, leads
  15. 0:31capabilities research at Epoch AI, where
  16. 0:33he and his team are developing new ways
  17. 0:35to track AI's rapidly evolving strengths
  18. 0:38and limitations. Here's Greg on why
  19. 0:40measuring and understanding AI's
  20. 0:42capabilities matters so much. Right now,
  21. 0:45AI capabilities
  22. 0:47large have gotten to such a point where
  23. 0:50what they're good at and bad at is
  24. 0:52starting to have real impact on the
  25. 0:54world, like on things we care about for
  26. 0:56other reasons. We're past the academic
  27. 0:59exam phase of understanding AI
  28. 1:02capabilities and for that we need high
  29. 1:04quality benchmarking, high quality
  30. 1:05evaluations. And one of my big questions
  31. 1:08is um just can AI come up with new
  32. 1:10ideas? If AI systems gets sort of
  33. 1:12superhuman at this that's a big deal for
  34. 1:15for the world both in terms of economic
  35. 1:18and scientific progress but also in
  36. 1:20terms of becomes harder to predict what
  37. 1:23AI systems are capable of. We use math
  38. 1:26as a test bed for this. I'm Sam
  39. 1:28Cherington and this is the Twimmel AI
  40. 1:30podcast. For over a decade, I've been
  41. 1:33exploring the ideas and innovation
  42. 1:35shaping the future of AI through
  43. 1:36conversations like this one that help
  44. 1:38you understand what's real, what's next,
  45. 1:41and what matters. Let's jump in.
  46. 1:52Let's talk a little bit about AI
  47. 1:53capabilities research broadly, how you
  48. 1:56pursue that research, and why you think
  49. 1:58it's important.
  50. 1:58>> We're getting to a point in the in the
  51. 2:01trajectory of AI overall where
  52. 2:05AI capabilities can have a real impact
  53. 2:08on the world. So
  54. 2:11maybe the transition could be
  55. 2:13characterized as school to work. Uh in
  56. 2:18maybe even up through Yeah. up through
  57. 2:202025, most ways that we tested what AI
  58. 2:25could do looked more like human exams
  59. 2:28like that you might give to you know to
  60. 2:30kids or to maybe advanced graduate
  61. 2:33students. I was going to say,
  62. 2:36>> right? Right. The the kid exams it is
  63. 2:38>> the bar exams, the medical exams
  64. 2:42>> and these p like I was perfectly good at
  65. 2:44these already end of 2024, beginning of
  66. 2:462025. So we really saw saw this
  67. 2:49transition happening over time to to be
  68. 2:52clear, but this is this is certainly the
  69. 2:55year where AI capabilities are impacting
  70. 3:01work activities that humans were engaged
  71. 3:04in. anyway uh for you know for their own
  72. 3:07purposes and so as these capabilities
  73. 3:11grow it's just important for us to keep
  74. 3:14tabs on them uh they grow very rapidly
  75. 3:17in some qualitative sense we see no no
  76. 3:20measurement we have shows any slowdown
  77. 3:23in what uh in how AI is getting better
  78. 3:26and better at any tasks that we're able
  79. 3:28to measure and so we think that just
  80. 3:32understanding the impact AI will have on
  81. 3:34the world. It's important to know what
  82. 3:36it can do, what it can't do, and keep a
  83. 3:38and keep tabs on how that how that
  84. 3:40trajectory is going.
  85. 3:41>> I want to push back on one thing you
  86. 3:43said in terms of we're not seeing any
  87. 3:47slowdowns
  88. 3:49that can be looked at broadly like in
  89. 3:52terms of maybe the number of, you know,
  90. 3:54benchmark data points that we're looking
  91. 3:56at. You know certainly we're seeing this
  92. 3:58broad performance but is it also true
  93. 4:02that like on a particular
  94. 4:05I feel like on a particular benchmark
  95. 4:07we're getting like the you know for for
  96. 4:09expected reasons when we went from you
  97. 4:12know GPG2 to GPG3 like we're taking
  98. 4:15these huge steps up these benchmarks and
  99. 4:18now like there's just a lot less room
  100. 4:20left and we're making more incremental
  101. 4:22progress. Do you see that as a a
  102. 4:25slowdown in AI performance or is would
  103. 4:27you characterize it differently?
  104. 4:29>> I would say well I'd say there's two
  105. 4:31issues here. One is on the benchmarks we
  106. 4:33have what does performance look like?
  107. 4:36Two is
  108. 4:38are there things our benchmarks don't
  109. 4:40measure that maybe there has been less
  110. 4:43or could even be negative uh improvement
  111. 4:47on and is there some reason why it's the
  112. 4:51things that are hard to measure where
  113. 4:53there's been the least progress. But on
  114. 4:55the first point because this is a very
  115. 4:56this is a much more cut and dried
  116. 4:58statistical point. We absolutely see
  117. 5:01continued progress on benchmarks like a
  118. 5:05benchmark that was challenging for GPT3
  119. 5:08uh you know many years whatever six
  120. 5:10years ago uh is now completely you know
  121. 5:14solved uh aced by GPT4 benchmarks
  122. 5:19challenging for GPT4 same for GPT5 now
  123. 5:21we've got GPT6 like so we do some
  124. 5:24statistical work to try to aggregate
  125. 5:26benchmarks over time Because the
  126. 5:29benchmarks that people used to track
  127. 5:32progress years ago, now every AI model
  128. 5:36off the shelf that that you've heard of
  129. 5:37that you might use is too good at them.
  130. 5:40Uh so people make new benchmarks and
  131. 5:41then you have to sort of do this like
  132. 5:43stitching of like okay well a model that
  133. 5:45was pretty darn good at this benchmark
  134. 5:48was not so good at this next one and you
  135. 5:50can use like a you can stitch those
  136. 5:52together and get this sort of composite
  137. 5:54score. We call it we we have a
  138. 5:56methodology for doing this. to call the
  139. 5:57epoch capabilities index. And that's a
  140. 6:01unified way of tracking benchmarks over
  141. 6:03time. And we really do see this thing
  142. 6:05continuing to go up uh you know pretty
  143. 6:08smoothly linearly and and that march of
  144. 6:12progress is uh is is steady. That's an
  145. 6:14important point that that's truly if um
  146. 6:18we we have we've made more benchmarks
  147. 6:21and AI systems that are harder that
  148. 6:23models start off at zero and then within
  149. 6:26a year or two years they're at 100% and
  150. 6:29we have to make a new benchmark to keep
  151. 6:30up. Uh and that's that's really been the
  152. 6:33story of the last uh certainly the last
  153. 6:35three years I I'll say with confidence
  154. 6:37there. It seems like that like
  155. 6:40normalizing across different benchmarks
  156. 6:43to
  157. 6:44produce a result that is coherent that
  158. 6:49sounds like a very challenging problem.
  159. 6:51like you you you describe the result as
  160. 6:55like you know kind of continued linear
  161. 6:57improvement and I imagine if that's like
  162. 7:01a constraint you can kind of map
  163. 7:04backwards to you know uh to factors that
  164. 7:08will show you that continuous linear
  165. 7:10improvement but you know starting from
  166. 7:13you know a set of disperate benchmarks
  167. 7:15that have their own kind of you know
  168. 7:18ranges and challenges and you know where
  169. 7:20AI you and where AI sits in them and and
  170. 7:23kind of mapping this across different
  171. 7:25generations of AI and coming up with a,
  172. 7:29you know, something that makes sense,
  173. 7:31you know, objectively at the end and
  174. 7:33then having that thing show continued
  175. 7:35linear progress. That sounds like um
  176. 7:42you almost like too good to be true and
  177. 7:44like suspicious that it is true.
  178. 7:46>> I I I yeah, I I agree. So, so let let me
  179. 7:49go into that a little because I think
  180. 7:50it's a it's a killer point to understand
  181. 7:53and it is surprising. So, I think you're
  182. 7:55very right to be surprised. In fact, if
  183. 7:58you look at AI benchmarks, they are all
  184. 8:02correlated with each other even if the
  185. 8:04benchmark the scores across models are
  186. 8:07correlated with each other even if the
  187. 8:09benchmarks are in nominally different
  188. 8:11domains. So maybe you don't expect you
  189. 8:15don't find this so surprising if it's
  190. 8:17like I've got one benchmark that's you
  191. 8:19know graduate chemistry questions and
  192. 8:22I've got another one that is coding
  193. 8:24puzzles and the models that are good on
  194. 8:27chemistry are also good on software
  195. 8:29engineering. Maybe that's not
  196. 8:30>> correlated largely by the the models
  197. 8:33themselves. Like as the models get
  198. 8:35better they're going to get better
  199. 8:36across these you know disparate
  200. 8:38benchmarks. But but this uh exactly but
  201. 8:40that's that's the surprising point. So
  202. 8:42all that you need to be true to see this
  203. 8:45linear trend over time is basically
  204. 8:47every time you know opus 4.5 opus 4.6
  205. 8:514.7 GPT 5.2 5.4 every time these come
  206. 8:55out if they get better on all benchmarks
  207. 8:58at once then you're going to see this
  208. 9:00linear tren this this big overtime
  209. 9:02linear trend like that's it's two sides
  210. 9:04of the same coin statistically. Why does
  211. 9:07this happen? Like this is a huge
  212. 9:09question. Like you're you're right to
  213. 9:10find that suspicious. And I would say
  214. 9:13there are two explanations that that we
  215. 9:15see. And it's an important question
  216. 9:18which of these explanations is more
  217. 9:20true. And of course the truth is some
  218. 9:22mix. But but let me let me give the the
  219. 9:25two explanations. One is the AI
  220. 9:29companies are making darn sure that
  221. 9:32their model gets better at all of these
  222. 9:35benchmarks.
  223. 9:36with each new generation of model. And
  224. 9:39the main mechanism you might imagine
  225. 9:40them doing that if they're like looking
  226. 9:42around, okay, people care about
  227. 9:44chemistry, people care about software,
  228. 9:46people care about operating, you know,
  229. 9:48answering my email, like they they would
  230. 9:51collect training data one form or
  231. 9:53another that would help the model learn
  232. 9:55how to do these how to do these tasks.
  233. 9:57And that's something of a
  234. 10:01um that's something of a you know very
  235. 10:03manual human in the loop process of just
  236. 10:05saying hey we're doing product
  237. 10:06development like they want this feature
  238. 10:08they want that feature our users want
  239. 10:09all these features and the way you do
  240. 10:11that in machine learning is you collect
  241. 10:12training data and throw it in I'm
  242. 10:14oversimplifying but that that's that um
  243. 10:17sort of like a shallow way to make it
  244. 10:19happen the so I'll call that shallow the
  245. 10:21other one is deep the deep way to make
  246. 10:23it happen is if the model generally is
  247. 10:26if the AI system generalize very very
  248. 10:29well. So yeah, you never trained it on
  249. 10:32chemistry, but you trained it a lot on
  250. 10:34math and software engineering and it
  251. 10:37just got super smart and with a modocum
  252. 10:40of chemistry basics, it was able to
  253. 10:43derive the rest from first principles
  254. 10:45sort of thing. This is like a deep form
  255. 10:47of generalization. Both of these
  256. 10:50happened to some extent like like kind
  257. 10:51of the revolution with GPT even GPT3
  258. 10:55around that era. This is like quite a
  259. 10:56while ago was oh if you train it to like
  260. 10:59predict the next word then it actually
  261. 11:01gets better at a wide wide range of
  262. 11:03languageoriented tasks that people had
  263. 11:06been trying to handle individually.
  264. 11:08Suddenly it just got got good at
  265. 11:10everything. And that wasn't because
  266. 11:11anyone was like oh let's have data
  267. 11:14specifically around uh sentiment
  268. 11:17analysis or or syntactic structural
  269. 11:20analysis. Like it wasn't anything like
  270. 11:22that. It was just give it give it this
  271. 11:23very general thing. But now we're on to
  272. 11:25this world where you have to like train
  273. 11:27it specifically. Right now there's a
  274. 11:29huge industry of collecting all this
  275. 11:31rich training data and it's really
  276. 11:32unclear how much generalization you're
  277. 11:35getting from you know in this deep way.
  278. 11:38Uh and and if and anyway that's that's
  279. 11:41sort of a big question because that deep
  280. 11:43generalization means you could really
  281. 11:45see if that happened in a in a rich way
  282. 11:47where you just didn't have to even be
  283. 11:49trained on something at all to gain
  284. 11:50capability in that area. then you might
  285. 11:53have a or like radically superhuman AI
  286. 11:56that would be um pretty hard to uh to
  287. 11:58understand like to to bound its
  288. 12:00capabilities.
  289. 12:02It's still a little counterintuitive
  290. 12:03that this would be that this would be
  291. 12:06linear and maybe you know maybe part of
  292. 12:09the question is like are we talking
  293. 12:11about
  294. 12:14you know linear across generations of uh
  295. 12:18of of model releases or linear within
  296. 12:21some regime like segment you know
  297. 12:23segmentally linear or
  298. 12:25>> the the thing I can say is we have like
  299. 12:28a statistical method so a benchmark
  300. 12:30score Like what are we talking about
  301. 12:32here? Our unit here is a score from 0 to
  302. 12:34100%. And typically the way these scores
  303. 12:38go just historically, empirically is
  304. 12:41like they follow like an S-curve. Like
  305. 12:42like they start out sort of bad, then
  306. 12:44they get better, they get better pretty
  307. 12:46rapidly, and then you can only get up to
  308. 12:48100%. So then they sort of level off.
  309. 12:50Takes them a while to get that last, you
  310. 12:52know, little chunk and then they're at
  311. 12:54100%. So you've you've got these like
  312. 12:56sigmoids, uh is is what that scurve is
  313. 12:59called. We've just got a statistical
  314. 13:00methodology for stitching sigmoids
  315. 13:03together and what comes out is a linear
  316. 13:06function. And this looks like like this
  317. 13:08is an artifact of the statistics to be
  318. 13:10clear. But this sure enough looks very
  319. 13:12regular in its own statistical analysis
  320. 13:15terms over across generations of of
  321. 13:19models. Like absolutely like it's been
  322. 13:21uh if anything it accelerated a little
  323. 13:23with the development of so-called um
  324. 13:26reasoning models around late 2024 where
  325. 13:28the models started to get good at math
  326. 13:30and and software engineering and things
  327. 13:33that require lots of logical reasoning.
  328. 13:35Uh maybe like a higher linear slope but
  329. 13:38uh but this is just like a that's a
  330. 13:39statistical artifact. You now have to
  331. 13:41ask, okay, well, if your score is like
  332. 13:44one model score is 150 and the next
  333. 13:46model score is 152, like what do those
  334. 13:49two points mean on your benchmark
  335. 13:51aggregate index?
  336. 13:53Now,
  337. 13:54>> I think part of part of what I'm what
  338. 13:56I'm asking myself as I hear this
  339. 13:59described as linear is does it say that
  340. 14:02in spite of the model the Frontier Lab
  341. 14:06vendor claims that you know Fable is
  342. 14:10like dramatically better than Opus or
  343. 14:13that GPT6 is like the step function uh
  344. 14:16you know improvement over GPT5.6.
  345. 14:19you know, linear implies to me that that
  346. 14:22it's kind of incremental and there's
  347. 14:24really been no step functions.
  348. 14:26>> This is um this is a really great
  349. 14:28question and yeah, we believe this is
  350. 14:30one of the big uses of this statsy tool
  351. 14:34we've got is you can try to find you
  352. 14:37know trend breaks and you can try to
  353. 14:40find accelerations. Uh and we mostly
  354. 14:43don't see that. Yeah. Like one way to
  355. 14:45put it that is everything just about is
  356. 14:48on trend. Fable uh Astra whatever is on
  357. 14:53trend but the trend is crazy. Like you
  358. 14:56gota you got to hold both of these
  359. 14:58together.
  360. 15:01>> Right. Right. Right. Right. Right. This
  361. 15:03is like this is sort of the zoomed out
  362. 15:04view epoch takes with with everything is
  363. 15:08a lot of this we think the engine of AI
  364. 15:10progress is heavily mediated by raw
  365. 15:14inputs like GPUs the training compute
  366. 15:17like training data maybe those are the
  367. 15:20big ones and then there's this like
  368. 15:21factor of the labs come up with
  369. 15:22algorithmic like innovations like
  370. 15:24research breakthroughs but we really
  371. 15:26think like a lot of it is coming from
  372. 15:28the scale up in the in the compute and
  373. 15:31the scale up in the training data So
  374. 15:33until you see the data and the GPU the
  375. 15:35data and compute supply chain like ramp
  376. 15:39uh significantly or discontinuously uh
  377. 15:43to be more precise you're not going to
  378. 15:44see you're not likely to see
  379. 15:47discontinuities on the model side.
  380. 15:49>> I think that's I think that's right. A
  381. 15:51big thing we're keeping our eye on to be
  382. 15:53clear is
  383. 15:55are there the are AI abilities emerging
  384. 15:58that could cause a discontinuity on the
  385. 16:00model side. So this is why people talk
  386. 16:03about recursive self-improvement or
  387. 16:05sometimes they call this a softwareonly
  388. 16:08intelligence explosion. And the point of
  389. 16:11these concepts is well sure right now
  390. 16:15there's a slow comparatively slow
  391. 16:17process for building more micro building
  392. 16:20more GPUs and collecting more human
  393. 16:22training data and whatnot. But if an AI
  394. 16:25system got really good on the
  395. 16:27algorithmic innovation side, could it do
  396. 16:30more with less or or do more with the
  397. 16:32with a fixed uh compute budget? And that
  398. 16:37could change the dynamics here. So the
  399. 16:38feedback loop would might no longer run
  400. 16:40through the physical world. We don't
  401. 16:43really it's not clear we see evidence of
  402. 16:44this yet. I'd say on balance we don't
  403. 16:47really see evidence of this yet but it's
  404. 16:49the sort of thing that has spooked the
  405. 16:52people inside the labs the the AI
  406. 16:54companies I mean saying like gosh this
  407. 16:57could be right around the corner that's
  408. 16:58like a bit of the fear and if that
  409. 16:59happens then the pace we're used to we
  410. 17:03might get that discontinuity
  411. 17:05>> have you developed a particular
  412. 17:06benchmark to identify and track this or
  413. 17:10is it more
  414. 17:13the way you look at the benchmarks that
  415. 17:15we've been talking about this
  416. 17:16statistical model that we've been
  417. 17:18talking about for an indication that you
  418. 17:21know there's some kind of you know
  419. 17:23exponential uh in the curve and and
  420. 17:26that's you know probably one of the
  421. 17:28likely causes you know for that kind of
  422. 17:30exponential or discontinuity
  423. 17:32>> all of the above. Uh we we want as many
  424. 17:35as many tools in our arsenal as as we
  425. 17:37can for this. So that to this statsy
  426. 17:40like acceleration detector is definitely
  427. 17:44a good uh is definitely a good tool. But
  428. 17:46we also sort of get you know down in the
  429. 17:49weeds with some of our specific
  430. 17:51evaluations and benchmarks saying I'll
  431. 17:53give you an example here. This is some
  432. 17:55forthcoming work we have uh that that
  433. 17:58I'm happy to to preview. We take a
  434. 18:01recent AI research breakthrough or
  435. 18:05innovation
  436. 18:07from humans and we take AI systems. We
  437. 18:12don't connect them to the internet and
  438. 18:13we choose AI systems that uh were
  439. 18:16developed and released just before this
  440. 18:18recent h whatever the recent human
  441. 18:20innovation is. And we basically say look
  442. 18:23here's the metric that that innovation
  443. 18:25the number that it makes go up. It
  444. 18:27improves this sort of efficiency. what
  445. 18:29whatever some some metric and we say AI
  446. 18:31system your goal is to improve this
  447. 18:33metric as much as you can but we don't
  448. 18:35tell it the innovation and so we see is
  449. 18:38it capable of replicating that human
  450. 18:41innovation that was highly relevant for
  451. 18:44AI you know research and development and
  452. 18:48so far we don't see them able to do this
  453. 18:50there's this nebulous idea called
  454. 18:52research taste which is like where's the
  455. 18:55good idea like what what experiment
  456. 18:57should I do next where Should I go
  457. 18:59hunting for uh big improvements more
  458. 19:02than the mundane improvements of just
  459. 19:04more compute, more data? Like where
  460. 19:06should I look for for a breakthrough?
  461. 19:08And AI systems don't seem great at doing
  462. 19:10this in um at least in an AI R&D
  463. 19:15context. Uh but but they're like getting
  464. 19:17a lot better at a lot of things that are
  465. 19:19adjacent to this uh to this area. So we
  466. 19:22really do think this is an important
  467. 19:24sometimes I almost call it a trip wire
  468. 19:26that we're you know the the future is
  469. 19:28foggy is some foggy landscape. We want
  470. 19:30to put a trip wire out there so that if
  471. 19:32AI ever does cross some threshold where
  472. 19:34it's able to do rapid improvements in AI
  473. 19:38algorithms themselves then uh we we want
  474. 19:41that trip wire to sound. we go, okay,
  475. 19:43gosh, watch out. Like maybe maybe
  476. 19:45something uh maybe we're about to see a
  477. 19:47big improvement big um you know, uptick
  478. 19:50in AI capabilities.
  479. 19:51>> And how does this relate to
  480. 19:54AI's performance on some of the math
  481. 19:57challenges? You know, certainly Navia
  482. 19:59Stokes has been uh in the news quite a
  483. 20:03bit recently. Prior to that, there were
  484. 20:04some Erdos results uh and others. like
  485. 20:08how do you think about the role that
  486. 20:10those math challenges play in this
  487. 20:13broader
  488. 20:14view of trying to understand AI's
  489. 20:17trajectory?
  490. 20:18>> Yeah, very good question. Uh I if I may,
  491. 20:21I would
  492. 20:23go rewind history just a just a bit here
  493. 20:26to um say some of the say the trajectory
  494. 20:29we've been on because once you you step
  495. 20:31back it's um it's it's wild. Uh so up
  496. 20:36until
  497. 20:40yeah up until
  498. 20:42fall of 2024 so we're talking two years
  499. 20:45ago uh grade school math was still a
  500. 20:49challenge for these LLM based AI systems
  501. 20:52the the kind of AI we're talking about
  502. 20:54here and all and the benchmark there was
  503. 20:57literally a benchmark of 8,000
  504. 20:59you know word problems so and so you
  505. 21:02know has so many apples that sell for
  506. 21:05this
  507. 21:06>> GSM8K that's the one
  508. 21:08>> and
  509. 21:10GPT4 which is 2023 had like done pretty
  510. 21:13well on this so people started looking
  511. 21:14at some like adding into the mix a
  512. 21:17benchmark just called math like four
  513. 21:19capital letters that um started pulling
  514. 21:22from some high school math competitions
  515. 21:25and there was like slow progress on this
  516. 21:26but but not a lot and then at the end
  517. 21:28fall of 2024 OpenAI comes out with the
  518. 21:3101 preview model full 01 by the end of
  519. 21:352024. And this showed a huge spike in uh
  520. 21:38in math competition capabilities. But
  521. 21:40we're still just talking like medium
  522. 21:42hard high school math competitions like
  523. 21:44something, you know, thousands of math
  524. 21:46nerd nerdy math 16-year-olds were were
  525. 21:49into across the across the US and and
  526. 21:54this became, you know, the new the new
  527. 21:56territory. My my company, Epoch, around
  528. 21:59that time released a benchmark called
  529. 22:01Frontier Math. Uh at the time that's all
  530. 22:04we called it. We now call it tiers one
  531. 22:05through four uh to distinguish from
  532. 22:08later versions I'll get to in a bit. Uh
  533. 22:11but the original frontier math consisted
  534. 22:13of problems that I would say go from
  535. 22:15advanced undergraduate to advanced grad
  536. 22:19student. So sort of early career
  537. 22:21research warm-ups almost. These are
  538. 22:24problems that are no longer thousands of
  539. 22:27uh thousands of high schoolers would
  540. 22:30tackle them. These require deep
  541. 22:32background knowledge in in various areas
  542. 22:36that that you know niche areas like
  543. 22:38we're talking a a problem from you know
  544. 22:42elliptic curves that you'd give to a
  545. 22:45grad student like shortly before they're
  546. 22:47going to embark on their novel thesis
  547. 22:49research just to get them familiar with
  548. 22:50the details of the of the niche that
  549. 22:52they're going to be trying to work in
  550. 22:54and make an original contribution to.
  551. 22:56Over the same period, we also saw high
  552. 22:58like so over 2025, we also saw high
  553. 23:01school math contests get completely
  554. 23:03aced. We got the gold medal on the
  555. 23:05International Math Olympiad. By the um
  556. 23:09even by the beginning of 2026, we saw
  557. 23:12most of these frontier math problems
  558. 23:14solved, though not all of them had had
  559. 23:16in fact been solved by an AI system at
  560. 23:19that point. And so just like even just
  561. 23:22here, this is a very rapid trajectory
  562. 23:25from o over the course of just a little
  563. 23:28over a year going from, you know, like
  564. 23:32hobbyist high schooler to graduate
  565. 23:36student like very competent kind of, you
  566. 23:39know, up there graduate student in terms
  567. 23:41of the problems they could solve. Really
  568. 23:43rapid AI progress.
  569. 23:46At the same time, we started to see
  570. 23:48people test AI systems on unsolved math
  571. 23:52problems. Uh problem like the the
  572. 23:54frontier was no longer something. This
  573. 23:57is what I mean. We'd gone from exam to
  574. 23:59on the job. Uh where it wasn't like
  575. 24:03here's your qualifying exam as a grad
  576. 24:05student anymore. It became more
  577. 24:06interesting to say, well, here's a small
  578. 24:08problem that maybe a mathematician, a
  579. 24:10professional thought about for an hour,
  580. 24:12didn't solve AI systems. Can you do it?
  581. 24:15And at that level we started this was
  582. 24:19all just this year like 2026 at the
  583. 24:21beginning of the year there were like
  584. 24:22starting to be claims of oh maybe here's
  585. 24:24a problem that was posed by a famous
  586. 24:26mathematician this is maybe the airdos
  587. 24:29problems there there's this just for
  588. 24:30background there's this very prolific
  589. 24:32mathematician Paul Erdos uh who proposed
  590. 24:35over the course of his life maybe like
  591. 24:38well over a thousand maybe a couple
  592. 24:39thousand problems that people still like
  593. 24:41haven't fully cataloged all the
  594. 24:43questions he asked some of those became
  595. 24:44central to to whole fields of math. Some
  596. 24:47of them were, you know, no one really
  597. 24:49paid much attention to, including Erdos
  598. 24:51himself. And AI started to notch a few
  599. 24:53solutions to some of the easier ones of
  600. 24:57those. Fast forward to May uh of this
  601. 25:02year and and an internal model probably
  602. 25:05something like the Astra model that we
  603. 25:07now have publicly available was able to
  604. 25:10solve a big one of those. This was sort
  605. 25:12of a historic moment. may the the
  606. 25:14so-called unit distance problem uh from
  607. 25:18uh which was a problem about how many
  608. 25:21points you can fit in the plane that are
  609. 25:23just one dis unit distance of one apart
  610. 25:25from each other sort of like very cool
  611. 25:27that you could frame it it's such a
  612. 25:28simple problem but there hadn't been
  613. 25:30progress much progress on this problem
  614. 25:32despite a lot of attempts uh since had
  615. 25:35had posed it you know decades and
  616. 25:37decades before and it was a serious
  617. 25:38problem it wasn't one of these no one
  618. 25:40had thought about it it was definitely
  619. 25:41had received a lot of attention and an
  620. 25:43AI system solved it. And it was really
  621. 25:45this this was like a aha moment of oh my
  622. 25:48goodness like AI can solve problems that
  623. 25:51humans uh have cared about just
  624. 25:53independently
  625. 25:55and we've seen a lot more of that since
  626. 25:57and the the you know with a sort of the
  627. 25:59the recent wow moment being the um
  628. 26:02solution of the millennium prize
  629. 26:04problem. This was you know the Navier
  630. 26:06Stokes problem. This was a collection of
  631. 26:08this these Millennium Prize problems, a
  632. 26:10collection of seven problems that um
  633. 26:13mathematicians sort of or at least one
  634. 26:15math institute set as like good goals
  635. 26:17for the new century. Like they were they
  636. 26:19were they're much older than the year
  637. 26:222000, but but they were sort of
  638. 26:24canonized in the year 2000 as hey like
  639. 26:26these are worthy targets of entire sub
  640. 26:28fields of mathematics and it would be a
  641. 26:30big deal and so one of them was was
  642. 26:32solved. So it's been a hell of a hell of
  643. 26:34a ride. The historical context is is
  644. 26:37super interesting. It, you know, we're
  645. 26:39all living through this trajectory, but
  646. 26:42to step back and see, you know, how much
  647. 26:44has happened in the past couple of years
  648. 26:46is pretty pretty nuts.
  649. 26:49>> Yeah. You you have to, as one of my
  650. 26:51colleagues said, you have to like not
  651. 26:52get frog boiled by it like like the the
  652. 26:55the rising and it's easy to
  653. 26:58>> I don't know. Anyway, um so so there is
  654. 27:01this interesting gap so far. Like you
  655. 27:04might hear the phrase jagged frontier
  656. 27:07where the AI capabilities are jagged in
  657. 27:10the sense that they'll be superhuman in
  658. 27:12some on some axis and some you know not
  659. 27:16superhuman still inferior in capability
  660. 27:18to human in in other a on some other
  661. 27:21axis and
  662. 27:24that uh remains the case so far in math
  663. 27:28in particular there's a couple things in
  664. 27:30hindsight the analysis by mathematicians
  665. 27:33of the solutions AI has found to these
  666. 27:36like big important previously unsolved
  667. 27:39math problems has it's always in the
  668. 27:43first place it's been nothing too super
  669. 27:46duper surprising maybe there was like a
  670. 27:49a direction that humans had overlooked
  671. 27:53or could have invested more in and AI
  672. 27:57sort of had the persistence may maybe
  673. 27:59this is a good a good word or or or the
  674. 28:02bravery or the tearity to like wade into
  675. 28:06these very in-depth intricate
  676. 28:08computations and calculations and sure
  677. 28:10enough came out the other side with with
  678. 28:12a solution and humans look back and are
  679. 28:14like ah that's not a crazy direction to
  680. 28:17try at all and maybe if we'd like gone
  681. 28:21in that direction we we we would have
  682. 28:22come up with this this isn't something
  683. 28:24that like would be surprising
  684. 28:26>> what I'm curious about is
  685. 28:31are we seeing the AI AI systems come up
  686. 28:33with a solution to the problem or an
  687. 28:37algorithm to create the solution to the
  688. 28:41problem. And and I think I'm trying to
  689. 28:43like ask a lot of different things here,
  690. 28:46you know. One is like creating
  691. 28:48creativity, one is maybe um levels of
  692. 28:52abstraction or the way it's like, you
  693. 28:54know, thinking abstracting, you know,
  694. 28:56around these problems. Um yeah, I'll
  695. 29:00kind of pause and let you react to that.
  696. 29:02Yeah, let let me let me throw out a
  697. 29:04couple things there. Let me know if you
  698. 29:06want me to expand on any of them. When
  699. 29:09humans solve these math problems, they
  700. 29:11they write up the solutions in p in like
  701. 29:14papers, proofs uh that you know this is
  702. 29:16true or this is not true. And those
  703. 29:18proofs are are not usually
  704. 29:21like not exactly algorithms. They're
  705. 29:22written in natural language. They have
  706. 29:24sort of they leave the reader to fill in
  707. 29:27some of the gaps. and uh AI systems are
  708. 29:30coming up with just the same kind of
  709. 29:32thing. Uh like like it's it's the same
  710. 29:34products. There's a lot of complaints
  711. 29:36about their writing style, but like the
  712. 29:38arguments, the substance of the
  713. 29:39arguments really are the thing we expect
  714. 29:42to see from humans. such that if a human
  715. 29:45came up with this and yeah, maybe
  716. 29:46cleaned it up a little, I don't know, uh
  717. 29:49would like would be lauded for having
  718. 29:51like made a very insightful connection
  719. 29:54between two different areas or having
  720. 29:57gotten a an intricate um argument just
  721. 30:01right and balanced different
  722. 30:03considerations to make everything come
  723. 30:04together to to land the proof. 100% like
  724. 30:08this is not a this is something we're
  725. 30:10past the point where we can say, "Well,
  726. 30:13AI systems are sort of like not really
  727. 30:16doing something that humans would, you
  728. 30:18know, be impressed by on human terms.
  729. 30:20No, AI systems are doing doing things.
  730. 30:22And is there a way to categorize or
  731. 30:26qualitatively
  732. 30:28are we finding that and I think you
  733. 30:30spoke to this a little bit like are you
  734. 30:32know is there a qualitative difference
  735. 30:34between you know the best human approach
  736. 30:37and the AI approach in a sense of like
  737. 30:40you know AI brute forced it you kind of
  738. 30:43spoke to this versus like pulled in
  739. 30:46things from some other domain and like
  740. 30:48created some elegant solution or are we
  741. 30:50past that point as well
  742. 30:51>> here I' I'd give more nuance. Uh we we
  743. 30:55um there's an emerging sense of at least
  744. 30:59for the moment what an AI shaped problem
  745. 31:03looks like. Maybe I could say AI has two
  746. 31:06big advantages right now. One is it
  747. 31:08knows everything like it knows all the
  748. 31:11prior work. It's got the literature
  749. 31:12memorized. I'm speaking loosely but that
  750. 31:15this is this is basically uh valid. Um
  751. 31:19the other is it's extremely patient and
  752. 31:22persistent. So if there is some
  753. 31:24intricate uh you know calculation you
  754. 31:28have to do I don't mean like numerically
  755. 31:29or algorithmically. When I say
  756. 31:31calculation I just mean lots of
  757. 31:33equations and you have to make all the
  758. 31:35pieces work out just as very high level
  759. 31:38conceptual logical kind of computation
  760. 31:41but still something very um intricate.
  761. 31:44Uh they just have the patience to to go
  762. 31:47for that. And if you take the Navier
  763. 31:49Stokes example, I think it's not a
  764. 31:51coincidence that um OpenAI's reported a
  765. 31:56system that that solved it included a
  766. 31:58swarm of thousands of different
  767. 32:00instances of AI systems trying lots of
  768. 32:03different things and sharing ideas where
  769. 32:06clearly it was able to like there's an
  770. 32:08element of brute force there. So there's
  771. 32:11degrees of brute force. This is where
  772. 32:12the nuance is. It wasn't like somehow uh
  773. 32:16you know you you asked it like it wasn't
  774. 32:20solving this in by any means in the
  775. 32:22dumbest most low-level way like we're
  776. 32:24we're past that point as well. They used
  777. 32:26to solve some problems in like kind of
  778. 32:28surprisingly low-level like grinded out
  779. 32:31kind of ways. Like that's not really
  780. 32:32what these look like anymore, but like
  781. 32:35the the highle ver a higher level
  782. 32:38version of that still is maybe what what
  783. 32:40you might what you might say here. So,
  784. 32:42and I like the frame of precise and
  785. 32:45intricate computations here. It's
  786. 32:47important I'll say one more thing about
  787. 32:49the Navier stroke solution. It's
  788. 32:51important to recognize that humans sort
  789. 32:53of put in place a lot of the structure
  790. 32:55through previous work that AI then went
  791. 32:58and as far as we can tell, I mean,
  792. 33:00humans are still figuring out exactly
  793. 33:01what goes into the AI solution, but as
  794. 33:04far as we can tell, it builds on a lot
  795. 33:06of prior human work. The AI didn't like
  796. 33:08build it up from the ground. that kind
  797. 33:11of theory building as mathematicians
  798. 33:13call it we still don't see that very
  799. 33:15much from AI systems last point I'll
  800. 33:18make before pausing on this is just uh
  801. 33:21remember its point in time like remember
  802. 33:23that trajectory we've seen very rapid uh
  803. 33:26so so mathematicians who are saying we
  804. 33:28have to adapt our profession to the fact
  805. 33:30that now you can push a button and get
  806. 33:32what used to have been a career making
  807. 33:34result like you know are also I think
  808. 33:37cautioning and we can expect that
  809. 33:39capabilities will freeze Right now we we
  810. 33:41see you know that that linear whether
  811. 33:43it's linear or exponential is a matter
  812. 33:45of perspective but it's going up and
  813. 33:48it's like we see no reason to expect it
  814. 33:50to stop there. So this is a big thing
  815. 33:52epoch is looking for. Can AI move from
  816. 33:56intricate computations plus lots of
  817. 33:58background knowledge to more of a
  818. 34:01developing uh new ideas more from whole
  819. 34:05cloth. Uh and and if so and that's a big
  820. 34:08thing we're watching for in our in our
  821. 34:09own math benchmarking uh activities.
  822. 34:12Yeah, it's interesting to think that
  823. 34:17at least the way I characterized this as
  824. 34:19brute force versus pulling in things
  825. 34:21from other domains,
  826. 34:24those are um you know those are not
  827. 34:28mutually exclusive. like for for an AI
  828. 34:31that has access to, you know, all of the
  829. 34:33research from every domain, you know, an
  830. 34:36approach that could be very fruitful is
  831. 34:38to brute force the exploration of
  832. 34:41adjacent domains and pull that into the
  833. 34:43solution to a given problem.
  834. 34:45>> That's it. I think a lot of what we see
  835. 34:48and maybe just to say the the other
  836. 34:50thing that would really supercharge them
  837. 34:53and bring them closer to full human
  838. 34:56parody or even human you know dominating
  839. 34:58human capabilities would be something
  840. 35:00like coming up with a new idea. Uh and
  841. 35:03maybe like in calculus the idea of the
  842. 35:07derivative or the integral of a function
  843. 35:09like that was something that was you
  844. 35:11know maybe mathematicians in the 1500s
  845. 35:14had been like grasping for this idea and
  846. 35:16then Newton or Libnets or whomever comes
  847. 35:18up with like hey here's like I can I can
  848. 35:22describe a new thing a new mathematical
  849. 35:24concept and that unlocks a lot for us
  850. 35:26that lets us solve problems we
  851. 35:29previously couldn't solve. And so that's
  852. 35:31the sort of thing we don't yet see AI
  853. 35:34doing. It's also the sort of thing
  854. 35:35that's rarer for humans to do. Like a
  855. 35:37lot of human math progress has been
  856. 35:49proven it or I studied one area and I
  857. 35:52noticed something and I thought it might
  858. 35:54be applicable to this nominally
  859. 35:55superficially different area like humans
  860. 35:57do that sort of stuff all the time
  861. 35:58that's 90%
  862. 36:00vaguely speaking of human mathematical
  863. 36:02work from from what I hear from
  864. 36:04mathematicians the new idea stuff is
  865. 36:06rarer so maybe we should expect it to be
  866. 36:08not the first thing AI gets good at but
  867. 36:11so far we don't see AI really you know
  868. 36:13AI really doing that yet
  869. 36:16>> how easy or difficult is it to define
  870. 36:21a new idea in in such a way that um you
  871. 36:26know we can easily distinguish them like
  872. 36:29you know if we think about kind of the
  873. 36:30evolution of this discourse around like
  874. 36:33is it AI creative you know initially it
  875. 36:35was like yeah it's very creative it
  876. 36:37created this like poem in you know in
  877. 36:40pirate like that's creativity that's you
  878. 36:42know that was a new idea uh but that's
  879. 36:44certainly not what we're talking about
  880. 36:46when we're talking about like advanc
  881. 36:48advancing science and advancing
  882. 36:50mathematics etc. Um
  883. 36:54you so h a how do we like how do we
  884. 36:58define that but also you know from your
  885. 37:02perspective as someone who's studying
  886. 37:03this is it like you know we've got this
  887. 37:06long line of of folks that are coming
  888. 37:08that says here's this AI with a new idea
  889. 37:10and you're like yeah well let me
  890. 37:11evaluate it against my new idea criteria
  891. 37:13no it's not really that or are we
  892. 37:17actually like is it pretty clear to
  893. 37:19everyone that the AI is not having new
  894. 37:21ideas and like we're waiting for the
  895. 37:23first person to come and say that AI has
  896. 37:25the new idea. Like how do you think
  897. 37:27about that whole space?
  898. 37:29>> Very very challenging nut to crack. Uh
  899. 37:32so we you know try to approach it from
  900. 37:35different angles.
  901. 37:37One thing you can try to say is
  902. 37:42I've got a problem. I don't know if a
  903. 37:44new idea will be required to crack it.
  904. 37:46And this could be a problem in math or
  905. 37:48in science or or any quantitative area.
  906. 37:52But humans have tried to crack it.
  907. 37:54Humans have tried to make you know this
  908. 37:56number go up. Whether that number is the
  909. 37:58efficacy of some you know drug at
  910. 38:00fighting some disease or something like
  911. 38:02that or some uh you know or it's a math
  912. 38:05result uh of some conjecture that the
  913. 38:08math have stumped mathematicians for a
  914. 38:10long time. And you can say, "Okay, well,
  915. 38:12I don't know if I I'm not sure how I'm
  916. 38:14going to measure a new idea, but I can
  917. 38:15at least tell if this problem has been
  918. 38:17solved or this number has been made to
  919. 38:20go up or whatever the the concrete
  920. 38:23metric is." And you can then you're you
  921. 38:26can at least check if that has happened.
  922. 38:28Now
  923. 38:30I was probably a year ago more
  924. 38:31optimistic that there would be a tighter
  925. 38:34relationship between math problem that
  926. 38:37humans have failed to solve and post
  927. 38:40hawk qualitative assessment of
  928. 38:42creativity.
  929. 38:44Uh but uh but you know that this is so
  930. 38:47this is sort of the combination the like
  931. 38:49you know double punch we're we're trying
  932. 38:52here. First you set a target that's hard
  933. 38:54on human terms. The humans care about
  934. 38:56the humans have tried to be creative to
  935. 38:58solve and then you do post talk review
  936. 39:01with experts and say what do you make of
  937. 39:03this solution? Why hadn't you found it?
  938. 39:05No offense uh you know what what is in
  939. 39:08the AI solution that humans had missed
  940. 39:10and so on. And this is risky like like
  941. 39:13risky in terms of humans being honest
  942. 39:15about it because you know your ego might
  943. 39:17be wrapped up in it or you might once
  944. 39:19you see the solution it feels more
  945. 39:21obvious in hindsight. So would you have
  946. 39:22described it like very very risky uh but
  947. 39:26uh but you still you try your best and
  948. 39:28you know you might say things like well
  949. 39:31uh how like here's an interesting
  950. 39:33approach that that we haven't
  951. 39:34operationalized in like in detail uh but
  952. 39:36it's it's a nice heristic I think. How
  953. 39:38big a hint would a human have needed to
  954. 39:42solve this in in hindsight? Uh so for
  955. 39:45example that unit distance a very uh
  956. 39:48prominent mathematician fields medalist
  957. 39:49Timothy Gowers uh did this sort of
  958. 39:52exercise on this unit distance problem
  959. 39:54once AI had solved it saying how big a
  960. 39:57hint and how like specific of a hint
  961. 40:00would I do I think a mathematician would
  962. 40:02have needed now in hindsight like to
  963. 40:05solve this problem. uh one one yeah and
  964. 40:08he came away with like an interesting
  965. 40:09exploration of well actually AI sort of
  966. 40:12disproved this conjecture as opposed to
  967. 40:13proving it true it said actually it's
  968. 40:15false even that is a huge hit humans had
  969. 40:17mostly thought it was true they were
  970. 40:19wrong but they'd been trying to prove
  971. 40:20that it was true instead of looking for
  972. 40:22a counter example so even just the hint
  973. 40:24of try to find a counter example is
  974. 40:27pretty big and then there was like one
  975. 40:28more piece and he was like I think
  976. 40:30though very hard to say we can't really
  977. 40:32do the counterfactual experiment I think
  978. 40:34with just that with like this sort of
  979. 40:36two-part hint. It would have been
  980. 40:38possible for like a human would have
  981. 40:39been like, "Oh, oh, okay, I know what to
  982. 40:41do now, blah, blah, blah." And like
  983. 40:42probably would have at least hastened a
  984. 40:44human solution. So, you can do things
  985. 40:46like that maybe still very qualitative,
  986. 40:48but you try to be honest with yourself
  987. 40:50and be rigorous and you could imagine
  988. 40:51doing a more rigorous version of this
  989. 40:53experiment.
  990. 40:54And anyway, things like that. But then
  991. 40:56the the the fact is math is still for
  992. 40:59the most part a um
  993. 41:02sort of discipline removed from physical
  994. 41:05real world impact. Uh not always but in
  995. 41:08surprise its impact is often hard to
  996. 41:10predict. There are domains where impact
  997. 41:12is not at all hard to like I'm in the
  998. 41:13chemistry lab and I'm trying to improve
  999. 41:15the yield of my synthesis process. And
  1000. 41:17if AI, you know, makes incremental
  1001. 41:20improvements, I expect incremental
  1002. 41:21yield. And if AI suddenly comes and
  1003. 41:23says, no, no, no, you're doing it all
  1004. 41:25wrong. Let me like redesign your whole
  1005. 41:26thing from scratch. I've got like a, you
  1006. 41:28know, maybe creative idea. Well, at some
  1007. 41:29point, you don't care if it's creative
  1008. 41:31or not in the qualitative sense. You
  1009. 41:32just care that you have saw a
  1010. 41:34discontinuity in the yield of your
  1011. 41:36chemical substance or whatever. And
  1012. 41:38that's, you know, that's uh
  1013. 41:41that's a bigger bigger impact for the
  1014. 41:43world. Anyway,
  1015. 41:44>> what I took from that is that we filter
  1016. 41:46on the problem and how hard humans have
  1017. 41:48kind of bang their head heads against
  1018. 41:50the problem. Then we look at the results
  1019. 41:54that AI has come up with and try to kind
  1020. 41:58of subjectively say, you know, is this
  1021. 42:01based on a new idea or is it, you know,
  1022. 42:03more brute force or something else? like
  1023. 42:06um but it's kind of imperfect and messy
  1024. 42:09and like we're not really sure but we're
  1025. 42:12we don't think that we've come up with
  1026. 42:15you know great examples of AIs coming up
  1027. 42:17with uh you know new quote unquote new
  1028. 42:22ideas one that's exactly right one
  1029. 42:25reference point I'd give is way back in
  1030. 42:27uh gosh well the the teens I forget
  1031. 42:31which year the the famous go playing
  1032. 42:33system the game of
  1033. 42:35Alph Go. Uh that
  1034. 42:36>> I had that thought earlier as we were
  1035. 42:38discussing this like if I remember
  1036. 42:40correctly the
  1037. 42:43the
  1038. 42:44you know part of the solution was like
  1039. 42:47oh that was an innovative move that no
  1040. 42:49human ever would have done.
  1041. 42:50>> There was a specific one. Yeah. Exactly.
  1042. 42:52>> There was a very specific move in there.
  1043. 42:54Yeah.
  1044. 42:54>> Yeah. They call it move 37 because it
  1045. 42:57was the 37th move in this game where uh
  1046. 43:01Yeah. Apparently, this was a very
  1047. 43:03surprising move. Like you said, no human
  1048. 43:05would have done it um in this particular
  1049. 43:08position in this in this game of Go. And
  1050. 43:12I think it's like encouraging for our
  1051. 43:14admittedly quite subjective program here
  1052. 43:17that Go experts looked at and were like,
  1053. 43:19"Oh my goodness, like that's not, you
  1054. 43:21know, that's that's new." And then it
  1055. 43:23also became clear that that was pivotal,
  1056. 43:25that that move gave the AI system
  1057. 43:28playing Go a decisive advantage. And so
  1058. 43:31that's sort of what we're looking for in
  1059. 43:33in math. And at least, you know, maybe
  1060. 43:36we have some hope that humans will be
  1061. 43:37able to spot it when it happens. Uh but
  1062. 43:40but yeah, that's that's what we haven't
  1063. 43:42seen quite yet. And again, the
  1064. 43:45trajectory is fast. The things are
  1065. 43:48always a little more qualitative than
  1066. 43:50you know that you know than they are or
  1067. 43:52or continuous than discreet. Uh but you
  1068. 43:55know, so anyway, this is this is where
  1069. 43:57we are now, which is kind of wild in its
  1070. 43:59own right. Earlier when we were talking
  1071. 44:00about frontier math, you kind of
  1072. 44:03characterize it as, you know, one to
  1073. 44:05four, I think, and then advanced and we
  1074. 44:08never circled back to
  1075. 44:10>> more. We've got a couple we've got a
  1076. 44:11couple uh two new sets of frontier math
  1077. 44:14problems. We ditched the tier system,
  1078. 44:16whatever.
  1079. 44:17>> Okay.
  1080. 44:17>> Uh so, so now um these are
  1081. 44:21I'll lump them together because it
  1082. 44:23really is the uh the this is our our
  1083. 44:26main tool. So we have uh you know our
  1084. 44:29latest iteration on frontier math uh
  1085. 44:31consists of a bunch of problems that
  1086. 44:33humans have tried and failed to solve.
  1087. 44:35They're curated to be of a special
  1088. 44:37interest to mathematicians where a and
  1089. 44:41some of them we've characterized even on
  1090. 44:42a bit of a scale from like this would be
  1091. 44:46moderately interesting or this would be
  1092. 44:49like maybe what that means is
  1093. 44:52two teams of mathematicians have tried
  1094. 44:54this problem. It's at least 10 years
  1095. 44:55old, but it's not central to any
  1096. 44:58research program. It's just something
  1097. 44:59that people would be happy to get an
  1098. 45:01answer to all the way up to like a major
  1099. 45:03breakthrough. Four tiers here.
  1100. 45:05Breakthrough is the highest. Uh none of
  1101. 45:07the breakthroughs have been solved yet.
  1102. 45:08But don't get me wrong, if Navier Stokes
  1103. 45:10had been in this problem set, it would
  1104. 45:12surely be a breakthrough there. Um
  1105. 45:15regardless of the fact that humans had
  1106. 45:17like done a lot of work and gotten maybe
  1107. 45:19within spitting distance of the of the
  1108. 45:21solution before AI finished it off. Um
  1109. 45:24but yeah, we really view this as just a
  1110. 45:26a way of finding AI um a way of finding
  1111. 45:32cases where AI might have had a creative
  1112. 45:34idea because we can tell easily whether
  1113. 45:38it solved these problems. This is maybe
  1114. 45:39the the the big piece of work on our
  1115. 45:41side was to make it easy to verify
  1116. 45:44automatically whether an AI system has
  1117. 45:46solved this problem. It's not obvious
  1118. 45:47when you think about it that no human
  1119. 45:49knows the answer. So, how do you tell
  1120. 45:51without a human looking if AI has has
  1121. 45:53found the answer, but there's a couple
  1122. 45:55uh techniques you can use for this? And
  1123. 45:56and we employ different techniques for
  1124. 45:58this. Um, it can go into that, but it's
  1125. 46:00something of a technical detail. And so,
  1126. 46:03when we when a new model comes out,
  1127. 46:05Astra came out, we hit go and it's like,
  1128. 46:08oh, it solved three new three new
  1129. 46:10problems from our list. Okay, let's go
  1130. 46:11look at the solutions. Let's share them
  1131. 46:13with mathematicians. Let's see what's
  1132. 46:14going on here. Um, and and then we we
  1133. 46:16analyze that. We with the help of
  1134. 46:18mathematicians analyze that for like
  1135. 46:20what kind of qualitatively what kind of
  1136. 46:22solution was this?
  1137. 46:23>> What does it mean when the models are so
  1138. 46:26good that we have to shift from you know
  1139. 46:28these benchmarks that we understand to
  1140. 46:30like this collection these collections
  1141. 46:32of problems that we just don't have the
  1142. 46:33answers to. It's definitely a phase
  1143. 46:35transition in the the science of AI
  1144. 46:38capabilities measurement which which is
  1145. 46:41you know what what I work on but but it
  1146. 46:43doesn't change the um it doesn't change
  1147. 46:46the game all that much. It just means we
  1148. 46:48have to look for harder problems. Uh I I
  1149. 46:51think I joke sometimes that even certain
  1150. 46:53forms of optimization problems like
  1151. 46:57allocating scarce resources efficiently
  1152. 46:59that's a benchmark of sorts that will
  1153. 47:02survive the singularity like even if AI
  1154. 47:04is like off going crazy running rampant
  1155. 47:07in the galaxy it'll still have to you
  1156. 47:08know it'll still have to decide how to
  1157. 47:11you know optimize. Uh so optimization is
  1158. 47:14uh is always going to be a benchmark
  1159. 47:16even if we're superhuman. Basically, it
  1160. 47:18used to just be we knew the answers
  1161. 47:19ahead of time. Now we don't. So, it
  1162. 47:21takes this extra trick of okay, how do
  1163. 47:23we tell when AI has gotten something
  1164. 47:25right, but this isn't uh this isn't
  1165. 47:28insurmountable. Especially like I was
  1166. 47:30saying with um you know, numbers in a
  1167. 47:32scientific context. It's intuitive that
  1168. 47:34okay, the best humans were able to do
  1169. 47:35was, you know, X, but we just say to AI,
  1170. 47:39hey, try to do Y greater than X and you
  1171. 47:42know, if it does, that's, you know,
  1172. 47:44those numbers go on and on. So, it's
  1173. 47:46it's no problem. But yeah, you're right
  1174. 47:48to sort of note the the phase shift and
  1175. 47:51to be somewhat surprised by it yet.
  1176. 47:53>> And is there a methodology, a concrete
  1177. 47:55methodology for assessing and awarding
  1178. 47:59partial credit? Is that part of the way
  1179. 48:00you think about uh benchmarks in this uh
  1180. 48:04you know postphase shift?
  1181. 48:06>> Yeah, it's a good question. It depends
  1182. 48:07on the benchmark. So sometimes it's
  1183. 48:09something like well you just want to
  1184. 48:11make this number go up and the higher
  1185. 48:12the better. So partial credit is very
  1186. 48:15continuous and very very natural in
  1187. 48:17those cases for math problems. Sometimes
  1188. 48:19it's just like well what we care about
  1189. 48:21is is you know if x does y always hold
  1190. 48:25like like is this is this always true
  1191. 48:27and then it's kind of sometimes humans
  1192. 48:30will find and publish partial results
  1193. 48:32for the AI systems. We usually just say,
  1194. 48:34you know, all or nothing, but we have a
  1195. 48:36large enough collection of such
  1196. 48:38conjectures that, you know, we hope to
  1197. 48:40have some, you know, smooth gradient of,
  1198. 48:42well, okay, this one got two more out of
  1199. 48:44100 or something. And so, so it uh, you
  1200. 48:47know, looks like a pretty smooth signal.
  1201. 48:49Um, that that's typically how we how we
  1202. 48:51approach these things. Uh, yeah, across
  1203. 48:54benchmarks. Yeah, maybe what I was
  1204. 48:56thinking was, you know, once you're in
  1205. 48:58the regime of very challenging unsolved
  1206. 49:01problems, is it possible that AI
  1207. 49:06systems are doing, you know, surprising
  1208. 49:08or novel or creative things but still
  1209. 49:10not quite getting them to the full uh,
  1210. 49:14you know, the full solution and do we
  1211. 49:15care about that? Do we want to to track
  1212. 49:18that and is that part of the way you've
  1213. 49:20constructed the benchmark?
  1214. 49:21>> It's a great question. It's not central
  1215. 49:23to how we construct a benchmark. It's
  1216. 49:25probably like it's certainly worth
  1217. 49:27tracking. The the problem is it's just
  1218. 49:28labor intensive. Uh like you could have
  1219. 49:30AI review the AI. We we do this
  1220. 49:32occasionally say like which problems did
  1221. 49:35it make partial progress on? What if we
  1222. 49:36let it think for longer on those
  1223. 49:38problems and nothing yet like it doesn't
  1224. 49:41seem good at understanding how close it
  1225. 49:43is to a solution. Humans aren't
  1226. 49:45necessarily very good at this either. I
  1227. 49:47did want to mention that while we've
  1228. 49:49focused on maybe the area where AI is
  1229. 49:53the very best like has come the farthest
  1230. 49:56the fastest math we do see lots of areas
  1231. 49:59where we also think it's important to
  1232. 50:02measure AI capabilities where it's um
  1233. 50:06where the progress is somewhat fuzzier
  1234. 50:08uh less clear and certainly less
  1235. 50:10superhuman or even in even in some well
  1236. 50:14may maybe still superhuman in some ways
  1237. 50:17But in are
  1238. 50:20less rarified.
  1239. 50:22>> What are some examples?
  1240. 50:24>> Couple couple examples.
  1241. 50:26One area that uh forthcoming work we're
  1242. 50:28excited about is just it's a question
  1243. 50:31everyone must be wondering can AI take
  1244. 50:33my job? Uh so our our our job at Epoch
  1245. 50:37uh often involves many things and uh all
  1246. 50:40sorts of research not just AI
  1247. 50:42benchmarking but other topics about
  1248. 50:44trends in AI and uh writing research
  1249. 50:47reports and doing data analysis making
  1250. 50:49infographics that are you know make a
  1251. 50:52point in an easy to digest way. uh we've
  1252. 50:55been collecting examples of these from
  1253. 50:58within epoch where the judgment at the
  1254. 51:02end of did the AI system manage to do
  1255. 51:04this you know epoch task um did is too
  1256. 51:09messy for traditional benchmarking. So
  1257. 51:12it wouldn't be obvious like how to say
  1258. 51:14did it get the right answer or not. It's
  1259. 51:16not about making a number go up. It's
  1260. 51:18not about satisfying some logical chain
  1261. 51:20of deductions. It's like here's the
  1262. 51:22infographic to go along with this
  1263. 51:24report. Like does it does it meet our
  1264. 51:27style guide and does it look good and
  1265. 51:29does it you know is it better or worse
  1266. 51:31than the uh the the human reference uh
  1267. 51:34that that we you know one of my
  1268. 51:36colleagues did. Um and here we see like
  1269. 51:40an interesting mix of uh some things I
  1270. 51:42think we we sort of all our experience
  1271. 51:45is like yeah if you ask it for a summary
  1272. 51:47of the literature it's going to do a
  1273. 51:48pretty good job. that's like very good
  1274. 51:49at that. But but there's other tasks
  1275. 51:50where it's um you know hard like
  1276. 51:53creating compelling visuals that follow
  1277. 51:56our style guide or this is a good one
  1278. 51:58coming up with a new um like uh short
  1279. 52:04form datadriven insight. Uh we call
  1280. 52:07these data insights. You can find them
  1281. 52:08on our website. And uh these like often
  1282. 52:12we try to boil down a single like
  1283. 52:14important observation into a single
  1284. 52:16chart and a single sentence. Uh and this
  1285. 52:19is like this is hard but we don't see
  1286. 52:21them doing a great job at this come up
  1287. 52:24with a new project for epoch to
  1288. 52:26undertake and do a prototype very
  1289. 52:28open-ended. not so good at this like
  1290. 52:31like these more open-ended tasks. We we
  1291. 52:33do see again my guess is if we we'll do
  1292. 52:36this go back in time some and see have
  1293. 52:38models gotten better at this my guess is
  1294. 52:40we'll see progress but the level
  1295. 52:43compared to like humans is still a
  1296. 52:45little um still a little meh. I want to
  1297. 52:48give one other example uh um which is
  1298. 52:51learning on the fly. This is a case
  1299. 52:54where this is another case relevant for
  1300. 52:56can AI take my job because you didn't
  1301. 52:58start you you weren't great at your job
  1302. 53:00on day one but as you practiced you know
  1303. 53:03you got better and humans you know have
  1304. 53:05this sort of you know they learn on the
  1305. 53:07fly uh we take a we like to measure this
  1306. 53:10for AI we have a a benchmark that's
  1307. 53:13taking a complicated board game we did a
  1308. 53:15human study for saying how many times do
  1309. 53:18humans have to play this board game in a
  1310. 53:19row before they sort of figure out how
  1311. 53:21it works and can get a very high score
  1312. 53:23on it. And we do the same with AI. We
  1313. 53:25say, "Play it once, take notes." Same AI
  1314. 53:28instance like, "Play it again, take
  1315. 53:30notes, play it again, take notes, and
  1316. 53:32see how you can do." AI systems have
  1317. 53:34gotten good at this board game, but not
  1318. 53:36by playing it repeatedly. It's more of a
  1319. 53:38intergenerational thing. Like, GPT 5.6
  1320. 53:42scores X out of the box and then stays
  1321. 53:44flat on its playthroughs. GPT6
  1322. 53:47scores much higher. It's like much
  1323. 53:48better at it, but it's still out of the
  1324. 53:50box and then it stays flat in its
  1325. 53:51playthroughs. it is still less than the
  1326. 53:53top human. So that this kind of uh on
  1327. 53:56the-fly learning is something we also
  1328. 53:57don't see AI systems as being good
  1329. 54:00enough at even to just pick up a board
  1330. 54:01game, let alone a messier uh job, you
  1331. 54:05know, real world job like that. So these
  1332. 54:07are sorts of things where we don't see
  1333. 54:08these capabilities and we think it's
  1334. 54:09important to try to catch when they
  1335. 54:12emerge. My gut on the the ladder, the AI
  1336. 54:16learning on the fly is that it's maybe a
  1337. 54:19harness solvable challenge versus a
  1338. 54:22model solvable challenge and maybe you
  1339. 54:25know no one's focused on that like with
  1340. 54:28the right set of memory structures and
  1341. 54:31like some kind of sidecar heristicy
  1342. 54:34thing like you know we can probably make
  1343. 54:37a big step improvement in
  1344. 54:40an AI's ability to like you know improve
  1345. 54:44game to game but it's not
  1346. 54:47uh it might not require dramatic model
  1347. 54:50changes. Do you do you think about that
  1348. 54:53and is is do you have you know in this
  1349. 54:57domain or in other domains like formal
  1350. 55:00or structured ways of thinking about you
  1351. 55:03know what uh you know model versus
  1352. 55:06harness you know granted this whole
  1353. 55:08model versus harness thing in and of
  1354. 55:10itself is you know practically brand new
  1355. 55:13um but it's proving to be important in
  1356. 55:17you know distinguishing you know where
  1357. 55:19the these improvements come from
  1358. 55:22>> Yeah, definitely. Uh
  1359. 55:26we've done some experiments on this for
  1360. 55:28the board game case in particular where
  1361. 55:30we've tried uh a very simple harness
  1362. 55:34versus the first party harnesses like
  1363. 55:36Clawude Code or Codeex. We've tried
  1364. 55:38multi- aent setups. We've tried like we
  1365. 55:42we give them a you know note-taking
  1366. 55:44tools and say you know make it clear
  1367. 55:46like okay you you're your your goal is
  1368. 55:48to learn the game so you can do really
  1369. 55:50well like in your final plays of the
  1370. 55:52game so explore take notes figure like
  1371. 55:54we try to prompt them well I I mean we
  1372. 55:57want to find good results if we can
  1373. 55:59because it's like that's a big deal for
  1374. 56:00the world the game as a test bed but
  1375. 56:03real jobs would also you know benefit
  1376. 56:06significantly risks would also go up if
  1377. 56:08this capability you know you test in the
  1378. 56:10lab. Is it a great hacker? You find
  1379. 56:12maybe it is, maybe it isn't, but you
  1380. 56:14find that it isn't. And yet it can learn
  1381. 56:16on the fly. Then okay, once you deploy
  1382. 56:17it, like look out. Um, so anyway, we we
  1383. 56:20try we've tried pretty hard and have
  1384. 56:22mostly not found anything that cracks
  1385. 56:25this uh that cracks this on the
  1386. 56:27learning. There's been some general
  1387. 56:29research about this as well, not just us
  1388. 56:31about uh you know people talk like there
  1389. 56:33was a people still use these so-called
  1390. 56:36skills where you like just write some
  1391. 56:38notes to an AI system of look when you
  1392. 56:41do work I ask you to do you're going to
  1393. 56:43need to use this tool or produce this
  1394. 56:45kind of output and like here's my notes
  1395. 56:47on how you could do that well so you
  1396. 56:49don't have to like learn it on the like
  1397. 56:51pick it up fresh every time so use these
  1398. 56:53skills and these show some improvement
  1399. 56:55but then there's like a plateau like
  1400. 56:57pretty quickly At
  1401. 56:59least in the research I'm familiar with
  1402. 57:01like you can get some improvements. My
  1403. 57:03sense is this is very loosely speaking
  1404. 57:06it's like lowanging fruit. Like if
  1405. 57:08you're telling it, oh, you're going to
  1406. 57:09have to navigate this like finicky
  1407. 57:11website that my company makes me use and
  1408. 57:14like just you're going to get stuck and
  1409. 57:15like you have to look in this menu and
  1410. 57:16that's where the button you're looking
  1411. 57:17for. Like straightforward advice that's
  1412. 57:20like obviously useful. But when it comes
  1413. 57:22to higher level strategic uh thinking,
  1414. 57:25kind of hard to bake that into the
  1415. 57:27harness where or or or like where where
  1416. 57:30the the the board game plays like like
  1417. 57:33has this um you you know sort of brings
  1418. 57:36this out where it's like there's not any
  1419. 57:39perfect advice that like you have to
  1420. 57:41find the advice for yourself in a sense
  1421. 57:44like you have to learn the game like so
  1422. 57:46I will say if we give it extremely
  1423. 57:47detailed strategy guides for here's how
  1424. 57:50you beat this like here's everything you
  1425. 57:52could possibly want. Like it's a cheat
  1426. 57:54by humans to like no one would play the
  1427. 57:55game that way. They do very they do very
  1428. 57:58well. They do much better than if they
  1429. 58:00basically have to find
  1430. 58:02>> do they right creating the strategy
  1431. 58:04guide
  1432. 58:04>> they can't create the strategy guide for
  1433. 58:06themselves which is like pretty
  1434. 58:07interesting. Uh so so like we do expect
  1435. 58:10eventually they'll get better at this
  1436. 58:12but it's again it's a big deal if that
  1437. 58:14if that happens because now it means
  1438. 58:17okay go back to that first benchmark
  1439. 58:20suite. Okay, they're not great at making
  1440. 58:22charts for us, uh, or coming up with
  1441. 58:25data insights for us, uh, you know,
  1442. 58:27without guidance, uh, but now put them
  1443. 58:30in on the team for a little while and
  1444. 58:33let them learn on the fly and maybe they
  1445. 58:35do become good. So you sort of see these
  1446. 58:36as a combination like if on our board
  1447. 58:39game where it's easy to measure they're
  1448. 58:41not getting better over time then it's
  1449. 58:42okay for us over in the real world like
  1450. 58:45human in the loop evaluation uh studies
  1451. 58:48for us to just do one try and say okay
  1452. 58:50well they weren't great and we don't
  1453. 58:51think they learn on the so that you know
  1454. 58:53but but really you know those go
  1455. 58:55together.
  1456. 58:55>> And is the board game a synthetic one
  1457. 58:58that you created for this challenge or
  1458. 59:01is it a a real board game that that
  1459. 59:03people play? The real board game, it's
  1460. 59:05called uh the first one we did this
  1461. 59:07with, we'll do this with more uh is
  1462. 59:09called Earthborne Rangers. Um it's
  1463. 59:11fairly obscure. If you go to like board
  1464. 59:13gamegeeek.com,
  1465. 59:15it's like 500th most popular or
  1466. 59:17something. So, it's not like up there.
  1467. 59:19There's a small devoted community to it.
  1468. 59:21Um and AI systems don't seem to know
  1469. 59:24that much about it. Like, they've heard
  1470. 59:26of it. They can give you some facts
  1471. 59:27about it. Uh but it doesn't seem like
  1472. 59:30something they've been trained on
  1473. 59:31specifically. It's impossible for us to
  1474. 59:33tell this with with high certainty with,
  1475. 59:35you know, high confidence, but but
  1476. 59:36that's sort of why we chose it. Um, it's
  1477. 59:39just hard to make these things from
  1478. 59:40scratch and know whether they're hard or
  1479. 59:42whether they're broken. Uh, you know, so
  1480. 59:44that's why we but but the risk which you
  1481. 59:46might have in mind is what if AI systems
  1482. 59:48like, you know, just know about it from
  1483. 59:50their training data. It's a risk on our
  1484. 59:52mind. It's just a risk. One thing we
  1485. 59:54will do in the future is we're moving on
  1486. 59:56to video games. Incidentally, it's
  1487. 59:58another area where AI is like very much
  1488. 1:00:00not like human reaction time and visual
  1489. 1:00:02and spatial reasoning again coming along
  1490. 1:00:05fast but not uh not really there yet. Um
  1491. 1:00:09that so video games you can always test
  1492. 1:00:11on a new game like a game that just came
  1493. 1:00:13out uh not in the training data pro
  1494. 1:00:15probably um or at least you know not as
  1495. 1:00:18robustly in the training data. Uh and so
  1496. 1:00:20we'll be testing can AI systems do well
  1497. 1:00:22on those? Do they do better on like, you
  1498. 1:00:24know, the older version of a video game
  1499. 1:00:27than like, you know, Grand Theft Auto 4
  1500. 1:00:30verse 5 or whatever, like when GTA 6,
  1501. 1:00:33whatever it is, comes out. Like, I don't
  1502. 1:00:35think we we'll be doing this one, but
  1503. 1:00:36but you know, are they worse at that?
  1504. 1:00:38>> Yeah. It's interesting when you when you
  1505. 1:00:41bring up video games, it immediately
  1506. 1:00:43calls to mind kind of the classical
  1507. 1:00:45reinforcement learning approaches, which
  1508. 1:00:48is kind of learning on the fly, but kind
  1509. 1:00:51of different.
  1510. 1:00:51>> Yeah. We we call the second sometimes in
  1511. 1:00:53context learning where the only like
  1512. 1:00:56lever the AI system really has is
  1513. 1:00:57managing its context and like look when
  1514. 1:01:00we give it the strategy guides and they
  1515. 1:01:02do much better clearly this is a
  1516. 1:01:03powerful lever but it's not as powerful
  1517. 1:01:06as updating the weights and like so yeah
  1518. 1:01:08I I mean like what's the big picture
  1519. 1:01:10here it's it's if you have this fast
  1520. 1:01:12loop of in context on the fly learning
  1521. 1:01:16uh then AI capabilities might sort of
  1522. 1:01:18improve much more rapidly than this
  1523. 1:01:20slower loop of uh collecting more
  1524. 1:01:23training data, doing more reinforcement
  1525. 1:01:24learning uh which which is done on
  1526. 1:01:27language models as well these days um
  1527. 1:01:30and then deploying a new model. Now AI
  1528. 1:01:32might automate that whole loop but it'll
  1529. 1:01:34still be slower than if it can just you
  1530. 1:01:36know work out its uh its own context on
  1531. 1:01:39the fly. Yeah.
  1532. 1:01:40>> Yeah. It it's interesting when you zoom
  1533. 1:01:42out on some of the the questions that
  1534. 1:01:45you focus on and we haven't talked
  1535. 1:01:47explicitly about it, but you have a blog
  1536. 1:01:49post where you kind of uh write out
  1537. 1:01:53these nine big hairy questions that
  1538. 1:01:56you're tracking. They kind of group into
  1539. 1:01:59like how do I assess where AI is now?
  1540. 1:02:02That's like can it do my job? Like how
  1541. 1:02:04is it, you know, competitively?
  1542. 1:02:07uh and you know where is the puck going
  1543. 1:02:10like can it learn on the fly you know
  1544. 1:02:12can it do research can it come up with
  1545. 1:02:14new ideas these are all it is kind of
  1546. 1:02:16where it is now but also like what what
  1547. 1:02:20you know is it positioned to like
  1548. 1:02:22dramatically
  1549. 1:02:24um I don't know self-improve or like you
  1550. 1:02:28know what's the trajectory
  1551. 1:02:29>> that's exa that's exactly right the um I
  1552. 1:02:33mean I think for for people thinking
  1553. 1:02:36more broadly about what they need to
  1554. 1:02:38know about AI that this is the second
  1555. 1:02:41topic is uh you know just as important
  1556. 1:02:44as the first uh but it's really
  1557. 1:02:49you know even if you take some of the
  1558. 1:02:51recent incidents like the hugging face
  1559. 1:02:52incident like the sort of think of two
  1560. 1:02:55axes
  1561. 1:02:56alignment of is the AI doing what I want
  1562. 1:02:59it to do holistically obviously hacking
  1563. 1:03:02into some other company's computers is
  1564. 1:03:04not well aligned uh and then
  1565. 1:03:05capabilities just like what can it get
  1566. 1:03:08done? uh can it hack into uh you know
  1567. 1:03:11reasonably well secured other companies
  1568. 1:03:13uh computers
  1569. 1:03:15um what that incident was was like an
  1570. 1:03:18early example very like I think pretty
  1571. 1:03:20robust example of well we don't have the
  1572. 1:03:23alignment thing solved and cap and
  1573. 1:03:25capabilities are getting so big so
  1574. 1:03:28strong that this is becoming a problem
  1575. 1:03:31uh and so paying attention to
  1576. 1:03:32capabilities like what we should expect
  1577. 1:03:35is becoming more important sort of
  1578. 1:03:37dayto-day today and we're just trying to
  1579. 1:03:39produce benchmarks that both say where
  1580. 1:03:42are we now exactly like you said like
  1581. 1:03:44what can it do now this is maybe
  1582. 1:03:47relevant for more you know mundane
  1583. 1:03:49economic utility or whatever as well as
  1584. 1:03:53uh you know do we see the the dynamics
  1585. 1:03:56of the system changing in a way that
  1586. 1:03:58might be alarming basically
  1587. 1:04:00>> you're kind of also thinking about it as
  1588. 1:04:01like first and second derivative of
  1589. 1:04:03capability
  1590. 1:04:04>> yeah I think that's right some of these
  1591. 1:04:06things make it uh the slope goes, you
  1592. 1:04:10know, much uh
  1593. 1:04:12faster uh hyperbolic growth or whatever
  1594. 1:04:15in some of these cases versus merely
  1595. 1:04:17just like, yep, it's uh it's growing and
  1596. 1:04:19it continues to grow. Both are
  1597. 1:04:21important. I I mean I think we've seen
  1598. 1:04:23the disruption we've seen to date has
  1599. 1:04:26come from these kind of unpredictable
  1600. 1:04:29thresholds like when you know GPT2 GPT3
  1601. 1:04:33suddenly you had the chat GPT moment and
  1602. 1:04:36you like oh this thing I want to talk to
  1603. 1:04:37this thing I might use this thing
  1604. 1:04:38instead of a search engine this thing is
  1605. 1:04:40useful to me for finding information and
  1606. 1:04:42then sudden and then like a you know a
  1607. 1:04:43little while years later whatever two
  1608. 1:04:46years later you had the it's getting
  1609. 1:04:48better slowly and steadily at helping me
  1610. 1:04:50with coding or with like using my
  1611. 1:04:52computer but suddenly
  1612. 1:04:53>> thinking models
  1613. 1:04:54>> with the thinking models then you had
  1614. 1:04:56the clawed code moment where it was like
  1615. 1:04:58oh suddenly this thing can actually do
  1616. 1:05:00software projects for me even if I'm not
  1617. 1:05:02a very not an engineer myself uh and
  1618. 1:05:05it's like hard to predict when these
  1619. 1:05:06thresholds will be crossed so it's
  1620. 1:05:08useful to still track but like the it's
  1621. 1:05:10useful to track the mundane capabilities
  1622. 1:05:12today and say look you know cyber is
  1623. 1:05:14maybe the latest example whereas like
  1624. 1:05:16look we see their ability in controlled
  1625. 1:05:18settings to do hacking going up uh
  1626. 1:05:22smoothly and at some point it's going to
  1627. 1:05:23cross some threshold where oh shoot it
  1628. 1:05:25can hack hugging face or whatever. Hard
  1629. 1:05:28to predict where those thresholds are
  1630. 1:05:29but still very useful to um to track
  1631. 1:05:32those capabilities. So so even the first
  1632. 1:05:34derivative as as you said version of
  1633. 1:05:36this we find very useful. We've got it's
  1634. 1:05:40a a fun benchmark we hope hope to
  1635. 1:05:42release uh maybe by the time this is
  1636. 1:05:44published we we'll we'll have it out on
  1637. 1:05:46a furniture assembly where uh we you we
  1638. 1:05:49build some IKEA furniture make some
  1639. 1:05:51mistakes and see if the AI system can
  1640. 1:05:52like just given photos like say oh like
  1641. 1:05:55wait you made a mistake in step back in
  1642. 1:05:57step four like stop go back um and
  1643. 1:06:00they've you know we see the same smooth
  1644. 1:06:02capabilities like Astra GPT6 Astra
  1645. 1:06:04pretty good at this actually uh and you
  1646. 1:06:07know but it wasn't and out of nowhere it
  1647. 1:06:08was like from a smooth and this is you
  1648. 1:06:10know important because we haven't seen
  1649. 1:06:13AI have much of an impact in physical
  1650. 1:06:15industry yet like it's been mostly
  1651. 1:06:17digital uh and even before robotics hit
  1652. 1:06:19the scene hits the scene we we might
  1653. 1:06:22expect that you know claude in your
  1654. 1:06:24glasses or something is saying like hey
  1655. 1:06:26you know I'm going to guide you through
  1656. 1:06:28repairing your car or whatever I'm going
  1657. 1:06:29to help you fix this machine in the
  1658. 1:06:30factory that broke so so you know we
  1659. 1:06:32want to track those capabilities too ju
  1660. 1:06:34just for the mundane impact anyway
  1661. 1:06:36impact all over the coming benchmarks.
  1662. 1:06:39Help us uh help us say so.
  1663. 1:06:41>> Awesome. Awesome. Well, Greg, thanks so
  1664. 1:06:43much for jumping on and sharing a bit
  1665. 1:06:45about what you're working on. It's super
  1666. 1:06:47interesting stuff and uh very
  1667. 1:06:49thoughtprovoking way of of thinking
  1668. 1:06:51about, you know, AI and where it is,
  1669. 1:06:54where it's going.
  1670. 1:06:55>> Thanks, Sam. Really appreciate it.
  1671. 1:06:57>> Awesome. Thank you.

About this transcript

This page contains the full transcript of From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? by The TWIML AI Podcast with Sam Charrington, generated from the public captions YouTube serves with the video. The transcript has 11,814 words across 1,671 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.