YouTube2Text

Jeff Dean: The 1% Rule for Building in AI — Transcript

by Y Combinator · 9,671 words · 1,435 segments · language en · Watch on YouTube

Full transcript

  1. 0:07All right. Should we go Should we get
  2. 0:09started, Jeff?
  3. 0:10>> Sure. Sounds great.
  4. 0:11>> All right. Jeff, welcome. And again,
  5. 0:13thank you so much for being here.
  6. 0:14Especially I just got a cold and thank
  7. 0:17you for being here.
  8. 0:17>> Yeah, I'm afraid I've lost my voice. I
  9. 0:19don't normally sound quite like this,
  10. 0:20but we'll we'll do what we can.
  11. 0:24>> So, um, you built map reduce, big table,
  12. 0:27tensorflow, the TPU, Gemini. We could
  13. 0:30spend a whole hour on all the things
  14. 0:32you've done, but what I love is that
  15. 0:34you're still making bold predictions in
  16. 0:36public. Last year, yes, last year in May
  17. 0:422025 at AI Ascent, you said that AI is
  18. 0:46at the level of a junior engineer.
  19. 0:50That was about a year ago. It's been How
  20. 0:55close are we to that prediction? Yeah, I
  21. 0:58mean I feel like uh the models have been
  22. 1:00getting a lot better at sort of
  23. 1:01agent-based longer running coding tasks
  24. 1:04and it seems pretty clear that they are
  25. 1:06now actually pretty capable and
  26. 1:08depending on exactly your definition of
  27. 1:10of junior engineer it seems pretty
  28. 1:12spot-on I would say.
  29. 1:15>> What did you underestimate from that
  30. 1:18prediction?
  31. 1:19Um I mean I think the
  32. 1:24the ability to do more and more complex
  33. 1:26tasks has been growing faster than I
  34. 1:28thought. Um and I also think uh outside
  35. 1:32of coding these these agent-based
  36. 1:34systems are are really starting to shine
  37. 1:36in other domains and I I think uh you
  38. 1:39know uh that's that's going to be an
  39. 1:41important trend in the future.
  40. 1:44>> So [snorts] give us another bold
  41. 1:47prediction. What do you think is going
  42. 1:48to be the 2027 edition?
  43. 1:51>> Uh I think you will see a lot more
  44. 1:54automation of uh ML systems themselves.
  45. 1:58um basically getting ML systems to
  46. 2:00improve their capabilities by running
  47. 2:03lots of experiments, breaking things
  48. 2:04down into subpros, you know, running
  49. 2:07those subpros in a tight automatic
  50. 2:09experimentation loop, putting the
  51. 2:11results together and being able to then
  52. 2:14uh you know, get some improved system uh
  53. 2:17out from that uh sort of fully automated
  54. 2:20problem decomposition and and automated
  55. 2:23experimentation that I think that's
  56. 2:24going to be really exciting. M
  57. 2:25>> I think that also applies not just to ML
  58. 2:28but also to other fields of science and
  59. 2:30engineering. Um basically anything where
  60. 2:32you can have a measurable objective uh I
  61. 2:35I think you can uh actually make a lot
  62. 2:37of progress these days.
  63. 2:40>> Now let's go back to a little bit in
  64. 2:42history. Back in way back in 2001 Google
  65. 2:47search used to run on hard drives.
  66. 2:49>> Yep. And you and Sanjay did the math and
  67. 2:53realized that at some point the whole
  68. 2:55search index would finally fit in all of
  69. 2:58the RAM of all the computers you had
  70. 3:02running
  71. 3:03and you made that radical realization
  72. 3:07and you basically in few days with
  73. 3:09Sanjay shipped in production a whole new
  74. 3:12search version that worked in RAM rather
  75. 3:14than hard drive and that was the thing
  76. 3:17that got Google to be so fast. Google
  77. 3:19searches.
  78. 3:21>> So history tends to remix.
  79. 3:25What is the it fits the memory moment
  80. 3:28right now in 2026 that everyone in this
  81. 3:31room is still
  82. 3:34should be thinking about and designing?
  83. 3:36>> Yeah. Yeah, I mean it's a little
  84. 3:38different, but I think uh you're going
  85. 3:40to see more and more
  86. 3:43uh uh high performance and um low energy
  87. 3:48uh inference hardware systems because I
  88. 3:51think everyone is now realizing that
  89. 3:53inference is the key to making you know
  90. 3:55these agent-based systems be available
  91. 3:58to more and more people and that latency
  92. 4:01is really important and that
  93. 4:03specialization of the hardware is a
  94. 4:05really key way you can make uh things
  95. 4:07that are more energy efficient and lower
  96. 4:10latency than more general purpose uh
  97. 4:12computational devices like say GPUs or
  98. 4:14TPUs
  99. 4:16>> because I think all everyone here is
  100. 4:18used to waiting for responses on on
  101. 4:21models. [clears throat]
  102. 4:22So
  103. 4:22>> waiting is no fun
  104. 4:25>> master speed.
  105. 4:27So you're saying what if we don't have
  106. 4:29to wait anymore?
  107. 4:30>> Yeah. I mean, I think we'll imagine what
  108. 4:32you could do with something where the
  109. 4:34latency is, you know, 50x better.
  110. 4:38>> Interesting thought.
  111. 4:40Now, what's one assumption that perhaps
  112. 4:436,000 people in this room hold that's
  113. 4:46already false about AI?
  114. 4:50>> Yeah. Uh, that's that's a good question.
  115. 4:51I mean I think um
  116. 4:55probably one thing is people don't quite
  117. 4:57realize how possible it is to have you
  118. 5:00know agent-based systems that can run
  119. 5:02not just for an hour or two hours on a
  120. 5:05problem you care about but for some
  121. 5:08problem domains and with highly capable
  122. 5:09models underlying them you can get them
  123. 5:12to run for days or weeks and do really
  124. 5:15really complicated tasks and I think
  125. 5:17that's you know starting some people are
  126. 5:20starting to see inklings of this but I
  127. 5:22don't think everyone has really
  128. 5:23internalized this and that's going to be
  129. 5:25really uh a pretty big deal.
  130. 5:28>> What's a particular task that you have
  131. 5:29run that has run for weeks? What what
  132. 5:32was it? What did the tell what did you
  133. 5:34tell the agents to solve? Yeah, I mean I
  134. 5:36think uh you can tell agents to uh go
  135. 5:40off and implement
  136. 5:43um you know completely new versions of
  137. 5:46software in different programming
  138. 5:47languages that might be you know have
  139. 5:49better safety properties or better
  140. 5:51performance properties uh that and then
  141. 5:53then they can go off and and actually do
  142. 5:55that in a you know pretty serious way.
  143. 5:58>> That's pretty cool.
  144. 6:00[clears throat] Now, one thing that
  145. 6:02you've been very well known for is
  146. 6:04you're really good at napkin math.
  147. 6:07Sounds funny. So, one of the stories
  148. 6:10about you is that back in uh 2013 when
  149. 6:14speech recognition started to work at
  150. 6:16Google, you did the nap napkin math
  151. 6:19where if every Google user used their
  152. 6:23phone and talked to it and used the
  153. 6:24speech recognition system for three
  154. 6:27minute just three minutes a day, you
  155. 6:29found that the system requires a Google
  156. 6:32server. you would have to double the
  157. 6:33fleet which would be really really
  158. 6:36expensive just to do speech translation.
  159. 6:39>> Yeah.
  160. 6:40>> And instead you basically built a custom
  161. 6:44ship and that was the origin story of
  162. 6:47the TPU.
  163. 6:48>> Yeah. Yeah. I mean I I sort of had done
  164. 6:51you know we were starting to see really
  165. 6:53good uh quality results on this the sort
  166. 6:57of deep learning based speech systems uh
  167. 6:59speech models we were training. um but
  168. 7:02they were computationally expensive
  169. 7:03compared to the old speech system but
  170. 7:05they haved the error rate. So that was
  171. 7:07like the equivalent of 20 years of
  172. 7:09advances in speech recognition in just a
  173. 7:11few months of like fiddling with the
  174. 7:13model and getting scaling it up a bit
  175. 7:15and getting better data. And so we
  176. 7:19started to get worried that if speech
  177. 7:20worked a lot better, people would use it
  178. 7:22more. And so that that back of the
  179. 7:24envelope calculation was really about
  180. 7:26that like well what if people start
  181. 7:29start to use speech recognition more to
  182. 7:31dictate emails or to talk to their phone
  183. 7:33or whatever. Um and yeah it turned out
  184. 7:37that um we realized that we needed some
  185. 7:41better solution than running on CPUs at
  186. 7:44the time. And so we came up with TPUs
  187. 7:47which are sort of very specialized for
  188. 7:50essentially low precision dense linear
  189. 7:52algebra which is at the heart of nearly
  190. 7:54all of the modern machine learning
  191. 7:56algorithms we we use today. And um if
  192. 8:00you build a specialized chip for low
  193. 8:02precision dense linear algebra and can't
  194. 8:04do anything else that turns out to be
  195. 8:06really useful for machine learning
  196. 8:07inference uh even though it can't run
  197. 8:09Chrome or Word or whatever. Uh, and so
  198. 8:12that system produced a chip a couple
  199. 8:14years later that was uh 30 to 80 times
  200. 8:19more energy efficient than CPUs and GPUs
  201. 8:22of the day and also much much lower
  202. 8:25latency like 20 to 30x lower latency
  203. 8:28>> which is incredible what the foundation
  204. 8:30that TPU has become today. No way you
  205. 8:32would have predicted that TPU would be
  206. 8:34so foundational now with transformer
  207. 8:36architecture which was invented way
  208. 8:38later before you actually invented the
  209. 8:40TPU. Yeah, I mean that's sort of why we
  210. 8:43built a general purpose linear algebra
  211. 8:46system, which is what a TPU is really.
  212. 8:48Um, because we knew ML algorithms were
  213. 8:51still evolving and you didn't want to
  214. 8:53over specialize, but you wanted to
  215. 8:55specialize enough that you got the
  216. 8:57dramatic performance benefits of we
  217. 9:00could have very big multiplier units. uh
  218. 9:03we could have you know high-speed memory
  219. 9:04we could have high-speed interconnect or
  220. 9:06later TPUs that like brought many many
  221. 9:09chips to bear on the same problem
  222. 9:11efficiently and u you know we've
  223. 9:14continued to scale those up and and
  224. 9:15improve their performance uh for over
  225. 9:18many many generations now
  226. 9:20>> incredible napkin math so
  227. 9:22>> what's good
  228. 9:23>> napkins are good
  229. 9:25>> so actually what's a good napkin math
  230. 9:28that everyone here who wants to be a
  231. 9:31future founder
  232. 9:32should run tonight to potentially build
  233. 9:35something as consequential as the TPU.
  234. 9:39>> Yeah, I mean, uh, it's always hard to
  235. 9:42say. Um, [clears throat] I think, uh,
  236. 9:46think about what problems you see in
  237. 9:49whatever it is you're thinking about,
  238. 9:51what what bottlenecks you see, and are
  239. 9:54there very different ways of thinking of
  240. 9:56the solutions to some of those problems
  241. 9:58that would get you, you know, an order
  242. 10:00of magnitude or two orders of magnitude
  243. 10:02better uh, performance or capability or
  244. 10:05whatever it is. Um, you know, because
  245. 10:08sometimes if you just squint at a
  246. 10:10problem and you think about not
  247. 10:13necessarily being anchored on exactly
  248. 10:14how that problem is solved today, but
  249. 10:16how you would solve it from first
  250. 10:18principles, you can come up with really
  251. 10:20good ideas that are, you know, maybe not
  252. 10:23what other people are thinking about.
  253. 10:25>> That's a good tip. [clears throat]
  254. 10:27>> No. Um, for everyone here who doesn't
  255. 10:29know, years ago, Jeff wrote a very
  256. 10:32famous list called the latency numbers.
  257. 10:36every engineer should know
  258. 10:38and these are numbers around for example
  259. 10:41how long a cache miss takes uh disk seek
  260. 10:46a network package traveling let's say
  261. 10:48from California to Netherlands
  262. 10:51um lots of numbers like this about
  263. 10:53distributed systems and systems
  264. 10:55engineering [clears throat]
  265. 10:56>> and it's been sort of taped and become
  266. 10:58the bible for a lot of distributed
  267. 10:59systems engineers
  268. 11:00>> okay yeah
  269. 11:01>> now fast forward that list is up for an
  270. 11:04update give us the AI edition for now
  271. 11:082026.
  272. 11:09>> Yeah, I mean I think if you looked at
  273. 11:10what is important in AI systems these
  274. 11:13days, you would want to know things like
  275. 11:17the bandwidth between you know your main
  276. 11:21memory system on your accelerator to the
  277. 11:23onchip memory to the um you know the
  278. 11:26multiplier unit or whatever. You want to
  279. 11:28know how much energy does it take to do
  280. 11:31a single multiplier operation.
  281. 11:34um you know uh what is the interconnect
  282. 11:36bandwidth between chips and how much
  283. 11:39does that uh how how many chips can you
  284. 11:42connect with that bandwidth and then if
  285. 11:43you go beyond that domain like what is
  286. 11:47the fall off in in uh network bandwidth
  287. 11:50when you need to talk to 10,000 strips
  288. 11:52instead of instead of uh 500 or
  289. 11:55something I think these are all really
  290. 11:57important numbers to to learn
  291. 11:59and and really affect how you think
  292. 12:02about solving particular kinds the
  293. 12:03problems.
  294. 12:04>> H [clears throat] and one interesting
  295. 12:07thing that I've heard you talk about is
  296. 12:08that nowadays the unit that you measure
  297. 12:12everything is energy.
  298. 12:15>> Yeah.
  299. 12:15>> You pointed out that doing a calculation
  300. 12:18or math costs about one pico.
  301. 12:21Uh but moving the data and doing data IO
  302. 12:24costs thousand times that.
  303. 12:26>> Yeah. Just bringing it in from HPM on an
  304. 12:28accelerator into the processor so it can
  305. 12:31actually compute on it. Yep.
  306. 12:33That gap kind of quietly decides what
  307. 12:36products are possible and how these
  308. 12:38algorithms in AI are built. So what are
  309. 12:42the kinds of problems that founders keep
  310. 12:44calling model problems but are in fact
  311. 12:47actually energy or data IO problems?
  312. 12:51Yeah, I mean I think the the example you
  313. 12:54raised of a thousandx difference in
  314. 12:56bringing mo moving data versus actually
  315. 12:59computing on it uh in in terms of energy
  316. 13:02is is a pretty significant one and it
  317. 13:04shapes a lot of aspects of what we do in
  318. 13:06machine learning. Um because if you
  319. 13:10didn't have that thousandx difference
  320. 13:11then you know you wouldn't have to do
  321. 13:14batching but you have to do batching of
  322. 13:16you know many examples or maybe many
  323. 13:18tokens at once in order to amortize that
  324. 13:21data movement [clears throat] so that
  325. 13:22you can uh you know not pay a thousandx
  326. 13:26slowdown but pay a 1000x divided by
  327. 13:28batch size uh energy cost. Um and you
  328. 13:32know for for really low latency batching
  329. 13:36is not really very good. Um so I think
  330. 13:39these [clears throat] kinds of things
  331. 13:40and the energy uh behind various
  332. 13:43decisions in the computer hardware we
  333. 13:45use really affects a lot of decisions we
  334. 13:48make in building higher level systems.
  335. 13:51>> A very concrete example is just how
  336. 13:53training models is done. There's this
  337. 13:55whole whole concept of batching the the
  338. 13:59data sets and running epochs. That's
  339. 14:01basically people perhaps may confuse
  340. 14:03that as a model problem, but it's really
  341. 14:05a systems data IO problem, right?
  342. 14:08>> Yeah. Yeah. I mean, you have to assemble
  343. 14:10batches to get better efficiency in your
  344. 14:13hardware. You know, ideally you might do
  345. 14:15batch size one training, but uh you
  346. 14:17know, it's um not as not as good in
  347. 14:21terms of efficiency. So people use re
  348. 14:24pretty large batches these days.
  349. 14:26>> Do you think uh it's possible for uh I
  350. 14:28know you're you're well known for uh
  351. 14:30taking off uh on a long week or weekend
  352. 14:33and coming up with this brilliant
  353. 14:34solution. Is there such things of Jeff
  354. 14:36going and working on it for a couple
  355. 14:38weeks and [clears throat]
  356. 14:38getting batch size equals one training
  357. 14:41done.
  358. 14:43>> Yeah, I've been thinking more about
  359. 14:44inference actually. So I think inference
  360. 14:46is a pretty interesting problem because
  361. 14:49you do want very low latency. You know
  362. 14:52training you don't necessarily need
  363. 14:53incredibly low latency. Um and I think
  364. 14:56there's a lot of room for specializing
  365. 14:58hardware more for inference than we are
  366. 15:00today.
  367. 15:02>> What are some of those interesting
  368. 15:03things that are on inference that you're
  369. 15:05really thinking a lot about?
  370. 15:07>> Um I mean just trying to minimize data
  371. 15:09movement. uh trying to think about
  372. 15:12incredibly
  373. 15:14uh low precision operations
  374. 15:17uh and maybe not supporting lots and
  375. 15:19lots of different kinds of precisions.
  376. 15:21Uh if you feel like you have a a good
  377. 15:25answer for what kinds of precision you
  378. 15:27need, maybe just build that into the
  379. 15:29hardware and and um not much else.
  380. 15:33which I think it brings down to a core
  381. 15:36[clears throat] analogy I heard from
  382. 15:38famous computer scientists that really
  383. 15:40the whole process of u AI is a big
  384. 15:44compression problem because in order to
  385. 15:48have the data to be f fully lossy and
  386. 15:50compress it and then restore it you
  387. 15:52basically need to understand it. Yeah, I
  388. 15:54mean if you truly understand the data,
  389. 15:57you should be able to compress it really
  390. 15:58well
  391. 15:59>> and now transformer architecture is
  392. 16:02basically one of the ways that has
  393. 16:04turned out to work really well.
  394. 16:07>> Yeah. Yeah, I would say
  395. 16:08>> working pretty well so far.
  396. 16:09>> Good work by my colleagues. [laughter]
  397. 16:12>> Yes. Now let's zoom out a bit. Um AI
  398. 16:16progress used to mean just better
  399. 16:18models. You could had more data trainer
  400. 16:20models with bigger parameters. But
  401. 16:23increasingly in the last years or so,
  402. 16:26it's everything around the model. Not
  403. 16:27just the model size and number of
  404. 16:29parameters or more data. It's everything
  405. 16:31around things like retrieval tools,
  406. 16:34memory, agent tools, and it might kind
  407. 16:37of get consolidated into what people
  408. 16:39call uh context engineering, right?
  409. 16:42[clears throat]
  410. 16:42>> Yeah. I mean I think uh the model is
  411. 16:45really only one piece of what you're
  412. 16:48trying to do which is build an overall
  413. 16:49system that can solve really interesting
  414. 16:52problems and that involves you know a
  415. 16:55model that knows how to use various
  416. 16:57tools. It maybe knows how to retrieve
  417. 16:59relevant information, maybe has a, you
  418. 17:02know, a history of other uh information
  419. 17:05that it has retrieved for past problems
  420. 17:08and it can put information into the
  421. 17:11context of the of the model. And the
  422. 17:14nice thing about that is that
  423. 17:15information is really clear to the
  424. 17:17model, unlike the training data the
  425. 17:19model was trained on where it's all kind
  426. 17:21of like trillions of tokens stirred
  427. 17:24together into a soup of of hundreds of
  428. 17:26billions or trillions of parameters, but
  429. 17:28it's all less clear than the actual
  430. 17:32context uh that the model sees directly
  431. 17:34for this particular problem or uses use
  432. 17:36case. And then I think being able to
  433. 17:39understand what tools are available,
  434. 17:43which ones are going to help me solve
  435. 17:44the help the model solve this next you
  436. 17:47know phase of the problem, how to
  437. 17:48decompose a problem into a sequence of
  438. 17:50of tool calls. Maybe trying multiple
  439. 17:52approaches to solve the problem and
  440. 17:55seeing which ones work and being able to
  441. 17:56evaluate that. you know this is the
  442. 17:58whole um you know orchestration of
  443. 18:02complex agent and multi-agent systems
  444. 18:04that I think is going to be more and
  445. 18:06more important and uh super exciting
  446. 18:08times I would say
  447. 18:10>> and I think the fun thing about this
  448. 18:11particular problem domain set is
  449. 18:13actually something that everyone in this
  450. 18:15room can actually do because before to
  451. 18:18train a model you needed incredible
  452. 18:19amount of resources incredible amount of
  453. 18:21access of to GPUs and data but for
  454. 18:24context engineering everyone here could
  455. 18:27do you have you just need the API to
  456. 18:29something like Gemini and then work on
  457. 18:31your own setup for your own retrieval
  458. 18:34your own tool calls and etc etc. So how
  459. 18:38does what are some tips for everyone
  460. 18:40here? How does everyone get better at
  461. 18:42and become exceptional at context
  462. 18:45engineering? Yeah, I mean I think uh
  463. 18:49[clears throat]
  464. 18:50a really good way to do it is to use
  465. 18:52these models and and sort of harnesses
  466. 18:55and tools and so on to try to solve
  467. 18:57problems and then some sometimes you can
  468. 19:00actually see where the models are
  469. 19:01failing. And often you can actually make
  470. 19:04the model work better and succeed at
  471. 19:07that kind of problem by not just
  472. 19:10adjusting the model parameters which is
  473. 19:11hard to do from the outside but from you
  474. 19:14know creating better guidelines for the
  475. 19:16model you know writing skills for the
  476. 19:19model to know how to use different tools
  477. 19:21that would be incredibly useful for
  478. 19:22solving this particular class of
  479. 19:24problem. And I think as you do that, you
  480. 19:28end up on this kind of improving
  481. 19:30self-improving of the setup that you're
  482. 19:33trying to use to to solve things. Uh,
  483. 19:36and you know that that's a really good
  484. 19:37way to get better at understanding what
  485. 19:41what additional information the model
  486. 19:43would want in order to become more
  487. 19:45capable.
  488. 19:46>> Can you give an example of uh some
  489. 19:48context engineering you personally have
  490. 19:50done? um I don't know skills you wrote
  491. 19:52tools that really made a huge different
  492. 19:54in your in your workflow. Yeah, I mean I
  493. 19:57guess uh Sanjay and I were working a few
  494. 19:59weeks ago and we you know we often do
  495. 20:03some amount of like uh performance
  496. 20:05improvement for very low-level libraries
  497. 20:08and we have a microbenchmark library
  498. 20:10we've written at Google where you can
  499. 20:12write microbenchmarks of how how long
  500. 20:15different kinds of operations take or
  501. 20:17how long does it take to populate this
  502. 20:18data structure whatever and sometimes
  503. 20:20those data structures are used on
  504. 20:22millions of processes across Google. So,
  505. 20:25it's actually pretty important to make
  506. 20:26sure they're high performance. And so,
  507. 20:28you can write microbenchmarks. Um, but
  508. 20:31then without an agent-based system, what
  509. 20:33you usually do is you measure what the
  510. 20:35current performance is on some
  511. 20:37benchmarks you care about. You make some
  512. 20:39modifications to improve the performance
  513. 20:42you hope. Then you rerun the the
  514. 20:45benchmarks, see where things improved.
  515. 20:48Um, you run a maybe a broader set of
  516. 20:49benchmarks, measure the cache footprint
  517. 20:52of things. And so we wrote a skill that
  518. 20:56basically taught the model how to do
  519. 20:57most of those things in in var in
  520. 21:00various sequences so that it could
  521. 21:01actually you know do self-improving uh
  522. 21:04benchmark measurement benchmark improve
  523. 21:06you know code changes measure the
  524. 21:08performance improvement and then iterate
  525. 21:10on that and that that seemed to work uh
  526. 21:12pretty well for some kinds of problems.
  527. 21:14And it really just is us giving the
  528. 21:18approach we would use as people to the
  529. 21:21model in a form that it could use.
  530. 21:23>> Wow, that seems very impressive. So
  531. 21:25you're saying you have this skill that
  532. 21:27if someone got access to it, it could do
  533. 21:30perform optimizations like Jeff Dean.
  534. 21:33Seems like the world would love this and
  535. 21:36is worth infinite amount of money to
  536. 21:38someone have access to this.
  537. 21:39>> Oh. Uh we actually published a document
  538. 21:41maybe a few months ago called
  539. 21:43performance hints that Sanjay and I
  540. 21:45wrote that's like a 30-page document
  541. 21:47about you know various kinds of
  542. 21:49performance tricks and some people have
  543. 21:51taken that and then given it in
  544. 21:53summarized form to various models and
  545. 21:55seen that they that model can now get
  546. 21:57you know better at uh per reasoning
  547. 21:59about performance issues in code.
  548. 22:01>> So you heard it all here. You could
  549. 22:02actually get your own optimize your own
  550. 22:05code like Jeff Dean if you take this
  551. 22:07this paper that you published when
  552. 22:08performance hints. Yep.
  553. 22:10>> It's all free available, so you should
  554. 22:11all try it.
  555. 22:12>> Very cool.
  556. 22:13>> Yeah.
  557. 22:14>> Now, you're talking about agents. Um,
  558. 22:16everyone here is probably building one
  559. 22:17or built one at some point. And I'm sure
  560. 22:20everyone has seen your agent go off the
  561. 22:23rail at perhaps like step 30 or 40. Like
  562. 22:26agents are great for like up to step, I
  563. 22:28don't know, 10 or something and then
  564. 22:29gets shaky at step 50. What do you think
  565. 22:32is the constraint today? Is it like
  566. 22:34context
  567. 22:36evaluators or just errors that compound
  568. 22:38because it's basically a openloop
  569. 22:40system?
  570. 22:41>> Yeah, I mean obviously we want agents to
  571. 22:44be able to run for very long periods of
  572. 22:46time because that's how they're going to
  573. 22:47solve more and more complicated
  574. 22:49problems. Um but as you as you observe
  575. 22:52today, you know, they sometimes stop
  576. 22:55working after, you know, 10 10
  577. 22:59interactions with the tools and so on.
  578. 23:01Um, and sometimes that's because the
  579. 23:04model is trying to do something it
  580. 23:05doesn't have a lot of experience doing.
  581. 23:07So it's been trained on a whole set of
  582. 23:09things and as soon as you get a little
  583. 23:11bit off the distribution of things it
  584. 23:13knows how to do then like most machine
  585. 23:15learning models it will you know its
  586. 23:18performance will suddenly will start to
  587. 23:20degrade and the farther you get off the
  588. 23:22comfort zone of what it knows how to do
  589. 23:25the the more likely it is to to not work
  590. 23:27as well. Um so there's a bunch of things
  591. 23:30you can do. So one is you know give the
  592. 23:33model skills and hints that kind of tend
  593. 23:35to keep it in in the uh sort of more
  594. 23:38brightly lit path of things it does know
  595. 23:40how to do. Um, I think you know having
  596. 23:43multi- aent systems where you have
  597. 23:45multiple agents trying different
  598. 23:47approaches and you can evaluate you have
  599. 23:49maybe another model or another agent
  600. 23:51that's evaluating which ones of those
  601. 23:53seem promising is another way to kind of
  602. 23:58in some sense search the path of pos
  603. 24:01search the space of possible solutions
  604. 24:03and stick to the ones that seem most
  605. 24:06promising and discard the ones that that
  606. 24:08didn't seem to work or maybe that went
  607. 24:10off the rails. or whatever. Um, and
  608. 24:13that's a very very useful general
  609. 24:15technique is you know inference time
  610. 24:18compute to perform search over plausible
  611. 24:20ways of solving the problem that can get
  612. 24:23much much higher performance or much
  613. 24:25more reliability in longunning agent
  614. 24:27flows.
  615. 24:30>> How are some ways you implemented this
  616. 24:32particular workflow for your agents
  617. 24:34internally?
  618. 24:35Yeah, I mean we have uh you know
  619. 24:37harnesses and then we have a whole set
  620. 24:39of skills uh particularly in the
  621. 24:42internal Google development environment.
  622. 24:43We have skills so that the agents can
  623. 24:45know how to use lots of our internal
  624. 24:48tooling for coding or for code reviews
  625. 24:50or for you know measuring performance or
  626. 24:54you know fetching log files. And um
  627. 24:57those are just skills that you can add
  628. 24:59to make the base model more capable even
  629. 25:01though it hasn't necessarily been
  630. 25:03trained on exactly the way that you know
  631. 25:06Google internal uh engineers would fetch
  632. 25:09log files from our you know proprietary
  633. 25:12system with the right kind of skill uh
  634. 25:15definition you can actually get it to
  635. 25:16work
  636. 25:17>> uh and that that improves the usefulness
  637. 25:19of the agents. Now let's talk about uh
  638. 25:23where startups can can win. This section
  639. 25:27is one that I personally care a lot
  640. 25:29about because also everyone here in this
  641. 25:31room needs to decide what to build in
  642. 25:33the future of your future founder. So
  643. 25:37the thing about Google is you co-design
  644. 25:39everything on the system from the
  645. 25:41processors to the products.
  646. 25:44um which are the layers that someone
  647. 25:47like Google would keep building and
  648. 25:49compounding being better and and where
  649. 25:52does a two three person team can still
  650. 25:55win?
  651. 25:57>> Yeah, I mean I think obviously Google
  652. 26:00and and our Gemini models and and our
  653. 26:02hardware infrastructure are really
  654. 26:04trying to build very general models that
  655. 26:06can do almost anything. But in in a lot
  656. 26:11of cases that means that we don't have a
  657. 26:14lot of attention on particular domains
  658. 26:16where perhaps a really well-designed
  659. 26:19surface that and maybe a model and set
  660. 26:22of skills or maybe a specialized model
  661. 26:25that uh isn't in sort of a general mix
  662. 26:28of of things that our models do well can
  663. 26:31actually have a significant advantage
  664. 26:33because you can build something
  665. 26:34delightful and you know really high
  666. 26:37accuracy. really high quality for a
  667. 26:40domain that you are really passionate
  668. 26:41about. And I think that's that's where
  669. 26:44you know the two or three people in a
  670. 26:46room uh building that that they're
  671. 26:49really excited about can have an
  672. 26:50advantage. Um but I I would also caution
  673. 26:54that the general models are definitely
  674. 26:57getting better at a broader and broader
  675. 26:59range of things. So you have to figure
  676. 27:00out, you know, is that thing you're
  677. 27:03working on, is that going to be a
  678. 27:04durable thing or do you think the models
  679. 27:07uh at the forefront are going to get
  680. 27:09better at that in the next six months or
  681. 27:1112 months or is it something they're not
  682. 27:13going to be able to do for a couple
  683. 27:15years or three years? And you know, you
  684. 27:17you want to weigh that as you're as
  685. 27:19you're deciding what to work on.
  686. 27:22>> So let's uh dive deeper into this. So
  687. 27:24the general models of course you're
  688. 27:26going to keep working on and keep making
  689. 27:27them all better.
  690. 27:29And how should the audience reason about
  691. 27:31what are those areas that uh it doesn't
  692. 27:35I mean h how should founder think about
  693. 27:38things to pick on and work on.
  694. 27:40>> Yeah. I mean I mean the most important
  695. 27:42thing is to pick something you're super
  696. 27:44excited about and want to build and you
  697. 27:46think would be useful in the world,
  698. 27:48right? So if you do that um that that's
  699. 27:51you're already way ahead uh than if you
  700. 27:54wake up and you're like oh I don't
  701. 27:55really want to do this or whatever or
  702. 27:57you're going to build something that is
  703. 27:59actually not that useful to to the world
  704. 28:01or to to many people. Um so I think
  705. 28:05that's the number one selection criteria
  706. 28:07I try to apply for what problem should I
  707. 28:10work on next. Um, second, I think you
  708. 28:14want to look at what the current more
  709. 28:16general models can do in that problem
  710. 28:19domain, right? You can you can test them
  711. 28:22with like, are they able to do this
  712. 28:23thing very well? And if they're
  713. 28:26completely failing, that's probably a
  714. 28:28good sign. If they're kind of able to do
  715. 28:30some of it but not very well, that's
  716. 28:33maybe not a great sign because that's a
  717. 28:35probably a a sign that the capability is
  718. 28:38starting to be present in those models
  719. 28:41and with more training data or larger
  720. 28:43scale models or or whatever it's likely
  721. 28:45to get better. So um you know look for
  722. 28:48something where the model succeeds 0% or
  723. 28:511% of the time not not 20%.
  724. 28:54>> How do you find those? I mean are those
  725. 28:56things effectively
  726. 28:58uh out of distribution from the training
  727. 29:00set and what exactly is the problem
  728. 29:02shape that fits that?
  729. 29:04>> Yeah, I mean I think uh sometimes it's
  730. 29:08uh a product that you build that might
  731. 29:11have access to particular kind of data
  732. 29:13that the underlying model might not the
  733. 29:15a general model. So it might be you're
  734. 29:18building something to help users
  735. 29:19organize all their own personal
  736. 29:21information and the model won't
  737. 29:22necessarily have access to that. And so
  738. 29:24there you can have a big advantage
  739. 29:26because all of a sudden your model has
  740. 29:28visibility or your product has
  741. 29:31visibility into important data. Um it
  742. 29:35could be some incredibly hard problem
  743. 29:38where if you get the right training data
  744. 29:40and you can train a more specific model
  745. 29:42than a general purpose one, you can
  746. 29:44actually do that in a very affordable
  747. 29:46way. you maybe it doesn't take that much
  748. 29:48compute to train a a niche model for
  749. 29:50this particular problem, but you can get
  750. 29:52something that's highly accurate. That
  751. 29:54can sometimes be a a really good uh
  752. 29:57building block for for solving a
  753. 30:00important problem that is maybe not
  754. 30:03handled very well by the general model.
  755. 30:04>> I think that's interesting. I think
  756. 30:06there are basically two paths. The first
  757. 30:07path is uh a little bit funny is uh you
  758. 30:10guys are organizing the world's
  759. 30:12information.
  760. 30:12>> Yeah,
  761. 30:13>> that's probably kind of well covered.
  762. 30:15>> Yeah. but organizing your personal
  763. 30:17information that's open which is funny.
  764. 30:21>> Yeah.
  765. 30:21>> And then the second path um you talked
  766. 30:24about more specialized models in certain
  767. 30:27domains. Can you tell us more about what
  768. 30:29are some of these domains?
  769. 30:30>> Yeah. I mean I think like if you look at
  770. 30:33uh my colleagues work on say alpha fold
  771. 30:35that was a very specific model for uh
  772. 30:38protein folding and it was highly
  773. 30:40successful
  774. 30:41um and was able to really handle that
  775. 30:44domain quite well so that all of a
  776. 30:46sudden you now have this amazing tool
  777. 30:48and model that can give you answers to
  778. 30:50questions about proteins and their
  779. 30:52structure um really effectively um but
  780. 30:55it's not a general model it's a very
  781. 30:58specific one and there are other I
  782. 31:01domains where that kind of approach can
  783. 31:03work really well. Uh maybe in material
  784. 31:05science or chip design or things like
  785. 31:07that that uh will enable you to leverage
  786. 31:11the capabilities of a very accurate but
  787. 31:14but niche model uh to do things that are
  788. 31:18hard today.
  789. 31:19>> That's a good example. So if some of you
  790. 31:21find a problem that's similar shape like
  791. 31:23alpha fold could be a good problem to
  792. 31:24work on. Now let's assume you found a
  793. 31:26problem to work on. We're going to talk
  794. 31:28a bit about how do you become a AI
  795. 31:30native founder? How do you really become
  796. 31:32good at it? Uh you in the past said that
  797. 31:35managing a fleet of agents, it's like 50
  798. 31:37or 100 agents is all about writing
  799. 31:40really good crisp design docs or specs.
  800. 31:45>> And how do people get good at that? What
  801. 31:47what do those look like? Yeah, I mean I
  802. 31:50think uh
  803. 31:54it's
  804. 31:55you you'll have a lot more success when
  805. 31:58working with your virtual agents if you
  806. 32:00can clearly specify what it is you want.
  807. 32:03And the clearer you are on what it is
  808. 32:04you want, the more the agent will have
  809. 32:07sort of guidelines and sort of rules of,
  810. 32:10you know, an outline of what it is
  811. 32:12trying to accomplish. Um whereas if you
  812. 32:15don't specify very much stuff, the agent
  813. 32:17has to sort of infer what it is you
  814. 32:19meant. And in many cases, it might infer
  815. 32:22things that are different than what you
  816. 32:23imagined. So we've always told computer
  817. 32:26scientists from the very beginning that
  818. 32:27really it's really important to specify
  819. 32:29what it is, what's the software that
  820. 32:31you're writing is trying to accomplish
  821. 32:34before then going and writing it. And so
  822. 32:36now we actually have agent-based systems
  823. 32:38that can do the writing, but the
  824. 32:40importance of specifying what what it is
  825. 32:42you want has actually gone up because
  826. 32:45before you'd be handing it off to a very
  827. 32:48intelligent human who maybe has context
  828. 32:50or can ask you follow-up questions. Um,
  829. 32:53and agents can sometimes do that, but I
  830. 32:54I think clear specifications is is a
  831. 32:58really good idea. Um, and to give you an
  832. 33:00example of a a a use of a coding agent
  833. 33:04that works extremely well is you can ask
  834. 33:07today's models to translate software
  835. 33:10from one computer language to another
  836. 33:12very effectively because in that case
  837. 33:15you actually have a incredibly detailed
  838. 33:18specification. You have the whole
  839. 33:19software that says what the system is
  840. 33:21supposed to do. And so if you have a
  841. 33:23Python implementation of something and
  842. 33:25you want a Go implementation of it, you
  843. 33:28know, that is something that the models
  844. 33:29seem incredibly capable at doing these
  845. 33:31days because you can it can sort of take
  846. 33:34all the tests that are in Python, make
  847. 33:36sure they pass in the Go version,
  848. 33:38translate the tests to Go, um you know,
  849. 33:41compare uh behavioral differences
  850. 33:43between the implementations until there
  851. 33:45aren't any um and be you know, highly
  852. 33:48effective because that spec is so clear.
  853. 33:51Hm. Now let's assume now every founder
  854. 33:53gets good at running hundreds of agents
  855. 33:55at the same time and all the code is
  856. 33:57written for them by the agents. What
  857. 33:59becomes the scarce skill?
  858. 34:03>> Yeah, I mean I think it's really having
  859. 34:07incredibly good taste in what you ask
  860. 34:09your agents to work on, right? That is
  861. 34:11the the crux of you know from my
  862. 34:15background uh a research problem. You
  863. 34:17know, a researcher can have all the
  864. 34:19tools and all the techniques, but often
  865. 34:22most of the battle is what problem are
  866. 34:25you gonna spend your time on? And if you
  867. 34:27pick the problem well and you succeed in
  868. 34:30in in solving it, that's way better than
  869. 34:33if you, you know, uh, delightfully
  870. 34:36execute a research investigation into a
  871. 34:39rather boring problem. And so that high
  872. 34:42level wisdom of what to work on, I think
  873. 34:44is incredibly important. And I think
  874. 34:46models are not necessarily going to be
  875. 34:47that good at it. So you're going to have
  876. 34:50people steering
  877. 34:54uh a lot of AI assisted computation in
  878. 34:57order to accomplish great things and
  879. 34:59more quickly. Um but that essence of of
  880. 35:02what it is you want your models to do is
  881. 35:05the the the key thing you should focus
  882. 35:07on.
  883. 35:08>> So let's talk a bit more about taste
  884. 35:09because it gets talked a lot about right
  885. 35:12now in this current era with agent
  886. 35:14coding. How do you exactly build taste
  887. 35:17and do that? I mean, yeah, that sounds
  888. 35:20so esoteric. How do you make it
  889. 35:22concrete?
  890. 35:23>> Yeah, I mean it it is a difficult thing.
  891. 35:26It's not like there's a measurable
  892. 35:28objective of of taste in a lot of cases.
  893. 35:31Um, I think some of it is from
  894. 35:33experience. You know, working on a lot
  895. 35:36of different problems in the past kind
  896. 35:38of teaches you about what kinds of
  897. 35:40problems might be interesting in the
  898. 35:42future or what kinds of things might be
  899. 35:45just barely possible by cobbling
  900. 35:48together these previous approaches and
  901. 35:51then some open problems you might have
  902. 35:53to work on in order to get to something
  903. 35:55kind of magical or or you know, highly
  904. 35:57useful. Um,
  905. 36:00another way you can get more experience
  906. 36:02for yourself is to just write down a
  907. 36:06bunch of things you think might be
  908. 36:07important in the next 12 months. And
  909. 36:10maybe you pick one of them to work on,
  910. 36:12but go back and evaluate in 12 months of
  911. 36:15these other things, which ones actually
  912. 36:18seemed important or which ones did other
  913. 36:20people in the world go out and and
  914. 36:22create and which ones did they did not
  915. 36:24seem to do yet. um that can give you a
  916. 36:27lot more samples for your own sort of
  917. 36:29taste creation uh capability. Um and and
  918. 36:33that's an important skill to have.
  919. 36:35>> I think a third way we were talking
  920. 36:37earlier was doing very crazy thought
  921. 36:41experiments.
  922. 36:42>> Oh yeah, that's another good way. I mean
  923. 36:44I think uh sometimes it's good to
  924. 36:49not take as a given things that most
  925. 36:52people seem to take as a as a given. Um
  926. 36:55so I was doing a crazy thought
  927. 36:56experiment with some colleagues the
  928. 36:58other day about you know
  929. 37:02for 60 years the whole silicon
  930. 37:06uh chip design industry uh design and
  931. 37:09fabrication industry have
  932. 37:12you know done tremendous work to make
  933. 37:15smaller and smaller scale transistors
  934. 37:17that are uh very low error rate right
  935. 37:22like because what what the assumption
  936. 37:24that we want is that every chip we
  937. 37:25manufacture of the same design should be
  938. 37:28identical to every other chip.
  939. 37:30>> You don't want any bits to flip.
  940. 37:31Everything
  941. 37:31>> no bits should flip. There's all kinds
  942. 37:33of things you there's all kinds of error
  943. 37:35margins built into you know memories
  944. 37:38have ECC memory these days. um you know
  945. 37:41at the at the macro scale we don't make
  946. 37:44that assumption when we're building
  947. 37:46large scale distributed systems right we
  948. 37:49we build reliable large scale
  949. 37:51distributed file systems out of
  950. 37:53unreliable parts right like individual
  951. 37:56discs can fail but your data should be
  952. 37:58safe and so we have mechanisms at a
  953. 38:00higher level to enable us to have um you
  954. 38:04know three copies of the data on three
  955. 38:06different machines and three different
  956. 38:07racks so that if any rack switch or
  957. 38:09individual ual machine or disk fails,
  958. 38:11you still have your data. We have read
  959. 38:13Solomon encoding techniques. Um, but we
  960. 38:17don't seem to do this at a really
  961. 38:19extreme level in the uh sort of
  962. 38:23transistor level scale of of the
  963. 38:26technology we're working on. So what
  964. 38:28would h basically a interesting thought
  965. 38:30experiment is what would happen if you
  966. 38:32tried to build a system out of
  967. 38:34transistors that might have you know 20
  968. 38:37errors per day.
  969. 38:38>> Oh my god. rather than one every million
  970. 38:40years, right? That would be a very
  971. 38:43different design point and might be
  972. 38:45might enable you to do really
  973. 38:46interesting things in the fabrication
  974. 38:48side of things. You have very different
  975. 38:50kind of design methodologies because if
  976. 38:52you want to get a signal from here to
  977. 38:54there, you and you have these super
  978. 38:56unreliable transistors. You might have
  979. 38:58very different ways of signaling. You
  980. 39:00might send it along multiple redundant
  981. 39:02paths uh in order to make sure that it
  982. 39:04gets along one of them. Um, and I think
  983. 39:07that would be a pretty interesting set
  984. 39:08of thought experiments. I'm not saying
  985. 39:10we should go do this, but you know,
  986. 39:12that's the kind of thing where you do
  987. 39:13want to, you know, occasionally question
  988. 39:17assumptions. Now, oftentimes these
  989. 39:20thought experiments don't work out
  990. 39:22because there are very good reasons
  991. 39:24that, you know, for the last 50 years,
  992. 39:26we've done this thing this way and not
  993. 39:28that way. But it it's good to kind of
  994. 39:30revisit those every so often.
  995. 39:32>> That is so wild. Well, I mean, it's
  996. 39:33starting to rhyme a lot with
  997. 39:35neuromorphic computing or the human
  998. 39:37brain and and how nature works.
  999. 39:39>> I mean, exactly like signals in our
  1000. 39:41brain are not especially reliable from
  1001. 39:43getting one place to another. And so, I
  1002. 39:46think in brains when there are really
  1003. 39:48important things you need to get from
  1004. 39:50one place to another, there are multiple
  1005. 39:51pathways that that enable you to sort of
  1006. 39:54do that.
  1007. 39:56>> What is uh in I mean, you have such an
  1008. 39:58impressive career. What is one of these
  1009. 40:01crazy assumptions that you threw out of
  1010. 40:02the window that actually built a
  1011. 40:04consequential system in the past?
  1012. 40:09>> Yeah, I mean I guess uh
  1013. 40:10>> that worked out actually.
  1014. 40:13>> Yeah, I mean I think uh well TPUs is a
  1015. 40:15good example like being able to
  1016. 40:17specialize hardware for a very niche
  1017. 40:20>> problem domain before that problem
  1018. 40:22domain seemed as important as it is
  1019. 40:24today uh is one thought experiment. Um
  1020. 40:28you know I think the
  1021. 40:30the origin of map produce is another
  1022. 40:33good example.
  1023. 40:34So we had worked the you know my Sanjay
  1024. 40:37and myself and a number of other
  1025. 40:39colleagues had worked on various
  1026. 40:41iterations of the crawling and indexing
  1027. 40:43system at Google and you know we'd sort
  1028. 40:46of written lots of hand parallelized
  1029. 40:49code with lots of checkpointing to make
  1030. 40:51sure it would be robust and reliable if
  1031. 40:53it was running on a 100 computers or a
  1032. 40:55thousand computers and some of those
  1033. 40:57died. Um, but that code tended to be
  1034. 41:02intermixed with the actually relatively
  1035. 41:04simple thing you often were trying to do
  1036. 41:06like I just want to like look at all the
  1037. 41:09contents of all the web pages and then
  1038. 41:11compute on the side a mapping from URL
  1039. 41:14to you know what language is this page
  1040. 41:16in the text of this page. Um, and it
  1041. 41:20would get obscured by all this kind of
  1042. 41:22other code for parallelization and
  1043. 41:24reliability. And so we sort of
  1044. 41:27remembered our training in functional
  1045. 41:29languages and realized we could squint
  1046. 41:31at those problems and developed this map
  1047. 41:33produce abstraction that you could have
  1048. 41:36above this implementation and then below
  1049. 41:39the implementation you could put all the
  1050. 41:42checkpointing and reliability mechanisms
  1051. 41:45into that lower level library that
  1052. 41:47everything could then build on. And so
  1053. 41:49that became a hugely successful way of
  1054. 41:52of dealing with very large scale
  1055. 41:53computations at Google in a robust and
  1056. 41:55reliable way. From that thought
  1057. 41:58experiment of like well if we squint at
  1058. 42:00it could we find lots of problems that
  1059. 42:01fit into this abstraction.
  1060. 42:03>> That's impressive. So this thought
  1061. 42:06experiment led you to create map reduce.
  1062. 42:08>> Yeah. Awesome.
  1063. 42:09>> Now let's go back to you talked a bit
  1064. 42:11about um about your interest right now
  1065. 42:15working on a lot of customized hardware.
  1066. 42:17So right now alpha chip
  1067. 42:19>> lays out chips. Now you also got alpha
  1068. 42:21evolve that proposes solutions,
  1069. 42:24>> evaluates them and keeps all the ones
  1070. 42:26that work. Seems like you're starting to
  1071. 42:27build all these system that can compound
  1072. 42:30and build AI that builds AI.
  1073. 42:32>> Yeah. I mean I think more generally
  1074. 42:34there's a there's this sort of
  1075. 42:39the foundation of the scientific method
  1076. 42:41of you propose an experiment you
  1077. 42:44implement what you need to run the
  1078. 42:46experiment and you evaluate the
  1079. 42:48experiment and then you get results from
  1080. 42:50that and I think there are more and more
  1081. 42:52problems that are now possible to
  1082. 42:55implement where that whole loop of
  1083. 42:58running you know not just a few
  1084. 43:00experiments but running many many
  1085. 43:01experiments because you're able to
  1086. 43:03automate that loop and make the latency
  1087. 43:05of that loop extremely low is going to
  1088. 43:08be really really important. It's going
  1089. 43:09to enable us to tackle you know lots of
  1090. 43:12different problem domains in science and
  1091. 43:14engineering and machine learning uh
  1092. 43:17model design itself and also in
  1093. 43:20engineering tasks like designing chips.
  1094. 43:23And so if you can actually do those
  1095. 43:25things in an automated way and have some
  1096. 43:28orchestration framework that can take
  1097. 43:31very high level objectives and break
  1098. 43:33them down into subpros and each of those
  1099. 43:35subpros can be one of these automated
  1100. 43:38loop that is exploring the best way to
  1101. 43:40solve that sub problem and then a
  1102. 43:43orchestration framework that can put
  1103. 43:46together subpros solutions into a you
  1104. 43:50know the overall solution for the higher
  1105. 43:52level problem that's going to be really
  1106. 43:54impactful and it's really really
  1107. 43:56important and I think it'll enable us to
  1108. 43:58do you know accelerate machine learning
  1109. 44:01progress it'll enable us to accelerate
  1110. 44:03science and enable us to accelerate
  1111. 44:05engineering and I think that's that's
  1112. 44:07going to be amazing
  1113. 44:09>> that sounds awesome I mean it sounds
  1114. 44:10like a lot of fields basically where you
  1115. 44:12can have very good evaluators and maybe
  1116. 44:15adjacent to basically things that can be
  1117. 44:17formally verified right those are ripe
  1118. 44:20for AI systems that can self-improve
  1119. 44:22Yeah, I think in a lot of cases
  1120. 44:25sometimes your evaluators need to be
  1121. 44:27made much faster. Mhm.
  1122. 44:28>> So as an example, my colleagues did some
  1123. 44:32work maybe a decade ago on um some uh
  1124. 44:36problems in quantum chemistry where
  1125. 44:38you're trying to understand the
  1126. 44:39properties of a particular molecule and
  1127. 44:41you can you know generate some molecule
  1128. 44:44configuration and then you want to
  1129. 44:45understand what properties it has. And
  1130. 44:48so you can run a very computationally
  1131. 44:50intensive density functional theory
  1132. 44:52simulator which is something that might
  1133. 44:54take like a a night of computation to
  1134. 44:57tell you the answer for one thing. Um
  1135. 45:00but what my colleagues did was
  1136. 45:04take a bunch of output from those
  1137. 45:07simulation runs the input molecule
  1138. 45:09configurations and the outputs of the
  1139. 45:11the expensive simulator and then use it
  1140. 45:14to train a neural approximation to the
  1141. 45:16simulator. So this is now a validation
  1142. 45:19device, but instead of it taking a
  1143. 45:22night, they made something that was
  1144. 45:25300,000 times faster.
  1145. 45:26>> Wow.
  1146. 45:27>> And nearly as accurate as running the
  1147. 45:28full scale simulator. So now that
  1148. 45:32completely changes how you would do
  1149. 45:33science, right? Because now you have 10
  1150. 45:35million things to screen. you know, you
  1151. 45:38could do that while you go to lunch
  1152. 45:40rather than it being a six-month
  1153. 45:42endeavor where you could try to scrape
  1154. 45:44together enough compute to to run all
  1155. 45:46these simulations. And I think there's a
  1156. 45:50lot of room in a lot of domains for much
  1157. 45:53faster validation models, possibly
  1158. 45:55learned valu validation models that can
  1159. 45:59uh you know get you a a approximation to
  1160. 46:02the true answer much much more rapidly.
  1161. 46:04And that changes how those experimental
  1162. 46:06loops can be thought of and how quickly
  1163. 46:08you can go around those loops.
  1164. 46:10>> What are some of the
  1165. 46:13spaces and problems that you're super
  1166. 46:14excited that this super sped up
  1167. 46:17scientific method is going to solve or
  1168. 46:20achieve? What particular problems or
  1169. 46:22spaces?
  1170. 46:23>> Yeah, I mean I think uh
  1171. 46:28well clearly machine learning itself is
  1172. 46:30one, right? So can we have a model that
  1173. 46:32is able to recursively self-improve
  1174. 46:35itself by running lots of experiments
  1175. 46:37and you know if you think about how
  1176. 46:39models are improved today in large
  1177. 46:41research teams you know what usually
  1178. 46:44happens is people think of some ideas
  1179. 46:46they run a bunch of smallcale
  1180. 46:48experiments they see if those small
  1181. 46:50scale experiments worked out well if so
  1182. 46:53they take the most promising ones of
  1183. 46:54those they try them at larger scale and
  1184. 46:56that gets then evaluated and then the
  1185. 47:00results get integr ated together into
  1186. 47:02you know a new recipe for your model. Um
  1187. 47:06but I think there's no uh you know real
  1188. 47:09impediment to making that be a much more
  1189. 47:11automated loop where the model itself
  1190. 47:15decides it's going to explore or maybe
  1191. 47:17with a nudge from some people uh at the
  1192. 47:20various highest level like oh why don't
  1193. 47:22you try some new ideas around model
  1194. 47:24architectures that incorporate this and
  1195. 47:27then it will go run lots of experiments
  1196. 47:30uh see which ones work and then those
  1197. 47:31will get incorporated at a much more
  1198. 47:33rapid rate and uh you know effectively
  1199. 47:36you want to optimize you know your
  1200. 47:39discoveries per unit of compute input.
  1201. 47:44>> Very cool.
  1202. 47:44>> Yeah.
  1203. 47:45>> Now going back to the room as all of you
  1204. 47:49will become at some point founders or
  1205. 47:51start your careers you will probably
  1206. 47:54collect lots of rejections. That will
  1207. 47:56happen. Uh it has happened to you too
  1208. 47:59Jeff. I mean there's a story that in
  1209. 48:022014
  1210. 48:04you with Jeff Hinton and Oral Fin wrote
  1211. 48:09a paper on distillation
  1212. 48:12>> which has to do with taking a big
  1213. 48:15teacher model to train a much smaller
  1214. 48:18and more efficient model that's a lot
  1215. 48:21cheaper to compute less model parameters
  1216. 48:24and it has become a trick that everyone
  1217. 48:27is using right now in industry. Yeah.
  1218. 48:30>> And
  1219. 48:32the thing is this paper got rejected at
  1220. 48:34Europe.
  1221. 48:36>> Yeah. I mean Yeah. I mean I think I
  1222. 48:39don't fault the program committee
  1223. 48:40because you know a lot of times a paper
  1224. 48:44gets three reviews and someone will look
  1225. 48:46at one of the reviewers will look at it
  1226. 48:48and in this case they said oh it's
  1227. 48:50unlikely to have significant impact.
  1228. 48:52>> Unlikely to have significant impact. But
  1229. 48:53you know I think you know when we wrote
  1230. 48:55the paper we actually saw this was a
  1231. 48:57super important problem because we knew
  1232. 49:00making cheaper highly capable models
  1233. 49:03from larger scale models was something
  1234. 49:04we desperately wanted to do because we
  1235. 49:07wanted to serve models to more and more
  1236. 49:09people in many different domains like
  1237. 49:11speech or vision. Um but you know
  1238. 49:13sometimes the reviewer maybe didn't have
  1239. 49:15that that experience because maybe
  1240. 49:16they're not thinking about you know
  1241. 49:19largecale AI services and are thinking
  1242. 49:22about you know is this a fundamental
  1243. 49:24advance um so so you know it gets
  1244. 49:27rejected every so often that's fine we
  1245. 49:29put it on archive people read it people
  1246. 49:31use it it's all good uh and you know we
  1247. 49:34do use it in making our flash models for
  1248. 49:37example from our larger scale pro model
  1249. 49:40that's partly why our flash models for
  1250. 49:41example in Gemini are so capable uh
  1251. 49:45relative to their size and and speed.
  1252. 49:47>> They're some of the best in the
  1253. 49:48benchmark for their model size class.
  1254. 49:50Yeah. Just impressive. And I think part
  1255. 49:52of the lesson is that even if you get
  1256. 49:54rejected, keep going.
  1257. 49:56>> Yeah. That's that's the lesson I would
  1258. 49:58distill from that.
  1259. 50:00[laughter]
  1260. 50:01>> Um no, I think the fun thing is that you
  1261. 50:04basically join when you when you join
  1262. 50:05Google as a 20 person startup back in
  1263. 50:081999.
  1264. 50:09Now, if you were to take the young Jeff
  1265. 50:12Dean from way back then to teleransport
  1266. 50:16him to now today.
  1267. 50:18>> Yeah.
  1268. 50:18>> In this era with your skills.
  1269. 50:20>> I'm feeling so vigorous and and young
  1270. 50:23now.
  1271. 50:24>> Um what would you do? Do you join a
  1272. 50:26frontier lab, start a company? I don't
  1273. 50:29know what what would you do? The c the
  1274. 50:34Jeff theme today 25-year-old Jeff theme.
  1275. 50:36>> Yeah. I mean,
  1276. 50:39it's always hard to say and it's a very
  1277. 50:41personal choice of what it is you want
  1278. 50:42to spend your time on. Um, to me, some
  1279. 50:46of the most important questions are,
  1280. 50:49are you going to work on something you
  1281. 50:52really care about, will you're working
  1282. 50:55on that? And if you're able to make
  1283. 50:57progress on it with a bunch of
  1284. 50:59colleagues you like working with uh if
  1285. 51:02you're able to make collectively solve
  1286. 51:04it or make progress on it, will that
  1287. 51:07make a difference in the world in some
  1288. 51:08positive way, right? Like will you
  1289. 51:10suddenly be able to do something and
  1290. 51:12offer that service to you know
  1291. 51:15partically
  1292. 51:19help biochemists or something or maybe
  1293. 51:21it's a broader thing. It'll help
  1294. 51:22programmers or it will help all
  1295. 51:24consumers. uh on the internet or or
  1296. 51:27other things. Um what you you know what
  1297. 51:32you should strive to do is to have
  1298. 51:33impact in the world that is positive and
  1299. 51:36to work with people you enjoy working
  1300. 51:38with and to you know uh work hard and
  1301. 51:42and do your best. Um so in terms of say
  1302. 51:46the particular trade-off you offered
  1303. 51:48joining a frontier lab versus say
  1304. 51:50starting a company with just one or two
  1305. 51:52or three of you you and your close
  1306. 51:54friends. Um, I think those are different
  1307. 51:57experiences, right? In a in a large
  1308. 51:59established organization, you have some
  1309. 52:02structure. You have lots and lots of
  1310. 52:04amazing colleagues who know lots of
  1311. 52:06things you don't. Um, you have lots of
  1312. 52:10interesting problems that uh you can
  1313. 52:12work on and h you already have a
  1314. 52:15platform for impact by your work, you
  1315. 52:18know, influencing lots and lots of
  1316. 52:20people in the world already. Um and then
  1317. 52:22as a very small startup,
  1318. 52:26you know, you have to have something
  1319. 52:28you're passionate about and there's a
  1320. 52:31lot of risk in taking on, you know,
  1321. 52:34working on that particular problem in a
  1322. 52:36way that uh you're going to succeed and
  1323. 52:38you're going to grow a, you know, an
  1324. 52:40endeavor in order to do that. But that
  1325. 52:42can also be incredibly rewarding, I
  1326. 52:44would imagine. So I I think um you know
  1327. 52:48it's really up to personal taste but but
  1328. 52:50at the very least regardless of what
  1329. 52:53path you take ask yourself if I work on
  1330. 52:56this problem and the best possible
  1331. 52:58outcome happens you know will the world
  1332. 53:01be a lot better in some way or will the
  1333. 53:03world go eh that's kind of cool but
  1334. 53:05whatever.
  1335. 53:06>> Uh that's not the kind of thing you
  1336. 53:08should spend your time on.
  1337. 53:10Now let's talk a bit about more about
  1338. 53:12that second path of working with people
  1339. 53:15that you really like in a small team.
  1340. 53:18You've been able to be an incredible
  1341. 53:20mentor and manager to many many
  1342. 53:22engineers and you've been able to build
  1343. 53:25huge systems and what are some some of
  1344. 53:29the lessons for everyone here on how to
  1345. 53:32get the most and how to work with smart
  1346. 53:34people or find smart people?
  1347. 53:36Yeah, I mean,
  1348. 53:39you always want to find people who have
  1349. 53:42really good skills in some some area
  1350. 53:45that's needed in, you know, a team
  1351. 53:47you're trying to form, whether that's
  1352. 53:50inside a company or uh starting a
  1353. 53:52company. Um, but you also want to find
  1354. 53:55people that are people you delight being
  1355. 53:59around, right? because you're going to
  1356. 54:00spend a lot of time around people
  1357. 54:02working on really hard problems and you
  1358. 54:05want people who are low ego that are
  1359. 54:08team players that you know have
  1360. 54:10complimentary skills to your own perhaps
  1361. 54:13um I always find working in a small team
  1362. 54:16where people know things that I don't
  1363. 54:18know and where maybe I have some skills
  1364. 54:20that other people don't have as much of
  1365. 54:22you know is super fun because you're
  1366. 54:24collectively building something or
  1367. 54:26working on something that none of you
  1368. 54:28could maybe do individually. ually, but
  1369. 54:30in the process of working on that, you
  1370. 54:34actually gain a lot of new knowledge and
  1371. 54:35new skills uh for yourself and so do
  1372. 54:38they. And you you kind of want to view
  1373. 54:41your engineering or research career as
  1374. 54:44you have an amazing tool belt of
  1375. 54:46techniques. And you always want to be
  1376. 54:48adding new tools to that tool belt
  1377. 54:50because you never know when you might
  1378. 54:53come across a problem where you need
  1379. 54:55these four specialized tools rather than
  1380. 54:57these three. And adding more tools makes
  1381. 55:00it more likely that the problems you you
  1382. 55:02encounter in the future will be solvable
  1383. 55:04by you.
  1384. 55:07>> Now, one last thing. I'm pretty sure
  1385. 55:09someone in this room or multiple people
  1386. 55:13will eventually build something as
  1387. 55:15consequential as you've done with map
  1388. 55:18reduce, TPU,
  1389. 55:21distillation, etc., etc. What problem do
  1390. 55:24you hope they would be working on? Oh
  1391. 55:28yeah. I mean I I think there's a lot of
  1392. 55:30interesting problems in the world and
  1393. 55:32I'll just rattle off a few. This is not
  1394. 55:34exhaustive because the world is a very
  1395. 55:36big place and full of problems. You know
  1396. 55:39I'm particularly excited about new
  1397. 55:41approaches to hardware. You know we that
  1398. 55:43thought experiment there was kind of you
  1399. 55:45know a
  1400. 55:47you know a indication of that or much
  1401. 55:50more efficient inference hardware. You
  1402. 55:52know, I think there are radically
  1403. 55:54different kinds of algorithms for
  1404. 55:56machine learning that might be much much
  1405. 55:58more data efficient than the approaches
  1406. 56:00we're using today. If you think about
  1407. 56:02our large scale models today, they
  1408. 56:04probably see a thousand times as much
  1409. 56:05data as a human does by the age of 18.
  1410. 56:09Yet, the human by the age of 18 is
  1411. 56:11better in a lot of things and, you know,
  1412. 56:13on par uh with those frontier models
  1413. 56:16that have seen way more data. So could
  1414. 56:17you come up with much more data
  1415. 56:20efficient systems that can learn
  1416. 56:22continuously learn from their own
  1417. 56:24actions? Uh continual learning is a
  1418. 56:26really interesting thing. I think multi-
  1419. 56:28aent interactions is an interesting
  1420. 56:30thing. Um you know I think you know
  1421. 56:35creating ways of having better discourse
  1422. 56:37among people in the world uh could be
  1423. 56:39interesting. Are there ways to have much
  1424. 56:41more civil conversations and you know
  1425. 56:44helping people meet other people are all
  1426. 56:46over the world that they should know
  1427. 56:48based on their interests. You know these
  1428. 56:50are kind of interesting things. I think
  1429. 56:52there there's lots of cool things in the
  1430. 56:54world and we should all go and strive to
  1431. 56:56make even cooler things occur.
  1432. 56:59>> That sounds wonderful. Thank you so much
  1433. 57:01Jeff Dane. That's all we have today.
  1434. 57:03>> Appreciate it.
  1435. 57:05>> Thank you all.

About this transcript

This page contains the full transcript of Jeff Dean: The 1% Rule for Building in AI by Y Combinator, generated from the public captions YouTube serves with the video. The transcript has 9,671 words across 1,435 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.