Jeff Dean: The 1% Rule for Building in AI — Transcript
Full transcript
- 0:07All right. Should we go Should we get
- 0:09started, Jeff?
- 0:10>> Sure. Sounds great.
- 0:11>> All right. Jeff, welcome. And again,
- 0:13thank you so much for being here.
- 0:14Especially I just got a cold and thank
- 0:17you for being here.
- 0:17>> Yeah, I'm afraid I've lost my voice. I
- 0:19don't normally sound quite like this,
- 0:20but we'll we'll do what we can.
- 0:24>> So, um, you built map reduce, big table,
- 0:27tensorflow, the TPU, Gemini. We could
- 0:30spend a whole hour on all the things
- 0:32you've done, but what I love is that
- 0:34you're still making bold predictions in
- 0:36public. Last year, yes, last year in May
- 0:422025 at AI Ascent, you said that AI is
- 0:46at the level of a junior engineer.
- 0:50That was about a year ago. It's been How
- 0:55close are we to that prediction? Yeah, I
- 0:58mean I feel like uh the models have been
- 1:00getting a lot better at sort of
- 1:01agent-based longer running coding tasks
- 1:04and it seems pretty clear that they are
- 1:06now actually pretty capable and
- 1:08depending on exactly your definition of
- 1:10of junior engineer it seems pretty
- 1:12spot-on I would say.
- 1:15>> What did you underestimate from that
- 1:18prediction?
- 1:19Um I mean I think the
- 1:24the ability to do more and more complex
- 1:26tasks has been growing faster than I
- 1:28thought. Um and I also think uh outside
- 1:32of coding these these agent-based
- 1:34systems are are really starting to shine
- 1:36in other domains and I I think uh you
- 1:39know uh that's that's going to be an
- 1:41important trend in the future.
- 1:44>> So [snorts] give us another bold
- 1:47prediction. What do you think is going
- 1:48to be the 2027 edition?
- 1:51>> Uh I think you will see a lot more
- 1:54automation of uh ML systems themselves.
- 1:58um basically getting ML systems to
- 2:00improve their capabilities by running
- 2:03lots of experiments, breaking things
- 2:04down into subpros, you know, running
- 2:07those subpros in a tight automatic
- 2:09experimentation loop, putting the
- 2:11results together and being able to then
- 2:14uh you know, get some improved system uh
- 2:17out from that uh sort of fully automated
- 2:20problem decomposition and and automated
- 2:23experimentation that I think that's
- 2:24going to be really exciting. M
- 2:25>> I think that also applies not just to ML
- 2:28but also to other fields of science and
- 2:30engineering. Um basically anything where
- 2:32you can have a measurable objective uh I
- 2:35I think you can uh actually make a lot
- 2:37of progress these days.
- 2:40>> Now let's go back to a little bit in
- 2:42history. Back in way back in 2001 Google
- 2:47search used to run on hard drives.
- 2:49>> Yep. And you and Sanjay did the math and
- 2:53realized that at some point the whole
- 2:55search index would finally fit in all of
- 2:58the RAM of all the computers you had
- 3:02running
- 3:03and you made that radical realization
- 3:07and you basically in few days with
- 3:09Sanjay shipped in production a whole new
- 3:12search version that worked in RAM rather
- 3:14than hard drive and that was the thing
- 3:17that got Google to be so fast. Google
- 3:19searches.
- 3:21>> So history tends to remix.
- 3:25What is the it fits the memory moment
- 3:28right now in 2026 that everyone in this
- 3:31room is still
- 3:34should be thinking about and designing?
- 3:36>> Yeah. Yeah, I mean it's a little
- 3:38different, but I think uh you're going
- 3:40to see more and more
- 3:43uh uh high performance and um low energy
- 3:48uh inference hardware systems because I
- 3:51think everyone is now realizing that
- 3:53inference is the key to making you know
- 3:55these agent-based systems be available
- 3:58to more and more people and that latency
- 4:01is really important and that
- 4:03specialization of the hardware is a
- 4:05really key way you can make uh things
- 4:07that are more energy efficient and lower
- 4:10latency than more general purpose uh
- 4:12computational devices like say GPUs or
- 4:14TPUs
- 4:16>> because I think all everyone here is
- 4:18used to waiting for responses on on
- 4:21models. [clears throat]
- 4:22So
- 4:22>> waiting is no fun
- 4:25>> master speed.
- 4:27So you're saying what if we don't have
- 4:29to wait anymore?
- 4:30>> Yeah. I mean, I think we'll imagine what
- 4:32you could do with something where the
- 4:34latency is, you know, 50x better.
- 4:38>> Interesting thought.
- 4:40Now, what's one assumption that perhaps
- 4:436,000 people in this room hold that's
- 4:46already false about AI?
- 4:50>> Yeah. Uh, that's that's a good question.
- 4:51I mean I think um
- 4:55probably one thing is people don't quite
- 4:57realize how possible it is to have you
- 5:00know agent-based systems that can run
- 5:02not just for an hour or two hours on a
- 5:05problem you care about but for some
- 5:08problem domains and with highly capable
- 5:09models underlying them you can get them
- 5:12to run for days or weeks and do really
- 5:15really complicated tasks and I think
- 5:17that's you know starting some people are
- 5:20starting to see inklings of this but I
- 5:22don't think everyone has really
- 5:23internalized this and that's going to be
- 5:25really uh a pretty big deal.
- 5:28>> What's a particular task that you have
- 5:29run that has run for weeks? What what
- 5:32was it? What did the tell what did you
- 5:34tell the agents to solve? Yeah, I mean I
- 5:36think uh you can tell agents to uh go
- 5:40off and implement
- 5:43um you know completely new versions of
- 5:46software in different programming
- 5:47languages that might be you know have
- 5:49better safety properties or better
- 5:51performance properties uh that and then
- 5:53then they can go off and and actually do
- 5:55that in a you know pretty serious way.
- 5:58>> That's pretty cool.
- 6:00[clears throat] Now, one thing that
- 6:02you've been very well known for is
- 6:04you're really good at napkin math.
- 6:07Sounds funny. So, one of the stories
- 6:10about you is that back in uh 2013 when
- 6:14speech recognition started to work at
- 6:16Google, you did the nap napkin math
- 6:19where if every Google user used their
- 6:23phone and talked to it and used the
- 6:24speech recognition system for three
- 6:27minute just three minutes a day, you
- 6:29found that the system requires a Google
- 6:32server. you would have to double the
- 6:33fleet which would be really really
- 6:36expensive just to do speech translation.
- 6:39>> Yeah.
- 6:40>> And instead you basically built a custom
- 6:44ship and that was the origin story of
- 6:47the TPU.
- 6:48>> Yeah. Yeah. I mean I I sort of had done
- 6:51you know we were starting to see really
- 6:53good uh quality results on this the sort
- 6:57of deep learning based speech systems uh
- 6:59speech models we were training. um but
- 7:02they were computationally expensive
- 7:03compared to the old speech system but
- 7:05they haved the error rate. So that was
- 7:07like the equivalent of 20 years of
- 7:09advances in speech recognition in just a
- 7:11few months of like fiddling with the
- 7:13model and getting scaling it up a bit
- 7:15and getting better data. And so we
- 7:19started to get worried that if speech
- 7:20worked a lot better, people would use it
- 7:22more. And so that that back of the
- 7:24envelope calculation was really about
- 7:26that like well what if people start
- 7:29start to use speech recognition more to
- 7:31dictate emails or to talk to their phone
- 7:33or whatever. Um and yeah it turned out
- 7:37that um we realized that we needed some
- 7:41better solution than running on CPUs at
- 7:44the time. And so we came up with TPUs
- 7:47which are sort of very specialized for
- 7:50essentially low precision dense linear
- 7:52algebra which is at the heart of nearly
- 7:54all of the modern machine learning
- 7:56algorithms we we use today. And um if
- 8:00you build a specialized chip for low
- 8:02precision dense linear algebra and can't
- 8:04do anything else that turns out to be
- 8:06really useful for machine learning
- 8:07inference uh even though it can't run
- 8:09Chrome or Word or whatever. Uh, and so
- 8:12that system produced a chip a couple
- 8:14years later that was uh 30 to 80 times
- 8:19more energy efficient than CPUs and GPUs
- 8:22of the day and also much much lower
- 8:25latency like 20 to 30x lower latency
- 8:28>> which is incredible what the foundation
- 8:30that TPU has become today. No way you
- 8:32would have predicted that TPU would be
- 8:34so foundational now with transformer
- 8:36architecture which was invented way
- 8:38later before you actually invented the
- 8:40TPU. Yeah, I mean that's sort of why we
- 8:43built a general purpose linear algebra
- 8:46system, which is what a TPU is really.
- 8:48Um, because we knew ML algorithms were
- 8:51still evolving and you didn't want to
- 8:53over specialize, but you wanted to
- 8:55specialize enough that you got the
- 8:57dramatic performance benefits of we
- 9:00could have very big multiplier units. uh
- 9:03we could have you know high-speed memory
- 9:04we could have high-speed interconnect or
- 9:06later TPUs that like brought many many
- 9:09chips to bear on the same problem
- 9:11efficiently and u you know we've
- 9:14continued to scale those up and and
- 9:15improve their performance uh for over
- 9:18many many generations now
- 9:20>> incredible napkin math so
- 9:22>> what's good
- 9:23>> napkins are good
- 9:25>> so actually what's a good napkin math
- 9:28that everyone here who wants to be a
- 9:31future founder
- 9:32should run tonight to potentially build
- 9:35something as consequential as the TPU.
- 9:39>> Yeah, I mean, uh, it's always hard to
- 9:42say. Um, [clears throat] I think, uh,
- 9:46think about what problems you see in
- 9:49whatever it is you're thinking about,
- 9:51what what bottlenecks you see, and are
- 9:54there very different ways of thinking of
- 9:56the solutions to some of those problems
- 9:58that would get you, you know, an order
- 10:00of magnitude or two orders of magnitude
- 10:02better uh, performance or capability or
- 10:05whatever it is. Um, you know, because
- 10:08sometimes if you just squint at a
- 10:10problem and you think about not
- 10:13necessarily being anchored on exactly
- 10:14how that problem is solved today, but
- 10:16how you would solve it from first
- 10:18principles, you can come up with really
- 10:20good ideas that are, you know, maybe not
- 10:23what other people are thinking about.
- 10:25>> That's a good tip. [clears throat]
- 10:27>> No. Um, for everyone here who doesn't
- 10:29know, years ago, Jeff wrote a very
- 10:32famous list called the latency numbers.
- 10:36every engineer should know
- 10:38and these are numbers around for example
- 10:41how long a cache miss takes uh disk seek
- 10:46a network package traveling let's say
- 10:48from California to Netherlands
- 10:51um lots of numbers like this about
- 10:53distributed systems and systems
- 10:55engineering [clears throat]
- 10:56>> and it's been sort of taped and become
- 10:58the bible for a lot of distributed
- 10:59systems engineers
- 11:00>> okay yeah
- 11:01>> now fast forward that list is up for an
- 11:04update give us the AI edition for now
- 11:082026.
- 11:09>> Yeah, I mean I think if you looked at
- 11:10what is important in AI systems these
- 11:13days, you would want to know things like
- 11:17the bandwidth between you know your main
- 11:21memory system on your accelerator to the
- 11:23onchip memory to the um you know the
- 11:26multiplier unit or whatever. You want to
- 11:28know how much energy does it take to do
- 11:31a single multiplier operation.
- 11:34um you know uh what is the interconnect
- 11:36bandwidth between chips and how much
- 11:39does that uh how how many chips can you
- 11:42connect with that bandwidth and then if
- 11:43you go beyond that domain like what is
- 11:47the fall off in in uh network bandwidth
- 11:50when you need to talk to 10,000 strips
- 11:52instead of instead of uh 500 or
- 11:55something I think these are all really
- 11:57important numbers to to learn
- 11:59and and really affect how you think
- 12:02about solving particular kinds the
- 12:03problems.
- 12:04>> H [clears throat] and one interesting
- 12:07thing that I've heard you talk about is
- 12:08that nowadays the unit that you measure
- 12:12everything is energy.
- 12:15>> Yeah.
- 12:15>> You pointed out that doing a calculation
- 12:18or math costs about one pico.
- 12:21Uh but moving the data and doing data IO
- 12:24costs thousand times that.
- 12:26>> Yeah. Just bringing it in from HPM on an
- 12:28accelerator into the processor so it can
- 12:31actually compute on it. Yep.
- 12:33That gap kind of quietly decides what
- 12:36products are possible and how these
- 12:38algorithms in AI are built. So what are
- 12:42the kinds of problems that founders keep
- 12:44calling model problems but are in fact
- 12:47actually energy or data IO problems?
- 12:51Yeah, I mean I think the the example you
- 12:54raised of a thousandx difference in
- 12:56bringing mo moving data versus actually
- 12:59computing on it uh in in terms of energy
- 13:02is is a pretty significant one and it
- 13:04shapes a lot of aspects of what we do in
- 13:06machine learning. Um because if you
- 13:10didn't have that thousandx difference
- 13:11then you know you wouldn't have to do
- 13:14batching but you have to do batching of
- 13:16you know many examples or maybe many
- 13:18tokens at once in order to amortize that
- 13:21data movement [clears throat] so that
- 13:22you can uh you know not pay a thousandx
- 13:26slowdown but pay a 1000x divided by
- 13:28batch size uh energy cost. Um and you
- 13:32know for for really low latency batching
- 13:36is not really very good. Um so I think
- 13:39these [clears throat] kinds of things
- 13:40and the energy uh behind various
- 13:43decisions in the computer hardware we
- 13:45use really affects a lot of decisions we
- 13:48make in building higher level systems.
- 13:51>> A very concrete example is just how
- 13:53training models is done. There's this
- 13:55whole whole concept of batching the the
- 13:59data sets and running epochs. That's
- 14:01basically people perhaps may confuse
- 14:03that as a model problem, but it's really
- 14:05a systems data IO problem, right?
- 14:08>> Yeah. Yeah. I mean, you have to assemble
- 14:10batches to get better efficiency in your
- 14:13hardware. You know, ideally you might do
- 14:15batch size one training, but uh you
- 14:17know, it's um not as not as good in
- 14:21terms of efficiency. So people use re
- 14:24pretty large batches these days.
- 14:26>> Do you think uh it's possible for uh I
- 14:28know you're you're well known for uh
- 14:30taking off uh on a long week or weekend
- 14:33and coming up with this brilliant
- 14:34solution. Is there such things of Jeff
- 14:36going and working on it for a couple
- 14:38weeks and [clears throat]
- 14:38getting batch size equals one training
- 14:41done.
- 14:43>> Yeah, I've been thinking more about
- 14:44inference actually. So I think inference
- 14:46is a pretty interesting problem because
- 14:49you do want very low latency. You know
- 14:52training you don't necessarily need
- 14:53incredibly low latency. Um and I think
- 14:56there's a lot of room for specializing
- 14:58hardware more for inference than we are
- 15:00today.
- 15:02>> What are some of those interesting
- 15:03things that are on inference that you're
- 15:05really thinking a lot about?
- 15:07>> Um I mean just trying to minimize data
- 15:09movement. uh trying to think about
- 15:12incredibly
- 15:14uh low precision operations
- 15:17uh and maybe not supporting lots and
- 15:19lots of different kinds of precisions.
- 15:21Uh if you feel like you have a a good
- 15:25answer for what kinds of precision you
- 15:27need, maybe just build that into the
- 15:29hardware and and um not much else.
- 15:33which I think it brings down to a core
- 15:36[clears throat] analogy I heard from
- 15:38famous computer scientists that really
- 15:40the whole process of u AI is a big
- 15:44compression problem because in order to
- 15:48have the data to be f fully lossy and
- 15:50compress it and then restore it you
- 15:52basically need to understand it. Yeah, I
- 15:54mean if you truly understand the data,
- 15:57you should be able to compress it really
- 15:58well
- 15:59>> and now transformer architecture is
- 16:02basically one of the ways that has
- 16:04turned out to work really well.
- 16:07>> Yeah. Yeah, I would say
- 16:08>> working pretty well so far.
- 16:09>> Good work by my colleagues. [laughter]
- 16:12>> Yes. Now let's zoom out a bit. Um AI
- 16:16progress used to mean just better
- 16:18models. You could had more data trainer
- 16:20models with bigger parameters. But
- 16:23increasingly in the last years or so,
- 16:26it's everything around the model. Not
- 16:27just the model size and number of
- 16:29parameters or more data. It's everything
- 16:31around things like retrieval tools,
- 16:34memory, agent tools, and it might kind
- 16:37of get consolidated into what people
- 16:39call uh context engineering, right?
- 16:42[clears throat]
- 16:42>> Yeah. I mean I think uh the model is
- 16:45really only one piece of what you're
- 16:48trying to do which is build an overall
- 16:49system that can solve really interesting
- 16:52problems and that involves you know a
- 16:55model that knows how to use various
- 16:57tools. It maybe knows how to retrieve
- 16:59relevant information, maybe has a, you
- 17:02know, a history of other uh information
- 17:05that it has retrieved for past problems
- 17:08and it can put information into the
- 17:11context of the of the model. And the
- 17:14nice thing about that is that
- 17:15information is really clear to the
- 17:17model, unlike the training data the
- 17:19model was trained on where it's all kind
- 17:21of like trillions of tokens stirred
- 17:24together into a soup of of hundreds of
- 17:26billions or trillions of parameters, but
- 17:28it's all less clear than the actual
- 17:32context uh that the model sees directly
- 17:34for this particular problem or uses use
- 17:36case. And then I think being able to
- 17:39understand what tools are available,
- 17:43which ones are going to help me solve
- 17:44the help the model solve this next you
- 17:47know phase of the problem, how to
- 17:48decompose a problem into a sequence of
- 17:50of tool calls. Maybe trying multiple
- 17:52approaches to solve the problem and
- 17:55seeing which ones work and being able to
- 17:56evaluate that. you know this is the
- 17:58whole um you know orchestration of
- 18:02complex agent and multi-agent systems
- 18:04that I think is going to be more and
- 18:06more important and uh super exciting
- 18:08times I would say
- 18:10>> and I think the fun thing about this
- 18:11particular problem domain set is
- 18:13actually something that everyone in this
- 18:15room can actually do because before to
- 18:18train a model you needed incredible
- 18:19amount of resources incredible amount of
- 18:21access of to GPUs and data but for
- 18:24context engineering everyone here could
- 18:27do you have you just need the API to
- 18:29something like Gemini and then work on
- 18:31your own setup for your own retrieval
- 18:34your own tool calls and etc etc. So how
- 18:38does what are some tips for everyone
- 18:40here? How does everyone get better at
- 18:42and become exceptional at context
- 18:45engineering? Yeah, I mean I think uh
- 18:49[clears throat]
- 18:50a really good way to do it is to use
- 18:52these models and and sort of harnesses
- 18:55and tools and so on to try to solve
- 18:57problems and then some sometimes you can
- 19:00actually see where the models are
- 19:01failing. And often you can actually make
- 19:04the model work better and succeed at
- 19:07that kind of problem by not just
- 19:10adjusting the model parameters which is
- 19:11hard to do from the outside but from you
- 19:14know creating better guidelines for the
- 19:16model you know writing skills for the
- 19:19model to know how to use different tools
- 19:21that would be incredibly useful for
- 19:22solving this particular class of
- 19:24problem. And I think as you do that, you
- 19:28end up on this kind of improving
- 19:30self-improving of the setup that you're
- 19:33trying to use to to solve things. Uh,
- 19:36and you know that that's a really good
- 19:37way to get better at understanding what
- 19:41what additional information the model
- 19:43would want in order to become more
- 19:45capable.
- 19:46>> Can you give an example of uh some
- 19:48context engineering you personally have
- 19:50done? um I don't know skills you wrote
- 19:52tools that really made a huge different
- 19:54in your in your workflow. Yeah, I mean I
- 19:57guess uh Sanjay and I were working a few
- 19:59weeks ago and we you know we often do
- 20:03some amount of like uh performance
- 20:05improvement for very low-level libraries
- 20:08and we have a microbenchmark library
- 20:10we've written at Google where you can
- 20:12write microbenchmarks of how how long
- 20:15different kinds of operations take or
- 20:17how long does it take to populate this
- 20:18data structure whatever and sometimes
- 20:20those data structures are used on
- 20:22millions of processes across Google. So,
- 20:25it's actually pretty important to make
- 20:26sure they're high performance. And so,
- 20:28you can write microbenchmarks. Um, but
- 20:31then without an agent-based system, what
- 20:33you usually do is you measure what the
- 20:35current performance is on some
- 20:37benchmarks you care about. You make some
- 20:39modifications to improve the performance
- 20:42you hope. Then you rerun the the
- 20:45benchmarks, see where things improved.
- 20:48Um, you run a maybe a broader set of
- 20:49benchmarks, measure the cache footprint
- 20:52of things. And so we wrote a skill that
- 20:56basically taught the model how to do
- 20:57most of those things in in var in
- 21:00various sequences so that it could
- 21:01actually you know do self-improving uh
- 21:04benchmark measurement benchmark improve
- 21:06you know code changes measure the
- 21:08performance improvement and then iterate
- 21:10on that and that that seemed to work uh
- 21:12pretty well for some kinds of problems.
- 21:14And it really just is us giving the
- 21:18approach we would use as people to the
- 21:21model in a form that it could use.
- 21:23>> Wow, that seems very impressive. So
- 21:25you're saying you have this skill that
- 21:27if someone got access to it, it could do
- 21:30perform optimizations like Jeff Dean.
- 21:33Seems like the world would love this and
- 21:36is worth infinite amount of money to
- 21:38someone have access to this.
- 21:39>> Oh. Uh we actually published a document
- 21:41maybe a few months ago called
- 21:43performance hints that Sanjay and I
- 21:45wrote that's like a 30-page document
- 21:47about you know various kinds of
- 21:49performance tricks and some people have
- 21:51taken that and then given it in
- 21:53summarized form to various models and
- 21:55seen that they that model can now get
- 21:57you know better at uh per reasoning
- 21:59about performance issues in code.
- 22:01>> So you heard it all here. You could
- 22:02actually get your own optimize your own
- 22:05code like Jeff Dean if you take this
- 22:07this paper that you published when
- 22:08performance hints. Yep.
- 22:10>> It's all free available, so you should
- 22:11all try it.
- 22:12>> Very cool.
- 22:13>> Yeah.
- 22:14>> Now, you're talking about agents. Um,
- 22:16everyone here is probably building one
- 22:17or built one at some point. And I'm sure
- 22:20everyone has seen your agent go off the
- 22:23rail at perhaps like step 30 or 40. Like
- 22:26agents are great for like up to step, I
- 22:28don't know, 10 or something and then
- 22:29gets shaky at step 50. What do you think
- 22:32is the constraint today? Is it like
- 22:34context
- 22:36evaluators or just errors that compound
- 22:38because it's basically a openloop
- 22:40system?
- 22:41>> Yeah, I mean obviously we want agents to
- 22:44be able to run for very long periods of
- 22:46time because that's how they're going to
- 22:47solve more and more complicated
- 22:49problems. Um but as you as you observe
- 22:52today, you know, they sometimes stop
- 22:55working after, you know, 10 10
- 22:59interactions with the tools and so on.
- 23:01Um, and sometimes that's because the
- 23:04model is trying to do something it
- 23:05doesn't have a lot of experience doing.
- 23:07So it's been trained on a whole set of
- 23:09things and as soon as you get a little
- 23:11bit off the distribution of things it
- 23:13knows how to do then like most machine
- 23:15learning models it will you know its
- 23:18performance will suddenly will start to
- 23:20degrade and the farther you get off the
- 23:22comfort zone of what it knows how to do
- 23:25the the more likely it is to to not work
- 23:27as well. Um so there's a bunch of things
- 23:30you can do. So one is you know give the
- 23:33model skills and hints that kind of tend
- 23:35to keep it in in the uh sort of more
- 23:38brightly lit path of things it does know
- 23:40how to do. Um, I think you know having
- 23:43multi- aent systems where you have
- 23:45multiple agents trying different
- 23:47approaches and you can evaluate you have
- 23:49maybe another model or another agent
- 23:51that's evaluating which ones of those
- 23:53seem promising is another way to kind of
- 23:58in some sense search the path of pos
- 24:01search the space of possible solutions
- 24:03and stick to the ones that seem most
- 24:06promising and discard the ones that that
- 24:08didn't seem to work or maybe that went
- 24:10off the rails. or whatever. Um, and
- 24:13that's a very very useful general
- 24:15technique is you know inference time
- 24:18compute to perform search over plausible
- 24:20ways of solving the problem that can get
- 24:23much much higher performance or much
- 24:25more reliability in longunning agent
- 24:27flows.
- 24:30>> How are some ways you implemented this
- 24:32particular workflow for your agents
- 24:34internally?
- 24:35Yeah, I mean we have uh you know
- 24:37harnesses and then we have a whole set
- 24:39of skills uh particularly in the
- 24:42internal Google development environment.
- 24:43We have skills so that the agents can
- 24:45know how to use lots of our internal
- 24:48tooling for coding or for code reviews
- 24:50or for you know measuring performance or
- 24:54you know fetching log files. And um
- 24:57those are just skills that you can add
- 24:59to make the base model more capable even
- 25:01though it hasn't necessarily been
- 25:03trained on exactly the way that you know
- 25:06Google internal uh engineers would fetch
- 25:09log files from our you know proprietary
- 25:12system with the right kind of skill uh
- 25:15definition you can actually get it to
- 25:16work
- 25:17>> uh and that that improves the usefulness
- 25:19of the agents. Now let's talk about uh
- 25:23where startups can can win. This section
- 25:27is one that I personally care a lot
- 25:29about because also everyone here in this
- 25:31room needs to decide what to build in
- 25:33the future of your future founder. So
- 25:37the thing about Google is you co-design
- 25:39everything on the system from the
- 25:41processors to the products.
- 25:44um which are the layers that someone
- 25:47like Google would keep building and
- 25:49compounding being better and and where
- 25:52does a two three person team can still
- 25:55win?
- 25:57>> Yeah, I mean I think obviously Google
- 26:00and and our Gemini models and and our
- 26:02hardware infrastructure are really
- 26:04trying to build very general models that
- 26:06can do almost anything. But in in a lot
- 26:11of cases that means that we don't have a
- 26:14lot of attention on particular domains
- 26:16where perhaps a really well-designed
- 26:19surface that and maybe a model and set
- 26:22of skills or maybe a specialized model
- 26:25that uh isn't in sort of a general mix
- 26:28of of things that our models do well can
- 26:31actually have a significant advantage
- 26:33because you can build something
- 26:34delightful and you know really high
- 26:37accuracy. really high quality for a
- 26:40domain that you are really passionate
- 26:41about. And I think that's that's where
- 26:44you know the two or three people in a
- 26:46room uh building that that they're
- 26:49really excited about can have an
- 26:50advantage. Um but I I would also caution
- 26:54that the general models are definitely
- 26:57getting better at a broader and broader
- 26:59range of things. So you have to figure
- 27:00out, you know, is that thing you're
- 27:03working on, is that going to be a
- 27:04durable thing or do you think the models
- 27:07uh at the forefront are going to get
- 27:09better at that in the next six months or
- 27:1112 months or is it something they're not
- 27:13going to be able to do for a couple
- 27:15years or three years? And you know, you
- 27:17you want to weigh that as you're as
- 27:19you're deciding what to work on.
- 27:22>> So let's uh dive deeper into this. So
- 27:24the general models of course you're
- 27:26going to keep working on and keep making
- 27:27them all better.
- 27:29And how should the audience reason about
- 27:31what are those areas that uh it doesn't
- 27:35I mean h how should founder think about
- 27:38things to pick on and work on.
- 27:40>> Yeah. I mean I mean the most important
- 27:42thing is to pick something you're super
- 27:44excited about and want to build and you
- 27:46think would be useful in the world,
- 27:48right? So if you do that um that that's
- 27:51you're already way ahead uh than if you
- 27:54wake up and you're like oh I don't
- 27:55really want to do this or whatever or
- 27:57you're going to build something that is
- 27:59actually not that useful to to the world
- 28:01or to to many people. Um so I think
- 28:05that's the number one selection criteria
- 28:07I try to apply for what problem should I
- 28:10work on next. Um, second, I think you
- 28:14want to look at what the current more
- 28:16general models can do in that problem
- 28:19domain, right? You can you can test them
- 28:22with like, are they able to do this
- 28:23thing very well? And if they're
- 28:26completely failing, that's probably a
- 28:28good sign. If they're kind of able to do
- 28:30some of it but not very well, that's
- 28:33maybe not a great sign because that's a
- 28:35probably a a sign that the capability is
- 28:38starting to be present in those models
- 28:41and with more training data or larger
- 28:43scale models or or whatever it's likely
- 28:45to get better. So um you know look for
- 28:48something where the model succeeds 0% or
- 28:511% of the time not not 20%.
- 28:54>> How do you find those? I mean are those
- 28:56things effectively
- 28:58uh out of distribution from the training
- 29:00set and what exactly is the problem
- 29:02shape that fits that?
- 29:04>> Yeah, I mean I think uh sometimes it's
- 29:08uh a product that you build that might
- 29:11have access to particular kind of data
- 29:13that the underlying model might not the
- 29:15a general model. So it might be you're
- 29:18building something to help users
- 29:19organize all their own personal
- 29:21information and the model won't
- 29:22necessarily have access to that. And so
- 29:24there you can have a big advantage
- 29:26because all of a sudden your model has
- 29:28visibility or your product has
- 29:31visibility into important data. Um it
- 29:35could be some incredibly hard problem
- 29:38where if you get the right training data
- 29:40and you can train a more specific model
- 29:42than a general purpose one, you can
- 29:44actually do that in a very affordable
- 29:46way. you maybe it doesn't take that much
- 29:48compute to train a a niche model for
- 29:50this particular problem, but you can get
- 29:52something that's highly accurate. That
- 29:54can sometimes be a a really good uh
- 29:57building block for for solving a
- 30:00important problem that is maybe not
- 30:03handled very well by the general model.
- 30:04>> I think that's interesting. I think
- 30:06there are basically two paths. The first
- 30:07path is uh a little bit funny is uh you
- 30:10guys are organizing the world's
- 30:12information.
- 30:12>> Yeah,
- 30:13>> that's probably kind of well covered.
- 30:15>> Yeah. but organizing your personal
- 30:17information that's open which is funny.
- 30:21>> Yeah.
- 30:21>> And then the second path um you talked
- 30:24about more specialized models in certain
- 30:27domains. Can you tell us more about what
- 30:29are some of these domains?
- 30:30>> Yeah. I mean I think like if you look at
- 30:33uh my colleagues work on say alpha fold
- 30:35that was a very specific model for uh
- 30:38protein folding and it was highly
- 30:40successful
- 30:41um and was able to really handle that
- 30:44domain quite well so that all of a
- 30:46sudden you now have this amazing tool
- 30:48and model that can give you answers to
- 30:50questions about proteins and their
- 30:52structure um really effectively um but
- 30:55it's not a general model it's a very
- 30:58specific one and there are other I
- 31:01domains where that kind of approach can
- 31:03work really well. Uh maybe in material
- 31:05science or chip design or things like
- 31:07that that uh will enable you to leverage
- 31:11the capabilities of a very accurate but
- 31:14but niche model uh to do things that are
- 31:18hard today.
- 31:19>> That's a good example. So if some of you
- 31:21find a problem that's similar shape like
- 31:23alpha fold could be a good problem to
- 31:24work on. Now let's assume you found a
- 31:26problem to work on. We're going to talk
- 31:28a bit about how do you become a AI
- 31:30native founder? How do you really become
- 31:32good at it? Uh you in the past said that
- 31:35managing a fleet of agents, it's like 50
- 31:37or 100 agents is all about writing
- 31:40really good crisp design docs or specs.
- 31:45>> And how do people get good at that? What
- 31:47what do those look like? Yeah, I mean I
- 31:50think uh
- 31:54it's
- 31:55you you'll have a lot more success when
- 31:58working with your virtual agents if you
- 32:00can clearly specify what it is you want.
- 32:03And the clearer you are on what it is
- 32:04you want, the more the agent will have
- 32:07sort of guidelines and sort of rules of,
- 32:10you know, an outline of what it is
- 32:12trying to accomplish. Um whereas if you
- 32:15don't specify very much stuff, the agent
- 32:17has to sort of infer what it is you
- 32:19meant. And in many cases, it might infer
- 32:22things that are different than what you
- 32:23imagined. So we've always told computer
- 32:26scientists from the very beginning that
- 32:27really it's really important to specify
- 32:29what it is, what's the software that
- 32:31you're writing is trying to accomplish
- 32:34before then going and writing it. And so
- 32:36now we actually have agent-based systems
- 32:38that can do the writing, but the
- 32:40importance of specifying what what it is
- 32:42you want has actually gone up because
- 32:45before you'd be handing it off to a very
- 32:48intelligent human who maybe has context
- 32:50or can ask you follow-up questions. Um,
- 32:53and agents can sometimes do that, but I
- 32:54I think clear specifications is is a
- 32:58really good idea. Um, and to give you an
- 33:00example of a a a use of a coding agent
- 33:04that works extremely well is you can ask
- 33:07today's models to translate software
- 33:10from one computer language to another
- 33:12very effectively because in that case
- 33:15you actually have a incredibly detailed
- 33:18specification. You have the whole
- 33:19software that says what the system is
- 33:21supposed to do. And so if you have a
- 33:23Python implementation of something and
- 33:25you want a Go implementation of it, you
- 33:28know, that is something that the models
- 33:29seem incredibly capable at doing these
- 33:31days because you can it can sort of take
- 33:34all the tests that are in Python, make
- 33:36sure they pass in the Go version,
- 33:38translate the tests to Go, um you know,
- 33:41compare uh behavioral differences
- 33:43between the implementations until there
- 33:45aren't any um and be you know, highly
- 33:48effective because that spec is so clear.
- 33:51Hm. Now let's assume now every founder
- 33:53gets good at running hundreds of agents
- 33:55at the same time and all the code is
- 33:57written for them by the agents. What
- 33:59becomes the scarce skill?
- 34:03>> Yeah, I mean I think it's really having
- 34:07incredibly good taste in what you ask
- 34:09your agents to work on, right? That is
- 34:11the the crux of you know from my
- 34:15background uh a research problem. You
- 34:17know, a researcher can have all the
- 34:19tools and all the techniques, but often
- 34:22most of the battle is what problem are
- 34:25you gonna spend your time on? And if you
- 34:27pick the problem well and you succeed in
- 34:30in in solving it, that's way better than
- 34:33if you, you know, uh, delightfully
- 34:36execute a research investigation into a
- 34:39rather boring problem. And so that high
- 34:42level wisdom of what to work on, I think
- 34:44is incredibly important. And I think
- 34:46models are not necessarily going to be
- 34:47that good at it. So you're going to have
- 34:50people steering
- 34:54uh a lot of AI assisted computation in
- 34:57order to accomplish great things and
- 34:59more quickly. Um but that essence of of
- 35:02what it is you want your models to do is
- 35:05the the the key thing you should focus
- 35:07on.
- 35:08>> So let's talk a bit more about taste
- 35:09because it gets talked a lot about right
- 35:12now in this current era with agent
- 35:14coding. How do you exactly build taste
- 35:17and do that? I mean, yeah, that sounds
- 35:20so esoteric. How do you make it
- 35:22concrete?
- 35:23>> Yeah, I mean it it is a difficult thing.
- 35:26It's not like there's a measurable
- 35:28objective of of taste in a lot of cases.
- 35:31Um, I think some of it is from
- 35:33experience. You know, working on a lot
- 35:36of different problems in the past kind
- 35:38of teaches you about what kinds of
- 35:40problems might be interesting in the
- 35:42future or what kinds of things might be
- 35:45just barely possible by cobbling
- 35:48together these previous approaches and
- 35:51then some open problems you might have
- 35:53to work on in order to get to something
- 35:55kind of magical or or you know, highly
- 35:57useful. Um,
- 36:00another way you can get more experience
- 36:02for yourself is to just write down a
- 36:06bunch of things you think might be
- 36:07important in the next 12 months. And
- 36:10maybe you pick one of them to work on,
- 36:12but go back and evaluate in 12 months of
- 36:15these other things, which ones actually
- 36:18seemed important or which ones did other
- 36:20people in the world go out and and
- 36:22create and which ones did they did not
- 36:24seem to do yet. um that can give you a
- 36:27lot more samples for your own sort of
- 36:29taste creation uh capability. Um and and
- 36:33that's an important skill to have.
- 36:35>> I think a third way we were talking
- 36:37earlier was doing very crazy thought
- 36:41experiments.
- 36:42>> Oh yeah, that's another good way. I mean
- 36:44I think uh sometimes it's good to
- 36:49not take as a given things that most
- 36:52people seem to take as a as a given. Um
- 36:55so I was doing a crazy thought
- 36:56experiment with some colleagues the
- 36:58other day about you know
- 37:02for 60 years the whole silicon
- 37:06uh chip design industry uh design and
- 37:09fabrication industry have
- 37:12you know done tremendous work to make
- 37:15smaller and smaller scale transistors
- 37:17that are uh very low error rate right
- 37:22like because what what the assumption
- 37:24that we want is that every chip we
- 37:25manufacture of the same design should be
- 37:28identical to every other chip.
- 37:30>> You don't want any bits to flip.
- 37:31Everything
- 37:31>> no bits should flip. There's all kinds
- 37:33of things you there's all kinds of error
- 37:35margins built into you know memories
- 37:38have ECC memory these days. um you know
- 37:41at the at the macro scale we don't make
- 37:44that assumption when we're building
- 37:46large scale distributed systems right we
- 37:49we build reliable large scale
- 37:51distributed file systems out of
- 37:53unreliable parts right like individual
- 37:56discs can fail but your data should be
- 37:58safe and so we have mechanisms at a
- 38:00higher level to enable us to have um you
- 38:04know three copies of the data on three
- 38:06different machines and three different
- 38:07racks so that if any rack switch or
- 38:09individual ual machine or disk fails,
- 38:11you still have your data. We have read
- 38:13Solomon encoding techniques. Um, but we
- 38:17don't seem to do this at a really
- 38:19extreme level in the uh sort of
- 38:23transistor level scale of of the
- 38:26technology we're working on. So what
- 38:28would h basically a interesting thought
- 38:30experiment is what would happen if you
- 38:32tried to build a system out of
- 38:34transistors that might have you know 20
- 38:37errors per day.
- 38:38>> Oh my god. rather than one every million
- 38:40years, right? That would be a very
- 38:43different design point and might be
- 38:45might enable you to do really
- 38:46interesting things in the fabrication
- 38:48side of things. You have very different
- 38:50kind of design methodologies because if
- 38:52you want to get a signal from here to
- 38:54there, you and you have these super
- 38:56unreliable transistors. You might have
- 38:58very different ways of signaling. You
- 39:00might send it along multiple redundant
- 39:02paths uh in order to make sure that it
- 39:04gets along one of them. Um, and I think
- 39:07that would be a pretty interesting set
- 39:08of thought experiments. I'm not saying
- 39:10we should go do this, but you know,
- 39:12that's the kind of thing where you do
- 39:13want to, you know, occasionally question
- 39:17assumptions. Now, oftentimes these
- 39:20thought experiments don't work out
- 39:22because there are very good reasons
- 39:24that, you know, for the last 50 years,
- 39:26we've done this thing this way and not
- 39:28that way. But it it's good to kind of
- 39:30revisit those every so often.
- 39:32>> That is so wild. Well, I mean, it's
- 39:33starting to rhyme a lot with
- 39:35neuromorphic computing or the human
- 39:37brain and and how nature works.
- 39:39>> I mean, exactly like signals in our
- 39:41brain are not especially reliable from
- 39:43getting one place to another. And so, I
- 39:46think in brains when there are really
- 39:48important things you need to get from
- 39:50one place to another, there are multiple
- 39:51pathways that that enable you to sort of
- 39:54do that.
- 39:56>> What is uh in I mean, you have such an
- 39:58impressive career. What is one of these
- 40:01crazy assumptions that you threw out of
- 40:02the window that actually built a
- 40:04consequential system in the past?
- 40:09>> Yeah, I mean I guess uh
- 40:10>> that worked out actually.
- 40:13>> Yeah, I mean I think uh well TPUs is a
- 40:15good example like being able to
- 40:17specialize hardware for a very niche
- 40:20>> problem domain before that problem
- 40:22domain seemed as important as it is
- 40:24today uh is one thought experiment. Um
- 40:28you know I think the
- 40:30the origin of map produce is another
- 40:33good example.
- 40:34So we had worked the you know my Sanjay
- 40:37and myself and a number of other
- 40:39colleagues had worked on various
- 40:41iterations of the crawling and indexing
- 40:43system at Google and you know we'd sort
- 40:46of written lots of hand parallelized
- 40:49code with lots of checkpointing to make
- 40:51sure it would be robust and reliable if
- 40:53it was running on a 100 computers or a
- 40:55thousand computers and some of those
- 40:57died. Um, but that code tended to be
- 41:02intermixed with the actually relatively
- 41:04simple thing you often were trying to do
- 41:06like I just want to like look at all the
- 41:09contents of all the web pages and then
- 41:11compute on the side a mapping from URL
- 41:14to you know what language is this page
- 41:16in the text of this page. Um, and it
- 41:20would get obscured by all this kind of
- 41:22other code for parallelization and
- 41:24reliability. And so we sort of
- 41:27remembered our training in functional
- 41:29languages and realized we could squint
- 41:31at those problems and developed this map
- 41:33produce abstraction that you could have
- 41:36above this implementation and then below
- 41:39the implementation you could put all the
- 41:42checkpointing and reliability mechanisms
- 41:45into that lower level library that
- 41:47everything could then build on. And so
- 41:49that became a hugely successful way of
- 41:52of dealing with very large scale
- 41:53computations at Google in a robust and
- 41:55reliable way. From that thought
- 41:58experiment of like well if we squint at
- 42:00it could we find lots of problems that
- 42:01fit into this abstraction.
- 42:03>> That's impressive. So this thought
- 42:06experiment led you to create map reduce.
- 42:08>> Yeah. Awesome.
- 42:09>> Now let's go back to you talked a bit
- 42:11about um about your interest right now
- 42:15working on a lot of customized hardware.
- 42:17So right now alpha chip
- 42:19>> lays out chips. Now you also got alpha
- 42:21evolve that proposes solutions,
- 42:24>> evaluates them and keeps all the ones
- 42:26that work. Seems like you're starting to
- 42:27build all these system that can compound
- 42:30and build AI that builds AI.
- 42:32>> Yeah. I mean I think more generally
- 42:34there's a there's this sort of
- 42:39the foundation of the scientific method
- 42:41of you propose an experiment you
- 42:44implement what you need to run the
- 42:46experiment and you evaluate the
- 42:48experiment and then you get results from
- 42:50that and I think there are more and more
- 42:52problems that are now possible to
- 42:55implement where that whole loop of
- 42:58running you know not just a few
- 43:00experiments but running many many
- 43:01experiments because you're able to
- 43:03automate that loop and make the latency
- 43:05of that loop extremely low is going to
- 43:08be really really important. It's going
- 43:09to enable us to tackle you know lots of
- 43:12different problem domains in science and
- 43:14engineering and machine learning uh
- 43:17model design itself and also in
- 43:20engineering tasks like designing chips.
- 43:23And so if you can actually do those
- 43:25things in an automated way and have some
- 43:28orchestration framework that can take
- 43:31very high level objectives and break
- 43:33them down into subpros and each of those
- 43:35subpros can be one of these automated
- 43:38loop that is exploring the best way to
- 43:40solve that sub problem and then a
- 43:43orchestration framework that can put
- 43:46together subpros solutions into a you
- 43:50know the overall solution for the higher
- 43:52level problem that's going to be really
- 43:54impactful and it's really really
- 43:56important and I think it'll enable us to
- 43:58do you know accelerate machine learning
- 44:01progress it'll enable us to accelerate
- 44:03science and enable us to accelerate
- 44:05engineering and I think that's that's
- 44:07going to be amazing
- 44:09>> that sounds awesome I mean it sounds
- 44:10like a lot of fields basically where you
- 44:12can have very good evaluators and maybe
- 44:15adjacent to basically things that can be
- 44:17formally verified right those are ripe
- 44:20for AI systems that can self-improve
- 44:22Yeah, I think in a lot of cases
- 44:25sometimes your evaluators need to be
- 44:27made much faster. Mhm.
- 44:28>> So as an example, my colleagues did some
- 44:32work maybe a decade ago on um some uh
- 44:36problems in quantum chemistry where
- 44:38you're trying to understand the
- 44:39properties of a particular molecule and
- 44:41you can you know generate some molecule
- 44:44configuration and then you want to
- 44:45understand what properties it has. And
- 44:48so you can run a very computationally
- 44:50intensive density functional theory
- 44:52simulator which is something that might
- 44:54take like a a night of computation to
- 44:57tell you the answer for one thing. Um
- 45:00but what my colleagues did was
- 45:04take a bunch of output from those
- 45:07simulation runs the input molecule
- 45:09configurations and the outputs of the
- 45:11the expensive simulator and then use it
- 45:14to train a neural approximation to the
- 45:16simulator. So this is now a validation
- 45:19device, but instead of it taking a
- 45:22night, they made something that was
- 45:25300,000 times faster.
- 45:26>> Wow.
- 45:27>> And nearly as accurate as running the
- 45:28full scale simulator. So now that
- 45:32completely changes how you would do
- 45:33science, right? Because now you have 10
- 45:35million things to screen. you know, you
- 45:38could do that while you go to lunch
- 45:40rather than it being a six-month
- 45:42endeavor where you could try to scrape
- 45:44together enough compute to to run all
- 45:46these simulations. And I think there's a
- 45:50lot of room in a lot of domains for much
- 45:53faster validation models, possibly
- 45:55learned valu validation models that can
- 45:59uh you know get you a a approximation to
- 46:02the true answer much much more rapidly.
- 46:04And that changes how those experimental
- 46:06loops can be thought of and how quickly
- 46:08you can go around those loops.
- 46:10>> What are some of the
- 46:13spaces and problems that you're super
- 46:14excited that this super sped up
- 46:17scientific method is going to solve or
- 46:20achieve? What particular problems or
- 46:22spaces?
- 46:23>> Yeah, I mean I think uh
- 46:28well clearly machine learning itself is
- 46:30one, right? So can we have a model that
- 46:32is able to recursively self-improve
- 46:35itself by running lots of experiments
- 46:37and you know if you think about how
- 46:39models are improved today in large
- 46:41research teams you know what usually
- 46:44happens is people think of some ideas
- 46:46they run a bunch of smallcale
- 46:48experiments they see if those small
- 46:50scale experiments worked out well if so
- 46:53they take the most promising ones of
- 46:54those they try them at larger scale and
- 46:56that gets then evaluated and then the
- 47:00results get integr ated together into
- 47:02you know a new recipe for your model. Um
- 47:06but I think there's no uh you know real
- 47:09impediment to making that be a much more
- 47:11automated loop where the model itself
- 47:15decides it's going to explore or maybe
- 47:17with a nudge from some people uh at the
- 47:20various highest level like oh why don't
- 47:22you try some new ideas around model
- 47:24architectures that incorporate this and
- 47:27then it will go run lots of experiments
- 47:30uh see which ones work and then those
- 47:31will get incorporated at a much more
- 47:33rapid rate and uh you know effectively
- 47:36you want to optimize you know your
- 47:39discoveries per unit of compute input.
- 47:44>> Very cool.
- 47:44>> Yeah.
- 47:45>> Now going back to the room as all of you
- 47:49will become at some point founders or
- 47:51start your careers you will probably
- 47:54collect lots of rejections. That will
- 47:56happen. Uh it has happened to you too
- 47:59Jeff. I mean there's a story that in
- 48:022014
- 48:04you with Jeff Hinton and Oral Fin wrote
- 48:09a paper on distillation
- 48:12>> which has to do with taking a big
- 48:15teacher model to train a much smaller
- 48:18and more efficient model that's a lot
- 48:21cheaper to compute less model parameters
- 48:24and it has become a trick that everyone
- 48:27is using right now in industry. Yeah.
- 48:30>> And
- 48:32the thing is this paper got rejected at
- 48:34Europe.
- 48:36>> Yeah. I mean Yeah. I mean I think I
- 48:39don't fault the program committee
- 48:40because you know a lot of times a paper
- 48:44gets three reviews and someone will look
- 48:46at one of the reviewers will look at it
- 48:48and in this case they said oh it's
- 48:50unlikely to have significant impact.
- 48:52>> Unlikely to have significant impact. But
- 48:53you know I think you know when we wrote
- 48:55the paper we actually saw this was a
- 48:57super important problem because we knew
- 49:00making cheaper highly capable models
- 49:03from larger scale models was something
- 49:04we desperately wanted to do because we
- 49:07wanted to serve models to more and more
- 49:09people in many different domains like
- 49:11speech or vision. Um but you know
- 49:13sometimes the reviewer maybe didn't have
- 49:15that that experience because maybe
- 49:16they're not thinking about you know
- 49:19largecale AI services and are thinking
- 49:22about you know is this a fundamental
- 49:24advance um so so you know it gets
- 49:27rejected every so often that's fine we
- 49:29put it on archive people read it people
- 49:31use it it's all good uh and you know we
- 49:34do use it in making our flash models for
- 49:37example from our larger scale pro model
- 49:40that's partly why our flash models for
- 49:41example in Gemini are so capable uh
- 49:45relative to their size and and speed.
- 49:47>> They're some of the best in the
- 49:48benchmark for their model size class.
- 49:50Yeah. Just impressive. And I think part
- 49:52of the lesson is that even if you get
- 49:54rejected, keep going.
- 49:56>> Yeah. That's that's the lesson I would
- 49:58distill from that.
- 50:00[laughter]
- 50:01>> Um no, I think the fun thing is that you
- 50:04basically join when you when you join
- 50:05Google as a 20 person startup back in
- 50:081999.
- 50:09Now, if you were to take the young Jeff
- 50:12Dean from way back then to teleransport
- 50:16him to now today.
- 50:18>> Yeah.
- 50:18>> In this era with your skills.
- 50:20>> I'm feeling so vigorous and and young
- 50:23now.
- 50:24>> Um what would you do? Do you join a
- 50:26frontier lab, start a company? I don't
- 50:29know what what would you do? The c the
- 50:34Jeff theme today 25-year-old Jeff theme.
- 50:36>> Yeah. I mean,
- 50:39it's always hard to say and it's a very
- 50:41personal choice of what it is you want
- 50:42to spend your time on. Um, to me, some
- 50:46of the most important questions are,
- 50:49are you going to work on something you
- 50:52really care about, will you're working
- 50:55on that? And if you're able to make
- 50:57progress on it with a bunch of
- 50:59colleagues you like working with uh if
- 51:02you're able to make collectively solve
- 51:04it or make progress on it, will that
- 51:07make a difference in the world in some
- 51:08positive way, right? Like will you
- 51:10suddenly be able to do something and
- 51:12offer that service to you know
- 51:15partically
- 51:19help biochemists or something or maybe
- 51:21it's a broader thing. It'll help
- 51:22programmers or it will help all
- 51:24consumers. uh on the internet or or
- 51:27other things. Um what you you know what
- 51:32you should strive to do is to have
- 51:33impact in the world that is positive and
- 51:36to work with people you enjoy working
- 51:38with and to you know uh work hard and
- 51:42and do your best. Um so in terms of say
- 51:46the particular trade-off you offered
- 51:48joining a frontier lab versus say
- 51:50starting a company with just one or two
- 51:52or three of you you and your close
- 51:54friends. Um, I think those are different
- 51:57experiences, right? In a in a large
- 51:59established organization, you have some
- 52:02structure. You have lots and lots of
- 52:04amazing colleagues who know lots of
- 52:06things you don't. Um, you have lots of
- 52:10interesting problems that uh you can
- 52:12work on and h you already have a
- 52:15platform for impact by your work, you
- 52:18know, influencing lots and lots of
- 52:20people in the world already. Um and then
- 52:22as a very small startup,
- 52:26you know, you have to have something
- 52:28you're passionate about and there's a
- 52:31lot of risk in taking on, you know,
- 52:34working on that particular problem in a
- 52:36way that uh you're going to succeed and
- 52:38you're going to grow a, you know, an
- 52:40endeavor in order to do that. But that
- 52:42can also be incredibly rewarding, I
- 52:44would imagine. So I I think um you know
- 52:48it's really up to personal taste but but
- 52:50at the very least regardless of what
- 52:53path you take ask yourself if I work on
- 52:56this problem and the best possible
- 52:58outcome happens you know will the world
- 53:01be a lot better in some way or will the
- 53:03world go eh that's kind of cool but
- 53:05whatever.
- 53:06>> Uh that's not the kind of thing you
- 53:08should spend your time on.
- 53:10Now let's talk a bit about more about
- 53:12that second path of working with people
- 53:15that you really like in a small team.
- 53:18You've been able to be an incredible
- 53:20mentor and manager to many many
- 53:22engineers and you've been able to build
- 53:25huge systems and what are some some of
- 53:29the lessons for everyone here on how to
- 53:32get the most and how to work with smart
- 53:34people or find smart people?
- 53:36Yeah, I mean,
- 53:39you always want to find people who have
- 53:42really good skills in some some area
- 53:45that's needed in, you know, a team
- 53:47you're trying to form, whether that's
- 53:50inside a company or uh starting a
- 53:52company. Um, but you also want to find
- 53:55people that are people you delight being
- 53:59around, right? because you're going to
- 54:00spend a lot of time around people
- 54:02working on really hard problems and you
- 54:05want people who are low ego that are
- 54:08team players that you know have
- 54:10complimentary skills to your own perhaps
- 54:13um I always find working in a small team
- 54:16where people know things that I don't
- 54:18know and where maybe I have some skills
- 54:20that other people don't have as much of
- 54:22you know is super fun because you're
- 54:24collectively building something or
- 54:26working on something that none of you
- 54:28could maybe do individually. ually, but
- 54:30in the process of working on that, you
- 54:34actually gain a lot of new knowledge and
- 54:35new skills uh for yourself and so do
- 54:38they. And you you kind of want to view
- 54:41your engineering or research career as
- 54:44you have an amazing tool belt of
- 54:46techniques. And you always want to be
- 54:48adding new tools to that tool belt
- 54:50because you never know when you might
- 54:53come across a problem where you need
- 54:55these four specialized tools rather than
- 54:57these three. And adding more tools makes
- 55:00it more likely that the problems you you
- 55:02encounter in the future will be solvable
- 55:04by you.
- 55:07>> Now, one last thing. I'm pretty sure
- 55:09someone in this room or multiple people
- 55:13will eventually build something as
- 55:15consequential as you've done with map
- 55:18reduce, TPU,
- 55:21distillation, etc., etc. What problem do
- 55:24you hope they would be working on? Oh
- 55:28yeah. I mean I I think there's a lot of
- 55:30interesting problems in the world and
- 55:32I'll just rattle off a few. This is not
- 55:34exhaustive because the world is a very
- 55:36big place and full of problems. You know
- 55:39I'm particularly excited about new
- 55:41approaches to hardware. You know we that
- 55:43thought experiment there was kind of you
- 55:45know a
- 55:47you know a indication of that or much
- 55:50more efficient inference hardware. You
- 55:52know, I think there are radically
- 55:54different kinds of algorithms for
- 55:56machine learning that might be much much
- 55:58more data efficient than the approaches
- 56:00we're using today. If you think about
- 56:02our large scale models today, they
- 56:04probably see a thousand times as much
- 56:05data as a human does by the age of 18.
- 56:09Yet, the human by the age of 18 is
- 56:11better in a lot of things and, you know,
- 56:13on par uh with those frontier models
- 56:16that have seen way more data. So could
- 56:17you come up with much more data
- 56:20efficient systems that can learn
- 56:22continuously learn from their own
- 56:24actions? Uh continual learning is a
- 56:26really interesting thing. I think multi-
- 56:28aent interactions is an interesting
- 56:30thing. Um you know I think you know
- 56:35creating ways of having better discourse
- 56:37among people in the world uh could be
- 56:39interesting. Are there ways to have much
- 56:41more civil conversations and you know
- 56:44helping people meet other people are all
- 56:46over the world that they should know
- 56:48based on their interests. You know these
- 56:50are kind of interesting things. I think
- 56:52there there's lots of cool things in the
- 56:54world and we should all go and strive to
- 56:56make even cooler things occur.
- 56:59>> That sounds wonderful. Thank you so much
- 57:01Jeff Dane. That's all we have today.
- 57:03>> Appreciate it.
- 57:05>> Thank you all.
About this transcript
This page contains the full transcript of Jeff Dean: The 1% Rule for Building in AI by Y Combinator, generated from the public captions YouTube serves with the video. The transcript has 9,671 words across 1,435 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.