Andrej Karpathy: Software Is Changing (Again) — Transcript
Full transcript
- 0:01Please welcome former director of AI
- 0:04Tesla Andre Carpathy.
- 0:07[Music]
- 0:11Hello.
- 0:14[Music]
- 0:19Wow, a lot of people here. Hello.
- 0:22Um, okay. Yeah. So I'm excited to be
- 0:24here today to talk to you about software
- 0:27in the era of AI. And I'm told that many
- 0:30of you are students like bachelors,
- 0:32masters, PhD and so on. And you're about
- 0:34to enter the industry. And I think it's
- 0:36actually like an extremely unique and
- 0:37very interesting time to enter the
- 0:38industry right now. And I think
- 0:41fundamentally the reason for that is
- 0:43that um software is changing uh again.
- 0:47And I say again because I actually gave
- 0:49this talk already. Um but the problem is
- 0:52that software keeps changing. So I
- 0:54actually have a lot of material to
- 0:55create new talks and I think it's
- 0:56changing quite fundamentally. I think
- 0:58roughly speaking software has not
- 1:00changed much on such a fundamental level
- 1:02for 70 years. And then it's changed I
- 1:04think about twice quite rapidly in the
- 1:06last few years. And so there's just a
- 1:08huge amount of work to do a huge amount
- 1:09of software to write and rewrite. So
- 1:12let's take a look at maybe the realm of
- 1:14software. So if we kind of think of this
- 1:16as like the map of software this is a
- 1:17really cool tool called map of GitHub.
- 1:20Um this is kind of like all the software
- 1:21that's written. Uh these are
- 1:23instructions to the computer for
- 1:24carrying out tasks in the digital space.
- 1:26So if you zoom in here, these are all
- 1:28different kinds of repositories and this
- 1:30is all the code that has been written.
- 1:31And a few years ago I kind of observed
- 1:33that um software was kind of changing
- 1:35and there was kind of like a new type of
- 1:37software around and I called this
- 1:39software 2.0 at the time and the idea
- 1:42here was that software 1.0 is the code
- 1:44you write for the computer. Software 2.0
- 1:46know are basically neural networks and
- 1:48in particular the weights of a neural
- 1:50network and you're not writing this code
- 1:53directly you are most you are more kind
- 1:55of like tuning the data sets and then
- 1:56you're running an optimizer to create to
- 1:58create the parameters of this neural net
- 2:00and I think like at the time neural nets
- 2:02were kind of seen as like just a
- 2:03different kind of classifier like a
- 2:04decision tree or something like that and
- 2:06so I think it was kind of like um I
- 2:09think this framing was a lot more
- 2:10appropriate and now actually what we
- 2:12have is kind of like an equivalent of
- 2:13GitHub in the realm of software 2.0 And
- 2:15I think the hugging face is basically
- 2:18equivalent of GitHub in software 2.0.
- 2:20And there's also model atlas and you can
- 2:22visualize all the code written there. In
- 2:24case you're curious, by the way, the
- 2:25giant circle, the point in the middle,
- 2:28uh these are the parameters of flux, the
- 2:30image generator. And so anytime someone
- 2:32tunes a on top of a flux model, you
- 2:34basically create a git commit uh in this
- 2:37space and uh you create a different kind
- 2:39of a image generator. So basically what
- 2:41we have is software 1.0 is the computer
- 2:43code that programs a computer. Software
- 2:452.0 are the weights which program neural
- 2:48networks. Uh and here's an example of
- 2:50Alexet image recognizer neural network.
- 2:53Now so far all of the neural networks
- 2:55that we've been familiar with until
- 2:56recently where kind of like fixed
- 2:58function computers image to categories
- 3:01or something like that. And I think
- 3:03what's changed and I think is a quite
- 3:05fundamental change is that neural
- 3:06networks became programmable with large
- 3:09language models. And so I I see this as
- 3:12quite new, unique. It's a new kind of a
- 3:14computer and uh so in my mind it's uh
- 3:18worth giving it a new designation of
- 3:19software 3.0. And basically your prompts
- 3:22are now programs that program the LLM.
- 3:25And uh remarkably uh these uh prompts
- 3:28are written in English. So it's kind of
- 3:30a very interesting programming language.
- 3:33Um so maybe uh to summarize the
- 3:36difference if you're doing sentiment
- 3:37classification for example you can
- 3:39imagine writing some uh amount of Python
- 3:42to to basically do sentiment
- 3:44classification or you can train a neural
- 3:46net or you can prompt a large language
- 3:47model. Uh so here this is a few short
- 3:50prompt and you can imagine changing it
- 3:51and programming the computer in a
- 3:52slightly different way. So basically we
- 3:54have software 1.0 software 2.0 and I
- 3:57think we're seeing maybe you've seen a
- 3:59lot of GitHub code is not just like code
- 4:01anymore. there's a bunch of like English
- 4:03interspersed with code and so I think
- 4:05kind of there's a growing category of
- 4:07new kind of code. So not only is it a
- 4:09new programming paradigm, it's also
- 4:10remarkable to me that it's in our native
- 4:12language of English. And so when this
- 4:14blew my mind a few uh I guess years ago
- 4:17now I tweeted this and um I think it
- 4:20captured the attention of a lot of
- 4:21people and this is my currently pinned
- 4:23tweet uh is that remarkably we're now
- 4:25programming computers in English. Now,
- 4:28when I was at uh Tesla, um we were
- 4:31working on the uh autopilot and uh we
- 4:34were trying to get the car to drive and
- 4:37I sort of showed this slide at the time
- 4:39where you can imagine that the inputs to
- 4:41the car are on the bottom and they're
- 4:43going through a software stack to
- 4:44produce the steering and acceleration
- 4:47and I made the observation at the time
- 4:48that there was a ton of C++ code around
- 4:51in the autopilot which was the software
- 4:521.0 code and then there was some neural
- 4:54nets in there doing image recognition
- 4:56and uh I kind of observed that over time
- 4:58as we made the autopilot better
- 5:00basically the neural network grew in
- 5:02capability and size and in addition to
- 5:05that all the C++ code was being deleted
- 5:08and kind of like was um and a lot of the
- 5:12kind of capabilities and functionality
- 5:14that was originally written in 1.0 was
- 5:16migrated to 2.0. So as an example, a lot
- 5:19of the stitching up of information
- 5:20across images from the different cameras
- 5:22and across time was done by a neural
- 5:24network and we were able to delete a lot
- 5:26of code and so the software 2.0 stack
- 5:29quite literally ate through the software
- 5:32stack of the autopilot. So I thought
- 5:34this was really remarkable at the time
- 5:35and I think we're seeing the same thing
- 5:37again where uh basically we have a new
- 5:39kind of software and it's eating through
- 5:40the stack. We have three completely
- 5:42different programming paradigms and I
- 5:44think if you're entering the industry
- 5:45it's a very good idea to be fluent in
- 5:47all of them because they all have slight
- 5:49pros and cons and you may want to
- 5:50program some functionality in 1.0 or 2.0
- 5:53or 3.0. Are you going to train
- 5:54neurallet? Are you going to just prompt
- 5:55an LLM? Should this be a piece of code
- 5:57that's explicit etc. So we all have to
- 5:59make these decisions and actually
- 6:00potentially uh fluidly trans transition
- 6:03between these paradigms. So what I
- 6:06wanted to get into now is first I want
- 6:09to in the first part talk about LLMs and
- 6:11how to kind of like think of this new
- 6:13paradigm and the ecosystem and what that
- 6:15looks like. Uh like what are what is
- 6:17this new computer? What does it look
- 6:18like and what does the ecosystem look
- 6:20like? Um I was struck by this quote from
- 6:23Anduring actually uh many years ago now
- 6:25I think and I think Andrew is going to
- 6:27be speaking right after me. Uh but he
- 6:29said at the time AI is the new
- 6:30electricity and I do think that it um
- 6:33kind of captures something very
- 6:34interesting in that LLMs certainly feel
- 6:36like they have properties of utilities
- 6:38right now. So
- 6:41um LLM labs like OpenAI, Gemini,
- 6:44Enthropic etc. They spend capex to train
- 6:47the LLMs and this is kind of equivalent
- 6:48to building out a grid and then there's
- 6:51opex to serve that intelligence over
- 6:53APIs to all of us and this is done
- 6:56through metered access where we pay per
- 6:58million tokens or something like that
- 7:00and we have a lot of demands that are
- 7:01very utility- like demands out of this
- 7:03API we demand low latency high uptime
- 7:06consistent quality etc. In electricity,
- 7:08you would have a transfer switch. So you
- 7:10can transfer your electricity source
- 7:12from like grid and solar or battery or
- 7:14generator. In LLM, we have maybe open
- 7:16router and easily switch between the
- 7:18different types of LLMs that exist.
- 7:20Because the LLM are software, they don't
- 7:23compete for physical space. So it's okay
- 7:25to have basically like six electricity
- 7:26providers and you can switch between
- 7:28them, right? Because they don't compete
- 7:29in such a direct way. And I think what's
- 7:31also a little fascinating and we saw
- 7:33this in the last few days actually a lot
- 7:36of the LLMs went down and people were
- 7:38kind of like stuck and unable to work.
- 7:41And uh I think it's kind of fascinating
- 7:42to me that when the state-of-the-art
- 7:43LLMs go down, it's actually kind of like
- 7:45an intelligence brownout in the world.
- 7:47It's kind of like when the voltage is
- 7:49unreliable in the grid and uh the planet
- 7:52just gets dumber the more reliance we
- 7:55have on these models, which already is
- 7:56like really dramatic and I think will
- 7:58continue to grow. But LLM's don't only
- 8:00have properties of utilities. I think
- 8:02it's also fair to say that they have
- 8:03some properties of fabs. And the reason
- 8:06for this is that the capex required for
- 8:09building LLM is actually quite large. Uh
- 8:12it's not just like building some uh
- 8:14power station or something like that,
- 8:15right? You're investing a huge amount of
- 8:17money and I think the tech tree and uh
- 8:20for the technology is growing quite
- 8:22rapidly. So we're in a world where we
- 8:24have sort of deep tech trees, research
- 8:26and development secrets that are
- 8:28centralizing inside the LLM labs. Um and
- 8:32but I think the analogy muddies a little
- 8:34bit also because as I mentioned this is
- 8:36software and software is a bit less
- 8:38defensible because it is so malleable.
- 8:40And so um I think it's just an
- 8:43interesting kind of thing to think about
- 8:44potentially. There's many analogy
- 8:46analogies you can make like a 4
- 8:48nanometer process node maybe is
- 8:49something like a cluster with certain
- 8:51max flops. You can think about when
- 8:53you're use when you're using Nvidia GPUs
- 8:54and you're only doing the software and
- 8:56you're not doing the hardware. That's
- 8:57kind of like the fabless model. But if
- 8:59you're actually also building your own
- 9:00hardware and you're training on TPUs if
- 9:02you're Google, that's kind of like the
- 9:03Intel model where you own your fab. So I
- 9:05think there's some analogies here that
- 9:06make sense. But actually I think the
- 9:08analogy that makes the most sense
- 9:09perhaps is that in my mind LLM have very
- 9:12strong kind of analogies to operating
- 9:15systems. Uh in that this is not just
- 9:17electricity or water. It's not something
- 9:19that comes out of the tap as a
- 9:20commodity. uh this is these are now
- 9:22increasingly complex software ecosystems
- 9:25right so uh they're not just like simple
- 9:28commodities like electricity and it's
- 9:30kind of interesting to me that the
- 9:32ecosystem is shaping in a very similar
- 9:33kind of way where you have a few closed
- 9:36source providers like Windows or Mac OS
- 9:38and then you have an open source
- 9:39alternative like Linux and I think for u
- 9:42neural for LLMs as well we have a kind
- 9:45of a few competing closed source
- 9:47providers and then maybe the llama
- 9:49ecosystem is currently like maybe a
- 9:51close approximation to something that
- 9:53may grow into something like Linux.
- 9:55Again, I think it's still very early
- 9:56because these are just simple LLMs, but
- 9:58we're starting to see that these are
- 9:59going to get a lot more complicated.
- 10:01It's not just about the LLM itself. It's
- 10:02about all the tool use and the
- 10:03multiodalities and how all of that
- 10:05works. And so when I sort of had this
- 10:07realization a while back, I tried to
- 10:09sketch it out and it kind of seemed to
- 10:11me like LLMs are kind of like a new
- 10:12operating system, right? So the LLM is a
- 10:15new kind of a computer. It's sitting
- 10:17it's kind of like the CPU equivalent. uh
- 10:19the context windows are kind of like the
- 10:21memory and then the LLM is orchestrating
- 10:24memory and compute uh for problem
- 10:26solving um using all of these uh
- 10:29capabilities here and so definitely if
- 10:32you look at it looks very much like
- 10:34operating system from that perspective.
- 10:36Um, a few more analogies. For example,
- 10:38if you want to download an app, say I go
- 10:41to VS Code and I go to download, you can
- 10:43download VS Code and you can run it on
- 10:46Windows, Linux or or Mac in the same way
- 10:50as you can take an LLM app like cursor
- 10:53and you can run it on GPT or cloud or
- 10:55Gemini series, right? It's just a drop
- 10:57down. So, it's kind of like similar in
- 10:59that way as well.
- 11:00uh more analogies that I think strike me
- 11:02is that we're kind of like in this
- 11:041960sish
- 11:05era where LLM compute is still very
- 11:09expensive for this new kind of a
- 11:10computer and that forces the LLMs to be
- 11:13centralized in the cloud and we're all
- 11:15just uh sort of thing clients that
- 11:18interact with it over the network and
- 11:20none of us have full utilization of
- 11:22these computers and therefore it makes
- 11:24sense to use time sharing where we're
- 11:26all just you know a dimension of the
- 11:28batch when they're running the computer
- 11:30in the cloud. And this is very much what
- 11:32computers used to look like at during
- 11:33this time. The operating systems were in
- 11:35the cloud. Everything was streamed
- 11:36around and there was batching. And so
- 11:39the p the personal computing revolution
- 11:41hasn't happened yet because it's just
- 11:42not economical. It doesn't make sense.
- 11:44But I think some people are trying. And
- 11:46it turns out that Mac minis, for
- 11:48example, are a very good fit for some of
- 11:50the LLMs because it's all if you're
- 11:52doing batch one inference, this is all
- 11:53super memory bound. So this actually
- 11:55works.
- 11:56And uh I think these are some early
- 11:58indications maybe of personal computing.
- 12:00Uh but this hasn't really happened yet.
- 12:02It's not clear what this looks like.
- 12:03Maybe some of you get to invent what
- 12:05what this is or how it works or uh what
- 12:08this should what this should be. Maybe
- 12:10one more analogy that I'll mention is
- 12:12whenever I talk to Chach or some LLM
- 12:14directly in text, I feel like I'm
- 12:16talking to an operating system through
- 12:18the terminal. Like it's just it's it's
- 12:21text. It's direct access to the
- 12:22operating system. And I think a guey
- 12:24hasn't yet really been invented in like
- 12:26a general way like should chatt have a
- 12:29guey like different than just a tech
- 12:31bubbles. Uh certainly some of the apps
- 12:33that we're going to go into in a bit
- 12:35have guey but there's no like guey
- 12:38across all the tasks if that makes
- 12:40sense. Um there are some ways in which
- 12:43LLMs are different from kind of
- 12:45operating systems in some fairly unique
- 12:47way and from early computing. And I
- 12:49wrote about uh this one particular
- 12:52property that strikes me as very
- 12:54different uh this time around. It's that
- 12:57LLMs like flip they flip the direction
- 12:59of technology diffusion uh that is
- 13:02usually uh present in technology. So for
- 13:05example with electricity, cryptography,
- 13:07computing, flight, internet, GPS, lots
- 13:09of new transformative technologies that
- 13:10have not been around. Typically it is
- 13:12the government and corporations that are
- 13:14the first users because it's new and
- 13:16expensive etc. and it only later
- 13:18diffuses to consumer. Uh, but I feel
- 13:20like LLMs are kind of like flipped
- 13:22around. So maybe with early computers,
- 13:24it was all about ballistics and military
- 13:26use, but with LLMs, it's all about how
- 13:29do you boil an egg or something like
- 13:30that. This is certainly like a lot of my
- 13:32use. And so it's really fascinating to
- 13:33me that we have a new magical computer
- 13:35and it's like helping me boil an egg.
- 13:37It's not helping the government do
- 13:38something really crazy like some
- 13:40military ballistics or some special
- 13:42technology. Indeed, corporations are
- 13:43governments are lagging behind the
- 13:45adoption of all of us, of all of these
- 13:47technologies. So, it's just backwards
- 13:48and I think it informs maybe some of the
- 13:50uses of how we want to use this
- 13:52technology or like where are some of the
- 13:53first apps and so on.
- 13:56So, in summary so far, LLM labs LLMs. I
- 14:01think it's accurate language to use, but
- 14:03LLMs are complicated operating systems.
- 14:06They're circa 1960s in computing and
- 14:08we're redoing computing all over again.
- 14:10and they're currently available via time
- 14:11sharing and distributed like a utility.
- 14:13What is new and unprecedented is that
- 14:16they're not in the hands of a few
- 14:17governments and corporations. They're in
- 14:18the hands of all of us because we all
- 14:20have a computer and it's all just
- 14:21software and Chaship was beamed down to
- 14:24our computers like billions of people
- 14:26like instantly and overnight and this is
- 14:28insane. Uh and it's kind of insane to me
- 14:30that this is the case and now it is our
- 14:33time to enter the industry and program
- 14:34these computers. This is crazy. So I
- 14:37think this is quite remarkable. Before
- 14:39we program LLMs, we have to kind of like
- 14:42spend some time to think about what
- 14:43these things are. And I especially like
- 14:45to kind of talk about their psychology.
- 14:48So the way I like to think about LLMs is
- 14:50that they're kind of like people
- 14:51spirits. Um they are stoastic
- 14:54simulations of people. Um and the
- 14:56simulator in this case happens to be an
- 14:58auto reggressive transformer. So
- 14:59transformer is a neural net. Uh it's and
- 15:02it just kind of like is goes on the
- 15:04level of tokens. It goes chunk chunk
- 15:06chunk chunk chunk. And there's an almost
- 15:08equal amount of compute for every single
- 15:10chunk. Um and um this simulator of
- 15:14course is is just is basically there's
- 15:16some weights involved and we fit it to
- 15:19all of text that we have on the internet
- 15:20and so on. And you end up with this kind
- 15:22of a simulator and because it is trained
- 15:24on humans, it's got this emergent
- 15:26psychology that is humanlike. So the
- 15:28first thing you'll notice is of course
- 15:30uh LLM have encyclopedic knowledge and
- 15:32memory. uh and they can remember lots of
- 15:34things, a lot more than any single
- 15:36individual human can because they read
- 15:37so many things. It's it actually kind of
- 15:39reminds me of this movie Rainman, which
- 15:41I actually really recommend people
- 15:43watch. It's an amazing movie. I love
- 15:44this movie. Um and Dustin Hoffman here
- 15:46is an autistic savant who has almost
- 15:49perfect memory. So, he can read a he can
- 15:51read like a phone book and remember all
- 15:53of the names and phone numbers. And I
- 15:55kind of feel like LM are kind of like
- 15:57very similar. They can remember Shaw
- 15:58hashes and lots of different kinds of
- 16:00things very very easily. So they
- 16:02certainly have superpowers in some set
- 16:04in some respects. But they also have a
- 16:06bunch of I would say cognitive deficits.
- 16:08So they hallucinate quite a bit. Um and
- 16:11they kind of make up stuff and don't
- 16:13have a very good uh sort of internal
- 16:15model of self-nowledge, not sufficient
- 16:17at least. And this has gotten better but
- 16:19not perfect. They display jagged
- 16:21intelligence. So they're going to be
- 16:22superhuman in some problems solving
- 16:24domains. And then they're going to make
- 16:26mistakes that basically no human will
- 16:27make. like you know they will insist
- 16:29that 9.11 is greater than 9.9 or that
- 16:32there are two Rs in strawberry these are
- 16:34some famous examples but basically there
- 16:36are rough edges that you can trip on so
- 16:38that's kind of I think also kind of
- 16:40unique um they also kind of suffer from
- 16:43entrograde amnesia um so uh and I think
- 16:46I'm alluding to the fact that if you
- 16:48have a co-orker who joins your
- 16:49organization this co-orker will over
- 16:51time learn your organization and uh they
- 16:54will understand and gain like a huge
- 16:55amount of context on the organization
- 16:57and they go home and they sleep and they
- 16:59consolidate knowledge and they develop
- 17:01expertise over time. LLMs don't natively
- 17:03do this and this is not something that
- 17:04has really been solved in the R&D of
- 17:06LLM. I think um and so context windows
- 17:09are really kind of like working memory
- 17:10and you have to sort of program the
- 17:12working memory quite directly because
- 17:13they don't just kind of like get smarter
- 17:15by uh by default and I think a lot of
- 17:17people get tripped up by the analogies
- 17:19uh in this way. Uh in popular culture I
- 17:22recommend people watch these two movies
- 17:23uh Momento and 51st dates. In both of
- 17:26these movies, the protagonists, their
- 17:27weights are fixed and their context
- 17:29windows gets wiped every single morning
- 17:32and it's really problematic to go to
- 17:34work or have relationships when this
- 17:35happens and this happens to all the
- 17:37time. I guess one more thing I would
- 17:39point to is security kind of related
- 17:42limitations of the use of LLM. So for
- 17:44example, LLMs are quite gullible. Uh
- 17:46they are susceptible to prompt injection
- 17:48risks. They might leak your data etc.
- 17:50And so um and there's many other
- 17:52considerations uh security related. So,
- 17:55so basically long story short, you have
- 17:57to load your you have to load your you
- 18:00have to simultaneously think through
- 18:01this superhuman thing that has a bunch
- 18:03of cognitive deficits and issues. How do
- 18:05we and yet they are extremely like
- 18:07useful and so how do we program them and
- 18:10how do we work around their deficits and
- 18:12enjoy their superhuman powers.
- 18:15So what I want to switch to now is talk
- 18:17about the opportunities of how do we use
- 18:18these models and what are some of the
- 18:20biggest opportunities. This is not a
- 18:22comprehensive list just some of the
- 18:23things that I thought were interesting
- 18:24for this talk. The first thing I'm kind
- 18:26of excited about is what I would call
- 18:29partial autonomy apps. So for example,
- 18:32let's work with the example of coding.
- 18:34You can certainly go to chacht directly
- 18:36and you can start copy pasting code
- 18:38around and copyping bug reports and
- 18:40stuff around and getting code and copy
- 18:42pasting everything around. Why would you
- 18:44why would you do that? Why would you go
- 18:45directly to the operating system? It
- 18:47makes a lot more sense to have an app
- 18:48dedicated for this. And so I think many
- 18:50of you uh use uh cursor. I do as well.
- 18:53And uh cursor is kind of like the thing
- 18:56you want instead. You don't want to just
- 18:57directly go to the chash apt. And I
- 18:59think cursor is a very good example of
- 19:01an early LLM app that has a bunch of
- 19:03properties that I think are um useful
- 19:06across all the LLM apps. So in
- 19:08particular, you will notice that we have
- 19:09a traditional interface that allows a
- 19:12human to go in and do all the work
- 19:13manually just as before. But in addition
- 19:16to that, we now have this LLM
- 19:17integration that allows us to go in
- 19:19bigger chunks. And so some of the
- 19:21properties of LLM apps that I think are
- 19:23shared and useful to point out. Number
- 19:25one, the LLMs basically do a ton of the
- 19:28context management. Um, number two, they
- 19:31orchestrate multiple calls to LLMs,
- 19:33right? So in the case of cursor, there's
- 19:34under the hood embedding models for all
- 19:36your files, the actual chat models,
- 19:39models that apply diffs to the code, and
- 19:41this is all orchestrated for you. A
- 19:43really big one that uh I think also
- 19:46maybe not fully appreciated always is
- 19:48application specific uh GUI and the
- 19:50importance of it. Um because you don't
- 19:53just want to talk to the operating
- 19:54system directly in text. Text is very
- 19:56hard to read, interpret, understand and
- 19:59also like you don't want to take some of
- 20:00these actions natively in text. So it's
- 20:03much better to just see a diff as like
- 20:05red and green change and you can see
- 20:06what's being added is subtracted. It's
- 20:08much easier to just do command Y to
- 20:10accept or command N to reject. I
- 20:11shouldn't have to type it in text,
- 20:13right? So, a guey allows a human to
- 20:15audit the work of these fallible systems
- 20:17and to go faster. I'm going to come back
- 20:20to this point a little bit uh later as
- 20:21well. And the last kind of feature I
- 20:23want to point out is that there's what I
- 20:25call the autonomy slider. So, for
- 20:27example, in cursor, you can just do tap
- 20:29completion. You're mostly in charge. You
- 20:31can select a chunk of code and command K
- 20:33to change just that chunk of code. You
- 20:36can do command L to change the entire
- 20:37file. Or you can do command I which just
- 20:40you know let it rip do whatever you want
- 20:42in the entire repo and that's the sort
- 20:44of full autonomy agent agentic version
- 20:46and so you are in charge of the autonomy
- 20:48slider and depending on the complexity
- 20:50of the task at hand you can uh tune the
- 20:53amount of autonomy that you're willing
- 20:54to give up uh for that task maybe to
- 20:57show one more example of a fairly
- 20:58successful LLM app uh perplexity um it
- 21:03also has very similar features to what
- 21:04I've just pointed out to in cursor uh it
- 21:07packages up a lot of the information. It
- 21:08orchestrates multiple LLMs. It's got a
- 21:10GUI that allows you to audit some of its
- 21:13work. So, for example, it will site
- 21:15sources and you can imagine inspecting
- 21:17them. And it's got an autonomy slider.
- 21:18You can either just do a quick search or
- 21:20you can do research or you can do deep
- 21:22research and come back 10 minutes later.
- 21:24So, this is all just varying levels of
- 21:25autonomy that you give up to the tool.
- 21:27So, I guess my question is I feel like a
- 21:30lot of software will become partially
- 21:32autonomous. I'm trying to think through
- 21:33like what does that look like? And for
- 21:35many of you who maintain products and
- 21:36services, how are you going to make your
- 21:38products and services partially
- 21:40autonomous? Can an LLM see everything
- 21:42that a human can see? Can an LLM act in
- 21:45all the ways that a human could act? And
- 21:47can humans supervise and stay in the
- 21:49loop of this activity? Because again,
- 21:50these are fallible systems that aren't
- 21:52yet perfect. And what does a diff look
- 21:54like in Photoshop or something like
- 21:56that? You know, and also a lot of the
- 21:58traditional software right now, it has
- 22:00all these switches and all this kind of
- 22:01stuff that's all designed for human. All
- 22:03of this has to change and become
- 22:04accessible to LLMs.
- 22:07So, one thing I want to stress with a
- 22:09lot of these LLM apps that I'm not sure
- 22:11gets as much attention as it should is
- 22:14um we we're now kind of like cooperating
- 22:16with AIS and usually they are doing the
- 22:18generation and we as humans are doing
- 22:20the verification. It is in our interest
- 22:22to make this loop go as fast as
- 22:24possible. So, we're getting a lot of
- 22:25work done. There are two major ways that
- 22:28I think uh this can be done. Number one,
- 22:30you can speed up verification a lot. Um,
- 22:32and I think guies, for example, are
- 22:34extremely important to this because a
- 22:36guey utilizes your computer vision GPU
- 22:39in all of our head. Reading text is
- 22:41effortful and it's not fun, but looking
- 22:43at stuff is fun and it's it's just a
- 22:45kind of like a highway to your brain.
- 22:47So, I think guies are very useful for
- 22:49auditing systems and visual
- 22:51representations in general. And number
- 22:53two, I would say is we have to keep the
- 22:56AI on the leash. We I think a lot of
- 22:58people are getting way over excited with
- 23:00AI agents and uh it's not useful to me
- 23:03to get a diff of 10,000 lines of code to
- 23:05my repo. Like I have to I'm still the
- 23:07bottleneck, right? Even though that
- 23:0910,00 lines come out instantly, I have
- 23:11to make sure that this thing is not
- 23:12introducing bugs. It's just like and
- 23:15that it's doing the correct thing,
- 23:16right? And that there's no security
- 23:17issues and so on. So um I think that um
- 23:22yeah basically you we have to sort of
- 23:25like it's in our interest to make the
- 23:28the flow of these two go very very fast
- 23:30and we have to somehow keep the AI on
- 23:32the leash because it gets way too
- 23:33overreactive. It's uh it's kind of like
- 23:35this. This is how I feel when I do AI
- 23:37assisted coding. If I'm just bite coding
- 23:39everything is nice and great but if I'm
- 23:40actually trying to get work done it's
- 23:42not so great to have an overreactive uh
- 23:44agent doing all this kind of stuff. So
- 23:47this slide is not very good. I'm sorry,
- 23:48but I guess I'm trying to develop like
- 23:51many of you some ways of utilizing these
- 23:53agents in my coding workflow and to do
- 23:55AI assisted coding. And in my own work,
- 23:58I'm always scared to get way too big
- 23:59diffs. I always go in small incremental
- 24:02chunks. I want to make sure that
- 24:04everything is good. I want to spin this
- 24:06loop very very fast and um I sort of
- 24:09work on small chunks of single concrete
- 24:10thing. Uh and so I think many of you
- 24:13probably are developing similar ways of
- 24:14working with the with LLMs.
- 24:17Um, I also saw a number of blog posts
- 24:19that try to develop these best practices
- 24:22for working with LLMs. And here's one
- 24:24that I read recently and I thought was
- 24:25quite good. And it kind of discussed
- 24:26some techniques and some of them have to
- 24:28do with how you keep the AI on the
- 24:29leash. And so, as an example, if you are
- 24:32prompting, if your prompt is vague, then
- 24:34uh the AI might not do exactly what you
- 24:36wanted and in that case, verification
- 24:38will fail. You're going to ask for
- 24:40something else. If a verification fails,
- 24:42then you're going to start spinning. So
- 24:43it makes a lot more sense to spend a bit
- 24:45more time to be more concrete in your
- 24:46prompts which increases the probability
- 24:48of successful verification and you can
- 24:50move forward. And so I think a lot of us
- 24:52are going to end up finding um kind of
- 24:54techniques like this. I think in my own
- 24:56work as well I'm currently interested in
- 24:57uh what education looks like in um
- 25:00together with kind of like now that we
- 25:01have AI uh and LLMs what does education
- 25:04look like? And I think a a large amount
- 25:07of thought for me goes into how we keep
- 25:09AI on the leash. I don't think it just
- 25:11works to go to chat and be like, "Hey,
- 25:13teach me physics." I don't think this
- 25:14works because the AI is like gets lost
- 25:16in the woods. And so for me, this is
- 25:18actually two separate apps. For example,
- 25:20there's an app for a teacher that
- 25:22creates courses and then there's an app
- 25:24that takes courses and serves them to
- 25:26students. And in both cases, we now have
- 25:29this intermediate artifact of a course
- 25:31that is auditable and we can make sure
- 25:32it's good. We can make sure it's
- 25:33consistent. and the AI is kept on the
- 25:35leash with respect to a certain
- 25:37syllabus, a certain like um progression
- 25:40of projects and so on. And so this is
- 25:42one way of keeping the AI on leash and I
- 25:44think has a much higher likelihood of
- 25:45working and the AI is not getting lost
- 25:47in the woods.
- 25:49One more kind of analogy I wanted to
- 25:51sort of allude to is I'm not I'm no
- 25:54stranger to partial autonomy and I kind
- 25:56of worked on this I think for five years
- 25:57at Tesla and this is also a partial
- 26:00autonomy product and shares a lot of the
- 26:01features like for example right there in
- 26:03the instrument panel is the GUI of the
- 26:05autopilot so it's showing me what the
- 26:07what the neural network sees and so on
- 26:09and we have the autonomy slider where
- 26:10over the course of my tenure there we
- 26:13did more and more autonomous tasks for
- 26:15the user and maybe the story that I
- 26:18wanted to tell very briefly is uh
- 26:21actually the first time I drove a
- 26:22self-driving vehicle was in 2013 and I
- 26:25had a friend who worked at Whimo and uh
- 26:27he offered to give me a drive around
- 26:29Palo Alto. I took this picture using
- 26:31Google Glass at the time and many of you
- 26:33are so young that you might not even
- 26:35know what that is. Uh but uh yeah, this
- 26:37was like all the rage at the time. And
- 26:39we got into this car and we went for
- 26:40about a 30-minute drive around Palo Alto
- 26:42highways uh streets and so on. And this
- 26:45drive was perfect. There was zero
- 26:46interventions and this was 2013 which is
- 26:49now 12 years ago. And it kind of struck
- 26:52me because at the time when I had this
- 26:54perfect drive, this perfect demo, I felt
- 26:56like, wow, self-driving is imminent
- 26:59because this just worked. This is
- 27:00incredible. Um, but here we are 12 years
- 27:03later and we are still working on
- 27:04autonomy. Um, we are still working on
- 27:07driving agents and even now we haven't
- 27:09actually like really solved the problem.
- 27:10like you may see Whimos going around and
- 27:12they look driverless but you know
- 27:14there's still a lot of teleoperation and
- 27:16a lot of human in the loop of a lot of
- 27:18this driving so we still haven't even
- 27:20like declared success but I think it's
- 27:22definitely like going to succeed at this
- 27:24point but it just took a long time and
- 27:26so I think like like this is software is
- 27:29really tricky I think in the same way
- 27:31that driving is tricky and so when I see
- 27:34things like oh 2025 is the year of
- 27:36agents I get very concerned and I kind
- 27:38of feel like you know this is the decade
- 27:41of agents and this is going to be quite
- 27:44some time. We need humans in the loop.
- 27:45We need to do this carefully. This is
- 27:47software. Let's be serious here. One
- 27:51more kind of analogy that I always think
- 27:52through is the Iron Man suit. Uh I think
- 27:56this is I always love Iron Man. I think
- 27:58it's like so um correct in a bunch of
- 28:01ways with respect to technology and how
- 28:02it will play out. And what I love about
- 28:04the Iron Man suit is that it's both an
- 28:05augmentation and Tony Stark can drive it
- 28:08and it's also an agent. And in some of
- 28:10the movies, the Iron Man suit is quite
- 28:11autonomous and can fly around and find
- 28:13Tony and all this kind of stuff. And so
- 28:15this is the autonomy slider is we can be
- 28:17we can build augmentations or we can
- 28:19build agents and we kind of want to do a
- 28:21bit of both. But at this stage I would
- 28:23say working with fallible LLMs and so
- 28:25on. I would say you know it's less Iron
- 28:29Man robots and more Iron Man suits that
- 28:31you want to build. It's less like
- 28:33building flashy demos of autonomous
- 28:35agents and more building partial
- 28:36autonomy products. And these products
- 28:39have custom gueies and UIUX. And we're
- 28:41trying to um and this is done so that
- 28:43the generation verification loop of the
- 28:45human is very very fast. But we are not
- 28:48losing the sight of the fact that it is
- 28:49in principle possible to automate this
- 28:51work. And there should be an autonomy
- 28:52slider in your product. And you should
- 28:54be thinking about how you can slide that
- 28:55autonomy slider and make your product uh
- 28:58sort of um more autonomous over time.
- 29:01But this is kind of how I think there's
- 29:02lots of opportunities in these kinds of
- 29:04products. I want to now switch gears a
- 29:06little bit and talk about one other
- 29:08dimension that I think is very unique.
- 29:09Not only is there a new type of
- 29:11programming language that allows for
- 29:12autonomy in software but also as I
- 29:15mentioned it's programmed in English
- 29:16which is this natural interface and
- 29:19suddenly everyone is a programmer
- 29:20because everyone speaks natural language
- 29:22like English. So this is extremely
- 29:24bullish and very interesting to me and
- 29:26also completely unprecedented. I would
- 29:28say it it used to be the case that you
- 29:29need to spend five to 10 years studying
- 29:31something to be able to do something in
- 29:32software. this is not the case anymore.
- 29:35So, I don't know if by any chance anyone
- 29:37has heard of vibe coding.
- 29:40Uh, this this is the tweet that kind of
- 29:42like introduced this, but I'm told that
- 29:44this is now like a major meme. Um, fun
- 29:46story about this is that I've been on
- 29:49Twitter for like 15 years or something
- 29:51like that at this point and I still have
- 29:53no clue which tweet will become viral
- 29:56and which tweet like fizzles and no one
- 29:58cares. And I thought that this tweet was
- 30:00going to be the latter. I don't know. It
- 30:01was just like a shower of thoughts. But
- 30:03this became like a total meme and I
- 30:05really just can't tell. But I guess like
- 30:06it struck a chord and it gave a name to
- 30:08something that everyone was feeling but
- 30:10couldn't quite say in words. So now
- 30:13there's a Wikipedia page and everything.
- 30:17This is like
- 30:18[Applause]
- 30:25yeah this is like a major contribution
- 30:27now or something like that. So,
- 30:30um, so Tom Wolf from HuggingFace shared
- 30:32this beautiful video that I really love.
- 30:34Um,
- 30:37these are kids vibe coding.
- 30:42And I find that this is such a wholesome
- 30:44video. Like, I love this video. Like,
- 30:46how can you look at this video and feel
- 30:48bad about the future? The future is
- 30:49great.
- 30:52I think this will end up being like a
- 30:53gateway drug to software development.
- 30:56Um, I'm not a doomer about the future of
- 30:59the generation and I think yeah, I love
- 31:02this video. So, I tried by coding a
- 31:04little bit uh as well because it's so
- 31:07fun. Uh, so bike coding is so great when
- 31:09you want to build something super duper
- 31:10custom that doesn't appear to exist and
- 31:12you just want to wing it because it's a
- 31:13Saturday or something like that. So, I
- 31:15built this uh iOS app and I don't I
- 31:18can't actually program in Swift, but I
- 31:20was really shocked that I was able to
- 31:21build like a super basic app and I'm not
- 31:23going to explain it. It's really uh
- 31:24dumb, but uh I kind of like this was
- 31:27just like a day of work and this was
- 31:28running on my phone like later that day
- 31:30and I was like, "Wow, this is amazing."
- 31:32I didn't have to like read through Swift
- 31:33for like five days or something like
- 31:35that to like get started. I also
- 31:38vipcoded this app called Menu Genen. And
- 31:40this is live. You can try it in
- 31:41menu.app. And I basically had this
- 31:44problem where I show up at a restaurant,
- 31:45I read through the menu, and I have no
- 31:46idea what any of the things are. And I
- 31:48need pictures. So this doesn't exist. So
- 31:51I was like, "Hey, I'm going to bite code
- 31:52it." So, um, this is what it looks like.
- 31:55You go to menu.app,
- 31:58um, and, uh, you take a picture of a of
- 32:01a menu and then menu generates the
- 32:03images and everyone gets $5 in credits
- 32:06for free when you sign up. And
- 32:08therefore, this is a major cost center
- 32:10in my life. So, this is a negative
- 32:13negative uh, revenue app for me right
- 32:16now.
- 32:17I've lost a huge amount of money on
- 32:19menu.
- 32:21Okay. But the fascinating thing about
- 32:23menu genen for me is that the code of
- 32:28the v the vite coding part the code was
- 32:30actually the easy part of v of v coding
- 32:32menu and most of it actually was when I
- 32:35tried to make it real so that you can
- 32:36actually have authentication and
- 32:37payments and the domain name and averal
- 32:39deployment. This was really hard and all
- 32:41of this was not code. All of this devops
- 32:44stuff was in me in the browser clicking
- 32:47stuff and this was extreme slo and took
- 32:49another week. So it was really
- 32:51fascinating that I had the menu genen um
- 32:54basically demo working on my laptop in a
- 32:57few hours and then it took me a week
- 32:59because I was trying to make it real and
- 33:01the reason for this is this was just
- 33:02really annoying. Um, so for example, if
- 33:05you try to add Google login to your web
- 33:07page, I know this is very small, but
- 33:09just a huge amount of instructions of
- 33:11this clerk library telling me how to
- 33:13integrate this. And this is crazy. Like
- 33:15it's telling me go to this URL, click on
- 33:17this dropdown, choose this, go to this,
- 33:19and click on that. And it's like telling
- 33:21me what to do. Like a computer is
- 33:22telling me the actions I should be
- 33:24taking. Like you do it. Why am I doing
- 33:26this?
- 33:28What the hell?
- 33:31I had to follow all these instructions.
- 33:33This was crazy. So I think the last part
- 33:36of my talk therefore focuses on can we
- 33:39just build for agents? I don't want to
- 33:41do this work. Can agents do this? Thank
- 33:44you.
- 33:46Okay. So roughly speaking, I think
- 33:48there's a new category of consumer and
- 33:50manipulator of digital information. It
- 33:53used to be just humans through GUIs or
- 33:55computers through APIs. And now we have
- 33:57a completely new thing and agents are
- 34:00they're computers but they are humanlike
- 34:02kind of right they're people spirits
- 34:04there's people spirits on the internet
- 34:05and they need to interact with our
- 34:06software infrastructure like can we
- 34:08build for them it's a new thing so as an
- 34:10example you can have robots.txt on your
- 34:12domain and you can instruct uh or like
- 34:15advise I suppose um uh web crawlers on
- 34:18how to behave on your website in the
- 34:19same way you can have maybe lm.txt txt
- 34:21file which is just a simple markdown
- 34:23that's telling LLMs what this domain is
- 34:25about and this is very readable to a to
- 34:28an LLM. If it had to instead get the
- 34:30HTML of your web page and try to parse
- 34:32it, this is very errorprone and
- 34:33difficult and will screw it up and it's
- 34:35not going to work. So we can just
- 34:36directly speak to the LLM. It's worth
- 34:38it. Um a huge amount of documentation is
- 34:41currently written for people. So you
- 34:42will see things like lists and bold and
- 34:45pictures and this is not directly
- 34:47accessible by an LLM. So I see some of
- 34:51the services now are transitioning a lot
- 34:52of the their docs to be specifically for
- 34:54LLMs. So Versell and Stripe as an
- 34:57example are early movers here but there
- 34:59are a few more that I've seen already
- 35:01and they offer their documentation in
- 35:04markdown. Markdown is super easy for LMS
- 35:06to understand. This is great. Um maybe
- 35:10one simple example from from uh my
- 35:12experience as well. Maybe some of you
- 35:14know three blue one brown. He makes
- 35:15beautiful animation videos on YouTube.
- 35:19[Applause]
- 35:23Yeah, I love this library. So that he
- 35:25wrote uh Manon and I wanted to make my
- 35:27own and uh there's extensive
- 35:30documentations on how to use manon and
- 35:32so I didn't want to actually read
- 35:34through it. So I copy pasted the whole
- 35:35thing to an LLM and I described what I
- 35:37wanted and it just worked out of the box
- 35:39like LLM just bcoded me an animation
- 35:41exactly what I wanted and I was like wow
- 35:43this is amazing. So if we can make docs
- 35:45legible to LLMs, it's going to unlock a
- 35:48huge amount of um kind of use and um I
- 35:51think this is wonderful and should
- 35:52should happen more. The other thing I
- 35:55wanted to point out is that you do
- 35:56unfortunately have to it's not just
- 35:57about taking your docs and making them
- 35:58appear in markdown. That's the easy
- 36:00part. We actually have to change the
- 36:01docs because anytime your docs say click
- 36:04this is bad. An LLM will not be able to
- 36:06natively take this action right now. So,
- 36:09Verscell, for example, is replacing
- 36:11every occurrence of click with an
- 36:13equivalent curl command that your LM
- 36:15agent could take on your behalf. Um, and
- 36:18so I think this is very interesting. And
- 36:19then, of course, there's a model context
- 36:21protocol from Enthropic. And this is
- 36:23also another way, it's a protocol of
- 36:24speaking directly to agents as this new
- 36:26consumer and manipulator of digital
- 36:28information. So, I'm very bullish on
- 36:29these ideas. The other thing I really
- 36:31like is a number of little tools here
- 36:33and there that are helping ingest data
- 36:36that in like very LLM friendly formats.
- 36:38So for example, when I go to a GitHub
- 36:40repo like my nanoGPT repo, I can't feed
- 36:42this to an LLM and ask questions about
- 36:44it uh because it's you know this is a
- 36:46human interface on GitHub. So when you
- 36:48just change the URL from GitHub to get
- 36:50ingest then uh this will actually
- 36:52concatenate all the files into a single
- 36:54giant text and it will create a
- 36:55directory structure etc. And this is
- 36:57ready to be copy pasted into your
- 36:59favorite LLM and you can do stuff. Maybe
- 37:01even more dramatic example of this is
- 37:03deep wiki where it's not just the raw
- 37:05content of these files. uh this is from
- 37:08Devon but also like they have Devon
- 37:10basically do analysis of the GitHub repo
- 37:12and Devon basically builds up a whole
- 37:14docs uh pages just for your repo and you
- 37:18can imagine that this is even more
- 37:19helpful to copy paste into your LLM. So
- 37:22I love all the little tools that
- 37:23basically where you just change the URL
- 37:24and it makes something accessible to an
- 37:26LLM. So this is all well and great and u
- 37:29I think there should be a lot more of
- 37:30it. One more note I wanted to make is
- 37:32that it is absolutely possible that in
- 37:35the future LLMs will be able to this is
- 37:38not even future this is today they'll be
- 37:39able to go around and they'll be able to
- 37:40click stuff and so on but I still think
- 37:42it's very worth u basically meeting LLM
- 37:46halfway LLM's halfway and making it
- 37:48easier for them to access all this
- 37:49information uh because this is still
- 37:51fairly expensive I would say to use and
- 37:54uh a lot more difficult and so I do
- 37:56think that lots of software there will
- 37:58be a long tail where it won't like adapt
- 38:00apps because these are not like live
- 38:02player sort of repositories or digital
- 38:04infrastructure and we will need these
- 38:06tools. Uh but I think for everyone else
- 38:08I think it's very worth kind of like
- 38:09meeting in some middle point. So I'm
- 38:11bullish on both if that makes sense.
- 38:14So in summary, what an amazing time to
- 38:17get into the industry. We need to
- 38:18rewrite a ton of code. A ton of code
- 38:20will be written by professionals and by
- 38:23coders. These LLMs are kind of like
- 38:25utilities, kind of like fabs, but
- 38:27they're kind of especially like
- 38:28operating systems. But it's so early.
- 38:30It's like 1960s of operating systems and
- 38:34uh and I think a lot of the analogies
- 38:36cross over. Um and these LMS are kind of
- 38:38like these fallible uh you know people
- 38:41spirits that we have to learn to work
- 38:43with. And in order to do that properly,
- 38:45we need to adjust our infrastructure
- 38:47towards it. So when you're building
- 38:48these LLM apps, I describe some of the
- 38:50ways of working effectively with these
- 38:52LLMs and some of the tools that make
- 38:54that uh kind of possible and how you can
- 38:57spin this loop very very quickly and
- 38:59basically create partial tunneling
- 39:00products and then um yeah, a lot of code
- 39:03has to also be written for the agents
- 39:04more directly. But in any case, going
- 39:07back to the Iron Man suit analogy, I
- 39:09think what we'll see over the next
- 39:10decade roughly is we're going to take
- 39:12the slider from left to right. And I'm
- 39:15very interesting. It's going to be very
- 39:17interesting to see what that looks like.
- 39:19And I can't wait to build it with all of
- 39:21you. Thank you.
About this transcript
This page contains the full transcript of Andrej Karpathy: Software Is Changing (Again) by Y Combinator, generated from the public captions YouTube serves with the video. The transcript has 8,255 words across 1,139 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.