Claude Code + Karpathy Autoresearch = The New Meta — Transcript
Full transcript
- 0:00An open-source project just dropped
- 0:01that, when you combine it with Claude
- 0:03code, literally becomes self-improving
- 0:05AI. This is not engagement farming or
- 0:07hype bait. This is a real repo that was
- 0:09just released by Andre Karpathy, who's
- 0:11widely renowned as one of the foremost
- 0:13voices in AI and machine learning
- 0:14research. And basically, what he did
- 0:16was, while training his model, he
- 0:18thought, "Why don't I just have my
- 0:19models train my models instead?" He
- 0:21built an elegant pipeline, which he's
- 0:23calling auto research, and essentially
- 0:25completely and fully automates the
- 0:27process of experimentation. He says it
- 0:29right here. The idea is to give an AI
- 0:31agent a small but real LLM training
- 0:34setup and just let it experiment
- 0:35autonomously overnight. It'll modify the
- 0:37code, train for 5 minutes, check if the
- 0:40results improved, keep or discard, and
- 0:42then just repeat. You wake up in the
- 0:43morning to a log of experiments and it
- 0:45hopefully a better model. Now, I'm not
- 0:48in machine learning training. I don't
- 0:50help make models more intelligent. What
- 0:52I do is I take models that other people
- 0:54have made, and then I use them for the
- 0:56purposes of making money. And so,
- 0:58immediately when this dropped, I started
- 1:00thinking about ways that I could apply
- 1:01this principle of auto research into my
- 1:03own life to improve, obviously, my own
- 1:05economic outcomes. And there are so
- 1:07many, it's not even funny. So, what I'm
- 1:08going to do is I'm going to run through
- 1:09some real practical examples in a
- 1:11moment, things that I'm actually doing
- 1:12in my own business that you can
- 1:13implement inside of Claude code. Then,
- 1:15I'm going to show you how to actually do
- 1:16it, so go through the step-by-step of
- 1:18setting up the repo and building some
- 1:20experimentation done totally
- 1:22autonomously for you. And at the end,
- 1:23you will have a fully automated,
- 1:25self-improving pipeline, just like Andre
- 1:27Karpathy here has done for his own
- 1:29machine learning training. So, here's
- 1:30one of many examples. I do a lot of cold
- 1:32email in my own business and then for
- 1:34clients. And cold email, in case you
- 1:35didn't know, is where you package up a
- 1:37really nice, sexy-sounding offer, and
- 1:39then you send it to people you've never
- 1:41met with the hopes that they take you up
- 1:42on it and then maybe convert. So, jump
- 1:44on a call with you, fill out a form,
- 1:45whatever. Now, the key metric in cold
- 1:47email is usually reply rate,
- 1:49specifically positive reply rate, but
- 1:51reply rate's easier for our purposes, so
- 1:52that's what I'm going to go with. And as
- 1:54you could see, most cold email software
- 1:55tracks this for you out of the box. So,
- 1:572.4% of people replied to this campaign.
- 2:002.5% of people replied to this campaign,
- 2:03and so on. Well, turns out that's all
- 2:05you need in order to build an automatic
- 2:06experimentation pipeline. You need some
- 2:08metric that you want to improve, and
- 2:10then you need some way factor, some
- 2:12thing you can modify to improve it. And
- 2:14so, what I have is my metric is reply
- 2:16rate, and then what I have is the thing
- 2:18that I can adjust is my cold email copy.
- 2:20So, what this looks like in practice for
- 2:21me is a folder called email optimizer.
- 2:25There are a bunch of additional code
- 2:26files here, configs, places where I'm
- 2:29storing the results of my data, and so
- 2:31on and so forth, and aren't super
- 2:32important. What is important is this
- 2:34file right over here called
- 2:35orchestrator.py.
- 2:37And this contains all of the prompts
- 2:39that I'm feeding into my orchestrator
- 2:41agent, who essentially is responsible
- 2:43for spinning up new cold email campaigns
- 2:45and then testing them against each other
- 2:46until I get better and better results.
- 2:48And it was as simple for me as literally
- 2:49copying the repo and making some slight
- 2:51adjustments. What I do is I tell it that
- 2:53that's inspired by Carbon the thesis
- 2:55auto research pattern. The core idea is
- 2:57an AI agent that runs experiments
- 2:58autonomously in a tight loop using an
- 3:00objective metric as a feedback signal. I
- 3:02have the architecture over here. I run
- 3:05the loop every 4 hours, and at the end
- 3:07of every 4 hours, I actually have better
- 3:10copy uh that's self-evolving over time
- 3:12based on the results from my previous
- 3:14test. And you can see some of the
- 3:15examples right over here. Anything with
- 3:17C is what we call a challenger. Anything
- 3:19with B is what's called baseline. And
- 3:21so, the model starts with a baseline
- 3:23type of copy, which we could see right
- 3:25over here. And then it makes slight
- 3:28modifications based off of what it knows
- 3:30to perform really well in cold email
- 3:32copy before testing it out. It runs the
- 3:34two side by side, and then automatically
- 3:36harvests based off of the results, aka
- 3:39the number of replies, and so on and so
- 3:40forth for both campaigns. Now, that part
- 3:42isn't the important bit. I mean, we've
- 3:43been optimizing cold emails for for many
- 3:45years at this point. The important bit
- 3:47is it then creates new copy based off of
- 3:50the learnings from previous experiments.
- 3:53As the models get better and better and
- 3:55better, they log all of their learnings
- 3:57to a resource.md
- 3:59that significantly improves future
- 4:01models abilities to make changes. And so
- 4:03here's a big list of things that this is
- 4:06essentially figured out, move reply rate
- 4:09up. And so in that way we get to push
- 4:11towards that direction over time. Now
- 4:13this has only been running for a few
- 4:14days now. Imagine this running for a
- 4:16year. Instead of optimizing on a basis
- 4:18of once every 4 hours, imagine if this
- 4:20optimized on a basis of once every 5
- 4:22minutes. Well, that's what my next leg
- 4:23of testing is going to do. We're going
- 4:25to be significantly improving the
- 4:26volume, pumping all this stuff out at
- 4:2810x the level, and then optimizing and
- 4:31iterating our results again fully
- 4:32autonomously. So that's just one
- 4:33example. I'm going to run you guys
- 4:35through a bunch of other use cases that
- 4:36you can apply auto research to, whether
- 4:38you're in machine learning engineering
- 4:40or whether you're just trying to improve
- 4:42the profitability of let's say
- 4:44paper click ad campaign. But first,
- 4:45let's make sure we all know how to
- 4:47actually use this thing. So, the way
- 4:49that auto research works, to make a long
- 4:51story short, is we start with an
- 4:53experiment. And just like in science,
- 4:55everything begins with some sort of
- 4:57hypothesis, okay? So my hypothesis might
- 5:00be, hey, if I make a slight adjustment
- 5:03to the copy of this campaign so that
- 5:06it's a little bit punchier, I think it's
- 5:08going to go well. You insert that using
- 5:10this little test.md. It's your goal,
- 5:12metric, and some high-level
- 5:13instructions. From there, the auto
- 5:16research agent will go through, employ
- 5:19the experiment usually using API calls.
- 5:21In my case, in the example we just saw,
- 5:24um to instantly, in Karpathy's specific
- 5:26example we saw, he's doing it through
- 5:29adjusting what are called hyper
- 5:30parameters. And then after that, we
- 5:32measure the results. Now, in order for
- 5:34us to make sure that this works, you
- 5:36know, the hypothesis isn't enough. We
- 5:37need some sort of metric that we're
- 5:38tracking. Now in my case, the metric was
- 5:41obviously pretty simple. It was reply
- 5:42rate. In Karpathy's case, it was pretty
- 5:44simple. It was something called
- 5:45validation loss. As long as you have
- 5:47that, you can then just pick the winner
- 5:50and then make a slight change before
- 5:51looping back. And depending on how tight
- 5:54this feedback loop is, you could
- 5:55theoretically do this in a minute or
- 5:57two. I mean, if he had more
- 5:59infrastructure when he's training his
- 6:00models, he could probably do in 5
- 6:01minutes what he does in one. And then in
- 6:03that way, you know, progress really,
- 6:05really quickly over to some, you know,
- 6:07desired goal. In my case, if I had more
- 6:09cold email infrastructure, I could do
- 6:10the same thing. So, at this point, scale
- 6:12is more or less all you need. This
- 6:13allows you to run hundreds of tests with
- 6:15literally zero human involvement. I
- 6:17mean, I'm not even in the loop anymore.
- 6:19And to be clear, like if I was in the
- 6:21loop, would I be making better decisions
- 6:23than the AI model? Like probably. I'd be
- 6:25a much more efficient optimizer. But
- 6:27that doesn't really matter because the
- 6:28reality is I take a lot more time to
- 6:31optimize than a model does. I also eat,
- 6:33sleep, have to go to the washroom, and
- 6:36do a variety of other things with my
- 6:37day. AI agents don't. You could very
- 6:39quickly and easily set this up on,
- 6:41again, an hourly loop and have this run
- 6:4324 times a day, whereas realistically,
- 6:46if you were to try and do it all
- 6:47yourself, you could only do it a couple
- 6:48times. And so in that way, whatever
- 6:50metric that you're tracking goes up over
- 6:52time, right? In my case, reply rates
- 6:54significantly go up. You know, test one,
- 6:56I might be at a 1.5%, test 12, I might
- 6:59be at a 2.7%. Before you know it, I
- 7:01reach literally like the optimal quality
- 7:04possible for my set of cold emails and
- 7:07then their audiences. And you can apply
- 7:08this, as mentioned, to a bunch of other
- 7:09strategies. So, what are those
- 7:11strategies? The requirement that you
- 7:13need is anything that has an objective
- 7:17metric you can track
- 7:18and an API or application programming
- 7:21interface that you can send a request to
- 7:24to get. Okay? So, some brief examples of
- 7:27this. Cold email copy. Obviously,
- 7:30fantastic. Why? Well, because in our
- 7:32case, we have the Instantly API. The
- 7:35Instantly API allows us to query
- 7:37metrics, and so I can give the agent the
- 7:39ability to call a quick tool, call up
- 7:41the Instantly API, see how the
- 7:43performance was relative to you know the
- 7:45the challenger in the base campaign.
- 7:47At the same time, you know, I have a
- 7:49very clear metric, which in my case is
- 7:51reply rate.
- 7:52Okay, how about landing pages? Let's say
- 7:55you're doing some form of CRO, which is
- 7:57conversion rate optimization, and you
- 7:59want to test to see how you can make
- 8:01your landing pages as efficient as
- 8:02humanly possible. Well, you can now
- 8:04completely automate it with auto
- 8:06research. What you do is you pick the
- 8:08metric that you want, which in our case
- 8:10would literally just be conversion rate,
- 8:13okay?
- 8:14And then if your website is hosted
- 8:15locally or it's hosted using some API or
- 8:18something like that, let's say a website
- 8:19builder like Wix or or or or WordPress
- 8:22or Webflow, what you can do is you can
- 8:24give it access to the API, and then you
- 8:27can say, "Hey, change this according to
- 8:29this resource of best practices that
- 8:31other agents have done. Make your
- 8:32change, test that for, I don't know, a
- 8:34day, depending on how much volume you
- 8:36have, and at the end
- 8:38consolidate the winner and then get rid
- 8:39of the loser." You can do the exact same
- 8:41thing for ad creatives, okay? What's the
- 8:43main thing that you want for ad
- 8:45creatives? Obviously, it's going to be
- 8:45some form of conversion rate as well,
- 8:47whatever specific type of conversion
- 8:49rate is, that's up to you, okay? But all
- 8:51you need to do is query some sort of
- 8:53API. Now, you know, a lot of these ad
- 8:56platforms like Facebook and Google
- 8:58basically already do this for you. Mind
- 9:00it, granted I don't think they do it
- 9:02anywhere near as effectively as you can
- 9:03with modern models like Opus 4.6 or GPT
- 9:065.4. What you can do is you can give it
- 9:08the API to call a specific ad resource,
- 9:11and then you could also just give it the
- 9:13metric to optimize for, which is CVR,
- 9:15and then it'll crush.
- 9:17How about some form of customer
- 9:18satisfaction for chatbot scripts? Maybe
- 9:21use some sort of customer satisfaction
- 9:23score, and then, you know, now you just
- 9:25adjust the main template, okay, that all
- 9:28customer service agents, whether human
- 9:30or AI, are are going off of. That's
- 9:32super simple and easy to do. How about
- 9:34product descriptions for some sort of
- 9:36e-comm? If you have like, I don't know,
- 9:38Amazon FBA or something like that, you
- 9:41know, maybe they don't necessarily have
- 9:42APIs, but maybe now you set up what's
- 9:45called Chrome DevTools MCP,
- 9:48give it a very tightly scoped list of
- 9:50steps that has to do to update the
- 9:51actual body of the landing page, and
- 9:54then based off of metrics like, I don't
- 9:56know, how many freaking dollars you've
- 9:58sold in the last little while, you can
- 9:59very quickly optimize and make your
- 10:01product landing page better and better
- 10:02and better. You know, in my case, I make
- 10:04a lot of YouTube content thesis. I could
- 10:06do this automatically with YouTube
- 10:07titles and the YouTube data analytics V3
- 10:09API. You know, you could optimize
- 10:11subject lines for your newsletters in
- 10:12the same way. You could optimize pricing
- 10:14pages the same way. You could optimize
- 10:15literally whatever you want. And so
- 10:17hopefully it's clear that at least for,
- 10:19you know, most sales and marketing
- 10:21purposes, and we're not even going into
- 10:22the back end here, Auto Research allows
- 10:24you to build a consolidated set of
- 10:27knowledge on what works, what doesn't,
- 10:29and then have that running in the
- 10:30background for you 24/7 with no human
- 10:32involvement. But how the heck does this
- 10:34actually work? Well, three simple steps.
- 10:36The first is we're going to clone the
- 10:37repo. I'll show you how to do that in a
- 10:38moment. We're then going to write some
- 10:40sort of test, okay? And you can call it
- 10:42test.md, you can call it whatever you
- 10:43want. But this only needs to include a
- 10:45goal, a metric, and a test method. And
- 10:48then you just give the agent the ability
- 10:50to run this on autopilot. In my case,
- 10:52I'm using a service called GitHub
- 10:54Actions, which allows me to store this
- 10:56in the cloud and then run this on
- 10:57regular intervals, like 4 hours. You can
- 10:59use GitHub Actions, you could use Modal,
- 11:02you could use a billion other providers,
- 11:03and I'll show you how to do all of that
- 11:04right now. So the first thing you need
- 11:05to do is you just need to get the Auto
- 11:06Research repo. Now I have a link, it's
- 11:09the top one or top two in the
- 11:11description, so click on that, you'll
- 11:13head over to this page. This will
- 11:14include all the information, including
- 11:16the Python training scripts, the project
- 11:18description, a bunch of other scripts,
- 11:20and then what he's calling program.md,
- 11:23where you provide the model everything
- 11:25that it needs in order to manage this
- 11:26whole research process. So you see in
- 11:28his case he says, "This is an experiment
- 11:29where you're going to do your own
- 11:30research. Work with the user to agree on
- 11:32this, create this, read that, verify
- 11:35this exists, and so on and so forth. And
- 11:37I want you to know this stuff is not
- 11:38super important. He's actually
- 11:39explicitly said that his prompt is
- 11:41probably pretty crappy and that it'd be
- 11:43very easy to make a better one. So, with
- 11:45all that in mind, now we need to go over
- 11:47to our agent. Next, head over to an
- 11:48integrated development environment or
- 11:50some sort of tool that allows you to run
- 11:52Claude code. In my case, I'm using
- 11:54what's called Antigravity. You guys
- 11:56could use Visual Studio Code. You guys
- 11:57could use like a hundred different apps,
- 11:58to be honest. Um by the way, if this
- 12:00seems like magic to you, you don't know
- 12:02what any of the buttons on the page are,
- 12:03I literally run through all of it in an
- 12:05extensive 4-hour Claude code course that
- 12:07even teaches you like what the different
- 12:08icons are and so on and so forth. Really
- 12:10holds your hand through it. So, just
- 12:12head to the top uh right-hand corner of
- 12:13the video for that. Anyway, assuming you
- 12:15have all this stuff open, we're going to
- 12:16want to create a new folder. So, I'm
- 12:18going to go here to open folder. Then
- 12:21I'm going to go new.
- 12:22I'm going to say Carpathia Auto Research
- 12:24Demo.
- 12:25And then I'll click create. Then I'm
- 12:27going to open this folder. Okay, and now
- 12:29in order to open up Claude code, I'm
- 12:30just going to double-click anywhere in
- 12:31here. Click on my little Claude code
- 12:33button. And then what I want to do is I
- 12:35basically want to clone this. So, I'll
- 12:37say, "Hey, clone this in the current
- 12:41working directory."
- 12:43What this is going to do is it's going
- 12:44to make an HTTP request over to GitHub
- 12:47and then clone this service, store that
- 12:50down below in a folder called Auto
- 12:51Research. And the reason why we're doing
- 12:53this is cuz we just want all of the
- 12:54context of this whole repo before we go
- 12:56ahead and actually define, you know,
- 12:58what it is that we're going to do on
- 12:59this. And so, what I want to do, just
- 13:01for a demonstration sake, is I'm just
- 13:02going to reproduce my cold email
- 13:03example. After that, I'm going to give
- 13:05myself a little bit of space. And now,
- 13:06because I have access to a voice
- 13:08dictation tool called uh WhisperFlow
- 13:09down over here, I'm just going to hold
- 13:11my FN key and then tell it what I want.
- 13:14Hey, I want you to use the context in
- 13:16the Auto Research folder to help me
- 13:18build a very similar idea, except
- 13:21instead of testing for validation loss
- 13:23and iterating on a machine learning
- 13:24model, I want you to do all of this, for
- 13:27cold email. The metric I'm interested in
- 13:29optimizing for is my reply rate. The
- 13:32platform I'm going to be doing all this
- 13:33stuff on is Instantly, and I'll give you
- 13:35the API credentials and everything that
- 13:37you need in a moment. And finally, the
- 13:39thing that you're going to change
- 13:40between one experiment and the other is
- 13:42going to be the copy of the cold emails.
- 13:45Finally, I want you to take all this and
- 13:47then put this on the cloud using GitHub
- 13:49Actions, so it runs once every hour and
- 13:51it has everything it needs to work on
- 13:54autopilot. Once I pasted that in, press
- 13:56enter.
- 13:57Now it's going to go through all of the
- 13:58auto research documentation. You know,
- 14:01it has a few things here that's probably
- 14:02not super important like this image
- 14:04which shows
- 14:05I don't know, the progress on Karpathy
- 14:07side. You can see here his baseline was
- 14:09validation BPB. It's some form of
- 14:12basically accuracy, how good the model
- 14:14is. Started up here, and then after just
- 14:16a few runs it got all the way down over
- 14:18here by adjusting various parameters.
- 14:20This is more or less everything that's
- 14:22going to occur except with our cold
- 14:23email. So it'll be like um, you know,
- 14:25invert this graph. Reply rate will start
- 14:27here, and then the idea is the reply
- 14:29rate will go up over time.
- 14:31Anyway, I'm going to let it run for
- 14:32however long it needs to before it does
- 14:34everything that it has to. And then at
- 14:35the end of it, we're going to have a
- 14:36fully functional auto research campaign.
- 14:38It's now asking me some questions. How
- 14:40should the system generate new email
- 14:42copy variants? So I'll say Claude. Do
- 14:45you already have campaigns running in
- 14:46Instantly or will this create everything
- 14:48from scratch? So I'm going to say from
- 14:49scratch, then click submit answers. And
- 14:52it's now going through and building the
- 14:53email-optimizer. For simplicity, I'm
- 14:56naming it similar to the other one just
- 14:58so you guys could see what's going on.
- 15:00Now it's actually building an Instantly
- 15:02client which will contain all of the API
- 15:04calls that it needs to make to Instantly
- 15:06to get the information. We also have the
- 15:08orchestrator. Now orchestrator, just for
- 15:10anybody that doesn't isn't inherently
- 15:13familiar with the language, is basically
- 15:15almost always going to be like your top
- 15:17top top level agent. And the idea is
- 15:19it's the orchestrator which
- 15:20orchestrates, okay, kind of like a
- 15:22conductor in a symphony or something,
- 15:24the function of a bunch of lower level
- 15:26agents
- 15:27or tools. And so in this case, what this
- 15:29orchestrator is doing, I'm going to try
- 15:31and draw a little blue cloud code logo.
- 15:34Didn't do a very good job there. But
- 15:35basically what's occurring is this is
- 15:37orchestrating any sub-agents that we'll
- 15:39need for maybe the purposes of writing
- 15:41copy.
- 15:42Um it'll orchestrate the calling of like
- 15:44the instantly
- 15:45API. It'll orchestrate the I don't know
- 15:48storing of documents and
- 15:51uh I don't know JSON results and
- 15:52obviously you could build in a database.
- 15:53You could do whatever the heck you want
- 15:54there.
- 15:55And so this is what is essentially going
- 15:57to be us just speaking to the
- 15:58orchestrator saying, "Hey man, here's
- 16:00what you are. You're an email optimizer
- 16:02orchestrator and you have access to all
- 16:04this stuff." Next, the utility scripts
- 16:06are just little one-off API calls like
- 16:08tools that allow it to do things like
- 16:09purge old leads, deploy in batch, test
- 16:12my parsers, and so on.
- 16:14>> [gasps]
- 16:14>> The config files here like baseline,
- 16:16resource, and in this case I fed it some
- 16:18additional documentation from a big
- 16:20course I did inside of Maker School that
- 16:22teaches people how to write good quality
- 16:24cold emails.
- 16:25This is just things that I have the
- 16:27ability to change, so I can change my
- 16:28baseline test. That's the first test the
- 16:30cold email optimizer will ever test
- 16:33against. I could change what's in
- 16:34resources, although obviously that's
- 16:36going to be added to.
- 16:37And then I also have little tokens over
- 16:39here so I can access the APIs. And then
- 16:42finally the GitHub actions workflow. All
- 16:43right, and I'm scrolling through here.
- 16:45It's doing the vast majority of the
- 16:46work, which is pretty nice.
- 16:48And in in this case it's actually
- 16:50creating some sub-agents to do it for
- 16:51me. If I open this up, you could see we
- 16:53actually have a bunch of data. So this
- 16:55is a .env.example,
- 16:57so um I'm going to ask it to basically
- 16:59set this up as a demo, meaning you guys
- 17:01can just pump in whatever the heck you
- 17:02want. And then also I'm going to add a
- 17:04Slack webhook. The reason why I'm doing
- 17:06that is because I basically just want it
- 17:08to be able to tell me how it's doing
- 17:09whenever it makes the changes. Now that
- 17:11it's doing some testing, we can
- 17:12basically go ahead. So just for
- 17:13demonstration purposes I'll say, "Great.
- 17:16Create a baseline and a challenger and
- 17:19show me dry run this is a demo. You can
- 17:22see what it's written over here as well.
- 17:24The way it works is it runs every hour
- 17:26via GitHub Actions Cron. This is a
- 17:28scheduling tool that triggers once per
- 17:30hour. There's three steps. It's going to
- 17:32harvest by collecting results from the
- 17:34previous experiment. It'll generate by
- 17:36creating a new challenger and then it'll
- 17:38deploy by creating the campaigns,
- 17:40drawing the leads from a pre-existing
- 17:41pool or database that I've given it and
- 17:44then finally activating everything. So
- 17:46you see the leads are over here. We have
- 17:49uh the usage. Then obviously we have
- 17:51like the big fat long scripts as well. I
- 17:53didn't have to write any of it. And this
- 17:54all follows very similar logic to what
- 17:56Carpathia was doing. It's just instead
- 17:58of doing this for like machine learning
- 18:00purposes, we're obviously doing this for
- 18:01financial purposes, better reply rates.
- 18:03Now because so much of this stuff occurs
- 18:05completely autonomously, I would
- 18:06recommend you always have a way to
- 18:08visualize or at least keep track of
- 18:10things as they go. And so what I've done
- 18:12is I've set up a little Slack uh ping
- 18:14via webhook that notifies me every time
- 18:16a new challenger or a baseline variant
- 18:19test is created. And so what happened is
- 18:21the other day we actually tested three
- 18:22different ones. You could see some of
- 18:24these tests were pretty small and pretty
- 18:25minor. We just made adjustments to the
- 18:27subject line and so on and so forth, but
- 18:29it stores things like the baseline and
- 18:31the challenger. And then whenever a
- 18:32harvest occurs, it tells us more or less
- 18:34which one won. Okay, and we now have the
- 18:35baseline which is a subject of quick
- 18:37question. Gives me a bunch of uh
- 18:39baseline copy here. So I actually wrote
- 18:41this initial first email.
- 18:43And then it's generating the challenger.
- 18:44The hypothesis is the baseline is too
- 18:46long. It buries the offer and it also
- 18:48lacks a specific CTA time. So it's going
- 18:50to try rewriting it to sub-75 words,
- 18:52leading with relevance, front-loading
- 18:54the risk reversal, and ending with a
- 18:55concrete time ask. You can see it's
- 18:57quite the significant change here. Hey
- 18:59first name, I drive PPC leads for a
- 19:01two-million-a-year dental marketing firm
- 19:02in Calgary. I've sent over 10 million in
- 19:03business to agencies like yours through
- 19:05cold outbound alone. Got a backlog of
- 19:07people wanting PPC right now a variety
- 19:08of verticals. I'd send you booked
- 19:09appointments and only charge if we hit a
- 19:11number you and I agree on beforehand.
- 19:13Zero risk on your end. Is this worth a
- 19:14quick call or even then gives a specific
- 19:16time. So, I mean it remains to be seen
- 19:18whether the challenger is going to be
- 19:19better than the the baseline, of course,
- 19:21but that's just part of the game. And
- 19:22you can see this is now actually been
- 19:23deployed. We have uh that same copy over
- 19:26here. And then if we go over to our
- 19:28baseline, we also have the baseline copy
- 19:30over here, which is that old cold email.
- 19:33And basically what's going to occur now
- 19:34is they're just going to test against
- 19:35each other until we figure out which one
- 19:37is better. Now, on net because this is
- 19:39an AI model we're working with, in my
- 19:41experience most challengers are not up
- 19:43to the task of the baseline. Usually the
- 19:45baseline is better because I wrote it,
- 19:47but eventually the challengers do become
- 19:50better and you start seeing significant
- 19:52improvements in the reply rate relative
- 19:54to the original, um which then you know
- 19:56makes them higher and higher and higher
- 19:57and higher over time. Then the
- 19:59challenger becomes the new baseline and
- 20:01then you just repeat. And so basically
- 20:02what this is, to be honest, is like the
- 20:04automation of like scientific
- 20:05experiments.
- 20:07Um
- 20:07you know, this is something where right
- 20:09now there's so much logistical overhead
- 20:11and friction involved in like running
- 20:13any sort of experiment, whether you're a
- 20:14marketer, a salesperson, somebody doing
- 20:16some back-end function or business or
- 20:17whatever. And this just eliminates all
- 20:19that friction. I no longer have to like
- 20:21copy and paste the leads. I no longer
- 20:22have to like do anything manually. Um
- 20:24it's all done via simple API calls. And
- 20:26now that we can put this thing on a
- 20:28loop, even though every time I run the
- 20:29orchestrator it's technically like a
- 20:30different agent, it has all the context
- 20:32from all of the previous runs, which
- 20:34allows it to grow more intelligent over
- 20:35time. I anticipate that eventually after
- 20:38something like 500 to maybe 1,000 runs,
- 20:41you'll probably have to consolidate some
- 20:42of the previous learnings so that that
- 20:44document doesn't get super long, but
- 20:46whatever you're using this for, whether
- 20:47landing pages, PPC, uh newsletter copy,
- 20:51whatever the heck, you know, SEO pages,
- 20:53hopefully you guys understand that as
- 20:54things get better, the new challengers
- 20:57are just going to be millions upon
- 20:59millions of times more profitable and
- 21:01efficacious than uh what your initial
- 21:03baseline was. Now, I should note there's
- 21:05There's other things that have to do in
- 21:06order to set this up completely, like
- 21:08for instance, I actually had to go grab
- 21:09my API keys. The way you do this on
- 21:10Instantly is pretty straightforward,
- 21:12settings, then you go integrations, then
- 21:13you go API keys down here, then you
- 21:15create an API key, say whatever the heck
- 21:17you want, select scopes all, and then
- 21:18actually copy it over. Um, you obviously
- 21:20need to do that as well with whatever AI
- 21:22model you're using to do the
- 21:23orchestration. In my case, I was using
- 21:25Claude Opus 4.6, so I just went over to
- 21:27Anthropic, got their API key. Um, and
- 21:28then, you know, if you have any other
- 21:29other services you want to use like
- 21:31GitHub for instance or whatever, you
- 21:33also have to push that up. But, agents
- 21:35will handle all that stuff for you. Just
- 21:37ask them, "Hey, you know, where do I go
- 21:38to get my API key? Okay, can you sign me
- 21:40in?" and so on and so forth, and you'll
- 21:41be good to go. The final thing I want to
- 21:43talk about are use cases that I would
- 21:44consider not ideal for some form of auto
- 21:47optimization. Um, in general, things
- 21:49that work really well are things that
- 21:51have fast feedback loops, okay? So, why
- 21:54did Karpathy's um, AI agent or like nano
- 21:57GPT loop work so well? Because it was
- 22:00literally a five-minute loop.
- 22:02You know, if you have a five-minute
- 22:03loop, technically speaking, that means
- 22:05that in 60 minutes, you could run 12
- 22:08experiments. And so, obviously, 12
- 22:10experiments is a lot of data, and
- 22:12assuming that, you know, you're running
- 22:13it on your own servers or whatever, you
- 22:14can just have that thing churn. Um, but
- 22:16basically, that means that your your
- 22:17iteration loop will be much faster
- 22:20because you'll be able to kind of draw
- 22:21like this as opposed to like this, you
- 22:23know?
- 22:24It's going to take a lot longer to
- 22:25figure out what works and what doesn't
- 22:27if you're all the way down here.
- 22:28Another good thing to keep in mind is
- 22:30you need a clear metric. So, in my case,
- 22:32reply rate was a fantastic metric. Why?
- 22:34That's objective. How many people
- 22:35actually my email campaigns?
- 22:37Click-through rate, very, very objective
- 22:39cuz, you know, obviously, this is people
- 22:41clicking through an email. All this
- 22:42stuff is automatically tracked. But, if
- 22:43you had something that was way fuzzier
- 22:46in addition to way slower, you know, the
- 22:48probability of you actually making this
- 22:50work uh, is is much lower because how do
- 22:52you subjectively measure like warmth,
- 22:55you know? You can't. It's like
- 22:57happiness. It's like you can't. What you
- 22:58have to do is you have to find proxies
- 23:00for all these things, which are usually
- 23:01like scales and metrics and analytics
- 23:04and so on and so forth.
- 23:05And then the third thing is you need
- 23:07some sort of API access to change the
- 23:08inputs. Um and if you don't have the API
- 23:10access, you could build some sort of
- 23:12Chrome DevTools or CLI based flow, but
- 23:14like you need to have that because if
- 23:16you don't, how the heck is the agent
- 23:17supposed to make any changes? What are
- 23:18they going to do? Just give you a list
- 23:19of changes to manually go in? You can do
- 23:21that, but that sort of defeats the whole
- 23:22purpose.
- 23:23Anyway, so what I'm going to do with
- 23:25this is I'm going to provide everything
- 23:26that you guys need, including the email
- 23:28optimizer repo, uh the Carpathy auto
- 23:31research GitHub repo, and everything
- 23:33else down below. Feel free to take a
- 23:35look at it, give it a click, explore it,
- 23:37and use it for your own use case.
- 23:38I'd be really interested to hear what
- 23:40you guys end up using this on.
- 23:42Um this is what all major labs that are
- 23:44working on machine learning models
- 23:45around the world are currently doing, by
- 23:47the way, in case it wasn't clear.
- 23:48They're constantly running many, many,
- 23:50many experiments behind the scenes
- 23:52overnight to like make their models
- 23:53better and so on and so forth. So, the
- 23:55fact that we're able to democratize that
- 23:56and now do that for ourselves, for our
- 23:58own businesses, and for our own own
- 23:59models is now like incredible. But I'd
- 24:02be really curious to hear what sort of
- 24:03use cases you guys have with us. And,
- 24:05you know, if it makes sense, I could
- 24:06compile a list of these use cases and
- 24:08then I could make another follow-up
- 24:09video that just goes through every
- 24:10single one and then even gives like real
- 24:12examples of them. Um because that'd be
- 24:13really dope.
- 24:15Aside from that, if you guys could do me
- 24:15a big solid, something like 73% of you
- 24:18are not subscribed to the channel, which
- 24:19really hurts cuz I try and make
- 24:20high-quality content for both
- 24:22subscribers and non-subscribers, but
- 24:23YouTube pushes my content way more
- 24:25heavily when that ratio improves. So, if
- 24:27if I gave you any value today
- 24:28whatsoever, please do click the
- 24:29subscribe button. I hate asking for it,
- 24:31but it just makes a difference on
- 24:32YouTube, so I'll do what works. If you
- 24:33guys want more on Claude Code and so on
- 24:35and so forth, definitely check that out.
- 24:37And yeah, I mean, I will catch all y'all
- 24:38in the next video. Thanks so much for
- 24:40watching, guys, per usual. See you.
About this transcript
This page contains the full transcript of Claude Code + Karpathy Autoresearch = The New Meta by Nick Saraev, generated from the public captions YouTube serves with the video. The transcript has 5,575 words across 803 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.