YouTube2Text

Claude Code + Karpathy Autoresearch = The New Meta — Transcript

by Nick Saraev · 5,575 words · 803 segments · language en · Watch on YouTube

Full transcript

  1. 0:00An open-source project just dropped
  2. 0:01that, when you combine it with Claude
  3. 0:03code, literally becomes self-improving
  4. 0:05AI. This is not engagement farming or
  5. 0:07hype bait. This is a real repo that was
  6. 0:09just released by Andre Karpathy, who's
  7. 0:11widely renowned as one of the foremost
  8. 0:13voices in AI and machine learning
  9. 0:14research. And basically, what he did
  10. 0:16was, while training his model, he
  11. 0:18thought, "Why don't I just have my
  12. 0:19models train my models instead?" He
  13. 0:21built an elegant pipeline, which he's
  14. 0:23calling auto research, and essentially
  15. 0:25completely and fully automates the
  16. 0:27process of experimentation. He says it
  17. 0:29right here. The idea is to give an AI
  18. 0:31agent a small but real LLM training
  19. 0:34setup and just let it experiment
  20. 0:35autonomously overnight. It'll modify the
  21. 0:37code, train for 5 minutes, check if the
  22. 0:40results improved, keep or discard, and
  23. 0:42then just repeat. You wake up in the
  24. 0:43morning to a log of experiments and it
  25. 0:45hopefully a better model. Now, I'm not
  26. 0:48in machine learning training. I don't
  27. 0:50help make models more intelligent. What
  28. 0:52I do is I take models that other people
  29. 0:54have made, and then I use them for the
  30. 0:56purposes of making money. And so,
  31. 0:58immediately when this dropped, I started
  32. 1:00thinking about ways that I could apply
  33. 1:01this principle of auto research into my
  34. 1:03own life to improve, obviously, my own
  35. 1:05economic outcomes. And there are so
  36. 1:07many, it's not even funny. So, what I'm
  37. 1:08going to do is I'm going to run through
  38. 1:09some real practical examples in a
  39. 1:11moment, things that I'm actually doing
  40. 1:12in my own business that you can
  41. 1:13implement inside of Claude code. Then,
  42. 1:15I'm going to show you how to actually do
  43. 1:16it, so go through the step-by-step of
  44. 1:18setting up the repo and building some
  45. 1:20experimentation done totally
  46. 1:22autonomously for you. And at the end,
  47. 1:23you will have a fully automated,
  48. 1:25self-improving pipeline, just like Andre
  49. 1:27Karpathy here has done for his own
  50. 1:29machine learning training. So, here's
  51. 1:30one of many examples. I do a lot of cold
  52. 1:32email in my own business and then for
  53. 1:34clients. And cold email, in case you
  54. 1:35didn't know, is where you package up a
  55. 1:37really nice, sexy-sounding offer, and
  56. 1:39then you send it to people you've never
  57. 1:41met with the hopes that they take you up
  58. 1:42on it and then maybe convert. So, jump
  59. 1:44on a call with you, fill out a form,
  60. 1:45whatever. Now, the key metric in cold
  61. 1:47email is usually reply rate,
  62. 1:49specifically positive reply rate, but
  63. 1:51reply rate's easier for our purposes, so
  64. 1:52that's what I'm going to go with. And as
  65. 1:54you could see, most cold email software
  66. 1:55tracks this for you out of the box. So,
  67. 1:572.4% of people replied to this campaign.
  68. 2:002.5% of people replied to this campaign,
  69. 2:03and so on. Well, turns out that's all
  70. 2:05you need in order to build an automatic
  71. 2:06experimentation pipeline. You need some
  72. 2:08metric that you want to improve, and
  73. 2:10then you need some way factor, some
  74. 2:12thing you can modify to improve it. And
  75. 2:14so, what I have is my metric is reply
  76. 2:16rate, and then what I have is the thing
  77. 2:18that I can adjust is my cold email copy.
  78. 2:20So, what this looks like in practice for
  79. 2:21me is a folder called email optimizer.
  80. 2:25There are a bunch of additional code
  81. 2:26files here, configs, places where I'm
  82. 2:29storing the results of my data, and so
  83. 2:31on and so forth, and aren't super
  84. 2:32important. What is important is this
  85. 2:34file right over here called
  86. 2:35orchestrator.py.
  87. 2:37And this contains all of the prompts
  88. 2:39that I'm feeding into my orchestrator
  89. 2:41agent, who essentially is responsible
  90. 2:43for spinning up new cold email campaigns
  91. 2:45and then testing them against each other
  92. 2:46until I get better and better results.
  93. 2:48And it was as simple for me as literally
  94. 2:49copying the repo and making some slight
  95. 2:51adjustments. What I do is I tell it that
  96. 2:53that's inspired by Carbon the thesis
  97. 2:55auto research pattern. The core idea is
  98. 2:57an AI agent that runs experiments
  99. 2:58autonomously in a tight loop using an
  100. 3:00objective metric as a feedback signal. I
  101. 3:02have the architecture over here. I run
  102. 3:05the loop every 4 hours, and at the end
  103. 3:07of every 4 hours, I actually have better
  104. 3:10copy uh that's self-evolving over time
  105. 3:12based on the results from my previous
  106. 3:14test. And you can see some of the
  107. 3:15examples right over here. Anything with
  108. 3:17C is what we call a challenger. Anything
  109. 3:19with B is what's called baseline. And
  110. 3:21so, the model starts with a baseline
  111. 3:23type of copy, which we could see right
  112. 3:25over here. And then it makes slight
  113. 3:28modifications based off of what it knows
  114. 3:30to perform really well in cold email
  115. 3:32copy before testing it out. It runs the
  116. 3:34two side by side, and then automatically
  117. 3:36harvests based off of the results, aka
  118. 3:39the number of replies, and so on and so
  119. 3:40forth for both campaigns. Now, that part
  120. 3:42isn't the important bit. I mean, we've
  121. 3:43been optimizing cold emails for for many
  122. 3:45years at this point. The important bit
  123. 3:47is it then creates new copy based off of
  124. 3:50the learnings from previous experiments.
  125. 3:53As the models get better and better and
  126. 3:55better, they log all of their learnings
  127. 3:57to a resource.md
  128. 3:59that significantly improves future
  129. 4:01models abilities to make changes. And so
  130. 4:03here's a big list of things that this is
  131. 4:06essentially figured out, move reply rate
  132. 4:09up. And so in that way we get to push
  133. 4:11towards that direction over time. Now
  134. 4:13this has only been running for a few
  135. 4:14days now. Imagine this running for a
  136. 4:16year. Instead of optimizing on a basis
  137. 4:18of once every 4 hours, imagine if this
  138. 4:20optimized on a basis of once every 5
  139. 4:22minutes. Well, that's what my next leg
  140. 4:23of testing is going to do. We're going
  141. 4:25to be significantly improving the
  142. 4:26volume, pumping all this stuff out at
  143. 4:2810x the level, and then optimizing and
  144. 4:31iterating our results again fully
  145. 4:32autonomously. So that's just one
  146. 4:33example. I'm going to run you guys
  147. 4:35through a bunch of other use cases that
  148. 4:36you can apply auto research to, whether
  149. 4:38you're in machine learning engineering
  150. 4:40or whether you're just trying to improve
  151. 4:42the profitability of let's say
  152. 4:44paper click ad campaign. But first,
  153. 4:45let's make sure we all know how to
  154. 4:47actually use this thing. So, the way
  155. 4:49that auto research works, to make a long
  156. 4:51story short, is we start with an
  157. 4:53experiment. And just like in science,
  158. 4:55everything begins with some sort of
  159. 4:57hypothesis, okay? So my hypothesis might
  160. 5:00be, hey, if I make a slight adjustment
  161. 5:03to the copy of this campaign so that
  162. 5:06it's a little bit punchier, I think it's
  163. 5:08going to go well. You insert that using
  164. 5:10this little test.md. It's your goal,
  165. 5:12metric, and some high-level
  166. 5:13instructions. From there, the auto
  167. 5:16research agent will go through, employ
  168. 5:19the experiment usually using API calls.
  169. 5:21In my case, in the example we just saw,
  170. 5:24um to instantly, in Karpathy's specific
  171. 5:26example we saw, he's doing it through
  172. 5:29adjusting what are called hyper
  173. 5:30parameters. And then after that, we
  174. 5:32measure the results. Now, in order for
  175. 5:34us to make sure that this works, you
  176. 5:36know, the hypothesis isn't enough. We
  177. 5:37need some sort of metric that we're
  178. 5:38tracking. Now in my case, the metric was
  179. 5:41obviously pretty simple. It was reply
  180. 5:42rate. In Karpathy's case, it was pretty
  181. 5:44simple. It was something called
  182. 5:45validation loss. As long as you have
  183. 5:47that, you can then just pick the winner
  184. 5:50and then make a slight change before
  185. 5:51looping back. And depending on how tight
  186. 5:54this feedback loop is, you could
  187. 5:55theoretically do this in a minute or
  188. 5:57two. I mean, if he had more
  189. 5:59infrastructure when he's training his
  190. 6:00models, he could probably do in 5
  191. 6:01minutes what he does in one. And then in
  192. 6:03that way, you know, progress really,
  193. 6:05really quickly over to some, you know,
  194. 6:07desired goal. In my case, if I had more
  195. 6:09cold email infrastructure, I could do
  196. 6:10the same thing. So, at this point, scale
  197. 6:12is more or less all you need. This
  198. 6:13allows you to run hundreds of tests with
  199. 6:15literally zero human involvement. I
  200. 6:17mean, I'm not even in the loop anymore.
  201. 6:19And to be clear, like if I was in the
  202. 6:21loop, would I be making better decisions
  203. 6:23than the AI model? Like probably. I'd be
  204. 6:25a much more efficient optimizer. But
  205. 6:27that doesn't really matter because the
  206. 6:28reality is I take a lot more time to
  207. 6:31optimize than a model does. I also eat,
  208. 6:33sleep, have to go to the washroom, and
  209. 6:36do a variety of other things with my
  210. 6:37day. AI agents don't. You could very
  211. 6:39quickly and easily set this up on,
  212. 6:41again, an hourly loop and have this run
  213. 6:4324 times a day, whereas realistically,
  214. 6:46if you were to try and do it all
  215. 6:47yourself, you could only do it a couple
  216. 6:48times. And so in that way, whatever
  217. 6:50metric that you're tracking goes up over
  218. 6:52time, right? In my case, reply rates
  219. 6:54significantly go up. You know, test one,
  220. 6:56I might be at a 1.5%, test 12, I might
  221. 6:59be at a 2.7%. Before you know it, I
  222. 7:01reach literally like the optimal quality
  223. 7:04possible for my set of cold emails and
  224. 7:07then their audiences. And you can apply
  225. 7:08this, as mentioned, to a bunch of other
  226. 7:09strategies. So, what are those
  227. 7:11strategies? The requirement that you
  228. 7:13need is anything that has an objective
  229. 7:17metric you can track
  230. 7:18and an API or application programming
  231. 7:21interface that you can send a request to
  232. 7:24to get. Okay? So, some brief examples of
  233. 7:27this. Cold email copy. Obviously,
  234. 7:30fantastic. Why? Well, because in our
  235. 7:32case, we have the Instantly API. The
  236. 7:35Instantly API allows us to query
  237. 7:37metrics, and so I can give the agent the
  238. 7:39ability to call a quick tool, call up
  239. 7:41the Instantly API, see how the
  240. 7:43performance was relative to you know the
  241. 7:45the challenger in the base campaign.
  242. 7:47At the same time, you know, I have a
  243. 7:49very clear metric, which in my case is
  244. 7:51reply rate.
  245. 7:52Okay, how about landing pages? Let's say
  246. 7:55you're doing some form of CRO, which is
  247. 7:57conversion rate optimization, and you
  248. 7:59want to test to see how you can make
  249. 8:01your landing pages as efficient as
  250. 8:02humanly possible. Well, you can now
  251. 8:04completely automate it with auto
  252. 8:06research. What you do is you pick the
  253. 8:08metric that you want, which in our case
  254. 8:10would literally just be conversion rate,
  255. 8:13okay?
  256. 8:14And then if your website is hosted
  257. 8:15locally or it's hosted using some API or
  258. 8:18something like that, let's say a website
  259. 8:19builder like Wix or or or or WordPress
  260. 8:22or Webflow, what you can do is you can
  261. 8:24give it access to the API, and then you
  262. 8:27can say, "Hey, change this according to
  263. 8:29this resource of best practices that
  264. 8:31other agents have done. Make your
  265. 8:32change, test that for, I don't know, a
  266. 8:34day, depending on how much volume you
  267. 8:36have, and at the end
  268. 8:38consolidate the winner and then get rid
  269. 8:39of the loser." You can do the exact same
  270. 8:41thing for ad creatives, okay? What's the
  271. 8:43main thing that you want for ad
  272. 8:45creatives? Obviously, it's going to be
  273. 8:45some form of conversion rate as well,
  274. 8:47whatever specific type of conversion
  275. 8:49rate is, that's up to you, okay? But all
  276. 8:51you need to do is query some sort of
  277. 8:53API. Now, you know, a lot of these ad
  278. 8:56platforms like Facebook and Google
  279. 8:58basically already do this for you. Mind
  280. 9:00it, granted I don't think they do it
  281. 9:02anywhere near as effectively as you can
  282. 9:03with modern models like Opus 4.6 or GPT
  283. 9:065.4. What you can do is you can give it
  284. 9:08the API to call a specific ad resource,
  285. 9:11and then you could also just give it the
  286. 9:13metric to optimize for, which is CVR,
  287. 9:15and then it'll crush.
  288. 9:17How about some form of customer
  289. 9:18satisfaction for chatbot scripts? Maybe
  290. 9:21use some sort of customer satisfaction
  291. 9:23score, and then, you know, now you just
  292. 9:25adjust the main template, okay, that all
  293. 9:28customer service agents, whether human
  294. 9:30or AI, are are going off of. That's
  295. 9:32super simple and easy to do. How about
  296. 9:34product descriptions for some sort of
  297. 9:36e-comm? If you have like, I don't know,
  298. 9:38Amazon FBA or something like that, you
  299. 9:41know, maybe they don't necessarily have
  300. 9:42APIs, but maybe now you set up what's
  301. 9:45called Chrome DevTools MCP,
  302. 9:48give it a very tightly scoped list of
  303. 9:50steps that has to do to update the
  304. 9:51actual body of the landing page, and
  305. 9:54then based off of metrics like, I don't
  306. 9:56know, how many freaking dollars you've
  307. 9:58sold in the last little while, you can
  308. 9:59very quickly optimize and make your
  309. 10:01product landing page better and better
  310. 10:02and better. You know, in my case, I make
  311. 10:04a lot of YouTube content thesis. I could
  312. 10:06do this automatically with YouTube
  313. 10:07titles and the YouTube data analytics V3
  314. 10:09API. You know, you could optimize
  315. 10:11subject lines for your newsletters in
  316. 10:12the same way. You could optimize pricing
  317. 10:14pages the same way. You could optimize
  318. 10:15literally whatever you want. And so
  319. 10:17hopefully it's clear that at least for,
  320. 10:19you know, most sales and marketing
  321. 10:21purposes, and we're not even going into
  322. 10:22the back end here, Auto Research allows
  323. 10:24you to build a consolidated set of
  324. 10:27knowledge on what works, what doesn't,
  325. 10:29and then have that running in the
  326. 10:30background for you 24/7 with no human
  327. 10:32involvement. But how the heck does this
  328. 10:34actually work? Well, three simple steps.
  329. 10:36The first is we're going to clone the
  330. 10:37repo. I'll show you how to do that in a
  331. 10:38moment. We're then going to write some
  332. 10:40sort of test, okay? And you can call it
  333. 10:42test.md, you can call it whatever you
  334. 10:43want. But this only needs to include a
  335. 10:45goal, a metric, and a test method. And
  336. 10:48then you just give the agent the ability
  337. 10:50to run this on autopilot. In my case,
  338. 10:52I'm using a service called GitHub
  339. 10:54Actions, which allows me to store this
  340. 10:56in the cloud and then run this on
  341. 10:57regular intervals, like 4 hours. You can
  342. 10:59use GitHub Actions, you could use Modal,
  343. 11:02you could use a billion other providers,
  344. 11:03and I'll show you how to do all of that
  345. 11:04right now. So the first thing you need
  346. 11:05to do is you just need to get the Auto
  347. 11:06Research repo. Now I have a link, it's
  348. 11:09the top one or top two in the
  349. 11:11description, so click on that, you'll
  350. 11:13head over to this page. This will
  351. 11:14include all the information, including
  352. 11:16the Python training scripts, the project
  353. 11:18description, a bunch of other scripts,
  354. 11:20and then what he's calling program.md,
  355. 11:23where you provide the model everything
  356. 11:25that it needs in order to manage this
  357. 11:26whole research process. So you see in
  358. 11:28his case he says, "This is an experiment
  359. 11:29where you're going to do your own
  360. 11:30research. Work with the user to agree on
  361. 11:32this, create this, read that, verify
  362. 11:35this exists, and so on and so forth. And
  363. 11:37I want you to know this stuff is not
  364. 11:38super important. He's actually
  365. 11:39explicitly said that his prompt is
  366. 11:41probably pretty crappy and that it'd be
  367. 11:43very easy to make a better one. So, with
  368. 11:45all that in mind, now we need to go over
  369. 11:47to our agent. Next, head over to an
  370. 11:48integrated development environment or
  371. 11:50some sort of tool that allows you to run
  372. 11:52Claude code. In my case, I'm using
  373. 11:54what's called Antigravity. You guys
  374. 11:56could use Visual Studio Code. You guys
  375. 11:57could use like a hundred different apps,
  376. 11:58to be honest. Um by the way, if this
  377. 12:00seems like magic to you, you don't know
  378. 12:02what any of the buttons on the page are,
  379. 12:03I literally run through all of it in an
  380. 12:05extensive 4-hour Claude code course that
  381. 12:07even teaches you like what the different
  382. 12:08icons are and so on and so forth. Really
  383. 12:10holds your hand through it. So, just
  384. 12:12head to the top uh right-hand corner of
  385. 12:13the video for that. Anyway, assuming you
  386. 12:15have all this stuff open, we're going to
  387. 12:16want to create a new folder. So, I'm
  388. 12:18going to go here to open folder. Then
  389. 12:21I'm going to go new.
  390. 12:22I'm going to say Carpathia Auto Research
  391. 12:24Demo.
  392. 12:25And then I'll click create. Then I'm
  393. 12:27going to open this folder. Okay, and now
  394. 12:29in order to open up Claude code, I'm
  395. 12:30just going to double-click anywhere in
  396. 12:31here. Click on my little Claude code
  397. 12:33button. And then what I want to do is I
  398. 12:35basically want to clone this. So, I'll
  399. 12:37say, "Hey, clone this in the current
  400. 12:41working directory."
  401. 12:43What this is going to do is it's going
  402. 12:44to make an HTTP request over to GitHub
  403. 12:47and then clone this service, store that
  404. 12:50down below in a folder called Auto
  405. 12:51Research. And the reason why we're doing
  406. 12:53this is cuz we just want all of the
  407. 12:54context of this whole repo before we go
  408. 12:56ahead and actually define, you know,
  409. 12:58what it is that we're going to do on
  410. 12:59this. And so, what I want to do, just
  411. 13:01for a demonstration sake, is I'm just
  412. 13:02going to reproduce my cold email
  413. 13:03example. After that, I'm going to give
  414. 13:05myself a little bit of space. And now,
  415. 13:06because I have access to a voice
  416. 13:08dictation tool called uh WhisperFlow
  417. 13:09down over here, I'm just going to hold
  418. 13:11my FN key and then tell it what I want.
  419. 13:14Hey, I want you to use the context in
  420. 13:16the Auto Research folder to help me
  421. 13:18build a very similar idea, except
  422. 13:21instead of testing for validation loss
  423. 13:23and iterating on a machine learning
  424. 13:24model, I want you to do all of this, for
  425. 13:27cold email. The metric I'm interested in
  426. 13:29optimizing for is my reply rate. The
  427. 13:32platform I'm going to be doing all this
  428. 13:33stuff on is Instantly, and I'll give you
  429. 13:35the API credentials and everything that
  430. 13:37you need in a moment. And finally, the
  431. 13:39thing that you're going to change
  432. 13:40between one experiment and the other is
  433. 13:42going to be the copy of the cold emails.
  434. 13:45Finally, I want you to take all this and
  435. 13:47then put this on the cloud using GitHub
  436. 13:49Actions, so it runs once every hour and
  437. 13:51it has everything it needs to work on
  438. 13:54autopilot. Once I pasted that in, press
  439. 13:56enter.
  440. 13:57Now it's going to go through all of the
  441. 13:58auto research documentation. You know,
  442. 14:01it has a few things here that's probably
  443. 14:02not super important like this image
  444. 14:04which shows
  445. 14:05I don't know, the progress on Karpathy
  446. 14:07side. You can see here his baseline was
  447. 14:09validation BPB. It's some form of
  448. 14:12basically accuracy, how good the model
  449. 14:14is. Started up here, and then after just
  450. 14:16a few runs it got all the way down over
  451. 14:18here by adjusting various parameters.
  452. 14:20This is more or less everything that's
  453. 14:22going to occur except with our cold
  454. 14:23email. So it'll be like um, you know,
  455. 14:25invert this graph. Reply rate will start
  456. 14:27here, and then the idea is the reply
  457. 14:29rate will go up over time.
  458. 14:31Anyway, I'm going to let it run for
  459. 14:32however long it needs to before it does
  460. 14:34everything that it has to. And then at
  461. 14:35the end of it, we're going to have a
  462. 14:36fully functional auto research campaign.
  463. 14:38It's now asking me some questions. How
  464. 14:40should the system generate new email
  465. 14:42copy variants? So I'll say Claude. Do
  466. 14:45you already have campaigns running in
  467. 14:46Instantly or will this create everything
  468. 14:48from scratch? So I'm going to say from
  469. 14:49scratch, then click submit answers. And
  470. 14:52it's now going through and building the
  471. 14:53email-optimizer. For simplicity, I'm
  472. 14:56naming it similar to the other one just
  473. 14:58so you guys could see what's going on.
  474. 15:00Now it's actually building an Instantly
  475. 15:02client which will contain all of the API
  476. 15:04calls that it needs to make to Instantly
  477. 15:06to get the information. We also have the
  478. 15:08orchestrator. Now orchestrator, just for
  479. 15:10anybody that doesn't isn't inherently
  480. 15:13familiar with the language, is basically
  481. 15:15almost always going to be like your top
  482. 15:17top top level agent. And the idea is
  483. 15:19it's the orchestrator which
  484. 15:20orchestrates, okay, kind of like a
  485. 15:22conductor in a symphony or something,
  486. 15:24the function of a bunch of lower level
  487. 15:26agents
  488. 15:27or tools. And so in this case, what this
  489. 15:29orchestrator is doing, I'm going to try
  490. 15:31and draw a little blue cloud code logo.
  491. 15:34Didn't do a very good job there. But
  492. 15:35basically what's occurring is this is
  493. 15:37orchestrating any sub-agents that we'll
  494. 15:39need for maybe the purposes of writing
  495. 15:41copy.
  496. 15:42Um it'll orchestrate the calling of like
  497. 15:44the instantly
  498. 15:45API. It'll orchestrate the I don't know
  499. 15:48storing of documents and
  500. 15:51uh I don't know JSON results and
  501. 15:52obviously you could build in a database.
  502. 15:53You could do whatever the heck you want
  503. 15:54there.
  504. 15:55And so this is what is essentially going
  505. 15:57to be us just speaking to the
  506. 15:58orchestrator saying, "Hey man, here's
  507. 16:00what you are. You're an email optimizer
  508. 16:02orchestrator and you have access to all
  509. 16:04this stuff." Next, the utility scripts
  510. 16:06are just little one-off API calls like
  511. 16:08tools that allow it to do things like
  512. 16:09purge old leads, deploy in batch, test
  513. 16:12my parsers, and so on.
  514. 16:14>> [gasps]
  515. 16:14>> The config files here like baseline,
  516. 16:16resource, and in this case I fed it some
  517. 16:18additional documentation from a big
  518. 16:20course I did inside of Maker School that
  519. 16:22teaches people how to write good quality
  520. 16:24cold emails.
  521. 16:25This is just things that I have the
  522. 16:27ability to change, so I can change my
  523. 16:28baseline test. That's the first test the
  524. 16:30cold email optimizer will ever test
  525. 16:33against. I could change what's in
  526. 16:34resources, although obviously that's
  527. 16:36going to be added to.
  528. 16:37And then I also have little tokens over
  529. 16:39here so I can access the APIs. And then
  530. 16:42finally the GitHub actions workflow. All
  531. 16:43right, and I'm scrolling through here.
  532. 16:45It's doing the vast majority of the
  533. 16:46work, which is pretty nice.
  534. 16:48And in in this case it's actually
  535. 16:50creating some sub-agents to do it for
  536. 16:51me. If I open this up, you could see we
  537. 16:53actually have a bunch of data. So this
  538. 16:55is a .env.example,
  539. 16:57so um I'm going to ask it to basically
  540. 16:59set this up as a demo, meaning you guys
  541. 17:01can just pump in whatever the heck you
  542. 17:02want. And then also I'm going to add a
  543. 17:04Slack webhook. The reason why I'm doing
  544. 17:06that is because I basically just want it
  545. 17:08to be able to tell me how it's doing
  546. 17:09whenever it makes the changes. Now that
  547. 17:11it's doing some testing, we can
  548. 17:12basically go ahead. So just for
  549. 17:13demonstration purposes I'll say, "Great.
  550. 17:16Create a baseline and a challenger and
  551. 17:19show me dry run this is a demo. You can
  552. 17:22see what it's written over here as well.
  553. 17:24The way it works is it runs every hour
  554. 17:26via GitHub Actions Cron. This is a
  555. 17:28scheduling tool that triggers once per
  556. 17:30hour. There's three steps. It's going to
  557. 17:32harvest by collecting results from the
  558. 17:34previous experiment. It'll generate by
  559. 17:36creating a new challenger and then it'll
  560. 17:38deploy by creating the campaigns,
  561. 17:40drawing the leads from a pre-existing
  562. 17:41pool or database that I've given it and
  563. 17:44then finally activating everything. So
  564. 17:46you see the leads are over here. We have
  565. 17:49uh the usage. Then obviously we have
  566. 17:51like the big fat long scripts as well. I
  567. 17:53didn't have to write any of it. And this
  568. 17:54all follows very similar logic to what
  569. 17:56Carpathia was doing. It's just instead
  570. 17:58of doing this for like machine learning
  571. 18:00purposes, we're obviously doing this for
  572. 18:01financial purposes, better reply rates.
  573. 18:03Now because so much of this stuff occurs
  574. 18:05completely autonomously, I would
  575. 18:06recommend you always have a way to
  576. 18:08visualize or at least keep track of
  577. 18:10things as they go. And so what I've done
  578. 18:12is I've set up a little Slack uh ping
  579. 18:14via webhook that notifies me every time
  580. 18:16a new challenger or a baseline variant
  581. 18:19test is created. And so what happened is
  582. 18:21the other day we actually tested three
  583. 18:22different ones. You could see some of
  584. 18:24these tests were pretty small and pretty
  585. 18:25minor. We just made adjustments to the
  586. 18:27subject line and so on and so forth, but
  587. 18:29it stores things like the baseline and
  588. 18:31the challenger. And then whenever a
  589. 18:32harvest occurs, it tells us more or less
  590. 18:34which one won. Okay, and we now have the
  591. 18:35baseline which is a subject of quick
  592. 18:37question. Gives me a bunch of uh
  593. 18:39baseline copy here. So I actually wrote
  594. 18:41this initial first email.
  595. 18:43And then it's generating the challenger.
  596. 18:44The hypothesis is the baseline is too
  597. 18:46long. It buries the offer and it also
  598. 18:48lacks a specific CTA time. So it's going
  599. 18:50to try rewriting it to sub-75 words,
  600. 18:52leading with relevance, front-loading
  601. 18:54the risk reversal, and ending with a
  602. 18:55concrete time ask. You can see it's
  603. 18:57quite the significant change here. Hey
  604. 18:59first name, I drive PPC leads for a
  605. 19:01two-million-a-year dental marketing firm
  606. 19:02in Calgary. I've sent over 10 million in
  607. 19:03business to agencies like yours through
  608. 19:05cold outbound alone. Got a backlog of
  609. 19:07people wanting PPC right now a variety
  610. 19:08of verticals. I'd send you booked
  611. 19:09appointments and only charge if we hit a
  612. 19:11number you and I agree on beforehand.
  613. 19:13Zero risk on your end. Is this worth a
  614. 19:14quick call or even then gives a specific
  615. 19:16time. So, I mean it remains to be seen
  616. 19:18whether the challenger is going to be
  617. 19:19better than the the baseline, of course,
  618. 19:21but that's just part of the game. And
  619. 19:22you can see this is now actually been
  620. 19:23deployed. We have uh that same copy over
  621. 19:26here. And then if we go over to our
  622. 19:28baseline, we also have the baseline copy
  623. 19:30over here, which is that old cold email.
  624. 19:33And basically what's going to occur now
  625. 19:34is they're just going to test against
  626. 19:35each other until we figure out which one
  627. 19:37is better. Now, on net because this is
  628. 19:39an AI model we're working with, in my
  629. 19:41experience most challengers are not up
  630. 19:43to the task of the baseline. Usually the
  631. 19:45baseline is better because I wrote it,
  632. 19:47but eventually the challengers do become
  633. 19:50better and you start seeing significant
  634. 19:52improvements in the reply rate relative
  635. 19:54to the original, um which then you know
  636. 19:56makes them higher and higher and higher
  637. 19:57and higher over time. Then the
  638. 19:59challenger becomes the new baseline and
  639. 20:01then you just repeat. And so basically
  640. 20:02what this is, to be honest, is like the
  641. 20:04automation of like scientific
  642. 20:05experiments.
  643. 20:07Um
  644. 20:07you know, this is something where right
  645. 20:09now there's so much logistical overhead
  646. 20:11and friction involved in like running
  647. 20:13any sort of experiment, whether you're a
  648. 20:14marketer, a salesperson, somebody doing
  649. 20:16some back-end function or business or
  650. 20:17whatever. And this just eliminates all
  651. 20:19that friction. I no longer have to like
  652. 20:21copy and paste the leads. I no longer
  653. 20:22have to like do anything manually. Um
  654. 20:24it's all done via simple API calls. And
  655. 20:26now that we can put this thing on a
  656. 20:28loop, even though every time I run the
  657. 20:29orchestrator it's technically like a
  658. 20:30different agent, it has all the context
  659. 20:32from all of the previous runs, which
  660. 20:34allows it to grow more intelligent over
  661. 20:35time. I anticipate that eventually after
  662. 20:38something like 500 to maybe 1,000 runs,
  663. 20:41you'll probably have to consolidate some
  664. 20:42of the previous learnings so that that
  665. 20:44document doesn't get super long, but
  666. 20:46whatever you're using this for, whether
  667. 20:47landing pages, PPC, uh newsletter copy,
  668. 20:51whatever the heck, you know, SEO pages,
  669. 20:53hopefully you guys understand that as
  670. 20:54things get better, the new challengers
  671. 20:57are just going to be millions upon
  672. 20:59millions of times more profitable and
  673. 21:01efficacious than uh what your initial
  674. 21:03baseline was. Now, I should note there's
  675. 21:05There's other things that have to do in
  676. 21:06order to set this up completely, like
  677. 21:08for instance, I actually had to go grab
  678. 21:09my API keys. The way you do this on
  679. 21:10Instantly is pretty straightforward,
  680. 21:12settings, then you go integrations, then
  681. 21:13you go API keys down here, then you
  682. 21:15create an API key, say whatever the heck
  683. 21:17you want, select scopes all, and then
  684. 21:18actually copy it over. Um, you obviously
  685. 21:20need to do that as well with whatever AI
  686. 21:22model you're using to do the
  687. 21:23orchestration. In my case, I was using
  688. 21:25Claude Opus 4.6, so I just went over to
  689. 21:27Anthropic, got their API key. Um, and
  690. 21:28then, you know, if you have any other
  691. 21:29other services you want to use like
  692. 21:31GitHub for instance or whatever, you
  693. 21:33also have to push that up. But, agents
  694. 21:35will handle all that stuff for you. Just
  695. 21:37ask them, "Hey, you know, where do I go
  696. 21:38to get my API key? Okay, can you sign me
  697. 21:40in?" and so on and so forth, and you'll
  698. 21:41be good to go. The final thing I want to
  699. 21:43talk about are use cases that I would
  700. 21:44consider not ideal for some form of auto
  701. 21:47optimization. Um, in general, things
  702. 21:49that work really well are things that
  703. 21:51have fast feedback loops, okay? So, why
  704. 21:54did Karpathy's um, AI agent or like nano
  705. 21:57GPT loop work so well? Because it was
  706. 22:00literally a five-minute loop.
  707. 22:02You know, if you have a five-minute
  708. 22:03loop, technically speaking, that means
  709. 22:05that in 60 minutes, you could run 12
  710. 22:08experiments. And so, obviously, 12
  711. 22:10experiments is a lot of data, and
  712. 22:12assuming that, you know, you're running
  713. 22:13it on your own servers or whatever, you
  714. 22:14can just have that thing churn. Um, but
  715. 22:16basically, that means that your your
  716. 22:17iteration loop will be much faster
  717. 22:20because you'll be able to kind of draw
  718. 22:21like this as opposed to like this, you
  719. 22:23know?
  720. 22:24It's going to take a lot longer to
  721. 22:25figure out what works and what doesn't
  722. 22:27if you're all the way down here.
  723. 22:28Another good thing to keep in mind is
  724. 22:30you need a clear metric. So, in my case,
  725. 22:32reply rate was a fantastic metric. Why?
  726. 22:34That's objective. How many people
  727. 22:35actually my email campaigns?
  728. 22:37Click-through rate, very, very objective
  729. 22:39cuz, you know, obviously, this is people
  730. 22:41clicking through an email. All this
  731. 22:42stuff is automatically tracked. But, if
  732. 22:43you had something that was way fuzzier
  733. 22:46in addition to way slower, you know, the
  734. 22:48probability of you actually making this
  735. 22:50work uh, is is much lower because how do
  736. 22:52you subjectively measure like warmth,
  737. 22:55you know? You can't. It's like
  738. 22:57happiness. It's like you can't. What you
  739. 22:58have to do is you have to find proxies
  740. 23:00for all these things, which are usually
  741. 23:01like scales and metrics and analytics
  742. 23:04and so on and so forth.
  743. 23:05And then the third thing is you need
  744. 23:07some sort of API access to change the
  745. 23:08inputs. Um and if you don't have the API
  746. 23:10access, you could build some sort of
  747. 23:12Chrome DevTools or CLI based flow, but
  748. 23:14like you need to have that because if
  749. 23:16you don't, how the heck is the agent
  750. 23:17supposed to make any changes? What are
  751. 23:18they going to do? Just give you a list
  752. 23:19of changes to manually go in? You can do
  753. 23:21that, but that sort of defeats the whole
  754. 23:22purpose.
  755. 23:23Anyway, so what I'm going to do with
  756. 23:25this is I'm going to provide everything
  757. 23:26that you guys need, including the email
  758. 23:28optimizer repo, uh the Carpathy auto
  759. 23:31research GitHub repo, and everything
  760. 23:33else down below. Feel free to take a
  761. 23:35look at it, give it a click, explore it,
  762. 23:37and use it for your own use case.
  763. 23:38I'd be really interested to hear what
  764. 23:40you guys end up using this on.
  765. 23:42Um this is what all major labs that are
  766. 23:44working on machine learning models
  767. 23:45around the world are currently doing, by
  768. 23:47the way, in case it wasn't clear.
  769. 23:48They're constantly running many, many,
  770. 23:50many experiments behind the scenes
  771. 23:52overnight to like make their models
  772. 23:53better and so on and so forth. So, the
  773. 23:55fact that we're able to democratize that
  774. 23:56and now do that for ourselves, for our
  775. 23:58own businesses, and for our own own
  776. 23:59models is now like incredible. But I'd
  777. 24:02be really curious to hear what sort of
  778. 24:03use cases you guys have with us. And,
  779. 24:05you know, if it makes sense, I could
  780. 24:06compile a list of these use cases and
  781. 24:08then I could make another follow-up
  782. 24:09video that just goes through every
  783. 24:10single one and then even gives like real
  784. 24:12examples of them. Um because that'd be
  785. 24:13really dope.
  786. 24:15Aside from that, if you guys could do me
  787. 24:15a big solid, something like 73% of you
  788. 24:18are not subscribed to the channel, which
  789. 24:19really hurts cuz I try and make
  790. 24:20high-quality content for both
  791. 24:22subscribers and non-subscribers, but
  792. 24:23YouTube pushes my content way more
  793. 24:25heavily when that ratio improves. So, if
  794. 24:27if I gave you any value today
  795. 24:28whatsoever, please do click the
  796. 24:29subscribe button. I hate asking for it,
  797. 24:31but it just makes a difference on
  798. 24:32YouTube, so I'll do what works. If you
  799. 24:33guys want more on Claude Code and so on
  800. 24:35and so forth, definitely check that out.
  801. 24:37And yeah, I mean, I will catch all y'all
  802. 24:38in the next video. Thanks so much for
  803. 24:40watching, guys, per usual. See you.

About this transcript

This page contains the full transcript of Claude Code + Karpathy Autoresearch = The New Meta by Nick Saraev, generated from the public captions YouTube serves with the video. The transcript has 5,575 words across 803 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.