YouTube2Text

7 INSANE loops you need to try right now — Transcript

by Matthew Berman · 2,850 words · 406 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Loops are emerging as the single biggest
  2. 0:02unlock for people building software with
  3. 0:04artificial intelligence right now. But
  4. 0:06most people don't even know what loops
  5. 0:09are. And so today, I'm going to tell you
  6. 0:11what loops are. I'm going to show you
  7. 0:13why they're valuable. And then I'm
  8. 0:14actually going to give you many specific
  9. 0:18use cases that you can use loops for
  10. 0:20today. So, what is a loop? A loop is a
  11. 0:24way to allow your AI coding agent to
  12. 0:27work autonomously towards a specified
  13. 0:30goal. The most important thing about
  14. 0:32loops is that it removes humans. That
  15. 0:36allows the agent to work much more
  16. 0:38quickly towards this defined goal. And
  17. 0:41if it sounds very theoretical, I am
  18. 0:43going to break it down. So, what is a
  19. 0:45loop more specifically? Well, you need
  20. 0:48two things. You need a trigger and you
  21. 0:51need a goal. With those two things, you
  22. 0:54can complete the loop. A trigger is what
  23. 0:57kicks off the loop. And there are three
  24. 0:59ways to kick off a loop. One, you can do
  25. 1:02so manually. You literally tell the
  26. 1:04agent, "Go do this loop." Two is
  27. 1:08schedule. You can schedule a loop to
  28. 1:09happen at a certain time of day or on a
  29. 1:12repeating schedule. And then three, you
  30. 1:14have actions. You can have the loop kick
  31. 1:17off based on some kind of action like
  32. 1:19opening a PR. Now, to fully remove the
  33. 1:22human, we wouldn't want to kick
  34. 1:23everything off manually, but sometimes
  35. 1:25it is required. All right. And for the
  36. 1:28goal, the goal can be basically one of
  37. 1:31two things. It can be verifiable or we
  38. 1:34can use LLM as a judge. So, if it's
  39. 1:37verifiable, it is something concrete,
  40. 1:39some specific number or some way to test
  41. 1:41it deterministically. If it is LLM as a
  42. 1:44judge, that means we're giving the model
  43. 1:47the ability to determine when it has
  44. 1:50reached the goal. Let me give you two
  45. 1:52examples. So, for verifiable, 100% test
  46. 1:55coverage in our code base is an example.
  47. 1:57That is something that we know for sure
  48. 2:00and we have a nice way to test against
  49. 2:03when it is true. And for LLM as a judge,
  50. 2:06one example would be refactor until
  51. 2:09satisfied. And the satisfaction just
  52. 2:12means you as the LLM gets to determine
  53. 2:16when we are satisfactorily
  54. 2:19refactored enough. All right, enough of
  55. 2:21the theoretical. Let me actually show
  56. 2:23you some examples. So, a lot of people
  57. 2:25talk about loops, but they don't
  58. 2:26actually give concrete use cases and I
  59. 2:28wanted to fix this. That is why I am
  60. 2:31launching the loop library. It is a free
  61. 2:34library. I'm basically taking all the
  62. 2:36loops that I use and the ones that I see
  63. 2:38other people use and putting them in a
  64. 2:40single place so you can see them. You
  65. 2:43can be inspired by them to create your
  66. 2:45own loops or you can simply copy them
  67. 2:47straight from here. It's free. I'm going
  68. 2:49to drop the link down below. So, let's
  69. 2:51go over it. This is definitely my
  70. 2:53favorite loop and it's going to show you
  71. 2:56exactly how loops work. This is the
  72. 2:58sub-50 ms page load loop. Let me click
  73. 3:01into it. And here we are. So, the
  74. 3:04objective of this loop is to get every
  75. 3:07single page load in my app under 50
  76. 3:11milliseconds. And so, that is the goal.
  77. 3:13It is a very concrete, well-defined goal
  78. 3:16which really makes building a loop
  79. 3:19easier. So, what I tell it is continue
  80. 3:22optimizing the code for speed after each
  81. 3:25significant change, measure page load
  82. 3:27performance across every page under the
  83. 3:30same repeatable test conditions.
  84. 3:32Continue until That's the loop. Continue
  85. 3:36until every page loads in under 50
  86. 3:39milliseconds. So, it is literally going
  87. 3:41to go through my entire application,
  88. 3:43every window, every page, every modal,
  89. 3:46load it. If it's above 50 milliseconds,
  90. 3:49it's going to continuously optimize it
  91. 3:51until it gets it under 50 milliseconds.
  92. 3:53Once it's done with one, it moves on to
  93. 3:56the next. That's the loop. That's the
  94. 3:58goal. But how do I actually do that? How
  95. 4:01do I actually kick it off? Well, the
  96. 4:03trigger in this case is me. I am the
  97. 4:06human and I'm going to manually kick off
  98. 4:08this loop. You can certainly set it on a
  99. 4:11schedule and you can even trigger it on,
  100. 4:14let's say, a PR open. So, every time you
  101. 4:17open a new PR, you also want to make
  102. 4:19sure that that new PR doesn't make the
  103. 4:22page load over 50 milliseconds. So,
  104. 4:25let's kick it off. So, we're going to
  105. 4:26click copy right here. All you have to
  106. 4:28do is paste it in. So, I have the prompt
  107. 4:30right there and then at the end or at
  108. 4:32the beginning, it doesn't matter, type
  109. 4:34{slash} goal. And this is a feature in
  110. 4:37Codex. Cloud Code also has a {slash}
  111. 4:40goal feature, but as soon as you have
  112. 4:41this {slash} goal, it's telling Codex to
  113. 4:44continue working until the condition is
  114. 4:47met, the condition of every page loads
  115. 4:49under 50 milliseconds. That's it. You
  116. 4:52just hit go and it might run for 10
  117. 4:54minutes, it might run for 10 hours. It
  118. 4:57will just continue to run until it meets
  119. 4:59the goal. And so, you do have to keep a
  120. 5:01close eye on it if you're under a token
  121. 5:03budget constraint. So, here it is in
  122. 5:05action. I sent this as a goal, look for
  123. 5:07more optimizations to make sure every
  124. 5:09page loads in under 50 milliseconds on
  125. 5:12production. It worked for nearly 50
  126. 5:14minutes. So, I'm treating this as a
  127. 5:16production performance goal. I'll first
  128. 5:18measure the real team's page request
  129. 5:20path and it basically, as you can see
  130. 5:22here, went through every single page and
  131. 5:25optimized it to load under 50
  132. 5:27milliseconds. Loops are the frontier of
  133. 5:30AI workloads and if you want to power
  134. 5:32them reliably and at production scale,
  135. 5:35use the sponsor of today's video,
  136. 5:38DigitalOcean. If you're running
  137. 5:39production inference, you're probably
  138. 5:41running into some of these problems.
  139. 5:42Your inference stack is too complex to
  140. 5:44operate, costs are unpredictable, and
  141. 5:47I'm spending more time managing the
  142. 5:49infrastructure than actually building
  143. 5:50the things to be on the infrastructure.
  144. 5:52And most teams find out the hard way
  145. 5:54that the hard part of building AI
  146. 5:56applications is not using the model,
  147. 5:59it's actually everything around the
  148. 6:01model. The operational overhead, the
  149. 6:03fine-tuning inference complexity, the
  150. 6:05costs that become harder to predict as
  151. 6:08you scale. And that's why I want to tell
  152. 6:10you about Digital Ocean, the partner of
  153. 6:11this video. Digital Ocean is designed to
  154. 6:14minimize the total cost of ownership by
  155. 6:17giving teams a simpler path to
  156. 6:18production AI. They provide
  157. 6:20infrastructure that is optimized for
  158. 6:23inference and a vertically integrated
  159. 6:25core cloud that provides efficiency at
  160. 6:28scale. Vertically integrated is the
  161. 6:30keyword. And with transparent
  162. 6:32usage-based pricing that makes costs
  163. 6:35easy to predict. So, if you want to
  164. 6:37spend less time managing your
  165. 6:38infrastructure and actually building the
  166. 6:39thing you're excited about, Digital
  167. 6:42Ocean is the way to go. So, go check it
  168. 6:44out. They've been a fantastic partner.
  169. 6:45I've actually been using Digital Ocean
  170. 6:47for well over a decade at previous
  171. 6:49companies, so I can vouch for them. Go
  172. 6:51check them out, link down below. Now,
  173. 6:53back to the video. Here's another loop
  174. 6:56that I really like. This is called the
  175. 6:57overnight docs sweep. Each night, review
  176. 7:00the codebase in full and make sure all
  177. 7:01documentation reflects the latest
  178. 7:03changes from the previous day. Update
  179. 7:05the documentation as needed, then open a
  180. 7:07pull request with those changes. So,
  181. 7:09what I am doing is I'm making sure we
  182. 7:12have complete documentation based on any
  183. 7:15changes we may have made. This is an
  184. 7:17example of LLM as a judge. There's no
  185. 7:20verifiable way to know if we have
  186. 7:22complete documentation coverage. There
  187. 7:24may be some ways that we can say, "Okay,
  188. 7:27as long as a piece of documentation
  189. 7:29covers this section of the code." But
  190. 7:31ultimately, what we're doing is saying,
  191. 7:32"Okay, LLM, you decide." So, how do we
  192. 7:35actually use this? Well, once again,
  193. 7:37just hit the copy button. We're going to
  194. 7:39come into CodeX. We're going to click
  195. 7:41this automations tab. We're going to
  196. 7:43create via chat. We're going to delete
  197. 7:45this portion. I don't know why they put
  198. 7:47that in there, but I want to set up an
  199. 7:48automation. Then we paste in what we
  200. 7:50just copied, and then each night review
  201. 7:52the code base in full. Hit go and let it
  202. 7:54run. And hopefully, it will set up an
  203. 7:56automation just like this. So, there we
  204. 7:58go. I'll set this up as a recurring
  205. 7:59automation. So, first I'm loading the
  206. 8:01automation tool rather than writing a
  207. 8:03one-off note. Perfect. So, this is a way
  208. 8:05to keep your documentation always up to
  209. 8:08date. It is awesome. And by the way, I
  210. 8:10created this website with here.now. So,
  211. 8:13shout out to here.now, the partner on
  212. 8:16the Loop Library. I created it, and I
  213. 8:19simply said, "Deploy to here.now." and
  214. 8:21it was done. It's so easy. Next is the
  215. 8:24architecture satisfaction loop. This is
  216. 8:27one that Peter Steinberg himself says he
  217. 8:30uses often. Here we go. Refactor until
  218. 8:34you are happy with the architecture.
  219. 8:36Here is the trigger and the goal all in
  220. 8:38one sentence. Refactor, which is what
  221. 8:40the loop is going to do, until you are
  222. 8:43happy with the architecture. Happy with
  223. 8:45the architecture is the goal. This is
  224. 8:47another example of LLM as a judge. We
  225. 8:50can even give it more guidance on what
  226. 8:52happy with the architecture means. We
  227. 8:54can say, "Be very strict about
  228. 8:57simplicity." or make sure every single
  229. 8:59line of code is dry. Then, after each
  230. 9:02significant step, live test the system,
  231. 9:05run auto review, and commit. Track
  232. 9:07progress in and then we give it a
  233. 9:09markdown file to track the progress.
  234. 9:11This is fantastic. So, it's tracking its
  235. 9:14loop as it's actually looping. Now, you
  236. 9:16can kick this off manually, or you can
  237. 9:18run it every night. So, let's say during
  238. 9:21the day you're deploying a bunch of
  239. 9:22code, and then every night you're just
  240. 9:23making sure that it's refactored, it's
  241. 9:25dry, and it looks really solid. So, very
  242. 9:29good way to keep your code base very
  243. 9:31clean. Next, uh another one of my
  244. 9:33favorites, the logging coverage loop.
  245. 9:36So, let's click into it. Basically, what
  246. 9:38this loop is going to do is make sure
  247. 9:39that we have thorough logging throughout
  248. 9:42our app. And there's another loop that
  249. 9:44builds off of this that I'm going to
  250. 9:46show you in a minute, which these two
  251. 9:48loops together, you can start to see how
  252. 9:50loops can become so powerful. So, this
  253. 9:52says review the system's logging and add
  254. 9:54missing coverage until every important
  255. 9:56path produces useful, tested logs. And
  256. 10:01again, this just makes sure that we have
  257. 10:03logging for everything. And this is
  258. 10:05going to be manually kicked off, and
  259. 10:07this is going to be LLM as a judge
  260. 10:09because it says every important path,
  261. 10:12and important is non-deterministic. It
  262. 10:15just means the LLM gets to decide what's
  263. 10:17important and what isn't. And by the
  264. 10:19way, if you want hands-on help with
  265. 10:20loops and other AI topics at your
  266. 10:23company, my team is offering free
  267. 10:25consulting sessions. I'm going to drop a
  268. 10:27link down below. We're only doing a few
  269. 10:29of these, so go apply if you're
  270. 10:31interested. We'd love to talk to you.
  271. 10:33All right, so now imagine this. You have
  272. 10:35full logging coverage, but what do you
  273. 10:37actually do with those logs? Well, I
  274. 10:39have another loop for you. This is
  275. 10:41called the production error sweep. Every
  276. 10:43single night, we're going to review our
  277. 10:46production logs for errors. If you find
  278. 10:48an actionable issue, trace it to its
  279. 10:50root cause, fix it, verify the fix, and
  280. 10:54open a pull request. Then ping me in
  281. 10:56Slack with the findings and PR link. If
  282. 10:59no actionable errors are present, ping
  283. 11:01me with that result instead. So, we are
  284. 11:04kicking off a loop every night, and the
  285. 11:06loop is looking for every error in the
  286. 11:08logs and will fix them one by one with
  287. 11:11the end goal being no more unaddressed
  288. 11:15errors in the logs. So, that is a very
  289. 11:17concrete goal for this loop. All right,
  290. 11:20here's another loop. Something
  291. 11:21incredibly important to any website
  292. 11:24owner, any app owner is SEO. And not
  293. 11:27only SEO, now GEO. So, here's the SEO
  294. 11:31GEO visibility loop. Run an SEO GEO
  295. 11:35audit across crawlability, indexation,
  296. 11:38page intent, titles, internal link
  297. 11:40structure data, source citations, and
  298. 11:43answer first content. Rank the gaps. I'm
  299. 11:46not going to read the whole thing. Fix
  300. 11:47the highest leverage issues, rerun the
  301. 11:49same crawl, and here's the loop. Repeat
  302. 11:52until no critical technical issues
  303. 11:55remain. Again, you might have one issue,
  304. 11:58you might have 50 issues. The point is
  305. 12:01we've now kicked off a loop that fixes
  306. 12:04all of them until no more issues are
  307. 12:07present. So, this is a really cool one
  308. 12:09to run, let's say, once a week. All
  309. 12:11right, here's one of my favorite and one
  310. 12:13of the most hand-wavy loops that I have,
  311. 12:15but listen to this. This is called the
  312. 12:17full product evaluation loop. Create n
  313. 12:20realistic scenarios covering every major
  314. 12:22capability. Before testing, define clear
  315. 12:24success criteria and choose a consistent
  316. 12:26evaluation method such as pass-fail
  317. 12:28checks or a scoring rubric. Run every
  318. 12:31scenario under the same conditions and
  319. 12:33record evidence for each outcome. Fix
  320. 12:35the underlying cause of anything that
  321. 12:37that does not meet the criteria, rerun
  322. 12:40the affected scenarios, and then rerun
  323. 12:43the complete test. Continue until every
  324. 12:45scenario meets the original quality bar.
  325. 12:47Now, a lot of you might be thinking,
  326. 12:49"Wow, that just sounds like tests,
  327. 12:51right? It's just like a test suite."
  328. 12:53Well, kind of, but this is actually
  329. 12:55non-deterministic. This is allowing the
  330. 12:57model to go through every single use
  331. 13:00case in your application, in your
  332. 13:03product, figure out if it's good enough,
  333. 13:06determined by the LLM, and update it if
  334. 13:08necessary. This one really does work. It
  335. 13:12takes like 12 hours at times or more,
  336. 13:15but it really does come up with very
  337. 13:17good optimizations. Now, you can also
  338. 13:19customize this for your specific app.
  339. 13:22So, for example, I'm building something
  340. 13:24right now that requires me asking a
  341. 13:25question of an LLM, and it providing a
  342. 13:28really accurate response with sources.
  343. 13:31So, I tell it, "Come up with 100
  344. 13:34different use cases, wide-ranging use
  345. 13:36cases, for asking the LLM questions, and
  346. 13:39judge whether the response is good
  347. 13:41enough. If it's not, iterate and improve
  348. 13:43it." So, I could keep going, but if you
  349. 13:45want to find all of the loops and any
  350. 13:47new ones that I discover, go check out
  351. 13:49the Loop Library. I'm going to drop a
  352. 13:50link down below. And once again, shout
  353. 13:52out to here.now for hosting the Loop
  354. 13:55Library. Okay, so there are two major
  355. 13:58caveats with loops that I have to tell
  356. 14:00you about. Number one is it's not for
  357. 14:03every problem yet. Designing a loop
  358. 14:06isn't always easy. Specifically, coming
  359. 14:09up with the goal for the loop is not
  360. 14:12easy. If something can be verified, like
  361. 14:15every page loads under 50 seconds, that
  362. 14:17is perfect for a loop. When we have to
  363. 14:20have the AI judge, LLM as a judge,
  364. 14:23whether a goal is met or not, that's
  365. 14:26when it becomes a little more brittle,
  366. 14:28because we are leaving taste and
  367. 14:31judgment up to the model. This becomes
  368. 14:35even more difficult when we're talking
  369. 14:36about building features. I've not really
  370. 14:39found a way to build features with
  371. 14:41loops. You cannot say, "Loop until we
  372. 14:44build a full permissioning system." I
  373. 14:47mean, you technically can, but I'm not
  374. 14:49doing it, because I don't know which
  375. 14:52direction the AI is going to go. I don't
  376. 14:54know what features it's going to build.
  377. 14:55I don't know when or how it's going to
  378. 14:57decide which features are worthwhile
  379. 15:00versus which are not. So, that makes it
  380. 15:03not great from day zero feature
  381. 15:05building. Now, one example of building a
  382. 15:08product from scratch using a loop is
  383. 15:11something I did where I told the model,
  384. 15:14as a goal, to clone Excel
  385. 15:17feature parity. And it was running for
  386. 15:21days and days and days until I finally
  387. 15:23stopped it. It actually opened up Excel
  388. 15:25on my computer, used computer use, and
  389. 15:29literally clicked through and made sure
  390. 15:31that it had feature parity. And yes, it
  391. 15:32was running for days before I finally
  392. 15:34stopped it. So, I do not recommend doing
  393. 15:36that. And that brings me to the second
  394. 15:39big caveat. Loops are very expensive.
  395. 15:42They are churning through tokens
  396. 15:45autonomously until they hit the goal.
  397. 15:47Some of these agents might run for 10
  398. 15:49minutes, some of them can run for days.
  399. 15:53So, for you token maxers out there,
  400. 15:55loops are fantastic. But for those of
  401. 15:57you who don't have an unlimited token
  402. 15:59budget, this might not work for you
  403. 16:03today. And by the way, if you like
  404. 16:04coding with loops, you might also like
  405. 16:06these four open-source projects that I
  406. 16:09reviewed that you can use right now.

About this transcript

This page contains the full transcript of 7 INSANE loops you need to try right now by Matthew Berman, generated from the public captions YouTube serves with the video. The transcript has 2,850 words across 406 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.