7 INSANE loops you need to try right now — Transcript
Full transcript
- 0:00Loops are emerging as the single biggest
- 0:02unlock for people building software with
- 0:04artificial intelligence right now. But
- 0:06most people don't even know what loops
- 0:09are. And so today, I'm going to tell you
- 0:11what loops are. I'm going to show you
- 0:13why they're valuable. And then I'm
- 0:14actually going to give you many specific
- 0:18use cases that you can use loops for
- 0:20today. So, what is a loop? A loop is a
- 0:24way to allow your AI coding agent to
- 0:27work autonomously towards a specified
- 0:30goal. The most important thing about
- 0:32loops is that it removes humans. That
- 0:36allows the agent to work much more
- 0:38quickly towards this defined goal. And
- 0:41if it sounds very theoretical, I am
- 0:43going to break it down. So, what is a
- 0:45loop more specifically? Well, you need
- 0:48two things. You need a trigger and you
- 0:51need a goal. With those two things, you
- 0:54can complete the loop. A trigger is what
- 0:57kicks off the loop. And there are three
- 0:59ways to kick off a loop. One, you can do
- 1:02so manually. You literally tell the
- 1:04agent, "Go do this loop." Two is
- 1:08schedule. You can schedule a loop to
- 1:09happen at a certain time of day or on a
- 1:12repeating schedule. And then three, you
- 1:14have actions. You can have the loop kick
- 1:17off based on some kind of action like
- 1:19opening a PR. Now, to fully remove the
- 1:22human, we wouldn't want to kick
- 1:23everything off manually, but sometimes
- 1:25it is required. All right. And for the
- 1:28goal, the goal can be basically one of
- 1:31two things. It can be verifiable or we
- 1:34can use LLM as a judge. So, if it's
- 1:37verifiable, it is something concrete,
- 1:39some specific number or some way to test
- 1:41it deterministically. If it is LLM as a
- 1:44judge, that means we're giving the model
- 1:47the ability to determine when it has
- 1:50reached the goal. Let me give you two
- 1:52examples. So, for verifiable, 100% test
- 1:55coverage in our code base is an example.
- 1:57That is something that we know for sure
- 2:00and we have a nice way to test against
- 2:03when it is true. And for LLM as a judge,
- 2:06one example would be refactor until
- 2:09satisfied. And the satisfaction just
- 2:12means you as the LLM gets to determine
- 2:16when we are satisfactorily
- 2:19refactored enough. All right, enough of
- 2:21the theoretical. Let me actually show
- 2:23you some examples. So, a lot of people
- 2:25talk about loops, but they don't
- 2:26actually give concrete use cases and I
- 2:28wanted to fix this. That is why I am
- 2:31launching the loop library. It is a free
- 2:34library. I'm basically taking all the
- 2:36loops that I use and the ones that I see
- 2:38other people use and putting them in a
- 2:40single place so you can see them. You
- 2:43can be inspired by them to create your
- 2:45own loops or you can simply copy them
- 2:47straight from here. It's free. I'm going
- 2:49to drop the link down below. So, let's
- 2:51go over it. This is definitely my
- 2:53favorite loop and it's going to show you
- 2:56exactly how loops work. This is the
- 2:58sub-50 ms page load loop. Let me click
- 3:01into it. And here we are. So, the
- 3:04objective of this loop is to get every
- 3:07single page load in my app under 50
- 3:11milliseconds. And so, that is the goal.
- 3:13It is a very concrete, well-defined goal
- 3:16which really makes building a loop
- 3:19easier. So, what I tell it is continue
- 3:22optimizing the code for speed after each
- 3:25significant change, measure page load
- 3:27performance across every page under the
- 3:30same repeatable test conditions.
- 3:32Continue until That's the loop. Continue
- 3:36until every page loads in under 50
- 3:39milliseconds. So, it is literally going
- 3:41to go through my entire application,
- 3:43every window, every page, every modal,
- 3:46load it. If it's above 50 milliseconds,
- 3:49it's going to continuously optimize it
- 3:51until it gets it under 50 milliseconds.
- 3:53Once it's done with one, it moves on to
- 3:56the next. That's the loop. That's the
- 3:58goal. But how do I actually do that? How
- 4:01do I actually kick it off? Well, the
- 4:03trigger in this case is me. I am the
- 4:06human and I'm going to manually kick off
- 4:08this loop. You can certainly set it on a
- 4:11schedule and you can even trigger it on,
- 4:14let's say, a PR open. So, every time you
- 4:17open a new PR, you also want to make
- 4:19sure that that new PR doesn't make the
- 4:22page load over 50 milliseconds. So,
- 4:25let's kick it off. So, we're going to
- 4:26click copy right here. All you have to
- 4:28do is paste it in. So, I have the prompt
- 4:30right there and then at the end or at
- 4:32the beginning, it doesn't matter, type
- 4:34{slash} goal. And this is a feature in
- 4:37Codex. Cloud Code also has a {slash}
- 4:40goal feature, but as soon as you have
- 4:41this {slash} goal, it's telling Codex to
- 4:44continue working until the condition is
- 4:47met, the condition of every page loads
- 4:49under 50 milliseconds. That's it. You
- 4:52just hit go and it might run for 10
- 4:54minutes, it might run for 10 hours. It
- 4:57will just continue to run until it meets
- 4:59the goal. And so, you do have to keep a
- 5:01close eye on it if you're under a token
- 5:03budget constraint. So, here it is in
- 5:05action. I sent this as a goal, look for
- 5:07more optimizations to make sure every
- 5:09page loads in under 50 milliseconds on
- 5:12production. It worked for nearly 50
- 5:14minutes. So, I'm treating this as a
- 5:16production performance goal. I'll first
- 5:18measure the real team's page request
- 5:20path and it basically, as you can see
- 5:22here, went through every single page and
- 5:25optimized it to load under 50
- 5:27milliseconds. Loops are the frontier of
- 5:30AI workloads and if you want to power
- 5:32them reliably and at production scale,
- 5:35use the sponsor of today's video,
- 5:38DigitalOcean. If you're running
- 5:39production inference, you're probably
- 5:41running into some of these problems.
- 5:42Your inference stack is too complex to
- 5:44operate, costs are unpredictable, and
- 5:47I'm spending more time managing the
- 5:49infrastructure than actually building
- 5:50the things to be on the infrastructure.
- 5:52And most teams find out the hard way
- 5:54that the hard part of building AI
- 5:56applications is not using the model,
- 5:59it's actually everything around the
- 6:01model. The operational overhead, the
- 6:03fine-tuning inference complexity, the
- 6:05costs that become harder to predict as
- 6:08you scale. And that's why I want to tell
- 6:10you about Digital Ocean, the partner of
- 6:11this video. Digital Ocean is designed to
- 6:14minimize the total cost of ownership by
- 6:17giving teams a simpler path to
- 6:18production AI. They provide
- 6:20infrastructure that is optimized for
- 6:23inference and a vertically integrated
- 6:25core cloud that provides efficiency at
- 6:28scale. Vertically integrated is the
- 6:30keyword. And with transparent
- 6:32usage-based pricing that makes costs
- 6:35easy to predict. So, if you want to
- 6:37spend less time managing your
- 6:38infrastructure and actually building the
- 6:39thing you're excited about, Digital
- 6:42Ocean is the way to go. So, go check it
- 6:44out. They've been a fantastic partner.
- 6:45I've actually been using Digital Ocean
- 6:47for well over a decade at previous
- 6:49companies, so I can vouch for them. Go
- 6:51check them out, link down below. Now,
- 6:53back to the video. Here's another loop
- 6:56that I really like. This is called the
- 6:57overnight docs sweep. Each night, review
- 7:00the codebase in full and make sure all
- 7:01documentation reflects the latest
- 7:03changes from the previous day. Update
- 7:05the documentation as needed, then open a
- 7:07pull request with those changes. So,
- 7:09what I am doing is I'm making sure we
- 7:12have complete documentation based on any
- 7:15changes we may have made. This is an
- 7:17example of LLM as a judge. There's no
- 7:20verifiable way to know if we have
- 7:22complete documentation coverage. There
- 7:24may be some ways that we can say, "Okay,
- 7:27as long as a piece of documentation
- 7:29covers this section of the code." But
- 7:31ultimately, what we're doing is saying,
- 7:32"Okay, LLM, you decide." So, how do we
- 7:35actually use this? Well, once again,
- 7:37just hit the copy button. We're going to
- 7:39come into CodeX. We're going to click
- 7:41this automations tab. We're going to
- 7:43create via chat. We're going to delete
- 7:45this portion. I don't know why they put
- 7:47that in there, but I want to set up an
- 7:48automation. Then we paste in what we
- 7:50just copied, and then each night review
- 7:52the code base in full. Hit go and let it
- 7:54run. And hopefully, it will set up an
- 7:56automation just like this. So, there we
- 7:58go. I'll set this up as a recurring
- 7:59automation. So, first I'm loading the
- 8:01automation tool rather than writing a
- 8:03one-off note. Perfect. So, this is a way
- 8:05to keep your documentation always up to
- 8:08date. It is awesome. And by the way, I
- 8:10created this website with here.now. So,
- 8:13shout out to here.now, the partner on
- 8:16the Loop Library. I created it, and I
- 8:19simply said, "Deploy to here.now." and
- 8:21it was done. It's so easy. Next is the
- 8:24architecture satisfaction loop. This is
- 8:27one that Peter Steinberg himself says he
- 8:30uses often. Here we go. Refactor until
- 8:34you are happy with the architecture.
- 8:36Here is the trigger and the goal all in
- 8:38one sentence. Refactor, which is what
- 8:40the loop is going to do, until you are
- 8:43happy with the architecture. Happy with
- 8:45the architecture is the goal. This is
- 8:47another example of LLM as a judge. We
- 8:50can even give it more guidance on what
- 8:52happy with the architecture means. We
- 8:54can say, "Be very strict about
- 8:57simplicity." or make sure every single
- 8:59line of code is dry. Then, after each
- 9:02significant step, live test the system,
- 9:05run auto review, and commit. Track
- 9:07progress in and then we give it a
- 9:09markdown file to track the progress.
- 9:11This is fantastic. So, it's tracking its
- 9:14loop as it's actually looping. Now, you
- 9:16can kick this off manually, or you can
- 9:18run it every night. So, let's say during
- 9:21the day you're deploying a bunch of
- 9:22code, and then every night you're just
- 9:23making sure that it's refactored, it's
- 9:25dry, and it looks really solid. So, very
- 9:29good way to keep your code base very
- 9:31clean. Next, uh another one of my
- 9:33favorites, the logging coverage loop.
- 9:36So, let's click into it. Basically, what
- 9:38this loop is going to do is make sure
- 9:39that we have thorough logging throughout
- 9:42our app. And there's another loop that
- 9:44builds off of this that I'm going to
- 9:46show you in a minute, which these two
- 9:48loops together, you can start to see how
- 9:50loops can become so powerful. So, this
- 9:52says review the system's logging and add
- 9:54missing coverage until every important
- 9:56path produces useful, tested logs. And
- 10:01again, this just makes sure that we have
- 10:03logging for everything. And this is
- 10:05going to be manually kicked off, and
- 10:07this is going to be LLM as a judge
- 10:09because it says every important path,
- 10:12and important is non-deterministic. It
- 10:15just means the LLM gets to decide what's
- 10:17important and what isn't. And by the
- 10:19way, if you want hands-on help with
- 10:20loops and other AI topics at your
- 10:23company, my team is offering free
- 10:25consulting sessions. I'm going to drop a
- 10:27link down below. We're only doing a few
- 10:29of these, so go apply if you're
- 10:31interested. We'd love to talk to you.
- 10:33All right, so now imagine this. You have
- 10:35full logging coverage, but what do you
- 10:37actually do with those logs? Well, I
- 10:39have another loop for you. This is
- 10:41called the production error sweep. Every
- 10:43single night, we're going to review our
- 10:46production logs for errors. If you find
- 10:48an actionable issue, trace it to its
- 10:50root cause, fix it, verify the fix, and
- 10:54open a pull request. Then ping me in
- 10:56Slack with the findings and PR link. If
- 10:59no actionable errors are present, ping
- 11:01me with that result instead. So, we are
- 11:04kicking off a loop every night, and the
- 11:06loop is looking for every error in the
- 11:08logs and will fix them one by one with
- 11:11the end goal being no more unaddressed
- 11:15errors in the logs. So, that is a very
- 11:17concrete goal for this loop. All right,
- 11:20here's another loop. Something
- 11:21incredibly important to any website
- 11:24owner, any app owner is SEO. And not
- 11:27only SEO, now GEO. So, here's the SEO
- 11:31GEO visibility loop. Run an SEO GEO
- 11:35audit across crawlability, indexation,
- 11:38page intent, titles, internal link
- 11:40structure data, source citations, and
- 11:43answer first content. Rank the gaps. I'm
- 11:46not going to read the whole thing. Fix
- 11:47the highest leverage issues, rerun the
- 11:49same crawl, and here's the loop. Repeat
- 11:52until no critical technical issues
- 11:55remain. Again, you might have one issue,
- 11:58you might have 50 issues. The point is
- 12:01we've now kicked off a loop that fixes
- 12:04all of them until no more issues are
- 12:07present. So, this is a really cool one
- 12:09to run, let's say, once a week. All
- 12:11right, here's one of my favorite and one
- 12:13of the most hand-wavy loops that I have,
- 12:15but listen to this. This is called the
- 12:17full product evaluation loop. Create n
- 12:20realistic scenarios covering every major
- 12:22capability. Before testing, define clear
- 12:24success criteria and choose a consistent
- 12:26evaluation method such as pass-fail
- 12:28checks or a scoring rubric. Run every
- 12:31scenario under the same conditions and
- 12:33record evidence for each outcome. Fix
- 12:35the underlying cause of anything that
- 12:37that does not meet the criteria, rerun
- 12:40the affected scenarios, and then rerun
- 12:43the complete test. Continue until every
- 12:45scenario meets the original quality bar.
- 12:47Now, a lot of you might be thinking,
- 12:49"Wow, that just sounds like tests,
- 12:51right? It's just like a test suite."
- 12:53Well, kind of, but this is actually
- 12:55non-deterministic. This is allowing the
- 12:57model to go through every single use
- 13:00case in your application, in your
- 13:03product, figure out if it's good enough,
- 13:06determined by the LLM, and update it if
- 13:08necessary. This one really does work. It
- 13:12takes like 12 hours at times or more,
- 13:15but it really does come up with very
- 13:17good optimizations. Now, you can also
- 13:19customize this for your specific app.
- 13:22So, for example, I'm building something
- 13:24right now that requires me asking a
- 13:25question of an LLM, and it providing a
- 13:28really accurate response with sources.
- 13:31So, I tell it, "Come up with 100
- 13:34different use cases, wide-ranging use
- 13:36cases, for asking the LLM questions, and
- 13:39judge whether the response is good
- 13:41enough. If it's not, iterate and improve
- 13:43it." So, I could keep going, but if you
- 13:45want to find all of the loops and any
- 13:47new ones that I discover, go check out
- 13:49the Loop Library. I'm going to drop a
- 13:50link down below. And once again, shout
- 13:52out to here.now for hosting the Loop
- 13:55Library. Okay, so there are two major
- 13:58caveats with loops that I have to tell
- 14:00you about. Number one is it's not for
- 14:03every problem yet. Designing a loop
- 14:06isn't always easy. Specifically, coming
- 14:09up with the goal for the loop is not
- 14:12easy. If something can be verified, like
- 14:15every page loads under 50 seconds, that
- 14:17is perfect for a loop. When we have to
- 14:20have the AI judge, LLM as a judge,
- 14:23whether a goal is met or not, that's
- 14:26when it becomes a little more brittle,
- 14:28because we are leaving taste and
- 14:31judgment up to the model. This becomes
- 14:35even more difficult when we're talking
- 14:36about building features. I've not really
- 14:39found a way to build features with
- 14:41loops. You cannot say, "Loop until we
- 14:44build a full permissioning system." I
- 14:47mean, you technically can, but I'm not
- 14:49doing it, because I don't know which
- 14:52direction the AI is going to go. I don't
- 14:54know what features it's going to build.
- 14:55I don't know when or how it's going to
- 14:57decide which features are worthwhile
- 15:00versus which are not. So, that makes it
- 15:03not great from day zero feature
- 15:05building. Now, one example of building a
- 15:08product from scratch using a loop is
- 15:11something I did where I told the model,
- 15:14as a goal, to clone Excel
- 15:17feature parity. And it was running for
- 15:21days and days and days until I finally
- 15:23stopped it. It actually opened up Excel
- 15:25on my computer, used computer use, and
- 15:29literally clicked through and made sure
- 15:31that it had feature parity. And yes, it
- 15:32was running for days before I finally
- 15:34stopped it. So, I do not recommend doing
- 15:36that. And that brings me to the second
- 15:39big caveat. Loops are very expensive.
- 15:42They are churning through tokens
- 15:45autonomously until they hit the goal.
- 15:47Some of these agents might run for 10
- 15:49minutes, some of them can run for days.
- 15:53So, for you token maxers out there,
- 15:55loops are fantastic. But for those of
- 15:57you who don't have an unlimited token
- 15:59budget, this might not work for you
- 16:03today. And by the way, if you like
- 16:04coding with loops, you might also like
- 16:06these four open-source projects that I
- 16:09reviewed that you can use right now.
About this transcript
This page contains the full transcript of 7 INSANE loops you need to try right now by Matthew Berman, generated from the public captions YouTube serves with the video. The transcript has 2,850 words across 406 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.