This Open Source Repo Solve Claude's #1 Problem — Transcript
Full transcript
- 0:00You cannot trust Claude to grade its own
- 0:02work, which is a huge problem, but one
- 0:04that this skill solves. It's called
- 0:06ClaudeX Loop, and the premise is simple.
- 0:08Instead of having Claude be responsible
- 0:10for planning, executing, and grading its
- 0:12own work, why don't we bring in Codex to
- 0:15also take a look at the plan and the
- 0:17execution of Claude and say, "Hey, this
- 0:20looks good, this doesn't. Here's what I
- 0:22think you should change." Because one of
- 0:23the big problems with every single AI
- 0:25model out there is that they look at
- 0:27their own work very favorably. When I
- 0:28ask Claude how good the plan is that it
- 0:30created, it's going to say, "This thing
- 0:32is awesome." So, it's important we have
- 0:34systems in place where we can bring in a
- 0:36second pair of eyes to look at what the
- 0:38first model built and say, "Thumbs up,
- 0:40thumbs down, and why." So, today I'm not
- 0:42only going to give you this skill, I'm
- 0:44going to break down how it works under
- 0:45the hood, and we'll do a quick demo.
- 0:47Now, the skill is broken out into four
- 0:48phases, and the idea is you invoke it
- 0:50before you add some sort of feature, or
- 0:52if you're starting a brand new
- 0:53greenfield project. And the general idea
- 0:55is whatever the first model comes up
- 0:57with as a plan or whatever it executes
- 0:59on, we wait until the second model takes
- 1:01a look at it and says, "Hey, this looks
- 1:02great before we move forward on big
- 1:04things." Now, for the first phase we're
- 1:06doing some reconnaissance. We're
- 1:07actually going to go out there on the
- 1:08web, Claude is going to scout out to see
- 1:09if the answers actually exist. We have
- 1:11the option to invoke deep research,
- 1:13which is the built-in dynamic workflow
- 1:15if we want to get really deep with the
- 1:17answers we're receiving. After that, we
- 1:19go into the interrogation phase. Think
- 1:20of that as an enhanced plan mode set of
- 1:22questions where we try to get on the
- 1:24same page with Claude before it creates
- 1:26its first plan. Secondly, this is where
- 1:27we bring in the second model. This is
- 1:29the review stage. So, Claude Code, after
- 1:32it comes up with a plan based on
- 1:33everything we've talked about and done
- 1:35our research on, it's going to create
- 1:36your standard plan.md. From there, Codex
- 1:39is going to review the plan in a
- 1:41read-only sandbox, and it's going to
- 1:43say, "Hey, either this is approved, or
- 1:45we need to revise X, Y, and Z." It's
- 1:48then going to send that answer with the
- 1:50revisions to Claude Code. Claude Code's
- 1:52going to take a look at it, say, "Mhm, I
- 1:54agree, I don't agree," and then it's
- 1:55going to send back its changes. This
- 1:58loop will continue up to five times.
- 2:00Now, I've never had it actually stall
- 2:02out at five times where we didn't reach
- 2:04an approved state. However, I have it
- 2:06stopped at five. You can easily change
- 2:08that with the skill so that it doesn't
- 2:10get stuck in some weird endless loop and
- 2:12you're just burning tokens forever and
- 2:13ever. This gives you like a very clear
- 2:15hard wall so you don't get stuck in that
- 2:16situation. Lastly, once Claude and Codex
- 2:19come to an agreement and a plan has been
- 2:20created, we then move into the build
- 2:22phase. Now, this skill gives you the
- 2:24ability to not just have Claude build.
- 2:26You can have Codex start building things
- 2:28as well. But, whatever model builds it,
- 2:30the second model is going to take a look
- 2:32at what it actually created again before
- 2:34it moves forward with anything. And
- 2:35again, under the hood, there's a lot of
- 2:36different variables you can tune. Like I
- 2:38mentioned before, we have the review
- 2:40section which is set to five. In the
- 2:42build section, when we have Codex take a
- 2:44look at what we built, I've knocked that
- 2:45down to two loops. Again, just so we
- 2:48don't get stuck in these sort of endless
- 2:50loop scenarios. And I haven't really run
- 2:52into issues where I felt like, "Oh, I
- 2:53needed to bump this up." But, you can
- 2:55play around with that. Further, you can
- 2:56play around with this more and actually
- 2:57like bring in a local model instead of
- 2:59Codex if that's also how you want to
- 3:01approach this. Now, if you use the
- 3:02previous version of the skill, which was
- 3:04called Grill Me Codex, the changes you
- 3:06want to pay attention to with the
- 3:07ClaudeX loop are a more enhanced
- 3:09interrogation mode so it's going deeper
- 3:11with its line of questioning. And in the
- 3:13execution phase, we've integrated Codex
- 3:15more so again we can get eyes on when it
- 3:17comes to the actual code that's been
- 3:19created. So, next step is the actual
- 3:21demo, but before we jump into that, a
- 3:23quick word from today's sponsor, me. So,
- 3:25I just released a completely updated
- 3:27version of my Claude Code Masterclass
- 3:29inside of Chase AI Plus and it is the
- 3:31number one way to go from zero to AI
- 3:33dev, especially if you don't come from a
- 3:35technical background. We focus on real
- 3:37use cases. This is updated every single
- 3:39week. So, if you're someone who wants to
- 3:40get better at this insane tool and you
- 3:42just don't come from a software
- 3:44development background, this is for you.
- 3:46So, if you want to check it out, there's
- 3:48a link to it in the pin comment. Now,
- 3:50for today's demo, we're going to use
- 3:51ClaudeX loop to recreate Calendly. If
- 3:53you don't know what Calendly is, it's a
- 3:55scheduling web app. You give it a link
- 3:57to people, they can see your calendar,
- 3:58they can pick times, it automatically
- 4:00creates either a Zoom link or a Google
- 4:02Meet. So, it's something that a lot of
- 4:04people use and actually pay for it, but
- 4:06I figured, "Hey, why don't we just
- 4:07create this ourselves and save us, you
- 4:09know, 10 bucks a month or whatever it is
- 4:10I'm paying." I should probably know
- 4:11that. So, what I'm going to do is I just
- 4:13invoke /claudesloop and I'm just going
- 4:16to give it a stream of consciousness and
- 4:17say something like,
- 4:19"I want to use Claude's Loop to
- 4:22essentially create our own version of
- 4:24Calendly. As of right now, I would
- 4:27probably just have it use Google Meet
- 4:29instead of Zoom,
- 4:31um but I'd want it tied with my calendar
- 4:33and essentially just recreate Calendly,
- 4:35all the major features." So, let's go
- 4:38ahead and do that. So, at the beginning,
- 4:39we're in phase zero, which is the
- 4:40research phase. So, it's saying, "Hey,
- 4:42how do you want to actually do this
- 4:43research? Either we do the web search,
- 4:45which is sort of your standard Claude's
- 4:46going to send out a few sub agents, or
- 4:48do we want to go full-blown deep
- 4:50research?" Now, with this skill, I have
- 4:52it pinned for Opus. So, as you know with
- 4:55deep research, if you do it on Fable and
- 4:56you just run /deepresearch, it's going
- 4:58to call Fable sub agents, which can kind
- 5:00of go nuts on your usage. So, as of now,
- 5:02I have it just do Opus off the bat just
- 5:04to sort of help you out. Again, you can
- 5:05totally do that.
- 5:07It's recommending web, but honestly, I'm
- 5:09going to have it do deep just to see
- 5:10what it comes back with. It's then going
- 5:11to show me the proposed deep research
- 5:13prompt and the questions it's actually
- 5:14trying to get answers to. So, it needs
- 5:16to know about Google Calendar, Meet,
- 5:18scheduling domain pitfalls, and
- 5:21generally the stack. So, if I approve
- 5:23that, we'll just say, "Go ahead and
- 5:24launch." And if I wanted to edit it,
- 5:26it's very simple to do. So, Claude
- 5:27finished doing deep research and then
- 5:29created an assumptions ledger, which is
- 5:30essentially a list of everything it
- 5:33assumes you want to happen for this
- 5:36project. Now, after this, it's going to
- 5:37have more questions for us where we can
- 5:39kind of branch down different paths, but
- 5:41these are the things where it's like,
- 5:42"Hey, I I I don't think there's a lot of
- 5:43question here." Now, you can either just
- 5:45give this a thumbs up and it's going to
- 5:47do all this and this is kind of get
- 5:48baked into the plan or you could say,
- 5:50"Hey, I actually want to change the
- 5:52stack or hey, I want to change how we're
- 5:53doing things with reminders and things
- 5:55of that nature." But for now, we're
- 5:56going to say, "Hey, this is confirmed."
- 5:57And then we're going to move into sort
- 5:59of the what it calls the load-bearing
- 6:01tier. You know, we all love the term
- 6:03load-bearing that's now been appended to
- 6:04everything Claude does. So now we're
- 6:06moving into those load-bearing
- 6:08questions. The first question is which
- 6:09Google account does your real calendar
- 6:10live on? For all of these questions, no
- 6:12matter your project, it's going to give
- 6:13you a series of questions alongside of
- 6:16recommendations. So if you have no clue,
- 6:18you can just go with recommended. But
- 6:19frankly, as a best practice, if you have
- 6:21no idea what Claude is telling you when
- 6:23it comes to like potential answers or
- 6:25like what is even going on, I highly
- 6:27suggest you go into this other section
- 6:29and basically just say something like
- 6:30explain this further. This is how you're
- 6:32actually going to get good when it comes
- 6:33to like AI and building things and not
- 6:35just being an accept monkey and hitting
- 6:36recommended and recommended over and
- 6:38over again. If you don't know, ask
- 6:40Claude to continue explaining it to you
- 6:41until you understand it. That's the only
- 6:43way you're actually going to learn. But
- 6:44kind of outside the scope of this video.
- 6:46So we're going to go with personal Gmail
- 6:48and I'm actually going to skip ahead to
- 6:50after I've answered all these questions
- 6:52because I think you kind of get how this
- 6:53phase works. Now, after we answer the
- 6:55load-bearing questions, then we go into
- 6:57some cosmetic decisions that again
- 6:59aren't really changing the base
- 7:01functionality of the application. In
- 7:03this scenario, it just lists them all
- 7:04out similar to the assumption ledger. So
- 7:07you can either say, "This is
- 7:08acceptable." and it's going to just move
- 7:10forward with it or you can say, "Hey,
- 7:11change one, change two, change
- 7:12whatever." I set it up this way so you
- 7:14can just or speed this process along.
- 7:16Now, after we answer those cosmetic
- 7:17questions, we then move into phase two.
- 7:19Claude is going to write that
- 7:20plan.markdown file and then it's going
- 7:23to get sent to Codex using GPT 5.6 soul.
- 7:26From there, they will have their back
- 7:27and forth for a maximum of five rounds
- 7:30or until they reach an accepted verdict.
- 7:32So Claude and Codex went back and forth
- 7:34for five rounds. At the beginning, there
- 7:35were 27 issues that were brought up in
- 7:38round one and we've gotten to round five
- 7:40and the issue is there's still a few
- 7:43more problems. They haven't come to an
- 7:44accepted verdict. I mentioned earlier
- 7:46how I never had this happen before, so
- 7:48I'm kind of glad it did here. So, we've
- 7:49hit five rounds, they haven't figured it
- 7:51out yet. There's still a few sort of
- 7:53just minor issues. What do you want to
- 7:55do? Well, you have options. You can
- 7:56either stop here, keep it as is, you can
- 7:58just accept the deadlock state, or you
- 8:00can extend it two rounds. So, that's
- 8:02what we're going to do and then we'll
- 8:04see what happens when they finally get
- 8:06on the same page. But, nice to know and
- 8:08for you to see that even though we have
- 8:10sort of these hard stops, for example,
- 8:11five rounds, you aren't stuck with that.
- 8:13If you get to that point, like we did
- 8:14here, you can very easily extend it. So,
- 8:16after seven rounds, they've come to an
- 8:18agreement and now it's going to ask you,
- 8:20"Who do you want to actually build this
- 8:21thing?" So, we can either have Claude
- 8:23build and Codex takes a look, or we can
- 8:25flip it and Codex builds and Claude
- 8:27takes a look. In certain situations,
- 8:29depending on what we're doing, it will
- 8:30also give you the option to do sort of a
- 8:32tandem build. So, let's say we were
- 8:33creating something that required like
- 8:35asset generation or something like GPT
- 8:37images too brought into the equation.
- 8:39Well, it will also say like, "Hey, let's
- 8:41bring in Codex for this specific portion
- 8:43of the project." But, in this case,
- 8:45we're going to do Claude builds and now
- 8:47it's going to start creating what it is
- 8:48called open book. So, Claude could
- 8:50finish creating the calendar demo, which
- 8:52is what we're looking at here. So, let's
- 8:53see if it actually works and then we'll
- 8:54go back and actually dive into what
- 8:57Codex added to this entire process. So,
- 8:59we have the intro call and then a
- 9:00working session. So, if we go to the
- 9:02intro call,
- 9:04we can then see a bunch of different
- 9:06times and schedules that are actually
- 9:08synced up to my Gmail calendar. So,
- 9:10let's say I'm picking something for the
- 9:1231st and I pick 10:00 a.m.
- 9:15You can just add my name.
- 9:18We'll put my email.
- 9:21And then we'll do confirm booking.
- 9:23And we can see the email got sent to us
- 9:25with a join link. And I can also see it
- 9:27on my calendar. I also had it create
- 9:29just a simple artifact kind of breaking
- 9:31down visually what Codex added to this
- 9:33entire process. We talked about this
- 9:35earlier, that back and forth phase where
- 9:37Claude had the plan and Codex made its,
- 9:39you know, adjustments or proposed
- 9:41adjustments. That went all the way to
- 9:42seven rounds and it started with 27
- 9:44issues and you can see each round it got
- 9:47lower and lower until round seven we
- 9:49finally had an approved verdict. And now
- 9:51here are some of the things that Codex
- 9:52actually brought to Claude's attention
- 9:54during the planning phase. Stuff like
- 9:56the double booking constraint could not
- 9:57compile. Things like OAuth connect flow
- 10:00having issues. There's also problems
- 10:02with concurrency that two reschedules of
- 10:04the same booking could both succeed and
- 10:06on and on and on. A lot of edge case
- 10:08related stuff. Now in the build phase
- 10:10Claude implemented that improved plan we
- 10:11had after the seven rounds and then we
- 10:13had a new Codex session pop up with
- 10:15completely fresh memory, new context
- 10:17window. It hadn't read the plan and then
- 10:19it read the code against the actual
- 10:21spec, what was created versus what was
- 10:23actually planned. So it came back with
- 10:2523 findings, 19 were accepted and fixed,
- 10:28four were rejected. So here are some of
- 10:30the things that Codex found in this
- 10:32build phase. The time grid drifted after
- 10:34every meeting. The management token was
- 10:36sitting in plain text. All day events
- 10:38blocked the wrong hours. And again, more
- 10:40and more things that again you could
- 10:43argue are edge cases but it would have
- 10:44taken Claude a long time to find them.
- 10:46So if Codex had never been in the room
- 10:48from the beginning, what would this have
- 10:49looked like? Well, we would have had
- 10:50broken features, we would have bookings
- 10:53that exist in the database and nowhere
- 10:54else and so on. Now, in reality, would
- 10:57this have been the case if we just
- 10:59relied on Claude? Probably at the
- 11:00beginning, we just would have taken some
- 11:02further iterations and we eventually
- 11:03would get to something that probably
- 11:05worked but with Codex, we found a lot of
- 11:07these issues in the planning phase so we
- 11:09didn't have to burn tokens and then burn
- 11:11more tokens after the fact and we were
- 11:14able to test these things with Codex,
- 11:16take a look at these with Codex before
- 11:18moving to production. So all in all this
- 11:21saves you a lot of time and money. So
- 11:23that's the Claude-X loop in action. If
- 11:25you want to get your hands on this, I'll
- 11:26put a link to it in the pinned comment
- 11:29and besides that, I'll see you around.
About this transcript
This page contains the full transcript of This Open Source Repo Solve Claude's #1 Problem by Chase AI, generated from the public captions YouTube serves with the video. The transcript has 2,706 words across 374 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.